Your AI demo is a lie (and how to make it real) | Arcade’s Alex Salazar

9 Sep 2025 · 1 h 2 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Dev Interrupted - Your AI Demo is a Lie (and how to make it real) | Arcade’s Alex Salazar

Episode Overview In this episode of Dev Interrupted, hosts Andrew Zigler and Ben Lloyd Pearson are joined by Alex Salazar, CEO of Arcade, to discuss the dichotomy between impressive AI demos and the complexities of creating production-ready AI systems. Salazar emphasizes the often underestimated challenges developers face when transitioning from flashy prototypes to secure, reliable AI applications.

Key Points Discussed

  • The gap between AI demos and real-world applications.
  • The importance of security and reliability in AI systems.
  • The "four demo killers" that lead to the failure of AI agent projects.
  • An innovative approach of increasing determinism to improve AI reliability.

---

Hosts' Introduction

  • Andrew Zigler & Ben Lloyd Pearson:
  • Welcomed back Ben after a summer break, discussing personal experiences with AI (including Ben's lighthearted failures while using AI for organizing Lego).
  • Transitioned into current industry news, setting the stage for the interview with Alex Salazar.

Industry News Highlights

  1. Google's Browser Ownership:
  2. Google avoids a breakup with Chrome after recent rulings regarding its monopoly.
  3. Discussion on how AI is changing browser dynamics, with potential for innovation in this area.
  1. Atlassian Acquires a Browser Company:
  2. Atlassian's purchase aims to create an AI-powered browser tailored for developers.
  3. Importance of rethinking browser functionality to enhance productivity.
  1. Coinbase's AI-Generated Code Claims:
  2. Discussion on Brian Armstrong's claim that 40% of daily code at Coinbase is AI-generated.
  3. Concerns around the implications of such metrics and their potential to mislead.
  1. JIRA Ticket Management:
  2. A critique of focusing on ticket closure rather than meaningful impact.
  3. Reference to an article by Sean Goedecke discussing the importance of aligning engineering efforts with organizational goals.

---

Main Discussion with Alex Salazar

AI Demos vs. Production Reality

  • Salazar highlights that a working demo represents only 1% of the journey toward a reliable AI system.
  • Common pitfalls include:
  • Inconsistency: Demos often lack rigorous evaluations that would be necessary in production environments.
  • Security Flaws: Real-world applications require stringent security measures that demos typically bypass.
  • Prohibitive Costs: What appears inexpensive in a demo may become costly when scaled.
  • High Latency: The complexity of AI systems can slow down operations, making them impractical for real-time applications.

Solutions to Demo Killers

  • Increasing Determinism:
  • Salazar’s team discovered that constraining AI capabilities (e.g., using intention-based tools like calculators or multiple-choice questions) can significantly reduce errors and enhance user experience.
  • Transitioning from open-ended systems to structured, deterministic workflows improves reliability.
  • Custom Workflow-Centric Tools:
  • The need for bespoke tools rather than relying solely on existing APIs.
  • Emphasis on building systems that prioritize security and workflow efficiency.

Development Strategies

  1. Start with Production in Mind: Always plan for the realities of deploying AI systems, not just for demos.
  2. Build Evaluation Processes: Create robust evaluation frameworks early to measure AI performance consistently.
  3. De-scope Projects: Focus on a single objective or workflow at a time to avoid overwhelming complexity.
  4. Utilize Modern APIs: Build on top of existing systems with modern APIs to simplify integration and reduce risk.

Company Culture and Team Dynamics

  • Salazar stresses the importance of hiring individuals who are agent native—those who have experience taking AI projects to production.
  • A successful team has members with diverse expertise that can cross-pollinate ideas and solutions, fostering an innovative environment.

---

Conclusion

  • The episode concludes with a call for engineers to engage practically with AI technology and learn through building.
  • Encouragement for teams to share insights and experiences to collectively navigate the evolving landscape of AI.

Resources Mentioned

  • Arcade: [Arcade.dev](https://www.arcade.dev)
  • Arcade YouTube Channel: [Watch examples and walkthroughs](https://www.youtube.com/@TryArcade)
  • LinkedIn for Alex Salazar: [Connect here](https://www.linkedin.com/in/alexsalazar/)

Additional Notes

  • The conversation highlights the critical shift from hype and demos to focused, practical engineering in AI development.
  • The episode is valuable for developers, engineers, and leaders in tech looking to bridge the gap between AI concepts and real-world applications.

---

Follow Dev Interrupted

  • Subscribe and follow for more insights on software engineering leadership and AI in development:
  • [Dev Interrupted Substack](https://devinterrupted.substack.com/)
  • [YouTube Channel](https://www.youtube.com/c/DevInterrupted)
  • [Twitter](https://twitter.com/DevInterrupted) and [LinkedIn](https://www.linkedin.com/showcase/dev-interrupted/)

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Welcome to Dev Interrupted. I'm your host Andrew Ziegler. And I'm your host, Ben Lloyd Pearson. This week, we're bringing Alex Salazar, CEO of Arcade.dev, on the show to discuss why a working AI demo is only 1 % of the journey. Because the gap between a flashy proof-of-concept and a secure production-ready system is huge. And giving AI less freedom is the surprising key to making it reliable and safe for your users. But first, Ben, welcome back. Yeah, I'm officially back. I've been out for a few weeks. You know, my wife and I, we welcomed our second son into the world, Arthur Gene Pearson. Had a really great summer break.

0:47Took some time off, spent it with my family. I feel like I came back with some lessons learned. You know, had a bit of an AI cleanse. It was really nice, actually. An AI cleanse? That doesn't sound like the bed I know. Yeah, I know, right? Yeah. Yeah, I practically have not touched AI at all, pretty much, since I left. There was one exception and it failed spectacularly, which is kind of hilarious. So one of the many projects I took on while I was out was I helped my four year old organize all of his Lego. And I had a couple of times where I just like asked Chad GPT, like, hey, what is this Lego that I don't recognize?

1:25And it just like it completely misidentified it and like was like not helpful whatsoever. whatsoever. And then another time, I think I sent it a picture of like a car I was building and was like, hey, help me design like a front end for this car. And it gave me some like elaborate ASCII diagrams that like made no sense whatsoever. And I'm like completely unconvinced that it was like a functional design at all. But, you know, I guess I have two lessons that I've learned from my time off. You know, not every problem needs to be solved with AI. And, you know, sometimes spending a little time disconnected from it all, like AI and just technology in general, is just a really good thing.

2:05So, you know, I encourage everyone out there, like just whether it's AI or it's just technology in general, like get disconnected a little bit, spend some time with your family or people you care about. It's been a really great summer for me. How was your summer, Andrew? You had a lot of guests on the show while it was out. It was really great to listen. Oh, yeah. You know, just while you've been out, we've, we just had a few folks you know come and keep your seat warm and and keep me company because i was by myself here and so uh you're huge thank you to all of the guests uh who did come on the show to to help fill in new segments i'm really looking forward to bringing more back in the future maybe giving bed more breaks to you know go uh play with more legos but i feel having heard your story i i kind of feel actually the opposite now i feel like we're going to be like the angel and the devil sitting on like our listeners shoulders like advocating these two sides of the world that we're all in right now because I mean gosh just last week I was on site all day at Code TV in Portland filming a vibe coding competition I was competing in so I just couldn't feel more different in terms of my AI usage and your story too is um is funny to me because I wonder now uh in my head like is the AI seeing Lego pieces or is it seeing a car it's kind of an interesting question yeah um yeah with that said if we got anyone from Lego who's listening to this I would love to help you all figure out how to train these models on lego to to help parents like me everywhere because wow i've got a whole bunch of organized lego now and no ai to help me make sense of it all so for now i have to rely on you know antiquated instruction manuals i guess classic yeah so with that said let's just get right back into it i'm excited to get into the news so what do we have on the the docket for today andrew oh gosh well we have some really interesting developments.

3:49One of the first ones that was big news for folks is Google avoiding a breakup with Chrome. So we've covered this before about how there was a ruling where Google had to sell Chrome and divest from its control of the browser in a monopoly ruling in U.S. courts, right? And this has been a really interesting change in the development here where they no longer have to give up that browser. So back when that happened, we all talked about how this was a shift in like the powers that be around browsers, which have a pretty defined market share at this point in terms of adoption, Chrome kind of eating all the rest of the pack.

4:28It put the future of browsers in question at a time when AI was already transforming all the tools we're using every day, including the one used by our guests last week, Zach Lloyd of Warp tackling the terminal, right? classic tools that are getting flipped on their head with AI and the browser is no different to that. So what did you think of this story, Ben? Yeah, I mean, this ruling, or I guess almost like a reversal of a previous ruling, it feels like a rug pull for the ages. Like at this point, I almost don't even know what to think or what to say about it because I feel like this isn't even going to be the end of it at this point.

5:04And it's been pretty crazy watching this drama unfold over the past few months because you know there really weren't a lot of buyers out there that made a whole lot of sense for chrome like it really didn't seem like there were going to be like anyone that materialized that that would pay what it's worth and would make sense to have that be the company versus you know google being the one that owns it like i think perplexity came out of the woodwork at one point and offered some some crazy amount of money for it it's just yeah kind of nuts so i mean my take is that you know instead of like just completely cleaving off chrome now there's like more of a nuanced restriction on what google can do with chrome moving forward so there's like restrictions on their ability to enter into like exclusivity agreements with like mobile phone producers they're also going to be required to share data about search rankings and basic user data with some of their competitors but i you know honestly i don't think this is we're anywhere near the end of this.

6:02I think this is going to get drug out for quite a while, and maybe even to a point where it starts to get disrupted by some of the more like AI native upstarts that we're seeing emerge in this space. Like I've thought for a while that browsers are probably going to get disrupted by AI because it's something that is, you know, it's an older technology that is ripe for disruption. And I think, you know, maybe before this really even takes effect, Chrome and Google may have an argument that they don't really have a monopoly because they they they're getting disrupted by these new technologies. And speaking of which, we have another news story that is right up this alley involving Atlassian buying an AI powered browser.

6:46So what's going on here, Andrew? Yeah, this is the latest, you know, immediately almost after this ruling just like days later or that really just like two days later, we hear a great announcement from Atlassian of tackling this exact problem that you and I are talking about, about the browser being ripe for disruption because it doesn't match the world that we live in and the way that we consume information. And it's even when you think about like how AI has eaten search, it makes sense that AI would then also start to eat the browser. What do you do in a browser? You're searching, you're looking for things.

7:17And so the latest announcement from Atlassian is they purchased the browser company and their bid to build the AI-powered browser of the future. And this is a browser that it will be power built for people that are developer minded and using new technologies. And so it's going to be embedded really closely with all of the things that we're already familiar with from the Atlassian ecosystem in terms of its breadth and its depth and its security. Right. And so this is a really interesting new development. I'm kind of curious to see how it's going to connect with everything else in the Atlassian world as well.

7:51Yeah, it seems like Atlassian is trying to pursue that like AI native web browser experience, you know, and the argument that they are making some of their promotional materials is that like browsers were kind of designed to be a consumption device. Like you consume a recipe off a cooking website or you consume an article that you're reading or you consume your email. And a lot of the productivity tools that we use were kind of tacked on to it as almost like an afterthought. It's not quite the right way to describe it, but it is kind of how they've been approached. And I absolutely love niche web browsers, so I wholly support anyone and everyone entering this space and trying to compete with it because I think there's plenty of room for competition.

8:35It is hard to compete when you have an established player like Chrome and Apple with Safari. But I love to see more companies in this space. But I have always kind of thought that browsers are like fundamentally not very well designed for professional work because they kind of treat all of our tools as separate entities rather than like a connected workflow that has to sort of work together. And, you know, just as an aside, I kind of think this is also a big part of why AI is often overhyped and is not as successful as people want it to be. Because we have such a strong tool disconnect, you know, whether it's just knowledge work in general, but also with software development.

9:16Like this is also a very common thing that happens. There's this tool disconnect that makes it, creates like sort of like this context chasm that's actually really tough for AI to bridge. Like humans have sort of adapted to get pretty good at bridging that. But AI really kind of struggles with that. And I think that's a part of it. So and I think Atlassian is the perfect type of company to like take on this type of challenge. And it makes sense for them to acquire a company like the Browser Company of New York as an opportunity to innovate, you know, versus like what we've seen some other companies in the AI space do where they use acquisition as a strategy to lock out.

9:59competition rather than as a strategy to actually like create new innovation. But this is, I mean, this is going to be an incredible challenge for them. You know, they're trying to rethink browsers for AI driven knowledge work. They're going to have to rethink user behaviors, like not just innovating new products. Like they're actually going to have to rethink how people interact with the web, which I think is really hard to change behaviors. You know, companies like Apple work really hard to make that happen. And not a lot of companies are successful at pulling that off. So, you know, I'm going to be watching it very closely.

10:37Like, I'm excited for it. I don't know, you know, what's going to come out of it. But, you know, maybe we'll try it out for Dev Interrupted at some point. Like, I think it's a pretty cool development. So let's move on to the next story, which is yet another executive from a tech company talking about how awesome their company is with AI. So what's going on here, Andrew? Yeah, so this is a tweet from Brian Armstrong, the CEO of Coinbase, who is no stranger to provocative comments on Twitter about how the practices at Coinbase or how he views his role as an entrepreneur, right? And so this is the latest in that installment.

11:16So this is him posting that about 40 % of daily code written at Coinbase is AI generated. And he wants to get it above 50 % by October. And he got a lot of negative lash back and attention from folks about this comment. And Ben, you may want to guide us through some of the things that you saw. Yeah, yeah, definitely a lot of negative attention over this one. And I think the problem here is that everyone needs to just stop associating the phrase AI generated with vibe coding. Like those are two very, very different statements when people say those things. So when you say 40, something like 40 % of our code is AI generated, it is probably not nearly as glamorous as you might think it is.

11:58You know, Andrew, like probably 40 % of the stuff that, well, at least before, before I went on break and went on an AI cleanse and haven't fully got back into AI. But probably 40 % of what I wrote or the words that I created in my computer were AI generated at some point, you know, but it's not like I'm writing like some opus using AI. It's a lot of like emails and Slack messages and like just routine toilsome work. So, yeah, a lot of what I've seen in these situations is AI doing this like routine grunt work rather than these like high profile features and capabilities. So don't view this as like, we've replaced all of our engineers and now AI is generating code.

12:39That's like the first point that I want to make. The second point is that, you know, while I think this is like an interesting metric to track as like a baseline of like, are our developers actually using AI to create stuff? I don't think it's a very good one to set goals around. And I really hope that he doesn't have this set as like some sort of long-term goal beyond October because, you know, somebody who used to work as a developer and got measured on, you know, effort-based productivity, you know, quote unquote productivity goals. I can think of at least a dozen ways to use AI to generate meaningless lines of code and get that number well above 50 % long before October.

13:21You know what? This week I could get that number up to 70, 80%. If that was the goal, we could be doing 100%. You know what? We could make all of our code AI generated. Yeah. You know, here's the problem with that is that it's arbitrary. Like ultimately, when you're measuring a percentage of all of your code and you're trying to see what percentage of it is AI generated versus not, you have to also remember that, you know, AIs are very verbose. They're loquacious. They love to add way more code than you probably ultimately need. So this can become a symptom of other problems within your organization and not representative of even like the total amount of code in your code base.

13:58Because if you want to get it above 50%, well, if you're creating on average 30 to 40 % more code baseline now and managing that much more, it's like that bloat is completely useless in a lot of ways. It's going to be hard to manage. And when you try to interconnect them all, it gets really difficult. And the thing is here is that like once you make this a metric that you want to attract and grow, like it becomes so easy to like gamify this to an extreme. Oh, my CEO wants as much AI generated code as possible. I'm going to just have this crazy huge output. And maybe that's what Brian wants, right?

14:39Maybe that's what he wants is more output from his engineers. but I think of the downstream implications of rotating on this metric as something as like a north star to track within your engineering work because I think it's going to leave nasty surprises for you in the future. I think that AI is best when it's used like an anvil to sharpen like a blade, you know, like you should really continue to make that blade as good and as sharp as possible and you shouldn't get distracted from what your goal is, which is like to make a really good blade, right? And so if you're ultimately just trying to make something really big, it's like, what are you doing here?

15:17Yeah, I get the impression that Armstrong is not the type of person that just wants volume. And part of that is because he has made comments multiple times in the past about how you can't just have more code, particularly through vibe coding. You do also, and particularly around things that do things like move money, you know, like Coinbase moves a lot of money and you can't like that's a really important thing to do. right you know so if you can't just have ai slop going into those systems completely but he has made a lot of comments around ai about how like they do encourage people to have really strong robust code reviews make sure that everything has appropriate checks in place and that there's always an appropriate human in the loop mechanism to review things so i kind of think it's a mixture of like like you know coinby is being pretty forward thinking in terms of like ai adoption mixed with Armstrong not quite articulating his opinion, you know, quite as well as maybe he could, coupled with the media just wanting to take comments like this and, you know, create controversy.

16:21It's a perfect storm. Yeah, exactly. I agree. Yeah, but he did also make some other statements recently when he sat down in an interview with Stripe president, John Collison, where he admitted that Coinbase had a directive recently where they gave all of their engineers a week to onboard with some AI tools that they had decided all their engineers had to start using. I think they had mentioned Cursor and maybe one other tool that they decided all their engineering had to start using. Gave them a week deadline to say, you need to onboard with this. And if you're not onboarded by the end of the week, or you don't have a good excuse, then you're fired.

17:02And they had a Saturday meeting where anyone who hadn't been onboarded and didn't have an excuse at that meeting got fired on the spot. And, you know, getting fired on a Saturday meeting aside, which I have lots of opinions on that I'll be for maybe another time. And, you know, maybe this is a harsh way to go about it, but I do actually, my hot take on this is I do actually kind of agree with his sentiment on this, even if I don't necessarily agree with the approach to it. Like we are reaching a stage where things like AI tooling is becoming a critical component of, of software development, of just knowledge work in general, but also particularly within software development.

17:43And if you finding yourself on a team or if you are you yourself or someone out there who is just refusing to be a part of this, you know, you need to have, you need to either be very conscious of why you're making that decision because there are valid reasons to not use AI in this day and age. You need to be very conscious of those reasons or you need to understand that you're making a decision that, you know, if you're an engineer that says, I'm not going to use an IDE or I'm not going to use a text editor as a part of my job or I'm not going to use CICD, people would question your competency as an engineer.

18:20And we're reaching that point with AI as well. and you know i may not agree with the approach that these execs take but i do somewhat agree with the sentiment at least but you know with that said i i kind of wish we would get beyond this whole like execs making vague statements about ai just to imply that they're like ahead of the curve uh you know what i really want to see is is to see the reality of this like show us what you're actually doing with the oh yeah stop telling us that you're just like writing all of this code with AI with these vague statements and like send your engineers out to events like have them share their knowledge like show us the cool stuff that they're doing if if they're actually generating this much code and it's valuable to us like share that knowledge with with the rest of the engineering community because we're all trying to learn right now you know and speaking of which Andrew you and I are going to be at an event this week the engineering leadership conference in San Francisco where I imagine among other things is we will be talking to all kinds of engineering leaders about how they're using AI within their organization.

19:25So if you're in San Francisco, if you're at this event, make sure you stop by. We'll have the Dev Interrupted booth there. We'll be recording some episodes. We'll be handing out some really cool swag. So, you know, we love doing this stuff. But yeah, if you're doing cool stuff in AI, it's time to start sharing what you're doing. Stop keeping it secret. Stop trying to just catch headlines and show us what you're doing. Yeah, I'm so excited to see you later. this weekend and get some face time in person with members of the dev interrupted community we're going to be there um and recording episodes even so stay tuned for some really great recorded episodes they're on site we're really excited to bring these lineup of guests for y 'all we got one more cool little story about everyone's favorite tool jira what do we have here andrew oh we're back in atlasia and my favorite so this is an article from sean godeki we love his engineering blog here at DevInterrupted.

20:17This is not the first time we've dropped one of these from Hacker News, making the front page here on DevInterrupted. This is an article about crushing Jira tickets. And what does it really mean for an engineer if they just hit it week after week after week, taking all those tickets from the beginning to the end of that Kanban board, right? And I love this article because it dives into why we're all making software in the first place. It's about shipping impact and not closing tickets. It's about sharing, or rather, It's about working together to find the real problems for your organization and solve them, but do so in a visible way where you have buy-in from your leaders and your impact of what you're doing is aligning with where the company is going.

20:58And this is really rooted in a really great anecdote that Sean shares in this article that I really resonated with. I love that he shared this really, you know, very raw moment from his career where he had taken a personal responsibility to maintain a high level of passing levels of these tests within his engineering work, right? Pushing it from 90 % to 100, made it completely green. That was his personal passion, his mission within that org. And ultimately, once he had achieved that, it was thrown aside effectively by leadership and by his manager. And it was seen as a waste of his time. And this was, you know, really upsetting for Shauna.

21:34Anyone who's had a passion project at work really resonates with that. I've been that person before where maybe you over-rotate on something that's really personally interesting to you or really aligned with other stuff that you want to be working on. But maybe not driving the highest impact within your organization. And so ultimately, this article gets at why, you know, when you go into JIRA and you're just like clearing out a week of sprint tickets that are falling on your plate, are you really solving for impact here? Or are there opportunities where you can actually look at what's important to your leaders, important to your manager, important to your, you know, your VP from your skip levels, and really then think outside the box to tackle the bigger problems that maybe there aren't tickets for yet.

22:21Well, first of all, Sean, if you're out there listening, I'm going to give you a hard truth. Tell me that you have bad JIRA hygiene without actually saying that you have bad JIRA hygiene. No, I'm joking. But I mean, if JIRA is an accurate representation of prioritizable work, closing tickets should actually represent some level of impact in an ideal world, right? But that's not the reality. We all live in the real world. JIRA is a mess. often nobody has the time to like infinitely groom their JIRA backlog to ensure that it's all properly organized and prioritized. Maybe we'll have AI bots that do that for us in the future, which we'll have a little bit on that in maybe just a moment.

23:07But, you know, I think this story is really more relevant than ever right now, because as AI accelerates our ability to do work, now we can accomplish more than ever. But that doesn't necessarily mean we're having a bigger impact than ever. If you're doing the wrong work, it just means you're doing more of the wrong thing than you've ever done before. And, you know, and I think now in this particular moment, in particular, the human being in the equation is still an incredibly critical element in determining what's important or what's the right thing to do in any one moment. And I say that with one big caveat i did see this really neat mcp prototype that you built andrew while i was out on my leave and that you were playing around with that where you were like asking clods like what should i prioritize today for my work and it was like going into some like jira queues and and get data and like coming up with like what it thought was like the most important stuff in the team's queue and giving you some recommendations and i thought that was actually really awesome and i was like okay well actually maybe ai can start to like tell us what's important now too yeah i think i think mcp is uh is pretty rad you know i agree with your perspectives on on sean and this article um or rather about like prioritizable work and what's in your jira really does it represent what's like is ultimately what is on your uh you know jira board should align with what is going to drive the most impact.

24:38But, you know, sadly, we don't all live in that perfect org. And oftentimes, those tickets that pile up don't match to it, right? But, you know, thinking about the thing like the MCP demo that I made, there was no ticket for that. There was, you know, a need within the organization that I identified and then kind of leapt at making something that was usable. And it wasn't perfect and it wasn't great. I didn't have all the answers. But no one working with MCP right now does, by the way. I wrote about that on Dev Interrupted last week on our Thursday send. And MCP, as y 'all know, we've talked about many times on the show.

25:15It's a really interesting thing to be messing around with. But if I talk about it anymore, I promise you, Ben, that Adam, you know, our editor, his head's probably already in his hands because of how long we've taken in this news segment. He's going to have so much of it on the floor. And so, folks, if you want to learn more about MCP, I would actually recommend that you join our webinar that's coming up on September 17th. I'm hosting that and giving a live MCP demo. And there's going to be info about that just after the jump. But before I do, Ben, do you have any last words from what we talked about today?

25:47No, it's great. Great to be back. Glad to help our audience navigate the AI-driven era once again. We'll bring back more guests, news hosts, I'm sure of it. And yeah, Let's keep moving forward. Great to be here. Well, great. Next up on Dev Interrupted, we have Forbes Cloud 100 Rising Stars. You know, it just announced Arcade.dev is among them. So stay tuned because after the break, we're sitting down with Alex Salazar. Are you investing in AI but struggling to see the real impact on your engineering team's productivity? Well, you're not alone. In a free 35-minute workshop that I'll be hosting with Linear B, we're going to show you how to translate AI metrics into business ROI, just like Expedia and Adobe have.

26:29And you'll learn a simple framework for understanding where AI is helping, where it's hurting, and where to focus your next investment. And besides, you're going to get a nice takeaway report on AI productivity as well. So don't miss out. The workshop's coming up on September 17th or 18th. Grab your slot and we'll see you there. Today, we're talking about how everyone is chasing gold with AI and how some people are digging and some people are selling the pickaxes. But Alex Salazar and Arcade, they're building the vault. Because as it turns out, AI that talks is easy, but AI that does something, you know, that's not too hard either.

27:10Where everything breaks down is when AI tries to do something securely. We've all seen this again and again, especially if you care about privacy, compliance, and customer trust. And Arcade isn't just handing out keys. They're helping developers build their own vaults so AI can act on behalf of users securely without blowing a hole in your whole security model. And if you're leading an engineering team and thinking about AI in production, you know, security is not an edge case. It's a foundational question that we're going to learn more about today. So Alex, thanks for joining us on Dev DevInterrupted.

27:45Thanks for having me. So let's start by talking about the problem that I've defined here. You know, we've talked a bit about AI and how it uses tools on DevInterrupted. You've been talking recently on LinkedIn about how a lot of AI projects, they look really great in demos and they're really flashy and they do this unique thing. But then you actually try to implement it with your team or take it on a larger implementation. And that's when something really hits the wall. You know, what happens there, Alex? Yeah. Look, I think one of the most disorienting parts of building agents versus normal software, the pre-agent software, is that it used to be that if you got a demo working, you were halfway there.

28:30And then you decided to like add some reliability and add some security. And it was all relatively straight shop. It was just labor. There wasn't a ton of risk. And that's not true at all. And so we started actually as an agent's company. We were trying to build a site reliability agent that could diagnose a problem in your environment. You'd get a alert from Datadog and we'd help you figure out what was going on. That was the original concept of the product. And we, like everybody else in agents, ran into this phenomenon where the demo is actually the easiest piece. The demo is like 1 % of the work.

29:00Once you get the demo working, 90 % of the work is getting it to be production-grade. Large language models are so malleable and they're so smart, generalizable. they can fool you right they can get a demo working relatively easy but going from that easy demo to something production grade it might be an entirely different product you might have to change everything and it's not obvious from the demo how close or far you are from production but you're typically really far from production so the things that get in the way the first one is consistency and accuracy you know when you have a demo you're not running evals if you know you're not looking to see like what percentage of the time the large language model is properly selecting the right next step in the workflow or the right tool call and the right parameters you're eyeballing and and if it happens enough times successfully you're like awesome demo and with the caveat that it's demo and whoever you show it to is like oh my god i can't believe you did that that's really cool only to do it once in order for everybody to be really happy right they produce it it needs to happen a lot you know if it's not if it's not to get everything right like north of 80 percent and actually more likely like 90 percent of the time depending on the use case the user experience is just going to suck and people aren't going to trust it people aren't going to use it just take security and safety out of it just users just think it sucks and so that's the first really really big one that's one of the biggest killers uh to demo production agents.

Read the full transcript

30:33The second one ends up being security and safety. And it's not even really just like a, it's not even really just like a, oh, you know, tinfoil hat, you know, CISO, large enterprise thing. Why can't JTPT send an email? It's 2025. They still can't send an email. It's a security thing, safety thing. And so, yeah, in a demo, you can give it like a Google app key, or you can hard code your credentials into some other service. There's all these hacks you can do in demo to make it work but that's never what you would do in production it might work for one single user but it would never work for a multi-user environment and so a lot of agents are blocked i mean why why isn't there a personal assistant agent today we've been talking about personal assistant agents for years and there still isn't one a lot of it is because how is the agent going to go talk to your gmail calendar the airlines how's it going to purchase something for you.

31:28Those problems are largely solved by what we built, but many of them are still aren't solved yet either. And the last one ends up being the last two are total costs. When you're doing a demo, you're not looking at the bill. You spent 50 bucks, no big deal. It was a demo. 50 bucks is nothing. But then when you extrapolate to like, you know, a hundred thousand users, suddenly like that's, you know, every day, the finance of the agent doesn't make any sense anymore. You'll lose money. You'll be negative margin. And then latency. People really underestimate the impact of latency. As you try and get the large lines of the more accurate and smarter and do things better, you start stuffing in more things in the context of my job.

32:05Well, that's going to drive the cost up, but it also drives up latency. And the more you do sequential thinking or any other kind of complex workflows, things really start to expand. And the smarter models are slower. And so all of a sudden, you have this really cool operation. Maybe it is really accurate. But it takes a minute or it takes, you know, 30 seconds or even 10 seconds, depending on what it is. Like, depending on the use case, the user may or may not find that viable. And so just, you know, so those are things we typically see chill projects. Yeah. And these constraints you've kind of defined are, they're interesting because they make for a very unique space where if you have like those three or four strong constraints around like why we don't have AI systems today, why they don't send emails, all of these things.

32:50like$50 for a demo versus, you know, millions of dollars for all your daily active users. Like when you have all of those kinds of constraints, you kind of get tempted to sacrifice one. And so that's where you get this demo phenomenon. You sacrifice security for the sake of putting on the veneer for the other kinds of functionality, the show like, look what this cool demo does, or look what we're building. But then when you actually go to build it, it becomes something totally different. And that's also because obviously the space is moving so quickly. It was a few months ago on this podcast, we had Kabir Bhatti from AWS, and they built an AI video generator tool.

33:25And he talked about that at length, about what they started with and then what they built over the course of the last year, and then what they ultimately brought to the market for mom-and-pop business owners. It probably had three different forms, ultimately. So that's where you get this rapid evolution. and a lot of times when we talk about this on Dev Interrupted, it seems like a team's expectation problem more than it is like, you know, all of those constraints, they're certainly solvable if you have the right expectations and you work about it in the right way. So, you know, what are some things that you see for teams that get it right, get it wrong about how they move past like the demo gloss and actually start solving?

34:08Yeah, so I'll tell you our story, which is ultimately how we pivoted into what is now the product, which is a tool execution and tool authorization system. So we started off as a SESSRE agent. And we were getting compounding error rates, which meant like every, you know, if you think of like a diagnostic flow, like why is service A slow? You're going to set the databases, you're going to set the servers, you're going to check to see if somebody committed code last night. There's all these checks. But like, it's almost an infinite space of things they might check. And the diagnostic like workflows, is the chains themselves can get really long.

34:43But every time you ask a large language model any question, you're rolling the dice. The power of a large language model is also its weakness, is that it's approximating this. And so, because it's approximation, it's non-deterministic. And so, it can get things wrong. It can hallucinate, you have errors. Every time you call it, you're rolling the dice, and if you roll a one, it fails. Well, if you've got really complex flows where you're calling it multiple times, you're rolling the dice more. And if you roll a one once, the whole chain fails. Right. And so we're running into these problems. That was problem one.

35:20They're called compounding error rates. And the second thing we were running into was, okay, well, we've given this agent super user access to Datadog and the servers and the databases. And it works. Right. Because it's great for a demo. But now, as we're going to go try and make this thing actually sellable, it's got to not do that. We have to really scope down resources and leverage existing permissing systems. And there was no way to do that. And so we ultimately had to invent all of this stuff for ourselves to make it work. And when we got it working and we started to show people a working demo, people were blown away.

35:54People who knew AI were blown away. They couldn't believe that we had really consistent data plots or that we were authenticating as on behalf of me into a service and properly scope permissions from within side of the agent. And then we realized, oh, this is a bigger product than the agent. So we ultimately pivoted and started showing and selling people that, which is what our key is now. yeah so my answer of like if i were building an agent today what would i do to get around all these things and to to really increase the odds of me succeeding obviously the issue is arcade but let's put that to the side so one i would this general software engineering best practices like you should be using stuff off the shelf to the degree that you can because there's so much to figure out and and build and learn that if you try and custom build everything from scratch you're going to run into a really big problem.

36:49Yeah. And so the easiest thing to start with is go grab an agent, go grab an agent framework and avoid trying to build your own. And yes, the learning curves are really steep, but they're really steep for a reason because when you custom build it yourself, you can fool yourself into thinking that you've got it because of the demo phenomenon. But all the frameworks are really complicated because they're building for production and they're not necessarily building for demo. And so by volume, we largely see LandGraph. And then behind it, we see OpenAI's Agent SDK. And there's a bunch of other ones.

37:23There's Pidantic, there's Mastra, there's ScrewAI. But pick one that you think fits your requirements in production and go and run. As you're picking those, I strongly recommend really thinking through what production is going to look like, not the demo. Optimize for the final state, not the demo because of this phenomenon, which I see a lot of people get stuck on. they'll optimize oh well look I deployed this one because it was so easy but they're optimistic for the demo and so they'll optimize for prod I think the other thing I would say is you've got to start with evals like you've got to build a muscle of evaluations right out of the gate because there's really no way you're going to get a production grade agent that's consistent enough if you're not really good with evaluations and there's a lot of products out there that can help you Planksmith there's Brain Trust we have a framework very specific for tool use and tool calling so the tools are out there you just need to set them up and use them and for people who are not native to aiml like traditional web developers this is very alien but it's really critical because we're all used to deterministic tests you know right integration tests unit tests and like how do you do a test on a function where there's an infinite number of correct answers an infinite number of incorrect correct answers and an infinite number of in the middle gray.

38:43Right. And you can't build those. Like there is no CI system that can support that. You need that. That's what evals are. And there's a different animal. The last thing I'll say on this topic, because I talk a lot about it to people is de-scope the living daylights out of your project. Even teams that have picked the right stacks and have built evals and all this stuff. But like the other thing that like where they hang themselves is they just, they try and do too many different things. and this stuff is complicated enough and new enough where you just increase your up significantly if you're trying to automate like one particular workflow first get the win and go and a bonus one is picking the project that you're going to go do like let's call it project selection like what is it you're going to work on what is it you're going to what's that what's that workflow you're going to automate man like that's half the battle and i would urge people to focus on places where there are modern APIs already available because if they don't already have modern APIs to access, the complexity goes through the roof.

39:50Like you try to build an agent that's browser-based, you can make it work. There are really good products out there like browser use and browser-based and a few others, but... Introducing a lot more complexity. A lot more complexity. And so if you're going to build an agent and you have the privilege of picking where you start, I would try and focus where there are modern APIs. Yeah. Okay. I love how you framed it about the eval situation, being so alien from writing normal tests. I think that's something that rings really true. That's something that I've had to adjust to as an engineer, moving from writing deterministic code to evaluating how my prompts or my workflows are doing.

40:27The idea of setting up the side-by-side benchmark is definitely new to a lot of folks, but it's a really critical part of building those tools. And it's just as important as the demo. I find whenever I work on those is starting from the eval, just like how you would start from tests or start from that more deterministic way is actually going to help frame what you're looking for a lot better. You know what success is earlier and you know failure immediately. I like too how you talk about your product moving because of opportunities, because of problems, because of constraints in the space, moving from like an SRE kind of agent, which an SRE is already solving a very complex, highly secure, very specific, like architectural infrastructural, you know, problem for teams, right?

41:11So then when you shift into like, we're just going to focus on solving the security, it's like, sounds like you taking your own advice, like what you just said about descoping. Because that security element that Arcade now is, was probably, or, you know, initially a very critical component of SRE Agent that was of the days of yore. And so you descoped yourself and you found what you're building now at Arcade. And just to bridge it into MCP a bit, because our listeners have, we talked about this a bit. We've had some guests even talk recently about MCP. We had CTO, co-founder of a first-of-its-kind MCP agency, who's like, they're out there making MCP tools for companies that are using these with agents right now.

41:50So there's a lot of interesting kind of conversations in this. We always are kind of going around security a bit. So I'm kind of curious to know, you know, we've talked a bit about the growing pains and about building with the tools. But, you know, what's your advice for folks that are experimenting with tool calls as part of their workflows that they're building with their teams, especially like in their engineering world that they live in? Yeah. So right now, tool use, tool calling is all the rage, right? It is the biggest bottleneck and biggest opportunity. in agents, which is crazy because we were deep in it before in school.

42:31Oh, so you're saying I was at those problem spaces. We were the OG. We all showed up. We were the OG. But so it's been really awesome to see the whole community kind of come to the same realizations that we had come to when we made the pivot and started to build the for-profit. That part's awesome. So let me contextualize it on why it's so important. Like why is it so exciting for everybody? But if you go back to what we just talked about, compounding error rates and consistency and accuracy, and if a model asks too many questions, you're going to hallucinate and have an error and all this eval stuff.

43:05The big insight we had in our agent, the reason we got it working was we ultimately, agents are non-deterministic, right? Lars Lang with Lars are non-deterministic. That is their power. That's also the risk. And so the big insight that we came to, like led us to Arcade was we dialed up determinism. There's, you know, prior, at least for us, prior to this big insight, we were all in on the large language model, figuring everything out. And that wasn't working because of error rates. And so when we, when we just had this big epiphany one day and we just started to like dial determinism back up, the analogy I give people is instead of letting a large language model just figure stuff out and go through its planning phases.

43:47and then, you know, in multiple series of planning phases, we were like, we're going to give it a discrete set of multiple choice questions. We're going to do, we're going to end at the TI-85 calculator and we're going to get all the questions. And it can only pick the butts during it. And when we, when we constrain the large language model, like to that degree, that's when things started working really well. And so the core insight that led us to Arcade, which is also the core insight that everyone's now having around tool use and MCD and all these different technologies is that if you dial up determinism and give the large language model these buttons and these buttons represent things that can connect to the outside world, you can build all the things that we all want to build but couldn't before.

44:36I see. So if I understand what you're saying correctly, it's instead of being like you are a free agent that can consider these tools as part of your thought and knowledge working process. Instead, it's like, okay, sit down. Here's your multiple choice exam. Each of these questions, there's five options, and you need to pick the best one from each. And that's actually a really fascinating kind of comparison to even how humans who are non-deterministic are evaluated. You could pick or use any kind of tool, but you want to know who's the best at it. You sit them all down in an exam, and you give them the same exact questions, and you see who picks them all correctly.

45:14So that's actually a really interesting way of kind of hijacking the same system that works for us as non-deterministic things. I mean, it's how all of us do math, right? You know, if I gave you, if I asked you to multiply two five-digit numbers together, like you'd probably get it wrong without a calculator or without a lot of time with pencil and paper, right? But if you're me, yeah. But even pencil and paper are tools. So if I just literally had you like do it in your head, you would probably get it wrong. So will a large language model. But if I give you a calculator, you're going to ace the exam.

45:49So will a large language model. And the impact's pretty dramatic. Like if you, if you give, you know, GPT 3.5 turbo, the oldest, worst tool calling model out there, and you hand it up like math tools, it will outperform a thinking model, like a brand new state of the art thinking model. and it'll do it much faster and much cheaper. And so the impacts are huge. And those are really simple operations. Now you think of something complicated where authentication and authorization is involved. Like, hey, go read my email. Go email Bob letting him know I'm five minutes late. Well, now what the large language we're going to have to figure out is actually almost impossible.

46:33Take hallucination rates out of complex workflows out of it. How's it going to handle an authorization flow, like an OAuth 2 flow, if you can't trust a model to hold keys, how are you going to, a security flow without you holding a secret is really difficult. We didn't know how to do it. And so, so how do you allow a large language model, an agent wrapping a large language model to go execute secure operations like reading your email? I mean, most things that matter require some degree of security. How are you going to do that? And the answer ended up being, you're going to do it inside the tool call, inside the button.

47:10okay so what does that mean for folks that are building now do they need to look at how they're building things differently so when you know let's let's pretend we're building a personal assistant agent and you say hey go check my email well the agent asking the agent to go figure out ola asking the agent to go infer like all the things that we need to go into that doesn't make any sense. Don't do that. That should be encapsulated in a tool. Whether it's MCP or RCA or the protocols don't matter as much. It's encapsulated in a tool. And that tool can be as deterministic as you want it to be. In many cases, they're completely deterministic.

47:50And so checking your email doesn't need a model to do anything. There are Google APIs and there's all kinds of stuff. And so you would, so when the model says, oh, I want to check email button, And after that, your function definition and your deterministic code is going to go through a standard OAuth flow. It's going to do all this different stuff. And it's going to go hit the right endpoints. It might be one endpoint. It might be multiple endpoints to go do the thing that you asked it to do. And now, here's the rub. This is where a lot of developers are getting stuck right now. They think, oh, well, Alex, what you're describing is just APIs.

48:27I'm just going to take the APIs that exist and put some natural language around them and I'm done, right? Unfortunately, that's not the way it works. So I'll give you a really good example. Let's say you say you want to reply to an email. Well, you're going to give the agent a reply to email button. So it's going to reply to email, right? Well, like, there is no reply to email, but there is no reply to email endpoint in Gmail. There's actually a series of operations you have to go do. A lot of it is in your own code. So it all has to still exist as a set block of API calls and you still need that orchestration, you know?

49:00Yeah, and they don't, and they're completely orthogonal to RESTful APIs. MCP and tools are ultimately APIs too. They're just different user, right? Like all the APIs that we know and love in the last 25 years, they're all, the user is really the developer who is wiring them into the code and they know exactly where they're going to use them and they know exactly how they're going to use them and it works every single time the same way. That's not how agents work. And so REST APIs typically are resource-based CRUD operations, create, read, update, delete. So almost every API you can think of is typically defined as a resource.

49:35And then most of the operations are create, read, update, and delete. An agent is different. They think in intention-based workflows. I'm going to reply to an email. I'm going to send an email. I'm going to go check to see if so-and-so replied. And those are orthogonal. And they might include zero API calls. They might include 100 API goals. They're just workflows versus resource-based front operations. And worse, RESTful APIs, because as we typically know them today, they're really what I refer to as inside out. Like a Google Drive API is going to describe Google Drive. It's going to explain Google Drive in a lot of detail.

50:14It has APIs to manipulate everything in Google Drive. your agent, let's say we're having a sales agent and it has an operation to brief you before a sales call. It's going to go check the CRM for communication. It's going to check Drive for brochures. It's going to go check your email for something else. It doesn't give a shit about Drive. What it cares about is getting the brochures. And so if the button is get brochures, that's all it cares about. All it cares about. But if the button instead is 37 different Google Drive operations, you're asking it to hallucinate. Because now you're asking it to not only hit the right button, but to hit the right sequence of buttons and figure out at inference in real time, like what the right sequence of buttons is to get to the right brochure.

51:01It can do it, but it's going to be slower, it's going to be more expensive, and you increase the odds that it's going to get it wrong when you don't need to. And so part of the problem, part of the paradigm shift for a lot of developers is it doesn't matter who gives me an MTP server. Like Google might tomorrow have the most beautiful SPS server that ever existed for Google Drive. You're still not the customer to build your own because you got to ground it in the domain and the use case of your agent because your agent's intention-based workloads. You're still going to need something between you and that or you're going to have to modify that to fit what you're looking for.

51:37Yeah. And I think right now, you know, that's something that people are only now starting to realize as they're starting to build tools. Because, yeah, it's super early. I mean, tool MCPs aren't going to have enough. But outside of very general purpose agents, which is the minority of the population of agents, you have to custom build your attention-based workflows. I want to shift the conversation a bit to learning a bit more about how your team is kind of like orienting itself around these problems and tackling these problems. Because like you've described in our conversation, it's very much a mindset shift in how you approach the problem, how you solve the problem with the constraints that are around you, but then also then how you test it.

52:16And ultimately, you build things by taking parts of other things and combining them in a secure way that ultimately specializes in on that intent. You talked about LLMs, they operate in intent versus APIs, they operate in actions or endpoints. right so when you are forming a team and you've and you brought people together to build arcade what does it look like on the inside of an ai native company and what are some things that you do there that you think that you do differently from maybe like more other companies yeah well i mean it's a great question so the first thing is who you hire it all starts with hiring and team makeup now because our product is ultimately like a service bus yeah we're executing just the get brochures button is is on the platform so we're like a service bus but a major feature in that service bus is that we're going to handle authentication and authorization on behalf of the agent right off probably their biggest feature so because we do that for our company in particular it's critical that we have the intersection of three really important skill sets for us you have be agent native and that's really hard to find there aren't a lot of people that don't build agents let alone they don't work at open ai and dropic and so my co-founder sam is one of those sam has implemented more than 100 agents or devras and so super lucky that we started this company together so somebody needs to be asian native and really understand how these things are go and go to prod have to really and this is this is the hard part not demo native but prod native which is a different thing.

53:56Yeah, that's definitely one of the big takeaways, I think, from this is moving the conversation from hype and from demo into prod. Yeah. But then once you get past agent nativeness, then it's really about what is it you're actually doing? And you got to be expert at that too. And so for us, as an example, we have to be expert at distributed systems because we're running all of these services, all these tools, and we have to be expert at all. And so, you know, half the team is out of Opta. we had to invent new ways of doing off to make all this stuff work right i think that's an important piece because you can't just like you can't just pattern match you can't just say oh well it worked like this last time we're gonna do the same over like it will rhyme but it won't be the same and so that intersection of expertise is what leads you to the right innovative solutions versus the old players just adding the word agents to something and then saying well now it's agent natives.

54:52Yeah, yeah. I mean, I think a lot of teams and folks are seeing that as well. But I want to double click on that first one you said being AI native or being like an agent native kind of worker. I think that's really interesting. So what do you think about that? What does that mean to you? I mean, that means you've put agents in production before. Okay, so I mean, like you've just made agents, you've taken the production, not necessarily that you are operating in an agentic way. If we're talking about an engine, like an engineering team. Yeah, like the right kind of like the mindsets or the skill shifts that you're seeing in your own team of like, oh, wow, like they stand out.

55:25They're a high performer and they're a high performer because of this. And a lot of times when I have these conversations with leaders like yourself, when we kind of peel that away a bit, it's because like, oh, they built this AI agent that does this thing for them or, oh, they captured their workflow in this way, which is one of those victories. So you have like. Yeah, yeah, for sure. I mean, like, I think right now everybody can, I'm going to sound like a hot tape, but people can fool themselves into thinking that they are, at least in an engineering capacity, that they're agent native because they use Kershaw or they use Windsor.

55:57That's an important skill set. Any engineer who's not deep in that stuff is going to be out of a job relatively soon. Because it's such an enormous lift. But it's not the same thing as having built an agent that went to production. Prod, not demo. Prod, not demo. Taking an agent to prod is a very different experience than having a conversation with Cloud Code. And that is what sets people apart right now in the engineering teams that are trying to build agents. Okay, that makes sense. So I'd be interested in knowing how that engineering group grows over time, the people who have built agents and taken them to prod right now.

56:36That's a small pool of people. It is a small pool of people, but you don't need a huge group because you still need everything else, right? And so, but what's great, if you have a really good engineering culture and everybody's in the same office, like us, I don't know how you do this in distributed teams. So I'll pass that to somebody else. But in a local team, the cross-pollination happens really fast. Yeah, it does. The person who is the agent expert is sitting next to the person who is the auth expert is sitting next to the person who is distributed systems expert. And they all start cross-pollinating.

57:06And all of a sudden, the distributed systems person is now starting to become a little bit more agent native every single day. and then all of a sudden they have an epiphany one day that even the agent builder didn't have. And vice versa. Like the auth person is cross-pollinating with the agent builder. And the operator's like, oh, well, maybe we did this way or that way. And so that's where the magic happens. For sure. We've talked about this with some guests before too. We had JJ Tang, CEO of Rootly, and we talked quite a bit about their knowledge sharing, working in the open, sharing team practices.

57:34This cross-pollinating effect that you're talking about is really critical for being a customer-oriented engineer because all of them have different aspects of what the customer is looking for or what the product needs to be. So the more you put them together, the better. I will say, in a distributed team myself, I will speak in defense of the distributed teams that we're also quite good at the cross-pollination. It's just that it happens a little differently. But I totally agree with you in terms of that being really important. And I know we're kind of starting to wrap up our conversation and come near the end of some of the stuff we've talked about.

58:07But you've captured a lot of really interesting thoughts about how teams are building tools, but also how they're orienting themselves around solving those problems. You know, is there any kind of lasting advice or that you'd like to give for if there may be that there may be there more that traditional engineering team who doesn't have that agent native engineer somewhere within them? You know, you can't cross pollinate from nothing. So how do you really start that culture? And what opportunities do you see for that kind of team? Yeah, I think the answer is both easy and extremely difficult. The answer is you have to start building with agents.

58:45And maybe the team's not ready to start building production-grade agents. Maybe it starts with hackathons. Maybe it starts with the leadership of the organization carving off time to letting the team play and learn. This stuff is really difficult. This stuff is, you're not going to pick it up in a day. it's going to take you months of tinkering and playing and building to run up against enough of the walls and learn because it's ultimately a new paradigm probably most listeners here don't remember the pre you know internet dates but a web developer was a alien creature to the traditional client server developer yeah right was an alien creature to like you know the mainframe developers, right?

59:30And so, you know, we take for granted how native most of us were to web, but it was a huge paradigm shift. And similarly right now, it's a huge paradigm shift. It's a whole new set of technologies, a whole new set of vendors, a whole new ecosystem, a whole new patterns. Testing's different. Workflows are different. Like, so the only way to learn it is to start doing it. And the books, you know, there might be books out there. They're great, but nothing, nothing is a replacement for hands-on experience. Yeah. You got to get out there. If you haven't built the demo, you got to build the demo. If you built the demo, maybe it's time to challenge yourself and think about how would that demo go to prod.

1:00:05I think that's a really powerful takeaway from this conversation. And when you build that demo, you should start by using Arcade to make it. So that when it goes to prod, it's easy. Yes. So, Alec, where can folks go to learn a little more about Arcade and how you're solving this problem that we've talked about? Yeah, thank you. So, yes, our website is arcade.dev. And so they should go there. We also have a YouTube channel where we have a ton of examples and walkthroughs on how to build pretty breakthrough agents that can go interact with the outside world. Things like Google and databases and your own APIs.

1:00:38We show you how to custom build your own tools in addition to using our out-of-the-box tools. We've done everything we can to make it as user-friendly for traditional developers as we possibly can. And I think we've done that. So I hope you come to our Q2 tips. Yeah, definitely. We'll put the links in the show notes so our listeners can go check it out, learn a little bit more about the problem space that you're working in. I think it's cool that y 'all also have examples up on YouTube. I'm definitely going to go check out the tool. Maybe there's a chance for us to collaborate on some of that stuff.

1:01:07I really appreciate you sitting down with us and walking us through some of the things that you've been building that's been top of mind. It's been really interesting for us. And for those that have been listening, if you've been listening to the conversation, you are one of those engineers who happens to have taken an AI agent to prod. We would love to hear about it. what you thought about today's conversation with Alex, but also just about your own experiences in doing that. So please drop a comment somewhere, anywhere that you're listening to this or find us on LinkedIn. We'd love to continue that conversation.

1:01:35And if you're not, you're still experimenting. Well, we want to hear from you too because we're all on the journey together. And the only way that we're going to get there safely and securely is by building it together. So thanks for joining us on Dev Interrupted and we'll see you next time.

1:01:54Be safe. Bye. Bye. Bye.

From the publisher

AI that talks is easy, but AI that acts securely is where everything breaks down. We're joined by Alex Salazar, CEO of Arcade, to confront the massive and often underestimated gap between a flashy AI demo and a production-ready system. Drawing from his team's own pivot from building agents to building the tools that secure them, he explains why a working demo is only 1% of the journey. Alex breaks down the four "demo killers" that cause most agent projects to fail: inconsistency, security flaws, prohibitive costs, and high latency.

Alex reveals the counterintuitive solution his team discovered: the key to making non-deterministic AI reliable is to dial up determinism. Learn why giving an AI a constrained set of intention-based tools - like a calculator or a multiple-choice test - dramatically reduces errors and solves critical security challenges that plague open-ended systems. He explains why you can't just wrap existing APIs and must instead build custom, workflow-centric tools for your agents. This is an essential listen for anyone who wants to build AI that doesn't just talk, but acts securely on behalf of your users.

Check out:

Follow the hosts:

Follow today's guest(s):

Referenced in today's show:



Support the show:

Offers:

More from Dev Interrupted

All 208 episodes
Your AI demo is a lie (and how to make it real)Dev Interrupted · 1 h 2 min
Listen in VO