99% Correct Is Still Failure: The Last Mile for Mission-Critical AI

9 Jun 2026 · 42 min · 17 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode argues that “99% correct” code generation is unacceptable for mission-critical systems (fighter jets, power grids, autonomous vehicles, medical hardware). It focuses on CodeMetal’s goal to close the “last mile” gap by providing provable assurance/verification for AI-assisted code translation and modernization, rather than relying on probabilistic outputs.

Guest backgrounds

Ryan Aitay, former CEO of Tableau; joins CodeMetal as president and COO. He references prior experience at Salesforce and Tableau and highlights collaboration with MIT/Lincoln Lab engineers (co-founder Peter Morales).

Key claims

Code-gen tools can generate code but can’t guarantee behavior equivalence at scale; verification is the hardest part. CodeMetal uses techniques like fuzzing, concolic testing, formal methods, and hardware-in-the-loop testing to prove correctness. “Assurance/provability” (not just human-in-the-loop) enables safe deployment and faster modernization.

Notable examples

Translating over a million lines of legacy C++ to Rust while preserving behavior on the same hardware; DoD-style modernization/simulation to avoid revalidation of stove-piped systems.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Importance of Accuracy in AI

0:00 to 0:16

Learn why 99% accuracy can still be a failure in mission-critical systems.

“It's great if you're, you know, 70, 80, 90, even 99 % correct, but 99 % correct is still failure when it comes to mission critical systems.”

CodeMetal and the Last Mile Challenge

0:30 to 1:14

Discussion on CodeMetal's mission to bridge the gap in AI for critical systems.

“AI can now write code faster than any human alive.”

Ryan Aitay's Journey and CodeMetal Vision

1:14 to 1:50

Insights from Ryan Aitay about his transition and vision for CodeMetal.

“Yeah, this is going to be an exciting one because this is a true gap, I think, in the market.”

Democratization of AI and Its Limits

1:50 to 2:26

Exploration of how AI is democratizing various functions but has limitations in critical areas.

“And just on a personal note, I was an OG Tableau user back in the day, like a complete fanboy of Tableau.”

Exploring Trustworthiness in AI

2:26 to 4:52

Ryan discusses the need for trustworthiness and verification in AI applications.

“So walk me through what you saw and what attracted you to this mission that CodeMetal is going on.”

Bridging the Gap Between Tools and Systems

5:04 to 7:13

Discussion on the integration of AI tools and legacy systems in mission-critical environments.

“Cause that's how I looked at the opportunity.”

Guaranteeing Reliability in AI Solutions

7:13 to 10:21

Ryan explains how CodeMetal ensures reliable AI outputs through rigorous testing.

“And, you know, it's, it's a lot, but it's also very, you know, it's a lot to unpack and understand, but it is a big opportunity.”

Human in the Loop and Its Role in Verification

10:21 to 14:01

Examination of the human in the loop concept and its implications for AI verification.

“I wonder too, cause you, you all always hear a lot of people's answer to this kind of problem you're talking about is human in the loop.”

Leveraging AI in Business Operations

14:01 to 17:20

Explore how businesses are using AI to enhance operations and decision-making.

“I'm curious, like how are you A, as an executive leveraging AI?”

Navigating the Nuances of AI Adoption

17:20 to 19:46

Discuss the importance of understanding AI's nuances and the risks involved.

“And I feel like the only way to get good at understanding the nuance is just by using AI a lot.”
Show all 17 chapters

Case Studies in AI Transformation

19:46 to 22:48

Learn about specific customer use cases and their impact on AI transformation.

“It's pretty mission critical that, you know, we're making sure that if there's a drone in the air, there's a tank or whatever it is that's being operated in a safe way that is dependable.”

The Importance of Assurance in AI Solutions

22:48 to 27:31

Understand the significance of assurance and accountability in mission-critical AI.

“How do we, you know, how do they get more time to think about, you know, what is the environment?”

Future of AI Accountability and Insurance

27:31 to 28:00

Delve into the future of accountability in AI and the emerging insurance markets.

“It's like the accountability, uh, aspect of it, because AI is not like a human to where it can be responsible and accountable.”

Accountability in AI

28:00 to 31:00

Explore the challenges of assigning accountability in AI technologies.

“I'm just curious, like, how do you think about accountability?”

Token Usage and Business Impact

31:00 to 37:00

Discuss the implications of token usage in AI on business effectiveness.

“get your take on you mentioned kind of token usage earlier but the uber cto i saw him the other day he basically said they blew through their entire AI budget.”

Evolving SaaS in an AI World

37:00 to 39:40

Learn how SaaS companies must adapt to integrate AI and meet new user demands.

“There's just a lot of uncertainty and you see that playing out in the stock price, for example.”

Advice for Business Leaders on AI

39:40 to 41:00

Gain insights on how business leaders should approach AI implementation.

“A lot of this, like the best way to learn these tools is to use them.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00It's great if you're, you know, 70, 80, 90, even 99 % correct, but 99 % correct is still failure when it comes to mission critical systems. A lot of these code gem tools are over here. The mission critical systems are over here and we have to help bridge that gap.

0:15Matt Paige:Welcome to the Talking AI Podcast, where we talk AI with both experts in the field and early adopters. I'm your host, Matt Paige, and we're here to demystify AI for you so you can get some value from it. Let's talk some AI. AI can now write code faster than any human alive. And most of the time, that's more than good enough. That's the magic powering the entire Vibe coding wave, the agentic coding craziness that we're experiencing right now. But there's a category of software where most of the time just doesn't cut it. The code running a fighter jet, a power grid, an autonomous vehicle, a piece of medical hardware.

0:51Matt Paige:When the code is wrong, the consequences aren't just a bug. They're a recall, an accident, a national security incident. And that's the gap that CodeMetal just raised$125 million to close. And that's why Ryan Aitay, the former CEO of Tableau, joined as president and COO to run point on what he calls safely delivering the last mile for mission-critical industries. We're going to get into what that actually means, where CodeMetal sits in this crowded market of cursor, clock code, and a host of other tools, and what an operator who ran one of the most iconic Perseed SaaS businesses thinks about this crazy world where agents are proliferating everywhere.

1:29Matt Paige:Ryan, welcome to Talking AI. It's great to be here, Matt. Thank you for having me. Yeah, this is going to be an exciting one because this is a true gap, I think, in the market. We use AI at our company 24-7. Our engineers are fully on board with it. But the point that you're solving here is critical. And just on a personal note, I was an OG Tableau user back in the day, like a complete fanboy of Tableau. But what I loved about Tableau is it wasn't just like the beautiful visualizations and charts, but it completely democratized analytics. And suddenly anybody in business could be their own data analyst.

2:10Matt Paige:And now we're living through something very similar with AI, where it's democratizing coding, design, video, writing. I mean, you name a function or purpose in a business and it feels like it's democratizing it or anybody can build. But you've stayed your next chapter on the idea that there's a category where democratization breaks down a little bit. So walk me through what you saw and what attracted you to this mission that CodeMetal is going on. Yeah. And by the way, thank you in my old capacity for being a Tableau community member because I'm still very connected to that experience. So look, I think, you know, coming into an environment which we're in now, which is this kind of AI world that we're in, I really just saw this large opportunity to work with really a lot of smart people, a lot of, you know, MIT, Lincoln Lab engineers, like our co-founder or sorry, founder and CEO, Peter Morales.

3:08and really it was you know i had all this great experience and i'm very grateful for my time at salesforce and tableau and and the things before that but it was like how do i take all these things that i've learned over you know 19 20 years and apply them to you know a new industry but also an industry and a company specifically like code metal where i could make a bigger impact what do i mean by that well it was you know make a difference make an impact because ultimately there's so many, there's a lot of AI noise, of course, it's like, you know, your podcast is called Talking AI. So it's like, there's a lot.

3:42We're part of the hype, man. We're part of the hype, but the hype's exciting, right? We're exciting to be in this environment, but there's this concept of like, how do we make AI more trustworthy? And, you know, because if you don't know enough, it could be dangerous. You know, of course there, and we'll talk probably about this today, but you know, my opportunity, I looked at it was like, how do we make an impact? There's this opportunity to make AI more trustworthy. And ultimately, I think there are a lot of people that don't understand it. So, you know, it's critical for our nation, for our military, you know, for our government, for mission critical companies, maybe that ship automobiles or airplanes or go down the list of things that we depend on every day as, you know, citizens.

4:27And if it's not done correctly or used correctly, like that could be problematic. And there are many ways and I can explain them, but like that, that could be problematic. And I think there's a lot of noise around like, well, is it going to take my job? What about is it safe for the things that I depend on every day? Like when I go get in my car or when I get on an airplane, like this is sort of the next chapter of, is it safe? Is it trustworthy? And our intent that CodeMetal is to really deliver that last mile.

4:52Matt Paige:Quick break in the pod. Our state of AI 2026 report just dropped and it breaks down what actually is changing in AI, what's hype and what leaders need to be paying attention to this year. You can grab it right now on our show notes or at hatchworks.com. Yeah. Cause that's how I looked at the opportunity. Yeah. But you're injecting this uncertainty into the world with this new, amazing thing we have AI and generative AI, but there's a lot of legacy systems that exist, these deterministic systems. So it's like, how do you graph those two together in a sense? I feel like that's part of it. And you talk about like the last mile for mission critical industries.

5:28Matt Paige:I'm curious, like, and this probably, you know, a bleeding edge of where is the last mile? Where does it start and where does it end? Or is it almost like a holistic view of how you build and leverage AI and the coding and building process? I mean, I think the last mile is relevant if we talk about things like code gen tools, and there are a lot of them today, the cloud codes, the codexes, the cursors, and they're all great, right? We know this, we use them ourselves. but when you get into and you can even ask the various AI tools today like you know can you can you generate code let's say can you translate c++ to rust and guarantee it'll be production ready and safe and they will all the tools will tell you well almost but not quite they can't guarantee it and that's the that's the tool itself right and we know that the hardest part is verification.

6:22And, you know, will it be correct? Have we tested every use case or edge case? Do we know every requirement? Is it secure? This is kind of like what I call the last mile, because there's just a lot of stuff that people forget at the end of the day that needs to happen. It's great if you're, you know, 70, 80, 90, even 99 % correct. But 99 % correct is still failure when it comes to mission critical systems. And so how do we, you know, our goal is to, you know, I think about like a lot of these code gem tools are over here, the mission critical systems are over here and we have to help bridge that gap.

6:59It's the assurance, if you will, like, how do I, how can I be assured that when I translate or modernize code, that it will actually work in a mission critical environment? And I think that is, that is really what we're trying to solve and we are solving. And, you know, it's, it's a lot, but it's also very, you know, it's a lot to unpack and understand, but it is a big opportunity.

7:19Matt Paige:Wait, it's funny. You mentioned like, if you ask AI today, if it could do, I forget the two, you mentioned Rust and something else. I feel like most of the time instead, it'll be like, heck yeah, I can do that. No big deal. In the end, you're like, you screwed XYZ up because you do have that sycophantic nature with AI where it will reassure you and say, yes, I got this, even if it doesn't have full context, which is scary. I think when you get to some of these, uh, you know, mission critical things like you're talking about, but you mentioned it though, like your teams, they use cloud code and some of these other tools.

7:54Matt Paige:So I think it'd be helpful. Like where you're not necessarily in competition with these tools. It's, it's almost like, um, tertiary or, uh, in support of these, these other tools so where if you look at the map of everything going on in this crazy world of ai coding tools and ai tools in general like where where does code metal sit in that uh that that hemisphere yeah um yeah i think at number one i think it's early uh in our in even though it feels like we're a few years into this we are but it is very early in the sense of like understanding what's possible and also, of course, identifying some of the risks and things that can be solved.

8:38We don't really see a lot of competitors, you know, at this point in terms of, you know, focused on verification, validation, assurance, the things that I've mentioned, making it trustworthy. You know, when I think about it, let me try to like, the thing I said before was C++ to Rust, right? So like, you could ask a question like, well, you know, and literally, if you go in and you say, like, go into chat GPT and say, like, can you do this in a production environment? Will you guarantee it'll work? It will say, well, yes, but like, there's always the but there. And so I think the scenario would be like, yes, you could use any code generation tool.

9:17You know, you could, they could like run an entire repository, it could be agentic, It could generate code PR. It could refactor some things. It's pretty good. But the problem becomes almost a behavior. So it's not really just a coding problem. It's more of a behavioral problem at scale. So if the tool succeeds, as I said before, 70 % of the time on a complex refactoring, you could still fail at some level and be a scenario where I think you run into challenges, right? So they can write code, but they cannot guarantee that, let's say a million lines of C++, or in this case, let's call it C++, legacy C++, that it will behave exactly like it did before once it's transformed.

10:09And that is the thing that we're solving. And there's a whole bunch of work that goes into making sure that it's guaranteed. But that is really what we're delivering. We're delivering a guarantee or an assurance that it will work. And that's the difference. Hmm.

10:22Matt Paige:That's interesting. I wonder too, cause you, you all always hear a lot of people's answer to this kind of problem you're talking about is human in the loop. Oh, we have a human in the loop here to verify all of this, which, you know, frankly, I think people get comfortable in that human in the loop, you know, verification decreases as, as time goes on. But I think that the bigger point is though, I mean, this is a process in a system like anything else and there's natural bottlenecks. And I think that the bottleneck is shifting in many ways to that human in the loop. So it's like, as this just completely progresses and scales, how do you think about human in the loop?

11:01Matt Paige:Because it's having that as the verification step, it almost feels like you have, you know, a game of Jenga. sits and you're at the very end of it where it's about to topple over and it's like yeah i mean i think and perhaps human in the loop you know there's i like to think of this as like you know well let's let's unpack a couple things so the first thing is like you know you've got ai makes everyone's jobs easier yes it used the correct way i agree i use it all the time i'm happy to talk about that as well and there's a human and you know let's say i've got an engineer and now they're not just writing code, let's say for said application, they can be doing other things.

11:38They can be thinking a bigger picture. They can be reviewing code. You know, is a, but what about the hardware, right? Because we're not just rewriting software. We're trying to connect it to hardware and do it in a production way. So we talk about our kind of our North Star here at CodeMetal, which is like, how do we, you know, basically make things as pro, like we want to make all things programmable, software, hardware, et cetera. Think of it like a, you know, if you've ever been in like a Tesla or a full self-driving experience where like you have the ability to ship over the air updates, you know, software impacting hardware, that's great.

12:16They can do it, but not everyone can do it. And so I think making sure that it works. Now, if I say, well, okay, let's just translate some old code and we'll just have a human review it. Well, what about this process of like, We need to analyze the code. We need to understand, you know, and this is, of course, some component. We do want to do test generation. So we do something called fuzzing, concolic testing. You go down the list of various things that we include in our suite to make sure that it checks every use case. It checks the source code. Is it behaving the same as it is in this new system?

12:49Is it checking every edge case? Is it checking for variables that aren't expected? So there's a, you know, and I'm going to get quickly out of my realm of expertise because I'm not an engineer. We should put a little caveat on this. But the point is, we're going a bit deeper to make sure that you can count on it ultimately at the end of the day. And that includes, in addition to human in the loop, it's hardware in the loop. So we want to be testing hardware continuously. It's a big part of like when I walked into one of our offices at first, I'm like, wow, there's all this like hardware. It's a different world, right?

13:20And so how do we make things run efficiently? you know from a small device to a mission critical device and do it in a safe way wait you lost me at was it fuzzing was that i've never heard of that term yeah i will you know think of it like um what are the different edge cases that could be a problem right and so it tests every different way and we use this thing called formal methods which is another word for math like there's a lot of this is my learning curve which is like okay what does this mean how do i you know trying to bring it all to uh english so that i like to say how can i make sure my parents can understand what i'm working on that's my uh my intent and

14:01Matt Paige:my goal as we talk about these things yeah that's a good north star but like this is separate but how amazing is it like that you have ai as this partner to like help guide you as you learn i think just just such a cool premise in general but you mentioned uh you mentioned like you're leveraging AI? Like I'm curious, companies right now, especially newer companies are growing in a totally different way because they have AI at the beginning versus existing companies going through their own innovators dilemma of how do we apply this to our existing business? I'm curious, like how are you A, as an executive leveraging AI?

14:43Matt Paige:I think that's super interesting for our audience, but also from a business context, how is AI being leveraged, whether it's in operations, it could be anything, sales, marketing, operations, name your function of the business, anything that kind of sticks out in your mind, either daily use or within the business? Yeah. You know, it's funny because we talk about this. We're a small company, you know, we're still under a hundred employees. But we, you know, we look at the opportunity, you know so we don't have like i was in theory one of the first business people non-engineers in the company yeah we have a lot more now but we can do a lot more than we could have you know if we had started this company in like you know 2000 or 2010 like even it's a it's a different world so i think you know we use you know whether it be you know even things like an m &a pipeline or you know recruiting and screening candidates more quickly like it accelerates a lot of what we can do or hey, give me an opinion.

15:43I wrote something, maybe I want to post it on LinkedIn or wherever I want to post it. I don't have a marketer yet. So hey, I need some, can you help me just like think through or give me perspective? But again, I'm not going to just take exactly what it says. I want to make sure that it's like verified. Again, it's the same premise. It's like, can I trust what is in here and how do I actually go through it? It's like, you know, in the past, you'd ask many different people before you maybe would make a decision. It's a similar thing. it accelerates that process. I do think that, you know, and of course there's, I think we found with the code generation tools is like, you know, if I'm building a simple application or a SaaS application or whatever it may be, maybe I can get started and advanced really quickly.

16:24If I'm, where we shift from that into a world where I just say like, let's be careful, let's make it sure it's trustworthy is we're gonna convert legacy code to modern code. Let's say it's a million lines, like I said before, it has to behave the same if it's gonna, let's say, you know, impact a, you know, an autonomous vehicle, right? Whether, especially in a military environment, like good on the, you can just see how this can spiral out of control very quickly. And so I think that every business right now needs to be using AI internally to run its operations, you know, like, Hey, I have built a cashflow statement.

17:00Hey, does this look right? Or maybe it can build it for you. Like, those are all great things, but you can't just like send it out. Like, I think that's the message I would want to tell people. Absolutely, if you're not using AI, the companies that aren't will lose to the companies that are, but be careful is sort of the intent, especially at this stage.

17:19Matt Paige:Yeah, there's definitely nuance about it. And I feel like the only way to get good at understanding the nuance is just by using AI a lot. I always go back to the Shopify CEO talking about, he has given a mandate for his team to use AI reflexively, just like it's just second nature. And then you do start to get, it's the nuance, it's the taste element of like, okay, I can use AI here or I know to push here. Like, I don't know about you, but I've been completely Claude code pilled over the past several, several months. And, you know, I was a very early open AI chat GBT, uh, adopter. And I went through this period of using both and now I've kind of shifted over.

18:02Matt Paige:I still use both, but I went through this exercise with our quarterly planning. I had it hooked up to our CRM and all the other data and had memory for everything we do, and I just had it build the full report for me. But to your point earlier, it wasn't just like, okay, there's the report go. Like I went through it and I found some interesting insights that I hadn't even noticed or thought of, which was just this really cool experience. in like another one we have it hooked up with again our crm and slack and whatnot and it's going through there and identifying okay what what prospects should we be following back up with that we haven't talked within a while or something's happened and it's just giving us this regular report with a you know a score attached to it and we now just have it like running and doing its thing so it's just once you see the possibilities it's like oh my god there's so much i can do with this.

18:58Yeah, I think that we're in the inning of perhaps, you know, or the stage of seeing all the possibilities, right? And then I think, you know, like we've identified code as one area where it is real and relevant and very, very useful. You know, I still think like, I'm still trying to work whether it's, you know, Claude or Picker. I'm like, hey, schedule a meeting. It's not quite my assistant yet. It's not quite perfect. It's getting better every day. You know, schedule a meeting, do this. It's like, well, I can't access your calendar or well, I can't do this or well, I scheduled it on the wrong day.

19:27So we're getting there, but I think it's education. It's just like what we're talking about here. Like, why am I doing this? It's because there's an opportunity to educate and make sure that, you know, not only impact the business, you know, or like, you know, we work a lot with various forms of our government, you know, it's pretty mission critical that things are done right. Right. It's pretty mission critical that, you know, we're making sure that if there's a drone in the air, there's a tank or whatever it is that's being operated in a safe way that is dependable. And that is a big opportunity.

19:59But not everyone understands this. So you can't just assume like, you know, if it sounds so good to be true, I always feel it probably is, right? That's a saying from the past. So just check your work. Yep.

20:09Matt Paige:It's maybe a theme, you know? What? 100%. And it definitely applies. I'd be curious though, like what you can tell us, like what are some cool use cases where you're starting to apply your approach, your methodology, your tooling with specific customers? Obviously, I'm sure some of the government ones you can't get this deep on, but are there any interesting ones that are worth noting? Yeah, I mean, maybe the one I mentioned before is a customer. It's one I can't mention the name yet soon, but the one where they basically had come to us and said, you know, we have over a million lines of legacy C++ code that they needed to basically translate and modernize, but using the same system, the same hardware that they've used.

20:59So it's kind of like, it's sort of like, I want to rewire a city, but I don't want the power to go out.

21:05Matt Paige:Yeah. It's okay, well, how do I do that without, I mean, it's impossible. So this is like an impossible task that all of a sudden we're like, well, we can do this in like a couple of weeks. It's not that hard. And we can guarantee it'll work, right? So a million lines of C++, you know, translated into Rust with zero issues in a very short timeframe is basically impossible. And now it is not impossible. I think maybe another example would be, this is the one I have to be even more vague with. I apologize. We can do this later when we can talk more about it. But, you know, think of like a Department of Defense or Department of War.

21:38They have bought software and, you know, various types of software that's been like, call it stove piped or disconnected. Some of it they can use, some of it they can't. Millions and millions of dollars has been spent on this. But at the end of the day, right, they want to make sure like, how do they modernize it? you know they can rewrite it perhaps but without like spending a bunch more money to revalidate that it's actually working so there's you know let's go through a process of like where i need to translate existing software into a modern maybe more maintainable scenario uh and i want to use that so i can do some simulation for war fighters right um pick your things right the code needs to be translated, needs to be provably correct.

22:28And then it needs to, we need to know that whatever goes in comes out and there's no need to revalidate it. So it's a, you know, again, I'm sorry I'm being vague, but like I, there's a lot of information that's sort of sitting in different places. We're bringing that all together and providing a simulation environment so they can run real world scenarios. And that, that is very powerful and really only something that we can do at this point.

22:49Matt Paige:Now that's, that's really interesting. Cause like you've, you've been in the business world for some time now like how many very long i'm very old yeah yeah how many transformation initiatives get proposed done and they just die on the vine because they become they're too complex they take too long by the time you get ready to do something things have changed like that's that's a very very real problem and you mentioned like tesla earlier there is this element of like modernization that potentially is possible now because like we talked about earlier the innovators dilemma a core part of that is just like there's so much built up infrastructure and nuance to things and ai can actually help accelerate i mean that's kind of what y 'all are doing in a sense i feel like in in a lot of ways yeah and i was trying to think of another one like you know i think robotics you know again and some of the use cases i give are more like perhaps military focused or defense focused but that's just a big big area of demand for us at the moment, but, you know, um, if you think about maybe a mid non-traditional type of environment where there's, uh, you know, let's say, uh, some kind of a tank or something that is autonomous, whatever it may be, right.

24:04Or drone or something like that. How do we, you know, how do they get more time to think about, you know, what is the environment? Like you see the environment, just like in a, in a self-driving car, you see the environment, but how do we get more time to react um and and how do i do that in a way that is safe and accurate like very important to be accurate in that environment so you know i want to compute let's say the the scene or sort of the visual faster so that that that the object can have a safer ride at higher speed gives the the, you know, the machine more time to react, that can only be done, you know, if they trust the output, right?

Read the full transcript

24:49And so if they don't trust the output, obviously, they're not going to do it. And so that it's, these are the types of things that are very, like, you know, very critical. That's why we say mission critical. We're not focused on the non mission critical things. These are like things that are very, very important. They, you know, failure can be catastrophic, and potentially any of these examples.

25:10Matt Paige:Yeah. And you mentioned a few times the word like guarantee where you can guarantee something. Like how do you, that word carries like a lot of weight. Like how do you think about that in the sense of your business and whatnot, especially, I mean, when the stakes are this high, you almost need that in a sense. Yeah, I mean, we think about it as assurance. You know, we try to talk to our customers about outcomes. you know how do you in a world where you know and and this i think is something that we could talk about for for a long time which is how do you price these things like how do you go through the process of you know making sure the customer sees value well when i can take a million lines of of you know do the impossible right which i said before is take a million lines of c++ and convert it to rust without changing behavior and it runs in the same system it's highly valuable and we can guarantee or prove that it will work, right?

26:07Prove is even a stronger word than guarantee, right? Because it's provable. So it's a verifiable proof that you're doing in these test environments and things like that.

26:16Matt Paige:It's not just like a satisfaction guarantee that put the stamp on it. It's good to go. Like there's something tangible behind it. Correct. And we can show those results, right? So we can say, hey, we will show you that it works. And that is where this becomes quite valuable. And if I can say, hey, this would have taken you, you know, I don't know, 50 to 100 developers for two years. Well, there's a cost to that. And then what are they not working on that they could be working on? We can just do this for you in a few weeks. It's a different world. And that's where I think that's why like code generation.

26:48And that's why I really like all these, you know, the open AIs, the Anthropics, et cetera, the cursors. They're all part of the same thing. They're just doing it a different, like a different stage, if you will. um we're very focused on provability assurance verification validation and delivering that last mile which i think is pretty important it's just maybe not on the top of everyone's mind because we're still in a mode of like oh i can i can write a blog post faster and i can write you know i can build a time and expense app like really quickly those are all great examples too um we're sort of over on this other spectrum which is like we want to make sure that we're safe and we want to make sure that there are no catastrophic issues that happen.

27:28That's why it's attractive to me.

27:31Matt Paige:I'd be curious to take on this. It's like the accountability, uh, aspect of it, because AI is not like a human to where it can be responsible and accountable. Um, but as things, as it becomes more mission critical and foundational, um, wait, where do you see that accountability element going? I could see like entire insurance markets starting to emerge around this. And there's, you know, obviously business insurance that already exists, but I feel like there's this new element that's, that's going to be introduced where, okay, well, the AI is at fault in some way. What happens then? I'm just curious, like, how do you think about accountability?

28:10Matt Paige:Uh, and this may not necessarily specific to CodeMetal, but I think it's going to be a bigger and bigger topic as we keep progressing. Yes. Um, it's a really interesting question. I think that, um, I feel like there, you know, there's two ways to just to look at this, which is like, there's probably a group. And again, there's a lot of media attention on everything we talk about with AI. There's probably a group that will say it's more dangerous, right? Cause it could make a mistake. Well, humans can make mistakes too, by the way. Right. We know this humans can hallucinate. Maybe we wouldn't call it that, but we know that you may get the wrong answer.

28:49but if we can prove that something works right because that's not really talked about it's just like well ai is amazing and it is but when we can prove it works with you know in our case we do it for mission critical environments but in and let's say any environment if we can prove that something can work using maybe it's what we do or some other technology you should be able to assign accountability to it it's just going to take time it's going to take education back to my earlier your point is like, we need to educate the risks, how it works, how you can prove it works and where you shouldn't use it.

29:26Like, you know, we, we know we shouldn't use it in certain scenarios. Like, for example, it's not great necessarily at, well, I don't want to go down that road because we'll get in other discussions, but there are things that's better at than things that is maybe not as good at is maybe the way to say that. Coding is one of them. It's great.

29:42Matt Paige:Totally. And I think one of the best analogies I've heard because it's probabilistic at its core and this this uh variability uh very um being able to verify it is one other element I haven't considered so my mind's kind of like opening to this new concept but think of this think of this like we call it verification and validation v and v you can't really it's easy to remember v and v like is it verified is it validated like this is how I look at a lot of these things I like that I like that, but the analogy that I've heard is okay. If I have a calculator and I put in two plus two and it gives me nine, I throw it away because it's broken.

30:19Matt Paige:It obviously is not working, but with AI, like if you get an incorrect answer, it doesn't mean it's broken. It may not have sufficient context or there's some other variable there. And the point you mentioned of guess who else, you know, is not always right. Humans, because we're kind of probabilistic beings in a sense too. We're working off the context, our nature, all these different factors that play into it and yeah i think it's just important that people view ai in that way and it gets back to the point you mentioned like it's going to make mistakes like you have to go into it knowing that in a sense yeah i mean i feel like we can all give a hundred examples of where it wasn't quite perfect but it was still pretty darn good yeah exactly exactly and so one other point i want to get your take on you mentioned kind of token usage earlier but the uber cto i saw him the other day he basically said they blew through their entire AI budget.

31:15Matt Paige:And for context, for those listening right now, we're in April. So if you're listening around Christmas time and thinking, oh, that's not too bad, we're in April. And he said they've blown through their entire AI budget and they have 5 ,000 engineers. 92 % are now using AI and agents in the coding process. I just feel like this concept of tokens is going to be such a critical part of the business conversation. the cost of goods or however you want to think about it. But what's your thoughts on token usage, how it applies, where it's going to go? Any thoughts on that, either from your business and use of AI or whatnot?

31:53I mean, I've seen many different spectrums, right? So whether it's the pursuit or usage model or tokenization and tokens, et cetera, I think is great. I think there's an opportunity and I think there's an evolution around whether it's called outcomes-based or value delivered. And I say that because like, let's just say that I, you know, I'm COO, I'm thinking about our usage of AI internally. Do we, you know, should we have a budget? I think it's more like, well, if you can demonstrate a business case or a use case that shows me that you can, you know, do more with less and be more, you know, have a better output, that's 100 % a win, right?

32:38Because it's not just about, it's like budget is one angle, but it's also hard if you're running a large company, having come from one to know like, what should I allocate and where do I want to spend? Yeah. I think in a world where you have to look at these AI initiatives as almost like, and I don't, never thought I would use the word again, but like a business transformational scenario, which is like, okay, what the C++ or S thing I mentioned before, if I'm going to like save, let's say, you know, five to$10 million, but also continue to be able to sell my hardware. And because I basically have been able to modernize it, that has a lot of sort of variables included in it that would impact my business.

33:20And I go back to like, I like to say, what we try to do is we help companies be compliant or safe is maybe another word. we help them go faster and then we help them ultimately be more open right because and those are things that you can't just say like oh well what's my budget to be safe compliant fast and open yeah i mean hopefully that's everything you're thinking about um and i think that that's really like what we're trying to unlock is ai can do a lot of these things but if done incorrectly it can destroy a company too we haven't seen enough of those examples yet but i'm pretty sure they're coming.

33:56She's like, oh, I bet the house on this particular thing using AI and, oh, it didn't work or, oh, there was some catastrophic thing that happened. We don't want that. We want to do those three things I mentioned and we want to be able to hopefully deliver. And when we're seeing, I mean, we have customers paying us based on the outcomes that we provide, but it's a different world, right? It's like a different way to think about what is the value? How do you calculate or you know sort of quantify what they would have spent otherwise and what is the value of being open and extensible right think about chip manufacturers right think about various languages to one particular vendor who i won't name that's dominant in others like how do we open the supply chain right that's a big opportunity as well which we're working on totally we got another v for the uh the two v's value so we got what verified uh validate and

34:50Matt Paige:values so we got the trifecta now. You're marketing, I like it. Exactly, exactly. Last thing I got for you. So obviously you ran Salesforce, massive company, like you're very adept in the whole SaaS ecosystem. But I'm curious, like how do you see that world evolving as it's not just humans using the tools? Like agents are becoming a big part of the users of the tools and I think a big part of it too is like, okay, well, how do we tailor our products and services to this new class of user, which is AI in a sense? Yeah. Um, I mean, like I mentioned, the other thing we talked about is there's definitely a lot of media attention and I think gen, you know, story generation, it's like SAS is dead or SAS is this or that.

35:39I still use these tools. Uh, and you know, it's not just because I work with Salesforce, but like we just got off a forecast call. We're using Salesforce, you know, there's use Tableau, like they use all these tools on a regular basis or maybe, you know, Workday or other tools, like they're not going away. They're in fact getting better in my mind. Now, does the market value them differently? I think they're trying to, because they don't exactly know how to value them in a new world, right? There are new, new things coming at them. You know, what's the long-term cashflow scenario? Like, of course, Like, you know, there's questions.

36:12But I look at these things as they're critical systems that are also evolving. I mean, having come from Tableau and seeing, you know, things like Tableau Next or Slackbot or other things, which are great technologies, they're evolving too, right? So I think the companies, and there are a lot of them that are all evolving, those are going to be great. The ones that aren't evolving, those are the ones that might be a little worried about. As it relates to, I think many of them are trying to switch their pricing models, etc. How the public markets value them, I guess it's hard to know what the public markets are going to do, in my opinion.

36:48Matt Paige:Yeah, well, I think your point's spot on because at the end of the day, stock price and valuations based on the prediction of or certainty of future cash flows. I think the market's just uncertain right now. It's not that it's not going to be there. There's just a lot of uncertainty and you see that playing out in the stock price, for example. But like the example, not to go deep on Salesforce, but you mentioned Slack, like it's getting almost a new rebirth because like this is this new conduit to talk to AI and agents. So there's this completely new use case that never existed before. That's like, oh, I think every business, I don't know if you have any thoughts on this, but I feel like every business needs to go through that exercise of, okay, what is enabled for us or how many users or AI use our products in a different way now that we have this new thing?

37:38I don't know if you have any thoughts or frameworks or approaches, or is it just, you know, get everybody together and see how they're using tools in unique ways. You mean internally? I just want to make sure I understand the question.

37:53Matt Paige:It could be internally or I think just your products and services in general. Like how could they be used in unique ways or, you know, existing companies? We have a lot of existing enterprise listeners and whatnot. And they're thinking like, okay, how do I evolve my product and where does my strategy go? And I think part of it's how do you tailor it to this new way of working in a sense? Yeah, I mean, I really, I think it just depends on what your goals are. At the end of the day, I mean, you know, we, we use, uh, like we use Salesforce, not just for sales, but we also of course think about it, like, where are we tracking our M and A opportunities or wherever it may be.

38:32So, um, what I, what you said about like the Slack scenario, I think is interesting because it now becomes like that interface for all of these, you know, whether it's CRM related things or things that are outside, you know, um, a lot of different use cases.

38:47Matt Paige:so last thing for you if a business leader you know taking one thing away from this conversation uh between like you know ai hype and what's real what's that one piece of advice you're giving to leaders listening right now that are worried excited confused whatever whatever analogy you want to give it yeah um you know i think um i'm going to go back to kind of a little bit of a thing that I think about a lot, which is, and I try to talk about with our team, which is there's a real opportunity to leverage AI in a way that can accelerate not just your company, but your personal goals, like at the end of the day, but you just have to know where it's best to use it and where it isn't.

39:31I think the biggest risk for people and or companies and or nations ultimately is to not do anything, right, or be afraid of it. A lot of this, like the best way to learn these tools is to use them. Try them all, right? Compare the results. I mean, I was doing this the other night where I was like, okay, I haven't hired all of the team members that I want yet. I'm trying to, but I need help with scheduling certain things and setting certain things up. And I've got my cloud code and that for cloud for work. I've got these other things running and I'm testing them all. Open AI, good on the list of various tools.

40:07Some do things better than others, right? And then I'm wondering, well, why doesn't this work? It should work. Right. Things that aren't quite perfect yet, but it evolves so quickly, which is great because if you check back in a week, it'll probably be doing a better job than it did last week. And so be aware, educate yourself, talk about these things, you know, listen to more stuff like your podcast here and others, I think are all important things like the learning curve is steep. and if you take a break from it or like, oh, I'm good, I don't need to use this stuff right now, I think those are the people and or the companies that I think are going to be perhaps behind and maybe not in position to take the lead.

40:47Matt Paige:Great advice. So listen to the podcast. If something doesn't work, come back in a week and try it again, which I think is very true. That's what I do. In today's age. Ryan, thanks for being on Talking AI. So where can people find you, learn more about Code Metal or I guess maybe it sounds like you have some openings as well. You're looking for some people. We're hiring a lot of people. Yes, we are rapidly growing. And if you're really thinking about the opportunity to make AI more trustworthy, I think we're a great place to be. I guess codemetal.com is kind of where we are. We don't yet have a conference or anything, but we'll be communicating more soon once we get our marketer on board.

41:28Awesome. Brian, thank you for talking AI. All right. Thank you. Have a great one, Matt.

41:33Matt Paige:Thanks for listening to the Talking AI Podcast. If you enjoyed the show, give us a follow or subscribe on your favorite podcast podcast. And don't forget to leave us a review. We love those. For more info on Talking AI, visit TalkingAIPodcast.com. Quick break in the pod. If you're listening to this podcast, chances are you've been thinking about how to actually use AI inside your business. And that's exactly why we built the AI Opportunity Finder. It's a free tool that helps you uncover high impact, tailored AI use cases based on your business, your goals, your pain points, and your industry. No fluff, no generic use cases, just real ideas that fit your business and the ranked by ROI potential.

42:13Matt Paige:It takes about three minutes to run and it's like having your own personal AI strategist for free. If you want to try it for free, check out the link in the show notes or go to hatchworks.com backslash AI dash opportunity dash finder.

From the publisher

AI can now write code faster than any human alive, and most of the time it's more than good enough. That's the magic powering the entire vibe coding wave. But there's a category of software where "most of the time" just doesn't cut it: the code running a fighter jet, a power grid, an autonomous vehicle, a piece of medical hardware. When that code is wrong, the consequences aren't a bug. They're a recall, an accident, a national security incident.

In this episode of Talking AI, Matt Paige sits down with Ryan Aytay, the former CEO of Tableau and now President and COO of CodeMetal, which just raised $125 million to close that gap. Ryan explains what he calls "the last mile" for mission-critical industries: the verification, validation, and provability layer that sits between AI-generated code and the systems where failure is catastrophic.

The conversation covers why 99% correct is still failure in defense and autonomous systems, how CodeMetal translated a million lines of legacy C++ to Rust in weeks (like rewiring a city without the power going out), and why the real problem isn't code generation, it's behavioral assurance at scale. Ryan also shares how he's using AI to run a sub-100-person startup, why the biggest risk for any company right now is doing nothing, and what an operator who lived through 19 years of per-seat SaaS at Salesforce thinks about outcomes-based pricing in the age of AI.

In this episode, you'll hear about:

Why every AI coding tool says "almost, but not quite" when asked about production-ready guarantees. The difference between code generation and behavioral assurance at scale. How CodeMetal translates legacy C++ to Rust with provable correctness in weeks, not years. The concept of V&V (verification and validation) and why it's the missing layer in AI code gen. Real use cases in defense, autonomous vehicles, and simulation environments. Why hardware in the loop matters as much as human in the loop. How a sub-100-person company uses AI across M&A, recruiting, marketing, and operations. Ryan's take on token economics, outcomes-based pricing, and the SaaS evolution. Why the biggest risk is inaction, not AI errors. What attracted Ryan to CodeMetal after 19 years at Salesforce and leading Tableau.

Key Moments

  • 02:47 — From Tableau fanboy to the trust gap in AI
  • 03:52 — Why Ryan left Salesforce/Tableau for CodeMetal
  • 05:55 — "Is it safe for the things I depend on every day?"
  • 06:45 — 99% correct is still failure for mission-critical systems
  • 08:20 — The sycophantic nature of AI: "Heck yeah, I can do that"
  • 09:22 — It's not a coding problem, it's a behavioral problem at scale
  • 11:22 — Human in the loop isn't enough: hardware in the loop
  • 14:30 — What is fuzzing? Formal methods explained in plain English
  • 16:02 — How a sub-100-person company leverages AI across every function
  • 18:19 — The Shopify mandate: using AI reflexively
  • 21:33 — Rewiring the city without the power going out: the million-line translation
  • 24:38 — Defense use cases: drones, autonomous vehicles, and simulation
  • 26:28 — "Prove is even a stronger word than guarantee"
  • 28:32 — Accountability and the coming wave of AI insurance
  • 32:54 — Token usage, the Uber CTO's blown budget, and outcomes-based pricing
  • 36:26 — SaaS isn't dead, it's evolving: Ryan's Salesforce/Tableau perspective
  • 40:08 — The biggest risk is doing nothing
  • 42:07 — Where to find CodeMetal (and they're hiring)

Key Links


Mentioned in this episode:

AI Opportunity Finder

Feeling overwhelmed by all the AI noise out there? The AI Opportunity Finder from HatchWorks cuts through the hype and gives you a clear starting point. In less than 5 minutes, you’ll get tailored, high-impact AI use cases specific to your business—scored by ROI so you know exactly where to start. Whether you're looking to cut costs, automate tasks, or grow faster, this free tool gives you a personalized roadmap built for action. 👉 Try it now at https://hatchworks.com/ai-opportunity-finder/

Free report from HatchWorks AI — State of AI 2026

What’s real in AI this year, what’s hype, and what leaders should prioritize — including production lessons, designing for agents, and governance. https://hatchworks.com/state-of-ai-2026/

More from Talking AI

All 84 episodes
99% Correct Is Still Failure: The Last Mile for Mission-Critical AITalking AI · 42 min
Listen in VO