AI Can Write Code. Why Isn’t Software Better?

28 Sep 2026 · 43 min · 18 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode argues that despite impressive AI coding agents, “software isn’t getting better” and real automation is missing. Instead of faster code generation, Diogo Almeida proposes embedding intelligence inside software via JEV, a new “primitive” that turns natural-language intent into probabilistic decisions usable by programs. Guests discuss reliability (uptime, determinism/consistency, robustness), why reliability beats flashy demos, and how this could enable “smart software,” new interfaces, and more useful SaaS—an “inverse saspocalypse” where apps gain capabilities rather than just faster PRs.

Guests

Ben Horowitz and Martin Casado (a16z; investors/tech leaders). Diogo Almeida (founder of TypeSafe AI; builds JEV; background in computer science/AI via Kaggle/NeurIPS/OpenAI/Google Brain; emphasizes pragmatism and “build prod, not God”).

Key claims/examples

OpenAI has tried automating customer service since 2020 with limited results; coding agents mostly produce “code like 10 years ago” and don’t add new software capabilities. Examples include password resets dominating “95%” support claims, and reliability definitions like “smart every time.”

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction of Diogo Almeida and JEV

1:08 to 2:10

Diogo Almeida discusses JEV and the vision for AI in software development.

“JEV is aimed at expanding what software itself can do, giving developers a new primitive for turning natural language intent into decisions that programs can actually use.”

The Vision Behind JEV

2:10 to 4:30

Diogo explains the goals of JEV and its approach to enabling automation in software.

“So I was actually asked for like an elevator pitch, which I tend to ramble on and I don't do well.”

Expanding the Capabilities of Software

4:30 to 7:10

A conversation on how JEV aims to enhance software capabilities and developer efficiency.

“But that code is the same thing a human being would have write.”

The Role of AI in Software Development

7:10 to 9:50

Discussion on how AI tools like JEV differ from traditional coding tools and their impact.

“Classifiers, they're designed to be useful.”

Diogo's Journey into AI

9:50 to 12:20

Diogo shares his unique background and journey into AI and software engineering.

“So I was a mathlete, but I actually never, oh man, this also is a little cringe.”

The Future of AI and Software

12:20 to 14:00

Discussion on the optimistic future of AI in software and its potential to create jobs.

“So you said something there that is so unusual in today's world, which is AI is really, really fun.”

The Automation Paradox

14:00 to 15:00

Exploring the disconnect between AI capabilities and real-world automation.

“That's one way to, that sounds much more ominous.”

Reflections on AI Development

15:00 to 18:00

A discussion on the evolution of AI capabilities and expectations over time.

“It's really, it's quite a kind of fascinating dichotomy.”

Challenges in AI Generalization

18:00 to 21:00

An analysis of AI's current limitations in automating productive tasks.

“by the generalization capabilities of RLHF.”

Exploring Automation in Industries

21:00 to 23:40

Insights into the types of problems easier to automate within different sectors.

“But we've been optimizing that judge instead of the automation part.”
Show all 18 chapters

The Importance of Reliability

23:40 to 26:20

The significance of reliability in AI systems and software development.

“looks like yeah I should be able to automate OpenAI has been trying to automate customer service since 2020.”

Future of Coding Agents

26:20 to 28:00

Speculating on the role of coding agents and their integration with advanced AI.

“So like not exactly determinism because I think determinism, it's useful for unit tests, but not real systems.”

The Future of Coding Agents and JEV

28:00 to 29:54

Exploring the implications of coding agents on software development and JEV usage.

“and like that's where a lot of software is today, right?”

SaaS Value in the AI Era

29:54 to 31:56

Discussing the impact of AI on SaaS companies and their future potential.

“like a 50th percentile architecture instead of a 60th because you want to move faster and have like codecs work overnight or something like that.”

Transforming Software Interfaces

31:56 to 33:36

Examining how AI could revolutionize software interfaces and user experience.

“Well, all the SaaS applications are going to all of a sudden get like dramatically more useful.”

Reimagining Software Capabilities

33:36 to 35:24

Discussing the limitations of AI in enhancing software capabilities and the potential for improvement.

“You know, there's just such a profound intuition here, which is if you use AI today to generate software, right?”

The Era of Probabilistic Programming

35:24 to 37:56

Exploring the future of probabilistic programming and its applications in AI.

“I'm not going to like over promise, under deliver that, but I will fight for that.”

Integrating AI with Existing Systems

37:56 to 40:55

Discussing the challenges and processes of integrating AI into current software systems.

“Like, I mean, we did this and we did this for the internet and we did this being friend to client server.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Where the fuck is all the automation? AI is so unbelievably smart, and yet it's so useless at all other stuff. It doesn't matter how much AI coding agents you use, the software actually isn't getting better. Maybe you're riding it faster, it's arguably getting worse. OpenAI has been trying to automate customer service since 2020. What I want instead is smart software. I want to expand what software itself can do, such that things that should be automatable can then be automatable. My favorite thing that you guys say is we built prod, not God. That's so good. Because if we had any other kind of like big lab leader, even if they had joy, they would cover it up.

0:38And then your view is so different. You're like, no, we're going to create a way better world. For nuanced reasons, I don't think we are on the path of RSI. In the SaaS-spocalypse story, AI can write software. But what if the bigger opportunity is putting intelligence inside the software itself? In this episode, Ben Horowitz and Martin Casado sit down with TypeSafe AI founder Diogo Almeida to talk about JEV and a different vision for how AI changes computing. Diogo argues that coding agents make it faster to produce the same kind of software we already have. JEV is aimed at expanding what software itself can do, giving developers a new primitive for turning natural language intent into decisions that programs can actually use.

1:22They get into why reliability matters more than impressive demos. what a new era of probabilistic programming could look like, and why AI might make existing software dramatically more useful, rather than simply replacing it. And beneath all of this is the question that drove Diogo to build TypeSafe in the first place. If AI is already this smart, where is all the automation? Today, we have the founder and leader of TypeSafe, Diogo, with us, who is a bit of a hero to both Martine and me. He is not only building a really interesting product, but creating what we think is a very important movement.

2:05So we're super excited about today. Welcome. Thank you for coming. Thank you. Maybe you can give us a brief on what is JEV, what is TypeSafe, why is it important? Is this a curse-friendly or no? Yeah. Oh, okay. What work are you talking about? Okay, okay, okay. Cool. So I was actually asked for like an elevator pitch, which I tend to ramble on and I don't do well. But like I realized my favorite elevator pitch for Jeff is where the fuck is all the automation? Like this is like so unbelievably tragic. Yes. So much intelligence. AI is so unbelievably smart. And yet so, not that I hate on chatbots or coding agents.

2:44I love them myself. But it's like so useless at all other stuff. And it's tragic. It's tragic that we have so much like diamond in the rough, but not polished for work. But TypeSafe is making AI for software. You know, we want to make AI powerful, not just for humans in the loop, but to actually build real software. And Jiv to us is our first model in this whole space to make it way, way better to like make automation. Yeah. And so it's been interesting because it's kind of caught fire in software world. So one of the things that made us go, what the hell is going on here is like every developer we know is calling us and going, oh, this is freaking awesome.

3:25It's great. It's fast. It's great. Everything's better. Then how does that, because everybody thinks of, well, we've got cloud code. We've got codecs. Don't we already have that? Like, what's the difference? And then how does that lead to real automation? Ooh, I wish I had like some slopped visuals because I have a favorite slopped visual for this. So I like Cloud Code and Codex. I love the description from Gary Tan and them. It's just-in-time software. Incredible way to describe what they're doing. It makes software on the fly, and you can program software in natural language, but it has the same expressive power as software.

4:01What I want instead is smart software. Instead of automating software engineering, I want to expand what software itself can do, such that things that should be automatable can then be automatable. and like in a more flowery language, I want to express things like intent. I want to expand the vocabulary of what we can do. And I can talk about like all sorts of like weird sci-fi things I want. But like programming is like hyper specifying like valuable things and then infinitely replicating them. It's so freaking cool. And I want to just make that more. Interesting. So one way to think about it is instead of kind of a tool that somewhat replaces a software engineer with a faster, maybe not even as good software engineer, what you're saying is, no, no, no, we're going to super empower the software engineers we have to write way, way better, more interesting things.

4:56Yeah, yeah. So this is, by the way, I just think so many people miss this point, and it's such a subtle point, and it's so important to actually tease it out, which is if you use something like Cloud Coder Codex, which is great, or Cursor, which is great, they write code. But that code is the same thing a human being would have write. Maybe it's better, maybe it's worse, but it's basically still code, just like code looked 10 years ago. And the thing with JEV is whether or not you're a cloud coder or a human, you have this new primitive, this new thing that you stick in your code that actually expands the power of software.

5:26So instead of writing code, it is something that you include in your code. Which, by the way, is interesting because it's this very powerful primitive, which would be great if you explain, but it's also a little bit different than how programmers think. For example, like it has this notion of probabilities or maybe... So an intelligent layer inside the software. Yeah, think of like a library that you can use natural language to describe what you want and you give it kind of a state machine and then it will choose what to do with some confidence levels, which we kind of haven't really had before.

5:58Ooh, there's a lot of tricks there. I will jump into one thing first, which is I love the first thing you said in the direction of where the fuck is all the automation. I love software so much. I wish I could be writing it all day. It's, would not recommend being a CEO to people, but whatever. And also it's wild that AI is so cool and software has been unchanged in 10 years. Yeah, exactly. You know, like that to me, like no one can square this together. And the most we can do is add like a little chat bot in the side sometimes that can take actions, but not all actions because some of the actions are not reliable enough.

6:32So I just want to give like that tiny aside. I love the point. I'm going to jump back to the point about like, this is a little bit of a different way to think about it. Yes, I think that machine native doesn't exactly match bits perfectly. And like, that's actually the art form that we are trying to do. In our onboarding on day one, I draw like the Venn diagram of like, what AI is good at, what is valuable in code, we're in the middle. So, you know, we don't output like extrapolated floats, for example, because like AI is just bad at that. But things like probabilities are not exactly novel. and it's similar to the, is Jev just a classifier argument?

7:09Jev is absolutely a classifier. Like classifiers are sick. Classifiers, they're designed to be useful. Yeah, they're designed to be useful. And actually it's the same interface as like some of those ML concepts. Of course. Because these came from like practical people who are trying to make systems work. And what I'm seeing is happening now is that Jev actually, my guess is that Jev probably is better than having an MLE team from 2019. making the stuff for you and you can just program it on the fly. Who knows what could be built? Because there were not that many good MLE teams in 2019 to build narrow things and to be able to collect data sets and measure it and all of that.

7:47And it is just the beginning. By the way, to this point, do you think there's a slider bar here where on one end it's language in, language out, like we have today. On the other end it's an existing imperative program. And then you can kind of move between the two? Or do you think this is the point in the design space, which is language in, kind of state machine out, which is going to solidify as a general purpose thing for programmers. Oh, that's a tricky one. So I will say the answer in my heart. The answer in my heart is that it is a slider. And actually, when I design for the properties we have, I might have made mistakes due to my personal preferences.

8:21But like intelligence per dollar is my North Star right now. And it could be wrong, just to be clear. Intelligence per second might be more valuable in the short term. But even like our interface, like calling the input state, this is intentional. Like it's to say. Oh, that's great. I didn't catch that. It's meant to be the inside of programs. Oh, I love that. So in my heart, because like, so we are really optimizing. A lot of the work I do is for even more complicated arrangements of the internals of program state. Can you put intelligence in there? I think that this is going to be an ever present battle to have.

8:51I, we're very intentional about our design. And also pragmatically, I think certain things happen. Like it's easier to make an AI at these milliseconds. so it'll be more like a database for a while than like a standard library thing. But I would love it to be a standard library thing too. Can I just pull back to you? Like what is the alchemy that creates a Diogo? I mean like you speak... When a man and a woman love each other. You speak like an AI researcher, you speak like a systems person, you speak like a programmer. And normally these things have been like not super overlapping and you're taking AI, which we've been pushing towards being a being and you're making it a programmer's tool.

9:29So maybe a little bit about your personal journey that... My history into AI is somewhat unorthodox. I was a mathlete. I was an award-winning mathlete. The way I describe it is I was good enough at... This is a cringe. I was good enough at math to get girls. So that's quite good. I didn't know that was a thing. Yeah, I did. You have to get quite good. And what kind of girls do you get when you're that good at math? That's an excellent... Whoa, that's... Our audience needs to know. Oh, no. We have to inspire the youth here. Don't do it. Don't do it. It's not worth it. Just be cool and chill and interesting.

10:05And don't overcompensate. Wow, I can't believe I said that. So I was a mathlete, but I actually never, oh man, this also is a little cringe. I never really liked math. I never really tried. I was just like big fish in Little Pond. And to me, math was always the path I was set on, but I hated it. Because it was always about winning competitions. but then computer science is actually a lot like math. It's basically like math, but cool and useful and fun and interesting. And I still love giving algorithms interviews. Is it the best thing for me to do? I don't know, but do I love it? Yes. And does it like allow me to like suss people out really well?

10:44Yes, it does. So I love computer science. I consider myself to be computer scientist much more before AI researcher, despite my history. And like what actually got me into it was I also won a Kaggle competition, not from sophisticated math, but from like just automating like the fuck out of it. You know, like just like more nested loops, more, you know, like it solved like a systems problem. Wow. You know, so that event eventually got me, like I was forced to speak at NeurIPS, normally an honor, but I hated it because I just wanted to be in the mines. Was that from the Kaggle team? Yes. Oh, wow.

11:20Yeah, actually the Kaggle host of it was Isabel Guillon, who was the co-inventor of the SVM. Actually, I think the first author of SVM. I'm not 100 % sure I'm the first author. And she just basically saw that I was like this person who really didn't fit into the research community and then adopted me and showed me, like it got me to meet all the AI people and that, you know, my career was just pushed into that direction. And from there, OpenAI? No, it was like a startup with Jeremy Howard. No kidding. I love Jeremy. Fantastic. Cool. And then Google Brain for a while. And then Retire for a while.

12:04And then eventually I was like just kind of tired of not doing anything. And I was like, you know what? Actually, AI is pretty damn fun. And I joined OpenAI because of that reason. And it worked out really well. Amazing. Really, really well. Incredible. So you said something there that is so unusual in today's world, which is AI is really, really fun. And then the company has such a different demeanor and view of AI than everybody else. My favorite thing that you guys say is we build prod, not God. Because if we had any other kind of like big lab leader, they'd be like trying to, even if they had joy, they would cover it.

12:53and then your view is so different. You're like, no, we're going to create a way better world, and it's going to be awesome, and there's going to be not only are there not going to be less jobs, there'll be more jobs, and there'll be way better jobs, and everybody's going to have a great time. And just being around you, you clearly believe that. So tell us about that. Because for us, TypeSafe, Jeff, it's more than a company. It's a whole movement towards a positive future. that most people in the AI world kind of don't like. Yes. Or they're not with it. I think they don't get it. Yes. You know, like it's just a classifier complaint.

13:33It's like an ML level concern while everyone else is having like a Jeff party. Because it's like, holy shit, like we can do all the things that we wanted to do. And I think if you don't like get developers, it'll be hard to understand what's really going on. So 100%, I agree with that. I do think that there's like a pretty negative world painted that I obviously disagree with. I think it really comes from this like, you know, mono model Kool-Aid that everyone believes. Right, right. One big brain to rule them all. That's one way to, that sounds much more ominous. Yes, yes. But that's what people hear.

14:13For sure, that's what people hear, yeah. But, you know, like, will that one, is that one brain really on the path to rule us all? Like, we have not automated really basic things that I don't think we want people to be doing. You know, like, there's lots of really, really basic stuff. And I think that, oh, man, it pains me when the world is discordant with the reality. And, like, part of the pain is, you know, on the where the fuck is all the automation. Like, how can we have AI be so freaking smart? and like there's so much so much financial incentive to automate stuff like yeah you could make an excuse for diffusion I don't buy it at all I shouldn't name names but like that obviously is not true part of the problem is like the discordance with the reality and the fact that AI has like so much potential is what made it really tragic for me that we had not released this so now it's like a little bit of a party for me but like I was afraid of AI and all dev users are like there's the happy AI the people on Jeff, and then there's the Morose AI, the people who are not.

15:17Yeah, yeah, yeah. It's really, it's quite a kind of fascinating dichotomy. It is really, well, I'll give you, and to your automation point, I had a funny conversation this morning with David George who runs our growth fund because we were talking about the new tools. I was like, have you tried the Muse thing? He's like, oh, it's awesome. I was like, what'd you do with it? He said, I finally canceled my New York Times subscription. And I was like, that is hard to do. But, you know, it's a kind of a, it's a very tip of the iceberg of the things that are horrible things to do that we need to automate.

15:57I think that if we were going to be really intellectually honest and we are really aiming for the North Star of automation, we cannot fall into the same anti-patterns that AI has fallen into, which is really focusing on outliers and demos, right? Like a lot of people ask me, like, what are your favorite use cases? And I'm like, I'm not sure if they work. I want them to work in the background such that like someone would trust that to run and not page them. And like people can build on top of that too. And like, you know. Composable, composable. Composable, but like other things like safe. It's a different type of safety where if you wanted to actually run with resources associated with it, with access to things, you need guarantees for that, or at least statistical guarantees.

16:39So it doesn't go wrong, break it, that type of thing? Well, I don't think our models will be doing that anytime soon unless someone does the software to do that, which would be very cool flex. I should figure out how to give credits for that. But not in a way that we're not responsible. Right, right, right. I'm just curious. How long has this intuition been percolating? Because I remember talking to you maybe in 2017. We did talk about that. Yeah. And then a lot of these ideas were in, you know, you were talking about data being important. You were talking about, like, you want to focus on the task.

17:17And like, but like, so I just like, you know, was this like, did you know that this was going to end up being a classifier? Or was this just an intuition that like, there's just kind of another way to view this entire kind of AI movement, you know? So actually a fun story about that chat in the talk from 2017. I think my talk was actually in a very similar theme. I think it was called something like AI modular in theory and flexible in practice, which is very software. So I'm a little bit consistent with that. I think that this really started right before ChatGPT. Like right when we released these things, I did not have intuition about this.

17:56And honestly, I was not even, I was very, very pleasantly surprised by the generalization capabilities of RLHF. When is this? Must be end of 2021, like fourth quarter of 2021. Like we were, it was really, really general. Like if you read the paper, it's unlike other papers that are like trying to prove their point. It was us actually, you know, scientific method-ish, trying to disprove, like, is it cheating? And, you know, my favorite query was, why is it important to eat socks before meditating? We'd made sure that was not on the internet beforehand. And, like, the models were able to, like, make plausible human-looking answers for this.

18:35And that, to us in the team, was the thing that clicked, like, this is not cheating, which you should always be afraid of cheating in ML. And then what really got me burnt was we released it. you know we did you know I'm obviously a big capabilities guy I did a lot to release that model I really thought that model had like a decent chance of being AGI and when it didn't that was like when my whole world came crashing down and I was like why? So you were kind of on the other train for a bit. Like the crazy train? Well no I'm just like RL generalizes like maybe we have AGI like RLHF generalizes pretty well.

19:16RLVR is the thing that doesn't generalize as well from what I've seen. And AGI in... Well, I was just saying more. I mean, like, you know, you were behind chat GPT. You were behind these early GPTs. That was a very different goal, which is like creating a chatbot that will talk to the human being was not a programmer's tool, et cetera. So I'm just wondering, like... Oh, well, actually, early, early, like 2020 OpenAI, when we talked about AGI, people used to describe it as Ilya in every if statement. So it's not, it's kind of like, but like part, we were talking about OpenAI culture. Part of it is that it's like intentionally vague.

19:55So it's a wide like tent so that everyone can be inside of it. But like, I am not, for nuanced reasons, I don't think we're on the path of RSI. And I still don't think we're in the path of RSI and I did then. I do think that what OpenAI defined as AGI is extremely doable. automating most of the world's economically valuable work actually sounds like, oh man, I don't, like there's a lot of work out there. A lot of it is very rote and simple and like by volume in order to be able to like outsource work, you need like simple instructions that like basic people can do. And as far as I can tell, the intelligence of that has been available in the models for like quite a while now.

20:37And like my, oh man, you know, chip on my shoulder is like, why is this not available? And then since our LHF, the industry just like kind of bifurcated into gigantic overpromise under deliver. I think GPT-3 was actually quite calibrated back in that day. But because humans evaluate how good the models are, it looks really good because they're the judge. But we've been optimizing that judge instead of the automation part. And that has been the missing thing. So I would say that it was really, really then that it like hit me, you know, like, why is this thing not more useful? And so you think that the measure that we should have is to what extent can you automate actual productive tasks?

21:17That the, when you say over-promise and under-deliver, that's the dimension in particular you're talking to. The ability to automate tasks. I, like, I think in my heart, it's like cool sci-fi. You know, and I think that, I think that that is the canary in the coal mine for cool sci-fi. Like, are you really telling me that math is solved or like even like two years ago, GPQA, that Google proof question answering is solved, but we still can't handle a drive through, right? Like, it's a very hard thing to hold in your head at once. And I think a lot of people don't have good answers to that. Can I just test one thing, which may not make sense, but I want to.

21:58I mean, isn't there an argument, though, that, like, the distribution of the real world is different than the digital world, right? It's heavy-tailed. There's a lot of exceptions. We don't have all the data. And, I mean, couldn't it be the case that the reason we're not doing productive stuff in the real world is just, like, we don't have the data for that distribution. We're not training on that distribution. and this is why it's just been basically relegated to like these lower dimensional manifolds like whatever math or code or... I don't entirely buy the data argument in my opinion. I do believe that there's a long tail for sure, like that would be kind of crazy to deny.

22:41And I don't think that in my like canary in the coal mine situation, we need to automate that long tail. Like I think that we need to be incredibly pragmatic on everything. and like building reliable software is always an investment, right? Like, you know, what were the three great virtues of a programmer? Laziness to not to do it again, hubris, and there was a third one. Yeah, no, I remember. This is from the Pearl days. Yeah, there was a third one. Very well, yeah. I wish I could remember it. But like, it's about like the laziness to like spend, you know, like 10 hours to do like the five minute task instantly and to never have to do it again.

23:18like it only make like it should be an ROI decision for people who like automate stuff like I would just like it to be automate a bowl and I think that people will just make like new kinds of work hence the Jev in Jevons new kinds of work once that stuff is doable but like as like a benchmark I feel like it's useful to see can we actually automate the stuff that it really really looks like yeah I should be able to automate OpenAI has been trying to automate customer service since 2020. You know, like, it's, you know, like, it's not. It's just pretty amazing. It's wild, you know, it's wild. I mean, inside companies, there's very little that's automated right now.

24:04And the projects haven't worked. Other than, programming has worked amazing. Can you maybe classify the types of problems you think that are easier to automate now? Because it was kind of interesting. So we've actually looked at support before the current generative wave. And it was interesting. You'd meet a company, and the company would say, we answer 95 % of all, like, you know, like, help desk calls. I'm like, that is so many. But then you actually look at the data. They're all the same. And you realize it's all password resets. And then, like, but if you did it by, like, uniqueness, it was only something like 50 % or something.

24:39So it just feels like when you're dealing with humans and natural systems, like, there's just kind of this very kind of, you know, like a long tail of exceptions. And so to what extent did like every, probably every hour I have somebody ping me and like, I'm using JEP for this new use case. I'm like, I had no idea, you know, like, you know. And so like, to what extent did you even predict like the broad range of use cases for it? Like, did you assume that was going to happen? And have you been surprised by that? Extremely surprised. Did not assume it would happen. This launch was like not something, like if anyone expected this, they are probably insane.

25:15Right? Like it is, I don't think someone could expect a chat GPT for developers because chat GPT was for, you know, like, you know, the normal users. And it's weird. I actually don't even know what percentage of the people who are part of the JF party are developers themselves. I can't imagine non-developers using it. I don't know how they would use it. But even my non-developer friends are just like part of the party and Twitter and memeing and everything like that. So number one, phenomenal number two um this will be hard to convey in this short message because like it's been like blood sweat and tears for years now like the amount i care about reliability is um it's it's a lot like reliability is what this thing is if you don't understand that it'll be very hard to make like a copycat that's benchmarked like it's i feel like every nine of reliability is going to be so valuable for everyone, even if it's not the most valuable thing market cap wise, because it will just enable new applications.

26:21And like we are fighting for like all sorts of like weird nines of reliability that like we don't even fully understand because we are just like, you know, like really getting this like electric motor of AI, of intelligence, like into people's like workstations and they can figure out what to do with it. What does reliability mean in this context is this is just like availability of the model or is it like i call the model and it returns the same thing or like how do i think about reliability yeah so for something that's inherently kind of stuck so uh not so much the former thing and the second thing is closer like i would describe the first thing as kind of like uptime or slas the second thing i would maybe call closer to determinism yeah something thirdly i would consider more like robustness so robustness i would kind of describe as similar intelligence every time.

Read the full transcript

27:11Oh, interesting. Yeah. So like not exactly determinism because I think determinism, it's useful for unit tests, but not real systems. Sure. Think about like if you add a UUID to a prompt, it should be the same because it's the same functionally, but it's not exactly deterministic. Right, right, right. I think that there's another layer of it that I don't really know what it's called yet. Like maybe this is what I would call like some form of intelligence, which is it doesn't have to be the similar function every time, but it needs to be smart every time. You know, like if you were in that situation, would this be an understandable thing for a human to think?

27:43Because a developer can program around that. And actually, to me, the highest honor of reliability will be to get to the point when people can program against Jev without making example queries. Like when you just trust it, you'll be in like perma flow state, just creating crazy software. and like that's where a lot of software is today, right? Like I don't think it's totally unrealistic, but I'm going to be fine. By the way, this is kind of a weird question. So like, I mean, feel free. Like if it's too weird, just feel free. But it occurs to me that actually the value of things like coding agents goes down if you have a primitive like this in a way, which is like, you could be like, you know, whatever.

28:24Some, you know, Codex builds all the software for me, but it doesn't actually use JEV. And so like the software itself that creates is somewhat limited or you can be like, okay, I as a human being, I will write the software without using a coding agent, but I've got this very generalized primitive that makes writing software easier. So do you feel like see a future where it's like the coding agent's using JEV and then you're telling the coding agents and then do you have like redundancy or do you feel it's like human symbolism? This is more of like a coding agent question than it is a JEV question.

28:51Oh yeah, for sure. My vibe is that I'm not in the coding mind as much as I'd like to be. So you two might be in there more than I am, which is sad. But my experience is that they are really good at syntax and really... They're bad at semantics. I would say incredibly bad at architecture. Yeah. So, like, to me, architecture is, like, the most human creative part of software. So I love using coding agents. I think that Jev is almost certainly not in distribution. That would be spooky if they trained on our user data. So it's probably not. But I think that when it is in distribution, I see no problem with like having it to do the syntax.

29:35And the thing with architecture is that maybe the models are actually like not just crap at architecture, but maybe they're 50th percentile at architecture. And if you don't know anything about architecture, it would be fine. So these are all like gray area trade-offs in order for you to navigate. And sometimes speed is the knob for your company or project to turn. Like you're willing to do a 50th, like a 50th percentile architecture instead of a 60th because you want to move faster and have like codecs work overnight or something like that. Actually, kind of along those lines, one of the interesting things or phenomenons in the market already is that, you know, when the coding agents came out, it was the SaaSpocalypse and all their values dropped through the floor.

30:15And then when Jeff came out, every SaaS company is like, this is the greatest thing ever. So explain that. I don't know what else to say, right? I think it's quite natural. In the Saspocalypse story, the story that I feel like has panned out really poorly is that software is very cheap and perhaps easy to replicate, which I think I could believe the former. I could not believe the latter because a lot of this stuff happens beneath the hood. I'm maybe overly a software fanboy here. Yeah, all of us. Okay, okay, okay. I didn't know where you might be the coding agent. We have a lot of legacy around that.

30:57Yeah. So I don't think that really panned out. So SaaS seems like maybe the markets don't agree, but I think SaaS is providing the same value it used to. Maybe the markets are just scared. But I think that SaaS will be one of the largest winners of the whole AI game. And I want to work really, really well with all the biggest, most boring, most in-the-know-of-user problem SaaS companies because I think that they are the best positioned to know what workflows to automate, what do people need, like that's what their bread and butter is, and to spend the big, like, you know, software is always a CapEx investment, but like you spend it ahead of time in order to make this experience even better that gets, you know, like distributed to all of that massive users.

31:42So I think that it's going to be, I'm not going to forecast anything about the financial markets, but I think as far as like a capabilities game goes, it's going to be like an inverse saspocalypse and I am so jazzed about it. I should make a name. Yeah, I got that. Yeah, you should have a name. Sassapalooza. Whoa! That sounds a little too fun. Well, all the SaaS applications are going to all of a sudden get like dramatically more useful. And by the way, you know, the kind of capital investment, like so much of a SaaS company's capital investment is actually getting to all the customers. And so if you've gotten to all the customers and then you make, you know, not just put a chatbot on your SaaS product, but actually make the software like way, way better.

32:30That's a hell of a thing. I don't know if this is a realistic dream or not, but I think that there's a world where like the multi-choice forms just disappear. You know, like I feel like they are, like they're always like something mapping natural language that usually the software already has into like a JEV-like output. and I think it's literally it's literally from the 80s it's like it's called we used to call it 4G LST fourth generation language yeah actually also I think this is from the 80s this might be an insult I was born then like I think that do what I mean is going to be like be taken to the absolute next level if I could shout out one JF application I don't know if it's reliable so I can't promise anything but it was so freaking cool someone was using like a voice to control your computer and it was basically constantly making decisions on like, is this a command or is it inserting text?

33:25Where's it inserting text? Like that sounds so unbelievably cool. That is cool. I feel like interfaces could just completely change and maybe we're going to have to make it cheaper and faster. Yeah, then you're at Star Trek. Oh, well. You know, there's just such a profound intuition here, which is if you use AI today to generate software, right? You're still creating the same software that you did before, but if you actually look at the average PR for a large company, it's like 10 lines, right? Seriously. I've been at Google. We actually did this study, so it's like 10 lines. So you're automating 10 lines.

34:03And by the way, those 10 lines are part of a learning from a customer or something. So you've kind of optimized something that's actually pretty minimal, but what it doesn't do is provide new capabilities to the software. It's kind of automating this thing which in the limit ends up being relatively minor, and now there's actually a new capability. And so it could just be the case that just software just actually gets better. And even before Jev, it didn't even occur to me that it doesn't matter how much AI coding agents you use, the software actually isn't getting better. Maybe you're writing it faster.

34:35It's arguably getting worse just because there's less oversight. So I think this is... And often more insecure. Yeah, for sure, for sure. But you actually now can make an argument Like, apps will have new functionalities as a result of this because there is this new primitive that you're providing. I mean, in a way, it speaks natural languages and it can reason, but it marries that to a state machine. If people take that as a takeaway, that would be the greatest compliment ever to what we are doing. I actually feel like it's almost too grand of a vision to expand beyond the three logic gates that we have into like, you know, our types are kind of like one of the same logic, but like one that's like a little brain in there.

35:19Like that would be the greatest compliment to like the type safe legacy. Cause like that is, that's a very non-trivial, huge thing for the world. I'm not going to like over promise, under deliver that, but I will fight for that. Yeah. I mean, listen, I mean, there's, I think pretty open questions to what, like how deep can this get as far as like, like really serious stuff, like state consistency or durability or like real systems level stuff where you actually need to like provide strong guarantees. And so 100 % this will change things like whatever, analyzing logs, analyzing emails, providing a UI, talking to the human, like that for sure.

35:53But like, you know, you could argue that over time this becomes like a smart database, you know. And also air traffic control system, which we really need. A little scary. Like I think automate the easy work before the hardware is always my philosophy. But I also think there's going to be like an entire era of probabilistic programming that's opened up. Like, my— By the way, you know there's a huge history of probabilistic programming. That basically died in, like, the 70s. I'm familiar with it. I actually think it's going to be, like, with the same— like, you could also call Jev, like, neuropsychology.

36:28So your co-founder, Eric, came from that background, he was telling me. Oh, cool. Oh, yes, yes, yes. He did a lot of biology. It goes up and down. But like what I mean is a more, I'm not a fan. My brand is pragmatism, incredible pragmatism. I'm not a fan of like biologically inspired stuff at all. It's never worked. Have you ever noticed that? I think it's never worked. It's useful to motivate crazy people to work on things for decades until it works. And then they refine it into like the engineering version. The story of AI, neural nets for sure. Yes, but like, you know, a lot of the stories about how it worked were not accurate, right?

37:07So like the hierarchical features of applications really did end up working because like otherwise Resnets wouldn't have worked. Longer story. I do think that it opens up like from a systems perspective. I'm not excited about this part because it's really, I'm excited for the world, not about me programming this because it sounds like really complicated. But I think that as we have like lots of intelligence at lots of like different cost and speed tradeoffs, the super systemsy types will be making tradeoffs at like, you know, like Jeb's going to be like a thousand times too smart for them. They just want like an approximate link to have an approximate guess to like optimistically route here and there.

37:43It's going to be like so crazy, this type of stuff that's available in the extreme systems. And the good news is we get to like rebuild systems again, which is great, right? We have a new, no, seriously, we have a new primitive. it's kind of a new way of thinking about doing software. Like, I mean, we did this and we did this for the internet and we did this being friend to client server. I mean, we do this periodically. And by the way, just because of the cybersecurity issues, we probably have to rebuild almost all the systems to just make them safe. I would think. I think it's pretty clear that there's not...

38:17Or at least the critical infrastructure, for sure. Yeah, yeah, yeah. Do you think about this more in terms of, like, apps, SaaS, analytics, or more in terms of, like, systems, foundations, or all the above? For what I would think of, or how— Yeah, just general application for this. When you think about, like, you're working on Jev, and you kind of envision that people are adapting it. Oh, man. You know, like, maybe do you even have an opinion? I have a little bit, and it's— So the way I think of it is a little like deep into the TCP guts. You know, like UDP, TCP, you know, like it's unreliable, too reliable.

39:00Speaking my language. Exactly. So when I think of AI, and this is why I care about intelligence per dollar, to be clear, when I think, and how I got to this conclusion, I work backwards from AI-based economic revolution. AI everywhere, sci-fi and everything, like all the software has AI all over the place. And I ask myself the question, what percentage of the calls to AI, I imagine it's like a function, which what percentage are like for human consumption where you need that style? And yeah, and it's going to be like many nines. And actually from that same question, how many will be at the first layer versus like deep in the guts?

39:36Right. And I think that it's going to be many nines in the guts, but it will start at the first layer. But like we need to, if you don't aim for the guts, that's weird. If you don't aim for the guts, it's going to take you a while to get there. I think people don't understand to what extent AI was kind of shipped to the night with software. Even if you try to embed AI in software, it kind of didn't behave, right? Because software doesn't really take natural languages. And you do all this weird stuff. You stick in the prompt, like, here's the JSON output that you want, and here's the schema, and it would never listen to it.

40:06And so what you ended up doing is just taking the output and giving it to a human. You're like, to hell with it, right? Or another LLM. That is what a while loop is, like the aging while loop, right? So it's like from first principles, it needs to be human in the loop, which is the chat. Or an agent, which is the while loop, because the natural language needs to be fed back into another. Yes, and I will say, I have watched this happen. There was almost like this kind of like five stages of grief. People would pick up AI and like, I'm going to use this within my software, right? And then it would go with whatever, denial, try to make it work, and anger.

40:38Then they go to acceptance, which is like, okay, never mind. I'm just going to give this to another element to a human being. So it's been very ships in the night. I think this is the first time I have seen when someone's like, actually, you can take an LLM, you can take AI, and you can actually map it to a state machine, and you can do that productively. And I hope so. I will not want to overpromise underdeliver as well. I don't know if it's ready for all the applications that have been overpromised. I really, really want it to, and my team will fight for that, obviously. We really, really care about reliability.

41:12we could have released so much sooner. I don't think people realize that. And I don't think that, honestly, I don't think that they will. Based on what I see at the Twitter discussion, I think people will never get it, but it'll just have that good vibe of how I can trust this. Well, it's the anti-frustration machine. It's a, I hope so. I hope, do what I mean, right? To me, that is about smoothness in the world, like having everything just move more smoothly together and interlink like gears. I actually have my whole AI utopia on different axes that I really, really want. And do what I mean is a huge part of this.

41:52Imagine if all technology just did what you mean. That's not sci-fi. Look at how smart AI is, right? Yeah. No, it's amazing. And maybe that's the thought to close on. Do what I mean. Yeah, I love it. Thank you, Diogo. This has been a great conversation. I really enjoyed it. thanks for listening to this episode of the a16z podcast if you like this episode be sure to like comment subscribe leave us a rating or review and share it with your friends and family for more episodes go to youtube apple podcast and spotify follow us on x a16z and subscribe to our sub stack at a16z.substack.com thanks again for listening and i'll see you in the next episode as a reminder the content here is for informational purposes only should not be taken as legal business tax or investment advice or be used to evaluate any investment or security and is not directed at any investors or potential investors in any a16z fund please note that a16z and its affiliates may also maintain investments in the companies discussed in this podcast for more details including a link to our investments please see a16z.com forward slash disclosures.

From the publisher

a16z’s Ben Horowitz and Martin Casado sit down with TypeSafe AI founder Diogo Almeida to ask a simple question: AI has become remarkably capable, so where is all the automation?

Diogo argues that coding agents may help us write software faster, but the software they produce still largely works the way software always has. TypeSafe is taking a different approach with Jev: putting intelligence inside software itself, so developers can build programs that reason about intent and make probabilistic decisions rather than simply generate text for a human to interpret.

They discuss why reliability is the key to making AI genuinely programmable, how this could open a new era of probabilistic software, and why established SaaS companies may be particularly well positioned to benefit. Ultimately, Diogo’s goal is straightforward: technology that can reliably “do what I mean.”


Resources:

Follow Diogo Almeida: https://x.com/CompleteSkeptic

Learn more about TypeSafe AI: https://typesafe.ai/

Follow TypeSafe AI: https://x.com/typesafeai

Follow Ben Horowitz on X: https://x.com/bhorowitz

Follow Martin Casado on X: https://x.com/martin_casado

Stay Updated:

Find a16z on YouTube: YouTube

Find a16z on X

Find a16z on LinkedIn

Listen to the a16z Show on Spotify

Listen to the a16z Show on Apple Podcasts

Follow our host: https://twitter.com/eriktorenberg

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

More from The a16z Show

All 489 episodes
AI Can Write Code. Why Isn’t Software Better?The a16z Show · 43 min
Listen in VO