In short
a16z Podcast Episode Notes
Episode Title Marc Andreessen and Amjad Masad: English As the New Programming Language
Episode Description Amjad Masad, founder and CEO of Replit, joins Marc Andreessen and Erik Torenberg to discuss the emergence of AI agents, the future of programming, and how software is evolving to create software. Key topics include the historical context of computing, the capabilities of AI agents, and the innovative approach of using natural language for programming.
---
Key Concepts
- The Evolution of Programming
- Historical Context: Discussion on major leaps in programming technology over decades, culminating in the rise of AI.
- AI Agents: AI systems that can plan, reason, and code independently for extended periods.
- Replit's Role
- User Experience: Simplifying coding through natural language. Users can describe what they want in plain English, and Replit's AI translates that into code.
- Focus on Ideas: By removing the complexity of development environments, Replit allows users to concentrate on their ideas rather than the technical details.
- Importance of Language in Programming
- End of Syntax: The idea that future programming may rely more on natural language than on traditional programming syntax.
- Global Language Support: Replit's AI can support multiple languages beyond English, catering to a global audience.
- AI and Reasoning
- Reinforcement Learning: How reinforcement learning has enhanced AI's ability to reason and solve problems over long contexts.
- Verification Loops: The integration of verification loops to maintain coherence in AI reasoning, enabling agents to perform complex tasks effectively.
- Challenges and Future of AI
- Coherence Over Time: Discussion on the duration AI agents can maintain coherent reasoning without errors.
- Diminishing Returns: Potential stagnation in AI capabilities and the debate over whether current models are hitting diminishing returns.
- Amjad Masad's Journey
- Background: His story from hacking university databases to founding Replit, emphasizing his passion for programming and innovation.
- Security Incident: A recount of how he navigated a hacking incident at his university, which ultimately led to a positive outcome and shaped his career trajectory.
---
Key Takeaways
- AI's Impact on Programming: The landscape of programming is changing, with AI tools enabling non-programmers to code through natural language descriptions.
- Future of Software Development: The next phase may involve users interacting with multiple AI agents simultaneously to streamline the development process.
- Cultural Reflection on Coding: The traditionalists in programming often resist changes brought about by new technologies, mirroring historical trends observed in the evolution of programming languages.
---
Conclusion The dialogue between Andreessen, Torenberg, and Masad highlights the transformative potential of AI in software development and the necessity of adapting to these advancements in technology. As programming becomes more accessible through natural language, the landscape of software engineering is poised for significant evolution.
---
Resources
- Follow Amjad Masad on X: [@amasad](https://x.com/amasad)
- Follow Marc Andreessen on X: [@pmarca](https://x.com/pmarca)
- Follow Erik Torenberg on X: [@eriktorenberg](https://x.com/eriktorenberg)
- Listen to the a16z Podcast: [Spotify](https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX) | [Apple Podcasts](https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711)
- Find more episodes and updates at [a16z.com](https://a16z.com)
---
Disclaimer: The content of this episode is for informational purposes only and should not be considered as legal, business, tax, or investment advice.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00We're dealing with magic here that we probably all would have thought was impossible five years ago or certainly 10 years ago. This is the most amazing technology ever, and it's moving really fast, and yet we're still, like, really disappointed. Like, it's not moving fast enough, and, like, it's, like, maybe right on the verge of falling out. We should both be, like, hyper-excited, but also on the verge of, like, slitting our wrists. Because, like, you know, the gravy train is coming to an end. Right. It is faster, but it's not at computer speed, right? Right. What we expect computer speed to be.
0:22It's sort of like watching a person work. It's like watching John Carmack on cocaine. The world's best programmer on a stimulant. On a stimulant. Yeah, that's right. Every few decades, programming takes a massive leap forward, and this might be the biggest one yet. In this episode, Mark Andreessen and I are joined by Amjad Massad, CEO and founder of Replit, to talk about how AI agents are changing what it means to code. We discuss the end of syntax, the rise of agents that can think and build software for hours, and how reinforcement learning and verification loops are pushing AI towards something that looks a lot like reasoning.
1:00And finally, Amjad shares his story from hacking his university database in Jordan to building one of the most powerful developer tools in the world. Let's get into it. So let's start with, let's assume that I'm a sort of a novice programmer. So maybe I'm a student, or maybe I'm just somebody who took a few coding classes and I've hacked around a little bit, or I don't know, I do Excel macros or something like that. But I'm like not, as well as I'm not like a master craftsman of coding. and somebody tells me about Replit and specifically AI and Replit. What's my experience when I launch in with what Replit is today with AI?
1:33Yeah, I think the experience of someone with no coding experience or some coding experience is largely the same when you go into Replit. The first thing we try to do is get all the nonsense away from setting up development environment and all of that stuff and just have you focus on your idea. So what do you want to build? Do you want to build a product? Do you want to solve a problem? Do you want to do a data visualization? So the prompt box is really open for you. You can put in anything there. So let's say you want to build a startup. You have an idea for a startup. I would start with a paragraph long kind of description of what I want to build.
2:05The agents will read that. It will just type in English. Standard English. You just type it in. I want to sell crepes online. So you just type in, I want to sell crepes online. It literally could be that four words or five words. Or it could be if you have a programming language you prefer or a stack you prefer, you could do that. but we actually prefer for you not to do that because we're going to pick the best thing for, we're going to classify the best stack for that request. If it's a data app, we'll pick Python, Streamlit, whatever. If it's like a web app, we'll pick JavaScript and Postgres and things like that.
2:37So you just type that. Or you can decide, you can say, and I want to do it in, I know Python or I'm learning Python in school and I want to do it in Python. That's right. The cool thing about Replet is we've been around for almost 10 years now and we built all this infrastructure. Replet runs any program language. So if you're comfortable with Python, you can go in and do that for sure. And then just again, I know this is obvious people have used it, but like I'm dealing in English. Yes. Go ahead. Yes, you're fully in English. I mean, just a little bit of sort of background here. Like when I came here and pitched to you like 10 years ago or whatever, seven years ago, what we were saying is we were exactly describing this future is that everyone would want to build software.
3:15And the thing that's kind of getting in people's ways is all the, what Fred Brooks called the accidental complexity of programming, right? They're like essential complexity, which is like, how do I bring my startup to market and how do I build a business and all of that. Accidental complexity is what package manager do I use and all of that stuff. We've been abstracting away that for so many years. And the last thing we had to abstract away is code. I had this realization last year, which is, I think we built an amazing platform, but the business is not performing. and the reason the business is not performing is that code is the bottleneck.
3:50Yes, all the other stuff is important to solve, but syntax is still an issue. Syntax is just an unnatural thing for people. So ultimately, English is the programming language. By the way, just to clarify, does it work with other world languages other than English at this point? Yes, you can write in Japanese, and we have a lot of users, especially Japanese. That tends to be very popular. So does it support these days? Does AI support every language, or do you still have to do custom work to craft a new language? No, most mainstream languages that has 100 million plus people who speak at AI is pretty good at it.
4:20Okay, yeah. Yeah, wow. So I did a bit of historical research recently for some reason. I just want to understand the moment we're in because it's such a special moment. It's important to contextualize it. And I read this quote from Grace Hopper. So Grace Hopper invented the compiler, as you know. At the time, people were programming machine code and that's what programmers do. That's what the specialists do. Yes. And she said, specialists will always be a specialist. They have to learn the underlying machinery of computers. But I want to get to a world where people are programming English. That's what she said.
4:53That's before Kurt Plathie, right? That's 75 years ago. And that's why I invented the compiler. And in her mind, like C programming is English. But that was just the start of it. You had C and then you go higher level, Python, JavaScript. And I think we're at a moment where it's the next step. instead of typing syntax, you're actually typing thoughts, which is what we ultimately want. And the machine writes the code. And the machine writes the code. Right, right. Yeah, I remember, you're probably not old enough to remember, but I remember when I was a kid, there were higher level languages by the 70s, like BASIC and so forth and Fortran and C, but you still would run into people who were doing assembly programming, assembly language, which by the way, you still do, like game companies or whatever still do assembly to get out.
5:31And they were hating on the kids that are doing BASIC. So the assembly people were hating on the kids doing BASIC, but there were also older coders who hated on the assembly programmers for doing assembly and not... and not doing direct machine code. Right. Not doing direct zero-in-one machine code. So people don't know. Assembly language is sort of this very low-level programming language that sort of compiles to actual machine code. Yeah. It's incomprehensible gibberish to most programmers. Even most programmers. You're writing in octal or something. You're writing like very, very close to the harbor.
5:54But even still, it's still a language that compiles to zeros and ones. Right. Whereas the actual real programmers actually wrote in zeros and ones. Yeah. And so there's always this tendency for the pros to be looked on the nose. Yeah, it does. And say the new people are being basically sloppy. They don't understand what's happening. They don't really understand the machine. And then, of course, what the higher level abstractions do is democratize. The absolute irony is I was part of the JavaScript revolution. I was at Facebook before starting Replit. And we built the modern JavaScript stack. We built React.js and all the tooling around it.
6:22And we got a lot of hate from the programmers that you should type vanilla JavaScript directly. And I was like, okay, whatever. And now that's mainstream. And now those guys that built their careers on the last wave we invented are hating on this new wave. People never change. Okay, got it. Okay, so you're typing English, I want to sell grapes online, I want to do this, I want to have a t-shirt, whatever the business is. Okay, what happens then? Yeah, and then a Repl.Agent will show you what it understood. So it's trying to build a common understanding between you and it. And I think there's a lot of things we can do better there in terms of UI, but for now I'll show you a list of tasks.
6:56I'll tell you, I'm going to go set up a database because you need to store your data somewhere. We need to set up Shopify or Stripe because we need to accept payments. And then it shows you this list and gives you two options initially. Do you want to start with a design so that we can iterate back and forth to get locked that design down? Or do you want to build a full thing? Hey, if you want to build a full thing, we'll go for 20, 30, 40 minutes and the agent will tell you, go here, install the app. I'm going to go set up the database, do the migrations, write the SQL, build the site. I'm going to also test it.
7:27So this is a recent innovation we did with Agent 3 is that after it writes the software, spins up a browser, goes around and tests in the browser and then any issue, it kind of iterates, kind of goes and fix the code. So we'll spend 20, 30 minutes building that. I'll send you a notification. I'll tell you the app is ready. So you can test it on your phone. You go back to your computer. You'll see, maybe you'll find a bug or an issue. You'll describe it to the agent. I'll say, hey, it's not exactly doing what I expected. Or if it's perfect, you're ready to go. And that's it. By the way, there's a lot of examples where people just get their idea in 20, 30 minutes, which is amazing.
8:01You just hit publish. You hit publish, couple clicks. You'll be up in the cloud. We'll set up a virtual machine in the cloud. the database is deployed, everything's done, and now you have a production database. So think about the steps needed just two or three years ago in order to get to that step. You have to set up your local development environment, you have to sign up for an AWS account, you have to provision the databases, the virtual machines, you have to create the entire deploying pipeline. All of that is done for you, and it's just, a kid can do it, a layperson can do it. If you're a programmer and you're curious about what the agent did, the cool thing about Replit, because we have this history of being an IDE, you can peel the layers.
8:41You can open the file tree and you look at the files. You can open GITs. You can push to GitHub. You can connect it to your editor if you want. You can open an Emacs. So the cool thing about Replit, yes, it is a Vibe coding platform that abstracts away all the complexities, but all the layers are there for you to look at. That was great. But let's go back to you said it. You say, I've got my IDE. You plug it in and it says, it gives you this list of things. And then when you describe it, you said, I'm going to do this, I'm going to do that. The I there in that case was the agent as opposed to the user.
9:06Yes. And so the agent lists the set of things that it's going to do, and then the agent actually does those things. Agent does those things. Okay. Yeah, that's a very important point. When we did this shift, we hadn't realized internally at Replit how much the actual user stopped being the human user, and it's actually the agent programmer. Right. So one really funny thing happened is we had servers in Asia. And the reason we had servers in Asia, because we wanted our Indian or Japanese users to have a shorter time to the servers. When we launched Agent, their experience got significantly worse.
9:38And we're like, what happened? Like, it's supposed to be faster. Well, it turns out it's worse. It's because the AIs are sitting in the United States. And so the programmer is actually in the United States. You're sending the request to the programmer, and the programmer is interfacing with a machine across the world. And so, yes, suddenly the agent is the programmer. Okay, so the new terminology agent is a software program that is basically using the rest of the system as if it were a human user, but it's not. It's a bot. That's right. It has access to tools such as write a file, edit a file, delete a file, search the package index, install a package, provision a database, provision object storage.
10:15It is a programmer that has the tools and interface. It has sort of an interface that is very similar to human programmer. And then, you know, we'll talk more about how this all works, but a debate inside the AI industry is with these, was kind of this, you know, this idea now of having agents that do things on your behalf and then go out, you know, go out and kind of accomplish missions. There's this, you know, kind of debate, which is, okay, how, like, obviously, you know, it's a big deal even to have an AI agent that can do relatively simple things. To do complex things, of course, is, you know, one of the great technical challenges of the last 80 years, you know, to do that.
10:46And then there's sort of this question of, like, Like, can the agent go out and run and operate on its own for five minutes, you know, for 15 minutes, for an hour, for eight hours? And meaning, like, sort of like, how long does it maintain coherence? Like, how long does it actually, like, stay in full control of its faculties and not kind of spin out? Because at least the early agents or the early AIs, if you set them off to do this, they might be able to run for two or three minutes. And then they would start to get confused and go down rabbit holes and, you know, kind of spin out. More recently, we've seen that agents can run a lot longer and do more complex tasks.
11:18Where are we on the curve of agents being able to run for how long and for what complexity tasks before they break? That's absolutely the main metric we're looking at. Even back in 2023, we've had the idea for software agents four or five years ago now. The problem every time we attempt them, the problem of coherence. You know, they'll go on for a minute or two and then they'll just, you know, compound in errors in a way that they just can't recover. And you can actually see it, right? Because they actually, if you watch them operate, they get increasingly confused and then, you know, maybe even deranged.
11:53Yeah, very deranged and they go into very weird areas and sometimes they start speaking Chinese and doing really weird things. But I would say sometime around last year, we maybe crossed a three, four, five minute mark. And it fell to us that, okay, we're on a path where long horizon reasoning is getting solved. And so we made a bet. And I tell my team— So long horizon reasoning meaning like dealing in like facts and logic in a sort of complex way. and the long horizon being over a long period of time. Yes. With many, many steps to a reasoning process. Yeah, that's right. So if you think about the way large language models work is that they have a context.
12:39This context is basically the memory, all the texts, all your prompts, and also all the internal talk that the AI is doing as its reasoning. So when the AI is reasoning, it's actually talking to itself. It's like, oh, now I need to go set up a database. Well, what kind of tool do I have? Oh, there's a tool here that says Postgres. Okay, let me try using that. Okay, I use that. I got feedback. Let me look at the feedback and read it and I'll read the feedback. And so that prompt box or context is where both the user input, the environment input, and the internal thoughts of the machine are all within.
13:15It's sort of like a program memory in memory space. And so reasoning over that was the challenge for a long time. That's when AI just like went off track. And now they're able to kind of think through this entire thing and maintain coherence. And there's now techniques around compression of contacts. So there's still, the contact length is still a problem, right? So I would say LLMs today, they're marketed as a million token length, which is like a million words almost. In reality, it's about 200 ,000 and then they start to struggle. So we do a lot of, we stop, we compress the memory. So if a memory, if a portion of the memory is saying that I'm getting all the logs from the database, you can summarize paragraphs of logs with one statement or the database setup.
14:06That's it, right? And so every once in a while we'll compress the context so that we make sure we maintain coherence. So there's a lot of innovation happened outside of the foundation models as well in order to enable that long context coherence. And what was the key technical breakthrough in the foundation models that made this possible, do you think? I think it's RL. I think it's reinforcement learning. So the way pre-training works is, you know, pre-training is the first step of training a large language model. It reads a piece of text. It covers the last words and tries to guess it. That's how it's trained.
14:42That doesn't really imply long context reasoning. It turns out to be very, very effective. It can learn language that way. But the reason we weren't able to move past that limitation is that that modality of training just wasn't good enough. And what you want is you want a type of problem solving over long contacts. So what reinforcement learning, especially from Codexution, gave us is the ability for the LLM to roll out what we call trajectories in AI. So a trajectory is a step-by-step reasoning chain in order to reach a solution. So the way, as I understand, reinforcement learning works is they put the LLM in a programming environment like Replit and say, hey, here's a code base, here's a bug in the code base, and we want you to solve it.
15:43Now, the human trainer already knows what the solution would look like. So we have a pull request that we have on GitHub, so we know exactly, or we have a unit test that we can run and verify the solution. So what it does is it rolls out a lot of different trajectories. They sample the model. And maybe one of those trajectories will reach, and a lot of them will just go off track, but one of them will reach the solution by solving the bug, and it reinforces on that. So that gets a reward, and the model gets trained that, okay, this is how you solve these type of problems. So that's how we're able to extend these reasoning chains.
16:17Got it. And how, as a two-part question is, how good are the models now at long reasoning? and I would say, and how do we know? Like, how is that established? There is a nonprofit called Meter that is measuring useful, has a benchmark to measure how long a model runs while maintaining coherence and doing useful things, whether it's programming or other benchmark tasks that they've done. And they put up a paper, I think, late last year that said every seven months, the minutes that a model can run is doubling. So you go from two minutes to, you know, four minutes in seven months. I think they've vastly underestimated that.
17:08Is that right? Vastly. It's doubling more often than seven months. So agent three, we measure that, you know, very closely. And we measure that in real tasks from real users. So we're not doing benchmarking. We're actually doing A-B tests and we're looking at the data that how users are successful or not. For us, the absolute sign of success is you made an app and you published it. Because when you publish it, you're paying extra money. You're saying this app is economically useful. I'm going to publish it. So that's as clear cut as possible. And so what we're seeing is in agent one, the agent can run for two minutes and then perhaps struggle.
17:43Agent two came out in February. It ran for 20 minutes. Agent three, 200 minutes. Some users are pushing it to like 12 hours and things like that, I'm less confident that it is as good when it goes to these stratospheres. But at like two, three hours timeline, it is really, it's insanely good. And the main innovation outside of the models is a verification loop. Actually, I remember reading a research paper from NVIDIA. So what NVIDIA did is they're trying to write GPU kernels using DeepSeek. And that was like perhaps seven months ago when DeepSeek came out. And what they found is that if we add a verify in the loop, if we can run the kernel and verify it's working, we're able to run DeepSeek for like 20 minutes.
18:33And it was generating actually optimized kernels. And so I was like, okay, the next thing for us, obviously as a sort of agent lab or like AppLay, our company, we're not doing the foundation model stuff, but we're doing a lot of research on top of that. And so, okay, we know that agents can run for 10, 20 minutes now or LLMs can stay coherent for longer. But for you to push them to 200, 300 minutes, you need a verifier in the loop. So that's why we spend all our time creating scaffolds to make it so that the agent can spin up a browser and do computer use style testing. So once you put that in the middle, what's happening is it works for 20 minutes, It spins up another agent, spins up a browser, tests the work of the previous agent, so it's a multi-agent system.
19:23And if it is, if it finds a bug, it starts a new trajectory and says, okay, good work. Let's summarize what you did the last 20 minutes. Now that plus the bug that we found, that's a prompt for a new trajectory. So you stack those on each other and you can go endlessly. So it's like setting up a marathon or like a relay race. as long as each step is done properly you could do in sort of an infinite number of steps that's right you can always compress the previous step into a paragraph and that becomes a prompt so it's an agent prompting the next agent right right right that's amazing so and then when an agent like when a modern agent like running on modern modern LMs that are trained this way when it let's say it runs for 200 minutes like when you watch the agent run is it like running is it like processing through like logic and tasks at the same pace that like a human being is or slower or faster it's actually I would say it is faster, but not that much significantly faster.
20:18It's not at computer speed, right? What we expect computer speed to be. It's like watching a person. Like if you watch that, if it's describing what it's doing, it's sort of like watching a person work. It's like watching John Carmack on cocaine work. The world's best programmer. Yeah. The world's best programmer on a stimulus. On a stimulus. Yeah, that's right. Working for you. Yeah, so it's very fast and you can see the file lifts running through. But every once in a while, it'll stop and it'll start thinking, I'll show you the reasoning. It's like, I did this and I did this. Am I on the right track?
20:51It kind of really tries to reflect. And then it might review its work and decide the next step or it might kick into the testing agent. Or, you know, so you're seeing it do all of that. And every once in a while, it calls a tool, for example, it stops and says, well, we ran into an issue, you know, Postgres 15 is not, compatible with this database ORM package that I have. Okay, this is a problem I haven't seen before. I'm going to go search the web. So it has a web search tool, go do that. And so it looks like a human programmer. And it's really fascinating to watch it. So one of my favorite things to do is just to watch the tool chain and the reasoning chain and the testing chain.
21:34Yeah, it is like watching a hyper-productive programmer. So, you know, we're kind of getting into here kind of the holy grail of AI, which is sort of, you know, generalized reasoning, you know, by the machine. So you mentioned this a couple of times with this idea of verification. So, so just for folks on the listening podcast who maybe aren't in the details, let me try to describe this and see if I have it right. So like just a, just a large language model, the way you would experience, you would have experienced with like chat GPT out of the gate two years ago or whatever would have been. It's like, it's incredible how fluid it is at language.
22:06It's incredible how good it is at like writing Shakespearean sonnets or rap lyrics. It's amazing how good it is at human conversation. But if you start to ask it like problems that involve like rational thinking or problem solving, all of a sudden, like you'd build math or the math, the whole show. And in the very beginning, it was, you could ask if you asked a very basic math problems that, you know, it would, it would not be able to do them. That's right. But then even when it got better at those, if you started to ask it to like, you know, it could maybe add two small numbers together, but it couldn't add two large numbers together.
22:31Or if it could add two large numbers, it couldn't multiply them. Yeah. And it's just like, all right, this is, and then it had this, there was this famous, the famous was the strawberry test, the famous strawberry test, which is how many R's are in the word strawberry? That's right. And there was this long period where it kept, it would just guess wrong. It would say there were only two R's in the word strawberry and it turns out there are three. So it was this thing, and so people were, and there was even this term that was being used, kind of the slur that was being used at the time was stochastic parrot.
22:57Yeah, I was thinking clanker. Well, clanker is the new slur. Clanker is just the full-on racial slur, I guess, AI is a species. But the technical critique was so-called stochastic parrot. Stochastic means random. So sort of random parrot, meaning basically that this thing was sort of a, large language models were like a mirage, where they were like repeating back to you things that they thought that you wanted to hear, but they didn't. And in a way, it's true in the pure pre-training LLM world. Right, for the very basic layer. But then what happened is, as you said, over the last year or something, there was this layering in of reinforcement learning.
23:31It's not new, crucially. it's AlphaGo, right? Describe that for a second. Yeah, so we had this breakthrough before and 2015 was the AlphaGo breakthrough, I think, 2015, 2016, where it is emerging of sort of, you know, you would know a lot better than me, the old AI debate between the connectionists, the people who think neural networks are the true sort of way of doing AI, and the symbolic systems, I think, or like the people that think that, you know, discrete reasonings, F-states, mints, knowledge bases, whatever, this is the way to go. And so there was a merging of these two worlds where the way AlphaGo worked is it had a neural network, but it had a Monte Carlo tree search algorithm on top of that.
Read the full transcript
24:17So the neural network would generate, would like generate a list of potential moves. And then you had a more discrete algorithm sort those moves and find the best based on just tree search, based on just trying to verify. Again, this is sort of a verifier in the loop, trying to verify which move might yield the best based on more classical way of doing algorithms. And so that's a resurgence of that movement where we have this amazing generative neural network that is the LLM. And now let's layer on more discrete ways of trying to verify whether it's doing the right thing or not. And let's put that in a training loop.
24:59And once you do that, the LLM will start gaining new capabilities, such as reasoning over math and code and things like that. Exactly right. Okay. And then that's great. And then the key thing there, though, for our elder work, for LLMs to reason, the key is that it be a problem statement that there is a defined and verifiable answer. That's right. Is that right? And so, and you might think about this as like, let's give a bunch of examples. Like in medicine, this might be like, you know, a diagnosis that like a panel of human doctors agrees with. or, by the way, or a diagnosis that actually, you know, solves the condition.
25:30In law, this would be a, you know, an argument that in front of a jury actually results in an acquittal or something like that. In math, it's an equation that actually solves properly. In physics, it's a result that actually works in the real world. I don't know, in civil engineering, it's a bridge that doesn't collapse, right? So there's always some test of practice. The first two do not work very well just yet. Like, I would say law and healthcare, they're still a little too squishy, a little too soft. It's unlike math or code. Like, the way they're training on math, they're using this sort of like a program language, provable language called Lean for proofs, right?
26:11So you can run a Lean statement. You can run a computer code. Perhaps you can run a physics simulation or a civil engineering sort of physics simulation. But you can't run a diagnosis. Okay. So I would say that... But you can verify it with human answers or not. Yeah, so that's a more HF in a way. So it is not the sort of autonomous RL, like fully scalable autonomous, which is why coding is moving faster than any other domain, is because we can generate these problems and verify them on the fly. But there's two, with coding, as anybody who's coded knows, there's coding, there's two tests, which is one is, does it code compile?
26:50Right. And then the other is, does it produce the right output? And just because it compiles doesn't mean it produces the right output. You tell me, but verifying that it's the correct output is harder. Yeah, so Sweebench is a collection of verified pull requests end states. So it is not just about compiling. So Sweebench is the main benchmark used to test whether AI is good at software engineering tasks. And we're almost saturating that. So last year, we're at like maybe 5 % early 24 or less. And now we're like 82 % or something like that with cloths on at 4.5. That's state of the art. And that's like a really nice hell climb that's happening right now.
27:36And basically, they went and looked on GitHub. They found the most complex repositories. They found bug statements that are very clear. And they found ProQuest that actually solved those bug statements with unit tests and everything. So there is an existing corpus on GitHub of tasks that the AIs can solve. And you can also generate them. Those are not too hard to generate what's called synthetic data. But you're right. It's not infinitely scalable because some human verifiers still need to kind of look at the task. But maybe the foundation models have found a way to have the synthetic training go all the way.
28:18Right. And then what's happening, I think, because what's happening is the foundation model companies are, in some cases, they're actually hiring human experts to generate new training data. Yes. So they're actually hiring mathematicians and physicists and coders to basically sit, and they're hiring human programmers, putting them on the cocaine. Yes. And having them, probably coffee. Yeah. And having them actually write code, and then write code in a way where there's a known result of the code running such that this RL loop can be trained properly. That's right. And then the other thing these companies are doing is they're building systems where the software itself generates the training data, generates the tests, generates the validated results, and that so-called synthetic training data.
28:58That's right. But again, those work in the very hard domains. It works to some extent in the software domains. And I think there's some transfer learning. You can see the reasoning work when it comes to tools like deep research and things like that. But we're not making as fast as progress in the more soft domains. So, softer domains, meaning domains in which it's harder or even impossible to actually verify correctness of result in a sort of a deterministic, factual, grounded, non-controversial way. Like if you have a chronic disease, you could have, you know, you have a POTS or, you know, whatever, EDS syndrome.
29:40And they're all clusters. Because it is the domain of abstraction. It is not as concrete as code and math and things like that. So I think there's still a long ways to go there. Right. So sort of the more concrete the problem. Like it's the concreteness of the problem that is the key variable, not the difficulty of the problem. Would that be a way to think about it? Yeah, I think the concreteness in a sense of can you get a true or false verifiable, I put it. Right. But like in any domain of human effort in which there's a verifiable answer, we should expect extremely rapid progress? Yes. Right.
30:11Okay. Absolutely. And I think that's what we're seeing. Right. And that for sure includes math. That for sure includes physics. For sure includes chemistry. For sure includes large areas of code. That's right. Right. What else does that include, do you think? Bio, like we're seeing with a protein folding. Like genomic, genomic. Yeah, yeah. Things like that. I think some areas of robotics, there's a clear outcome. But it's not that many, I mean, surprisingly. Well, it depends on your point of view. Some people might say that's a lot. And then you mentioned the pace of improvement. So what would you expect from the pace of improvement going forward for this?
30:48I think we're ripping on coding. I think it's just, it's going. I think it's going to be like what we're working on with Agent 4 right now is by next year, we think you're going to be sitting in front of Replit and you're shooting off multiple agents at a time. You're like planning a new feature. So I want a social network on top of my storefront. And another one is like, hey, refactor the database in your running parallel agents. So you have about five, 10 agents kind of working in the background and they're merging the code and taking care of all of that. But you also have a really nice interface on top of that, that you're doing design and you're interacting with AI in a more creative way, maybe using visuals and charts and things like that.
31:36So there's a multimodal angle of that interaction. So I think, you know, creating software is going to be such an exciting area. And I think that the lay person will be as good as what a senior software engineer that works at Google is today. So I think that's happening very soon. But, you know, I don't see them, and I'd be curious about your point of view, but like my experience between as a sort of a, you know, on the, let's say, healthcare side or more, you know, write me an essay side or more creative side, I haven't seen as much of a rapid improvement as what we're seeing in code. So I think code is going to go to the moon.
32:18Math is probably as well. Some, you know, scientific domains, bio, things like that, those are going to move really fast. Yeah. So there's this weird dynamic. Let's see if you agree with this. And Eric also, I'm curious your point of view on this. Like, there's this weird dynamic that we have, and we have this in the office here a lot, and I also have this with, like, the leading entrepreneurs a lot, which is this thing of, like, like, wow, this is the most amazing technology ever, and it's moving really fast, and yet we're still, like, really disappointed. And, like, it's not moving fast enough, and, like, it's, like, maybe right on the verge of stalling out.
32:48And like, you know, we should both be like hyper excited, but also on the verge of like slitting our wrists. Because like, you know, the gravy train is coming to an end. And I always wonder, it's like, you know, on the one hand, it's like, okay, like, you know, not all, I don't know, ladders go to the moon. Like, just because something, you know, looks like it works or, you know, doesn't mean it's going to, you know, be able to scale it up and have it work, you know, to the fullest extent. You know, so like, it's important to like recognize practical limits and not just extrapolate everything to infinity.
33:13On the other hand, like, you know, we're dealing with magic here that we, I think, probably all would have thought was impossible five years ago or certainly 10 years ago. Like, I didn't, you know, like, I, you know, I got my CS degree in the late 80s, early 90s. I never, I didn't think I would live to see any of this, right? Like, this is just amazing that this is actually happening in my lifetime. But there's a huge bet on AGI, right? Like, whether it's the foundation models, I think, you know, now the entire U.S. economy is sort of a bet on AGI. and there are crucial questions to ask whether are we on track to AGI or not.
33:43Because there are some ways that I can tell you it doesn't seem like we're on track to AGI because there doesn't seem to be transfer learning across these domains that are, you know, significant, right? So if we get a lot better at code, we're not immediately getting better at like generalized reasoning. We need to go also, you know, get training data and create RL environment for bio or chemistry or physics or math or law. And this has been the sort of point of discussion now in the AI community after the Dworkish and Richard Sutton interview where Richard Sutton kind of poured this cold water on the bitter lesson.
34:25So everyone was using this essay that he wrote called The Bitter Lesson. The idea is that there are infinitely scalable ways of doing AI research. And any time you can pour more compute and more data and get more performance out, that's the ultimate way of getting to AGI. And some people interpreted that interview that perhaps he's doubtful that we're even on a better lesson path here. And perhaps the current training regime is actually very much the opposite in which we are so dependent on human data and human annotation and all of that stuff. So I think that I agree with you. I mean, as a company, we're excited about where things are headed.
35:15But there's a question of like, are we on track to AGI or not? And be curious what you think. So, and, you know, I think, you know, Ilya Sescover makes a specific form of this argument, which is basically like we're just literally running out of training data. It's a fossil fuel argument. Right. Like we've slurped all the training data. Fundamentally, we've slurped all the data off the internet. That is where almost all the data is at this point. There's a little bit more data that's in like, you know, private dark pools somewhere that we're going to go get. But we have it all. And then, right, we're in this business now trying to generate new data.
35:42But generating new data is hard and expensive, you know, compared to just like slurping things off the internet. So there are these arguments. You know, having said that, you know, you get a definitional question here really quick, which are kind of a rabbit hole. But having said that, like you mentioned transfer learning. So transfer learning is the ability of the machine to be an expert in one domain and then generalize that into another domain. My answer to that is like, have you met people? And how many people do you know are able to do transfer learning? Not many. Not many, right? Well, because there's quite the opposite, actually.
36:09The nerdier they are in a certain domain, they kind of, you know, often they have blind spots. We joke about how everyone's just retarded in one area, or they make some massive mistake, and don't trust them on this, but on this other topic. Right. Yeah, well, and this is a well-known thing among, for example, public intellectuals. So this happens, there's actually been whole books written about this on so-called public intellectuals. So you get these people who show up on TV, and they're experts. And what happens is, they're like an expert in economics, right? And then they show up on TV, and they talk about politics, and they don't know anything about politics.
36:36Right? Or they don't know anything about medicine, or they don't know anything about the law, or they don't know anything about computers. You know, this is the Paul Gruckman talking about how the internet is going to be no more significant than the fax machine. Fax, yeah. He's a brilliant economist. He has no idea how a computer works. Is he a brilliant economist? Well, at one point. At one point. At one point. Let's get, even if, even if he's a brilliant, well, this is the thing, like, what does that mean? Should a brilliant economist be able to extrapolate, you know, the internet is a good question.
37:02But the point being, like, even if he is a brilliant, you know, or take anybody, oh, by the way, or like Einstein's like actually my favorite example. I think you'd agree Einstein was a brilliant physicist. He was a Stalinist. Like, he was a socialist and he was a Stalinist. And he was like, he thought, like, Stalin was fantastic. Well, he's already out, so. Yeah, okay, all right. True socialism. All right, Einstein. You know, I'll take your word for it. But, like, once he got into politics, he was just, like, totally loopy. Or, you know, even right or wrong, he just sounded like all of a sudden like an undergraduate lunatic, like somebody in a dorm room.
37:36Like, there was no transfer learning from physics into politics. Like, he was right or wrong. He didn't, there was no, there was clearly, there was nothing new in his political analysis. It was the same routine bullshit you get out of, you know. Yeah, so in a way, the argument you're making is like, we may be already at human level AI. I mean, perhaps the definition of AGI is something totally different. It's like above the human level that something that truly generalizes across domains. It's not something that we've seen. Yeah, like we've idealized, yeah, as I said, we've, we've, and in your look, we should, we should shoot big, But we've idealized a goal that may be idealized in a way that, like, number one, it's just, it's like so far beyond what people can do that it's no longer a relevant comparison to people.
38:16And usually AGI is defined as, you know, able to do everything better than a person can. And it's like, well, okay, so if doing everything better than a person can, it's like if a person can't do any transfer learning at all, doing even a little bit, a marginal bit might actually be better. Or it might not matter just because no human can do it. And so, therefore, you just stack up the domains. There's also this well-known phenomenon in AI. Typically, this works the other way, which is a phenomenon in AI. AI engineers always complain about, scientists always complain about, which is the definition of AI is always the next thing that the machine can't do.
38:45And so the definition of AI for a long time was kind of beat humans at chess. And then the minute it could beat humans at chess, that was no longer AI. That was just like, oh, that's just boring. That's computer chess. It became an entire decision. Yeah, it's computer chess. It's just like boring, and now it's an app on your iPhone, and nobody cares, right? And it's immediately... The Turing test was the next one. The Turing test. And then we passed it and nobody... This is a really big deal. There was no celebration. There was no parties. That's exactly right. There was no party. For 80 years, the Turing test, I mean, they made a movie about it, like the whole thing.
39:11That was the thing. And like, we blew right through it. And nobody even registered it. Nobody cares. It gets no credit for it. We're just like, ah, it's still a complete piece of shit. Like I said, yeah. Right. And so there's this thing where, so the AI scientists are used to complaining, basically, that they're always being judged against the next thing as opposed to all the things they've already solved. But that's maybe the other side of it, which is they're also putting out for themselves an unreasonable goal, and then doing this sort of self-flagellation kind of along the way. And I kind of wonder, yeah, I wonder kind of which way that cuts.
39:42Yeah, yeah. It's an interesting question. Like, I started thinking about this idea of, like, it doesn't matter whether it's truly AGI, and the way I define AGI is that you put in an AI system in any environment and efficiently learns. You know, it doesn't have to have that much prior knowledge in order to kind of learn something, but also can transfer that knowledge across different domains. But, you know, we can get to like functional AGI. And what functional AGI is, is just, yeah, collect data on every useful economic activity in the world today and train an LLM on top of that or train the same foundation model on top of that.
40:21And we'll go, we'll target every sector economy and you can automate a big part of labor that way. So I think, yeah, I think we're on that track. For sure. You tweeted after GPT-5 came out that you were feeling the diminishing returns. What were you expecting and what needs to be done? Do we need another breakthrough to get back to the pace of growth? What are your thoughts? I mean, this whole discussion is sort of about that. And my feeling is that, you know, GPT-5 got good at verifiable domains. It didn't feel that much better at anything else. The more human angle of it, it felt like it regressed.
40:57and like you had this sort of Reddit pitchfork sort of movement against Sam and OpenAI because they felt like they lost a friend. GPT-4.0 felt a lot more human and closer, whereas GPT-5 felt a lot more robotic, very in its head, kind of trying to think through everything. And so I would have just expected like when we went from GPT-2 to 3, it was clear it was getting a lot more human. It was a lot closer to our experience. You can feel like it's actually all it gets me. Like there's something about it that understands the world better. Similarly, three to four, four to five didn't feel like it was a better overall being as it were.
41:46But is that, is that, is that, is that a, is the question there like, is it emotionality? Is it... Partly emotionality, but again, partly, like, I like to ask models, like, very controversial things. Can it reason through, I don't know how deep we want to go here, but, like, what happened with World Trade 7? Right, right. Sure. It's an interesting question, right? Like, I'm not putting out a theory, but, like, it's interesting. Like, how did it, you know, and can it think through controversial questions in the same way that it can go think through a coding problem? And there hasn't been any movement there.
42:30Like, all the reasoning and all of that stuff. And not just that, you know, that's a cute example, but like COVID, right? Like, you know, the origins of COVID. You know, go, you know, dig up GPT-4 or other models and go to GPT-5. you're not going to find that much difference of, okay, let's reason together. Let's try to figure out what was the origins of COVID. Because it's still an unanswered question, you know? And I don't see them making progress on that. I mean, you play a lot with them. Do you feel like... Yeah, I use it differently. I don't know. Maybe I have different expectations. The way I...
43:01My main use case actually is sort of PhD and everything at my beck and call. And so I'm trying to get it to explain things to me more than I'm trying to like, you know, have conversations with it maybe. Maybe I'm just unusual with that. And that gets better. Well, so what I found specifically is a combination of like GPT-5 Pro plus deep reasoning or like Rock 4 Heavy, like the highest end models like that. You know, they now basically generate 30 to 40 page, you know, essentially books on demand on any topic. And so anytime I get curious about something, maybe it's my version of it, but it's something like, I don't know.
43:35Here's a good example. When an advanced economy puts a tariff on a raw material or on a finished good, like who pays? You know, is it the consumer? Is it the importer? Is it the exporter? Or is it the producer? And this is actually a very complicated, it turns out a very complicated question. It's a big, big, big thing that economists study a lot. And it's just like, okay, you know, who pays? And what I found, like, for that kind of thing is it's outstanding. Well, but it's outstanding at sort of going out of the web, getting information, synthesizing it. Correct. It gives me a synthesized 20, 30, 40 page.
44:05It basically tops out of 40 pages of PDF. But I can get up to 40 pages of PDF, but it's a completely coherent, and as far as I can tell, for everything I've cross-checked, a completely, like, world-class... Like, if I hired, you know, for a question like that, if I hired, like, a great, you know, econ PhD postdoc at Stanford who just, like, went out and did that work, like, it would maybe be that good. But then, of course, the significance is it's like, you know, at least for, this is true for many domains, you know, kind of PhD and everything, and so... But this is synthesizing knowledge, not trying to create new knowledge.
44:36Well, but this gets to the sort of, you know, of course, you get into the angels dancing on the head of a pen thing, which is like, what's the difference? How much new knowledge ever actually is there anyway? What do you actually expect from people when you ask them questions? And so what I'm looking for is like, yes, explain this to me in like the clearest, most sophisticated, most complex, most like complete way that it's possible for somebody, you know, for a real expert to be able to explain things to me. And that's what I use it for. And again, as far as I can tell from the cross-checking, like I'm getting, you know, like almost like basically a hundred out of a hundred, like I don't even think I've had an issue in months where it's like had a problem in it.
45:11And it's like, yeah, you can say, yeah, synthesizing is supposed to create new information, but like it's generating a 40-page, it's basically generating a 40-page book. That's amazing. That's like incredibly like fluid. It's, you know, it's, you know, the logical coherence of the entire, like it's a great, like if you evaluated a human author on it, you would say, wow, that's a great author. Yeah. You know, are people who write books, you know, creating new knowledge? Well, yeah. Well, sort of not because a lot of what they're doing is building on everything that came before them as synthesizing and combining.
45:41But also like a book is a creative accomplishment, right? And so, yeah, one of the things I'm interested in, I'm hoping AI could help us solve this, just like how confusing the information ecosystem right now. You know, everything feels like propaganda. Like it doesn't feel like you're getting real information from anywhere. So I really want an AI that could help me reason from first principles about what's happening in the world for me to actually get real information. And maybe that's an unreasonable sort of ask of the AI researchers, but I don't think we have made any progress there. So maybe I'm over-focused, maybe I'm over-focused on arguing with people as opposed to trying to get down to the line truce.
46:22But, well, here's the thing I do a lot with this is I just say, like, take a provocative point of view and then steel man the position. Take your COVID thing. So I often pair these. Steel man the position that it was a lab leak and then steel man the position that it was natural origins. And again, is this creativity or not? I don't know. But what comes back is 30 pages each of like, wow. That is the most compelling case in the world I can imagine with everything marshaled against it. The argument is structured in the most possible way. Part of the reason that's not happening is because it stopped being taboo to talk about human origin.
46:52When it was taboo, the AIs would talk down to you. It's like, oh, you're a conspiracy theorist. And so there's a, you know, a period of time. And so it takes something truly controversial. And they actually, they can't reason about it because of all the RLHF and all sorts of limitations. And as you know, I won't pick those specific ones here, but like there are certain big models that will still lecture you. That you're a bad person for asking that question. But, you know, like I was just, some of them are just like really, really open now to, you know, being able to do these things. And then, yeah, so, okay.
47:28Yeah, so, okay. So, yeah, so there's this, yeah, so basically, like, ultimately what you're looking for, like, the ultimate thing would be, if there's something that's like, I don't think anybody's really defined this really well, because it's not, because, again, it's like, conventional, all the conventional definitions of AGI are, like, basically comparing to people. Yeah. And there it's always like, you know, the conventional explanations of AGI always, for me, strike me a lot, like the debate around like whether a self-driving car works or not, which is, does a self-driving car work because it's a perfect driver or does it work because it is better than the human driver?
48:00And better than the human driver, I think, is actually quite, you know, just like with the chess thing and the go thing. I actually think like that, that's like a real thing. And then there's the like, is it a perfect driver, which is, you know, what they're obviously the self-driving car companies are working for. But then I think you're looking for something beyond the perfect driver. You're looking for the car who knows where to go. So I'm of two minds. So one mind is the sort of practical entrepreneur. And I just, I have so many toys to play with, to build. Like stop AI progress today and Repl.o will continue to get better for the next five years.
48:31There's so much we can do just on the app layer and the infrastructure layer. So, you know, but I think that it will, you know, the foundation models will continue to get better as well. And so it's a very exciting time in our industry. The other mind is more academic because as a kid, I've always been interested in the nature of consciousness, nature of intelligence. I was always interested in AI and reading the literature there. And I would point to the RL literature. So Richard Sutton, there's another guy, I think co-founder of DeepMind, Shane Lagg, wrote a paper trying to define what AGI is.
49:07and in there I think that the definition of HI I think is the original perhaps correct one which is efficient continual learning like if you truly want to build an artificial general intelligence that you can drop in any domain you can drop in a car without that much prior knowledge about cars and within you know how long does it take a human to learn how to drive within months being able to drive a car really well. Generalized skill acquisition, generalized understanding acquisition, generalized reasoning acquisition. And I think that's the thing that would truly change the world. That's the thing that would give us a better understanding of the human mind, of human consciousness.
49:54And that's the thing that will propel us to the next level of human civilization, I think. So on a civilization level, I think that's a really deep question. question, but separate from the economy and the industry, which is all exciting. But there's an academic aspect of it that I'm really... And so what odds, if we're on CalSheet today, what odds do we place on that? I'm kind of bearish on true AGI breakthrough because what we built is so useful and economically valuable. So in a way... Oh, good enough. Good enough is the enemy. Yeah, yeah. Do you remember that essay? Worse is better. Worst is better.
50:33Worst is better. So there's like a local, there's like a trap. There's like a local maximum trap. That's right. We're in a local maximum trap. Local maximum trap where it's, because it's good enough for so much economically productive work, it relieves the pressure in the system to create the generalized answer. Yes. And then you have the weirdos like Rich Sutton and others that are still trying to go down that path and maybe they'll succeed. But there's enormous optimization energy behind the current thing that we're hell climbing on this like local maximum. Right, right, right. Right. And the irony of it is everybody's worried about like the, you know, the gazillions of dollars going into building out all this stuff.
51:06And so the most ironic thing in the world would be if the gazillions of dollars are going into the local maximum. That's right. As opposed to a counterfactual world in which they're going into solving the general problem. But it's also potentially irrational. Like maybe the general problem is actually, you know, not within our lifetimes. Who knows? How much further do you think, like, do you think we squeeze most of the juice out of LLMs in general then? Or are there any other research directions that you're particularly excited about? Well, that's the thing. I think the problem is there aren't that many.
51:37I think the breakthroughs in RL are incredibly exciting, but we also knew about them now for like over 10 years where you marry generative systems with TreeSearch and things like that. But there's a lot more to go there. And I think, again, the original minds behind reinforcement learning are trying to go down that path and try to kind of bootstrap intelligence from scratch. Carmack is going down that path. As far as I understand, Carmack, you guys may be invested, but they're not trying to go down the LLM path. So there are people that are trying to do that, but I'm not seeing a lot of progress or outcome there.
52:15But I watch it kind of from far. Although, you know, for all we know, there's already a bot on X somewhere. Maybe. You never know. It might not be a big announcement. It might just be, you know, one day there's just like a bot on X that starts winning all the arguments. Yeah, it could be. Or a user of Reddit and all of a sudden it's generating incredible software. Okay, let's spend our remaining minutes. Let's talk about you. So, yeah, take us and start from the beginning with your life. And how did you get from being born and being in Silicon Valley? Okay. In two minutes. Yeah. I'm joking. Yeah.
52:51I got introduced to computers very, very early on. And so for whatever reason, so I was born in Amman, Jordan. And for whatever reason, my dad, who was just a government engineer at the time, decided that computers were important. And he didn't have a lot of money, took out a dad, bought a computer. It was the first computer in our neighborhood. First computer of anyone I know. and I just, one of my earliest memories, I was six years old, just watching my dad unpack this machine and sort of open up this huge manual and kind of finger type CD, LS, MKDIR and like I would, you know, be behind his shoulder and just like watching him, you know, type these commands and seeing the sort of machine kind of respond and do exactly what he's asked it to do.
53:41Popping Tylenol. Popping Tylenol. Exactly. autism activated. Of course. You have to. You have to. Exactly. What kind of what kind of computer was it? It was an IBM as far as I remember. IBM PC. It was IBM PC. So what year was this about? 1993. 1993. Okay. So did it have Windows at that point? No. It didn't have Windows. It was right before Windows. Right before Windows. I think Windows had been out but you would It was an add-on. It was an add-on. You wouldn't boot it up. So I think we bought the disk for Windows, and you had to kind of bootload it from the disk, and it will open Windows, and you can click around.
54:24It wasn't that interesting because there wasn't a lot on it. So a lot of time I just spent it in DOS, and writing batch files and opening games and messing around with that. But it wasn't until Visual Basic that I started. So after Windows 95 that I started making real software. and the first idea I had was I used to be a huge gamer so I used to go to these LAN gaming cafes and play Counter-Strike and I would go there and the whole thing is full of computers but they don't use any software to run their business. It was just like people running around just like writing down your machine number how much time you spend on it and how much did you pay and kind of tapping your shoulders like hey you need to pay a little more for that and I asked them like why don't you just build a piece of software that allows me to log in and have a time or whatever.
55:12And I was like, yeah, we don't know how to do that. And I was like, okay, I think I know how to do that. So I spent, I was like 12 or something like that. I spent like two years building that and then went out and tried to sell it and was able to sell it and was making so much money. I remember McDonald's opened in Jordan around the time when I was 13, 14. I took my entire class at McDonald's. It was very expensive, but I was bawling. I had all this money and I was showing off. And so that was the first business that I created. And then when it came to... And at the time, I started learning about AI, you know, reading sci-fi and all of that stuff.
55:54And when it came time to go to college, I didn't want to go to computer science because I felt like coding is on its way to get automated. I remember using these wizards. Do you remember wizards? I remember wizards. Wizards, basically. It's like extremely crude, early bots that generate code. That generate code, yeah. Yeah, and I remember you could type in a few things, like, here's my project, here's what it does, whatever, and then click, click, click, and it just scaffold a lot of code. I was like, oh, I think that's the future. Coding is such a... It's almost solved. Yeah, it's solved. Why should I go into coding?
56:26I was like, okay, if AI can do the code, what should I do? Well, someone needs to build and maintain the computers. And so I went to the computer engineering and did that for a while. But then rediscovered my love for programming, reading program essays on Lisp and things like that, and started messing around with Scheme and programming languages like that. But then I found it incredibly difficult to just like learn different programming languages. I didn't have a laptop at the time. And so every time I go to like wanting to learn Python or Java, I would go to the computer lab, download gigabytes of software, try to set it up, type a little bit of code, try to run it, you know, run into missing DLL issue.
57:10And I was like, man, this is so primitive. Like at the time, it was 2008, something like that. You know, we had Google Docs, we had Gmail. You could like open the browser and partly thanks to you and be able to kind of use software on the Internet. And I thought the web is the ultimate software platform. everything should go on the web. Okay, who's building an online development environment? And no one. And it felt like I found a$100 bill on the floor of Grand Central Station. Surely someone should be building this. But no, no one was building this. And so I was like, okay, I'll try to build it.
57:49And I got something done in a couple hours, which was a text box. You type in some JavaScript, and there's a button that says eval. you click eval and evaluates it shows you in an alert box. So one plus one, two. I was like, oh, I have a programming environment. I showed it to my friends, people started using it, I added a few additional things like saving the program. I was like, okay, all right, there's a real idea here. People love it. And then again, it took me two or three years to actually be able to build anything because the browser can only run JavaScript. And it took a breakthrough at the time Mozilla had a resource project called mScripten that allowed you to compile different programming languages like C, C++ into JavaScript.
58:36And for the browser to be able to run something like Python, I needed to compile C Python into JavaScript. So I was the first to do it in the world. So built, contributed to that project and built a lot of the scaffolding around it. And we, my friends and I compiled and I was like, okay, we did it for Python, let's do it for Ruby, let's do it for Loa. And that's how the emergence of the idea for Replit came, is that when you need a REPL, you should get it, you should REPL it. And so a REPL is the most primitive programming environment possible. So I added all these programming languages, and again, all this time, my friends were using it and excited about it.
59:15And I was on GitHub at the time, and just my standard thing is when I make a piece of software, I just open source it. And so I was open sourcing all the things. I was years building just like this underlying infrastructure to be able to just run code in the browser. And then it went viral. It went viral in Hacker News. And it coincided with the MOOC era. So massively online courses. Udacity was coming online. Coursera. And most famously, Code Academy. So Code Academy was the first kind of website that allowed you to code in the browser interactively and learn how to code. And they built a lot of it on my software that I was open sourcing all the way from Jordan.
59:51And so I remember seeing them on Hacker News and they were going super viral. I was like, hey, I recognize this. What are you using? And so I left a Hacker News comment. I was like, oh, you're using my open source package. And so they reached out to me. They're like, hey, we'd like to hire you. I was like, I'm not interested. I want to start a startup. I want to start this thing called Replit. And they're like, well, no, you should come work with us. We can do the same stuff. And I kept saying no. I was like, okay, I'll contract with you. They were paying me$12 an hour. I was really excited about it.
1:00:20Back from Amman. Um, but they came out to their, to their credit. They came out of Jordan to recruit me and spent a few, a few days there. And then, uh, you know, I kept saying no. And in the end, they gave me an offer. I can't refuse. Um, and they got me an O-Win visa, came to the United States. That's when you moved. So when, when, when was the first, cause you were born, what year? 1987. 87. What was the first year that you could remember where you had the idea that you might not live your life in Jordan and you might, you might actually move to the U.S.? Uh, when I watched Pirates of Silicon Valley.
1:00:49Is that right? Okay. Got it. All right. maybe 98 or 99. I don't know when it came out. That might be a good place to go. Yeah. Is it worth telling the hacker story? Because there's a version of the world where you didn't actually, like if that changed differently, maybe you wouldn't have gone to America. Right, right. Yeah, so in school, I was programming the whole time. So I just want to start businesses. I just like, I'm exploding with ideas all the time. And like the reason Replit exists is because I have ideas all the time. I just want to go type them on the computer and like build them. So I wasn't going to school.
1:01:19It was like incredibly boring for me. And part of the reason why Replet has a mobile app today is because I always wanted to program under the desk. It's like just two things. And so at school, they kept failing me for attendance. So I would get A's, but I just didn't show up. And so they would fail me. And so I felt it was incredibly unfair. And all my friends were graduating. Now this year, this was like 2011. I've been like for six years in college. It should be like a three or four year. And I was like incredibly depressed. I really wanted to be in Silicon Valley. And so I was like, oh, what if I change my grades?
1:01:57There we go. In the university database. And so I went into my parents' basement and implemented the polyphasic sleep. Are you familiar with that? I am. Leonardo da Vinci's polyphasic sleep. I didn't hear from Leonardo da Vinci. I heard it from Seinfeld. Because there's an episode where John Kramer goes on polyphasic sleep. 20 minutes every four hours. Yes, 20 minutes every four hours. And yes, and this somehow is going to work well. Yeah, and hacking, if you've ever done anything. As the meme goes, this has never worked for anybody else, but it might work for me. Yes. And a lot of what hacking is, is that you're coming up with ideas for like finding certain security holes and like writing a script and then running that script and that script will take like 20, 30 minutes to run and so you'll take that you know 20 30 minutes to sleep and go on so i spent two weeks just going mad like trying to hack into the university database and uh finally i found um a way i found a sequel injection somewhere on the site uh and i found a way to like be able to edit the the records but i didn't want to risk it so i went to my neighbor who's going to the same school uh i think till this day no one caught him but i went to him and i said um hey uh i have this way to change grades.
1:03:17Like, would you want to be my guinea pig? And I was honest about it. I was like, I'm not going to do it. Are you open to do it? He's like, yeah, yeah, yeah. They call it human trials. This is how medicine works. So we went and changed his grades and he went and pulled his transcript and the update wasn't there and went back to the basement. Well, it turned out that I had access to the slave database. I didn't have access the ambassador database. So find a way through the network, privileged escalation. It was an Oracle database that had a vulnerability and then found the real database. And then I just, you know, did it for myself, changed the grades and went and pulled my transcripts.
1:04:02And sure enough, it actually changed, went and bought the gown, went to the graduation parties, did all that and we're graduating and then one day I'm at home it's like maybe 6 or 7 p.m. I get a you know the telephone at home rings I'm gonna I'm gonna ring sound yes well hello and he's like hey this is the university registration system and I knew the guy that run it he's like look you know we're having this problem the system's been down all day and it keeps coming back to your record there's an anomaly in your record where you have a passing grade, but you're also banned from that final exam of subject.
1:04:47I was like, oh, shit. Well, it turns out the database is not normalized. So typically, when they ban you from an exam, the grade resets to 35 out of 100. But apparently, there's a Boolean flag. And by the way, all the column names in the database are single letters. That was the hardest thing. It's security by obscurity. And it turns out there's a flag that I didn't check. So when you go over attendance, when you don't attend and they want to fail you, they ban you from the final exam. So I changed the grades and that created an issue and brought down the system. So they were calling me. And I thought at the time, I was like, you know, I could potentially lie and it'll be a huge issue.
1:05:30Or I just like, I'll just fess up. I'll just fess up. So I said, hey, listen, look, yeah, I might know something about it. Hey, let me come tomorrow and kind of talk to you about what happened. So I go in and I open the door. It's the deans of all the schools. It's like computer science, computer science. They were all working on it for like days because it's like, it's a very computer heavy, you know, university. And it was like a problem. And they're all kind of really intrigued about what happened. And so I pull up a whiteboard and started explaining what I did. and everyone was engaged. I gave them a lecture, basically.
1:06:06Your oral exam for your PhD. So you were like, this is great. They were really excited. And I think it was endearing to them. I was like, oh, wow, this is a very interesting problem. And then I was like, okay, great, thank you. And I was like, hey, wait, wait, we don't know what to do with you. Do we send you to jail? And I was like, hey, we have to escalate to the university president. and he was a great man and I think he gave me a second chance in life and I went to him and I, you know, I explained the situation. I said, like, I'm really frustrated. I need to graduate. I need to get on with my life.
1:06:44I've been here for six years and I just can't sit in school. The stuff I already know, I'm a really good programmer. And he gave me a Spider-Man line at the time. It's like, with great power comes great responsibility and you have a great power. And it really affected me and I think he was right at the moment. And so he said, well, we're going to let you go, but you're going to have to help the system administrators secure the system for the summer. I was like, happy to do it. And I show up and all the programmers there hate me. They hate my guts. And they would lock me out. Like I would see them, they would be outside.
1:07:19I would knock on the door and no one would listen. It was like, they don't want to let me in. I tried to help them a little bit. They weren't collaborative. And so I was like, all right, whatever. And so it came time for me to actually graduate. It was the final project. And one of the computer scientists came to me and he said, look, I need to call a favor. I was a big part of the reason we kind of let you go and we didn't kind of prosecute you. So I want you to work with me on the final project. And it's going to be around security and hacking. I was like, no, I'm done with that shit. I just want to build programming environments and things like that.
1:07:55And he's like, no, you have to do it. I was like, okay. So I thought I'd do something more productive. So I wrote a security scanner that I was very proud of, that kind of crawls a different side, that tries to do SQL injection and all sorts of things. And actually, my security scanner found another vulnerability in the system. And so I went to the defense, and he's like, you need to run this security scanner live and show that there's a vulnerability. And I didn't understand what was going on at the time, but I just, okay. So I give the presentation about how the system works, and I was like, oh, let's run it.
1:08:25And it showed that there's security vulnerability. Okay, let's try to get a shell. So the system automatically runs all the security stuff and gets you a shell. And then the other dean that turned out, he was giving the mandate to secure the system. And now I started to realize I'm a pawn in some kind of rivalry here. And his face turned red and it's like, no, it's impossible. We secure the system. You're lying. I was like, you're accusing me of lying. All right, what should we know? Should we know your salary or your password? What do you want me to look up? And I was like, yeah, I look up my password.
1:09:04So I look up his password. And it was like gibberish. It was encrypted. And I was like, oh, that's not my password. See, you're lying. I was like, well, there's a decrypt function that the programmers put in there. So I do decrypt and it shows his password. And it was something embarrassing. I was like, I forgot what it was. And so he gets up, really angry, shakes my hand and leaves to change his password. I was able to hack into the university another time. Luckily, I was able to graduate, give them software, they secured the system. But yeah, later on, I would realize that, yeah, he wanted to embarrass the other guy, which is why I was in the middle.
1:09:40Economic politics. Well, I think the moral of the story is if you can successfully hack into your school system and change your grade, you deserve the grade and you deserve to graduate. I think so. And just for any parents out there, or children out there, you can cite me as the most. You can say that I'm proud of me as the moral authority in this. One maybe lesson I think that is very relevant for the AI age, I think that the traditional sort of more conformist path is paying less and less dividends. And I think kids coming up today should use all the tools available to be able to discover and chart their own paths.
1:10:19Because I feel like just listening to the traditional advice and doing the same things that people have always done is just not as, it's not working out as much as we'd like. Yeah, that's right. Thanks for coming to the podcast. Thank you, man. Fantastic. Thanks for listening to this episode of the A16Z Podcast. If you liked this episode, be sure to like, comment, subscribe, leave us a rating or review, and share it with your friends and family. For more episodes, go to YouTube, Apple Podcasts, and Spotify. Follow us on X at A16Z and subscribe to our Substack at a16z.substack.com. Thanks again for listening, and I'll see you in the next episode.
1:11:01As a reminder, the content here is for informational purposes only. It should not be taken as legal business, tax, or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any A16Z fund. Please note that A16Z and its affiliates may also maintain investments in the companies discussed in this podcast. For more details, including a link to our investments, please see a16z.com forward slash disclosures.
From the publisher
Amjad Masad, founder and CEO of Replit, joins a16z’s Marc Andreessen and Erik Torenberg to discuss the new world of AI agents, the future of programming, and how software itself is beginning to build software.
They trace the history of computing to the rise of AI agents that can now plan, reason, and code for hours without breaking, and explore how Replit is making it possible for anyone to create complex applications in natural language. Amjad explains how RL unlocked reasoning for modern models, why verification loops changed everything, whether LLMs are hitting diminishing returns — and if “good enough” AI might actually block progress toward true general intelligence.
Resources:
Follow Amjad on X: https://x.com/amasad
Follow Marc on X: https://x.com/pmarca
Follow Erik on X: https://x.com/eriktorenberg
Stay Updated:
If you enjoyed this episode, be sure to like, subscribe, and share with your friends!
Find a16z on X: https://x.com/a16z
Find a16z on LinkedIn: https://www.linkedin.com/company/a16z
Listen to the a16z Podcast on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX
Listen to the a16z Podcast on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711
Follow our host: https://x.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Stay Updated:
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Podcast on Spotify
Listen to the a16z Podcast on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
