In short
Podcast Episode Summary
Podcast Title Latent Space: The AI Engineer Podcast
Episode Title Why Every Agent Needs Open Source Cloud Sandboxes
Episode Description In this episode, Vasek Mlejnsky from E2B discusses the critical role of sandboxes for AI agents, highlighting the growth of E2B in the last two years, its extensive usage across Fortune 500 companies, and the implications of evolving AI workflows.
---
Key Topics Discussed
- Introduction to E2B and its Origins
- Transition from DevBook to E2B.
- Development of interactive documentation to improve developer experience.
- Initial sandbox technology from DevBook evolved into E2B's current offering.
- Growth of E2B
- E2B's rapid adoption by ~50% of Fortune 500 companies.
- Generation of millions of sandboxes weekly.
- Shift from individual projects to comprehensive solutions for AI agents.
- Challenges in Early Development
- Building an agent cloud and the initial hurdles.
- Limitations of early LLMs and models.
- Balancing growth against model capabilities.
- Use Cases for E2B Sandboxes
- Data analysis and visualization.
- Executing arbitrary code from model outputs.
- Running evaluations on code generation.
- Reinforcement learning applications.
- Technological Insights
- Discussion on the LLM Operating System (LLMOS) landscape.
- Differences in programming language usage (JavaScript vs. Python) on E2B.
- Comparison of AI VMs and traditional cloud infrastructure.
- Future Plans for E2B
- Development of higher-level agent frameworks and toolkits.
- Addressing limitations of chat-based interfaces and future agent capabilities.
- Vision for E2B as a full lifecycle infrastructure for LLMs.
- Technical Specifications
- Overview of sandbox technical specs, including usage-based billing.
- Discussion on pricing AI based on value delivered versus token usage.
- Features like forking, checkpoints, and parallel execution in sandboxes.
- Market Landscape and Competitors
- Perspectives on the broader LLMOS landscape.
- Insights on how E2B differentiates itself from competitors.
- Discussion on the importance of model agnosticism.
- AI Agent Development and Deployment
- The journey from code execution to app deployment.
- Roadmap for supporting LLM development.
- The need for seamless integration of AI-generated applications.
- Hiring and Future Directions
- Current hiring needs at E2B.
- The importance of being close to users and iterating quickly.
- Discussion on relocating to San Francisco and its strategic advantages.
---
Key Takeaways
- E2B is pivotal in providing sandbox environments necessary for AI agents, enabling a wide range of applications from data analysis to reinforcement learning.
- The podcast highlights the critical need for infrastructure that supports the rapid evolution of AI technologies and workflows.
- Future plans for E2B include expanding functionalities and moving towards becoming an all-encompassing platform for AI development and deployment.
- The conversation emphasizes the importance of real-time user feedback in shaping product development.
---
Conclusion This episode sheds light on the innovative landscape of AI tooling and how E2B is positioning itself as a leader in providing essential infrastructure for AI agents. The discussion reflects the evolving nature of AI development and the critical role that open source and cloud technologies play in this space.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:28Hey everyone, welcome to the Latent Space Podcast. Yeah, and my official legal first name is Vaclav, but everyone pronounces it as Vaclav, which I hate. So it's just Vasek. Okay, awesome. We're both invested in E2B in different ways. But you and I go back the furthest. I just realized three years ago when you were working on DevBook and you were interested in that developer experience angle and somehow you pivoted to E2B. Maybe you want to tell that story. Yeah, so Thomas, my co-founder, who's our CTO, So we've been interested in DevTools for quite a long time, like six, eight years. And before DevBook, there was like a bunch of iterations.
1:08After there were like different iterations and pivots of DevBook, we just like stopped renaming things at some point and just like went with DevBook. The one you are talking about was interactive documentation for developers. So basically the idea was we wanted to, instead like you as a developer, when you come to a tools docs website, instead of reading about everything and then Googling it and trying it in your coding editor, the idea was to give you interactive experience in the browser. So you would have like pre-made interactive guides, playgrounds, you could try things right away. And the company, the owner of the docs would prepare the experience for you because you would also be trying everything in the browser.
1:53they would now see if you get stuck anywhere, what you are doing, you know. So you get a very valuable onboarding analytics. We actually built an interactive playground for Prisma. I think it's still up. They are still using it. The last time I checked a couple months ago. So it was like interactive guides and playground that you could try out Prisma without having to manually set up all the databases and everything. So you could just try Prisma right away. and that was like the very very first version of actually our infrastructure that we are offering basically sandboxes it was sandboxes it literally was sandboxes the same technology but just completely unscalable uh so then that somehow in 2024 ish turned to e2b 2023 i think march 23 and uh we were pretty burnt out thomas and i we were working from prague from the czech republic lake from my apartment.
2:49Nothing was really moving, no growth. And GPT 3.5 came out. Like really first model, kind of good-ish with CodeGen. So we took a like, let's take 10 days break, two weeks break from DevBook. And because like everyone was trying things with AI, it was very clear like this is something where the future might go. So we wanted to just like from out of curiosity, try things out. We wanted to build like a dev in kind of like thing. The first idea we had was like, let's automate our work. Because with every project we were starting, there was a set of tools you always wanted to integrate in your backend.
3:33Like Stripe for SaaS business, like Stripe, Analytics, Slack notifications, emails, sending our emails. and so we gave the agent tools to run code and we needed some kind of sandbox. We were like, yeah, that's a good coincidence. We have sandbox from DevBook and we posted about it on Twitter. Basically, the agent actually pulled GitHub repository, wrote code, started the server, tested everything and at that point, I think we deployed it to Railway. Railway had like the best DX and it was easiest to just plug it into the agent. I tweeted about it and I think Greg Brockman like retweeted it like hey like I don't remember exactly what he said I need to find a tweet but it was around the time where OpenAI was retweeting or OpenAI co-founders were retweeting all the things that people are doing like to show like what you can do with GPT 3.5 and people like it had like half a million views after a few days and so we were with Tomas we were like we gotta do something like people are interested in it so we just open sourced it the repository and it was named the organization was like the AI company like we had no name I kind of wish like we could have like a stick like legally stick with that we just came up with a name we was like E2B because you take English you convert it to bits that's how it started we started building community around it and like two days later we were like let's Focus on the sandbox part, not on the agent part.
5:12We actually had a bunch of hypotheses behind it, like why it might be more interesting than building the agent. And then you had the small developer, run small developer on E2B. That was a little bit later. Yeah, when was the... This is in 2023 as well. When you first launched E2B, it was kind of like an AI agent's cloud. What did you call it? In 23, the hypothesis was like agents, code gen agents especially, will need some kind of environment to run the code. And the same way developer needs a laptop or something, you know. So, but we struggle a lot with how the product should actually look like and what's like a go-to market, the first version.
5:55And so the high-level idea was like, we will host your agent, everything from actually deploying it to then monitoring it and it will also have this environment running the code. One of the first test project was taking Sean's small agent project and deploying it inside our sandbox and just giving the agent tools to pull the GitHub repositories, work on that, do a PR and then post it on GitHub, the PR on GitHub, which was very, very popular. like there was clear there was something very interesting it's amazing how much people went with that it was literally not meant to do that it was meant to do Chrome extensions yeah and I think you really you had good insight that you let the agent plan the work in a markdown file basically and write down the spec which I think even now when you look at deep research agents that's sort of often times what they are doing they plan everything these days I would say they should do more structured output than markdown but you know that's that's an implementation detail.
7:01So that was taking off but it attracted a little bit different audience than we wanted so because it attracted a lot people who nowadays would be using tools like Lovable for example you know they just wanted to build projects for them and of course like it wasn't really working at the time like like it worked for like one simple website but the moment you wanted something more complex like it was was really hard. I've been also reflecting on why I stopped working on it. And there was actually a, so when I built it, it was with Cloud3, the new Cloud3 launch. It was trying to utilize the 100K context of Cloud3.
7:41And I used the same project to do this, to try to repeat the demo that I made for myself a month afterwards. And it wasn't anywhere as smart. So Cloud3 got dumber. But it looks like I made up the demo or something, but no, literally, I just reran the same code and it just was not as capable. And I mean, I think to some extent, this is like RNGsus hurting me or it's the type that it's like a different month. So the model is different or something. But I think also basically people had this vision of what they wanted and then they tried to do it in reality and they couldn't do it because the models are not ready.
8:16So a lot of feedback we got. At some point, people thought like small agents is from us because we had a website, like deploy small agent, but we were like, no, that's like Sean's work. He should go to his repository, submit an issue. No, they could clone it. It was like, you know, a few hundred lines of code. I did it however you want. For us, it was more like a test if the environment, the sandbox is useful for the agent. If it can sort of scale and it can work, which worked well, but then it took us another, I would say, six months to actually find the right go-to-market strategy. which was code interpreting for us.
8:58So typically AI data analysis, data visualization, inside like a headless Jupyter type of notebook and environment. Specifically Jupyter? It doesn't need to be Jupyter. The important part is that you don't need to explain the model and the model doesn't need to care about how to keep the state of the program running. So it was, especially with our earlier models, I think the models are now smarter, but they kept producing like code snippets and thought they can reference to previous variables and functions definitions, which if you had only normal code execution, like you run the code snippet and then you finish, it wouldn't work.
9:40So you need some kind of REPL environment. And what was, I think, especially early on, like Python was the language that the models were probably the best or one of the first languages is where the models were working very well. Looking back at it, there was a really strong pool from people just wanting to visualize data and talk with their data. Python was really good at it and it mixes well with Jupyter type of environment because you get charts out of the box. You can also support interactive charts. Jupyter itself isn't the right environment to actually do it because all the technical problems that you will run into once you start doing it on scale.
10:25So it gets slower and slower. And actually, what we are coming into is we are building our own thing internally just to support LMS specifically. It's got your own runtime. At what point did you start going from just code interpreter, run code, to expand? Because now you have people doing RFT, you have computer use and all of those things. When were the models ready for people to start using it? You demoed at one of our first events as well. and I think there's always this lag between the infrastructure that you build and the capabilities of the model. When did you go from just code interpreter to start saying, okay, now it's time to do computer use, now it's time to do RFT?
11:04That was probably end of 24, start of 25. So when you look even at our data, 24 we are growing, we are growing good, but 25 is up to the right. And so it feels like at 24 people are figuring out these agents and building them and trying them. And 25 is everyone moving them into production and finding more and more use cases. Around end of 24, start of 25, we started seeing things like using Sandbox for reinforcement learning type of use cases or using Sandbox for computer use, which was very interesting when Anthropic launched their computer use. We had like a desktop version of a Sandbox that was sitting in our GitHub repository for six months.
11:50We were like, okay, this is probably interesting, but no model can actually use it. So when Anthropic announced it, it was like, we have something here we can show you. Using with like lavables and Blitz's type of products also like started using the sandbox for more than just like run code snippet, like data analysis. And then deep research agents, that That has been something really big in the last few months. How does deep research agents... Are you referring to Manus? Yeah, Manus, for example. Deep research I typically think of as a web search heavy task. It doesn't really use a code interpreter in any way.
12:33I have some idea that Manus uses E2B in an interesting code interpreter way to do deep research. What's the difference? Yeah, I think it's a good idea to stop thinking about the sandbox just for code interpreting and more about like a runtime, code runtime for the LLM or the agent. The use case for the sandbox, it's a very horizontal in a sense that it can cover everything from the agent needs to create a file, make a to-do list. It uses a browser, not inside the sandbox. It's like separate. You have browser use, being used for research on the website, but then you download the data somewhere.
13:12You need to transform the data you want to do data analysis, you want to actually write a small app, you want to create an Excel sheet. So it's the same way you are using your laptop as a human that is the same useful for the agent. So you can think about it as more like a dev box and at the same time the agent that's using it is also like a very, very good developer, good accountant, good like slides creator researcher and so you are just basically giving it tools to let it do the job even better and faster. Yeah and when you say up and to the right I just want to share some numbers from the investor updates.
13:59So March 24 which we're talking about. I don't know. You shared this. You shared this publicly. But yeah March 24 you were doing 40 ,000 sandboxes. You want to say how many you've done last month? March 2025? 15 million. Yeah. So in one year, you've gone from 40 ,000 to 15 million. And yeah, I think you can kind of see the slope, especially from like the Sonnet 3.7 release. And I think this is like an interesting model versus infrastructure. I think there's this usual like VCEO, we should invest in like the tools, you know, the picks and shovels instead of the application layer. But I think this is the first time where the infrastructure is lagging the applications.
14:45That's a good point. I think 24 was all about the agent couldn't use the whole sandbox. And now, yeah, sometimes we are actually catching up with some features for the LMs that they need more than what we have at the moment. Yeah, I guess what we're doing here, you are another one of the LLMOS companies that we are talking to. We also did one with BrowserBase and I'm not sure who else would qualify under that term. Yeah. So like basically like you don't specifically yourself use LLMs internally in your products, but you enable others to work to augment their LLMs with your infrastructure. Yeah.
15:30I don't know any reflections on just the general LLMOS landscape. You have other competitors, like, you know, how is this evolving? How do you position in it? A lot of people are saying, if you are a GPT wrapper in 23, you were in a really bad position because all the value will be captured by the AI labs. It's good to be a GPT wrapper because you get all the advantages from a new model. You just switch it. I mean, the just is not so simple. You probably have evals. You need to change the prompt a little bit. But I would say it's increasingly easier to switch models. So we need to think about it the same way.
16:07like our users are switching models a lot. We need to be agnostic to the LMs. Oftentimes, like people want to deploy us in their cloud or on-premise. So that's also something very important. I think a good analogy here is sort of, like technologically, it's kind of, you want to be the Kubernetes of the world for the agent, but with much better DX and easier to use. One thing I'm thinking about also is like, what is valuable real estate to occupy in the LMOS. And I have this spectrum, I think on the browser-based episode I was talking about, either you can focus on browser emulation or you can focus on the VM or you can do a custom Python sandbox like Modal does.
16:54Would you say that you are the most general of all of them? Yes, it's very general, but that's not really historically how you want to market it because people don't know what to do with it. Okay. So we started as, when we had our website in 24, it said something like computer, like cloud computer for AI. People didn't just understand what to do with it. So we had to literally show them, first you change it to code interpreting because that's something people knew from OpenAI. And then you just show them very, very specific use case. And you use that to get an early traction and an early set of users.
17:32And we spent a lot of time, I mean, Teresa from our team spent a lot of time, especially her, on educating the market, you know, and developers. So, and I think this is the space, the AI space is you sometimes have to like show developers what they might need. And you kind of have to like trust your God, like this might go in this direction, maybe like this could make sense and show them like what they can build with it. Because it's very hard to imagine what you can build with things that you don't even have, right? So why would you need forking sandboxes or checkpointing sandboxes? How is that useful?
18:07Well, it turns out if you're building some Monte Carlo type of a thing, search for an agent, it's very useful. But to actually go for a developer who might not be in the AI deep research type of thing, it's not obvious. So you want to be agnostic, you want to be channel, but you want to also show people very clear use cases how they can use you. and over the time they then get educated and realize, okay, there's more use cases, understand the platform, they start coming up with their own ideas, but onboard people with general use cases is very hard. At least that's what we learned the hard way.
18:47Yeah, and you also don't really tailor to the more DevOps infrastructure person, which I think a lot of the other sandboxes are like, oh, we have a GVisor runtime and we have all these different terms that you don't really know if you're the AI engineer. type. Yeah, that's exactly true that our user isn't like an infra engineer. Even ML engineer usually isn't our type of a user. It's like AI engineer, you know, like from the definition you have, I still remain very bullish web developers and JavaScript world, TypeScript world. Even though we have a ton of usage from Python, but there's so many web developers and it's easier and easier to use LLMs.
19:28I really think that you need to cater to these type of developers and make things simple for them. And if they want to dive in, they can. Chances are they don't even want to. They want to focus on building. Yeah, product developers, not infrastructure developers. It's very interesting. Yeah, I mean, this whole GPT wrapper versus model lab thing, it's not like it's better to work on wrapper over model labs because the model labs people are making a lot of money. It's just that there is room for wrappers and the wrappers also do make money. and I think people were not seeing that in 2023. Yeah. And this is what we saw.
20:02It's not binary. It's not either or, yeah. Yeah, I think, and then just a quick check. Do you know like the rough percentage between JavaScript and Python? Slightly less JavaScript. More Python still. So like from number of downloads of our SDK per month, it's like 250 ,000 JavaScript, close to around half a million Python. Yeah. Something around that. Two to one. Yeah. Yeah, interesting. I mean, if the use case is really for code interpreting, generating charts and all these, then Python wins. Yeah, yeah, yeah. Exactly. Python wins for that. But once you go into more like, for example, building apps, generated apps, then it's JavaScript winning, right?
20:45Because you probably have, you have frameworks, all the swells, next.js, view.js frameworks, which is like JavaScript thing. It's interesting because I would think that it doesn't really matter what code you want the LLM to produce. It wouldn't dictate what kind of user is using us. But if you think about it, when developers are building AI data analysis, it's typically Python developer. When developers building a type of V0 like a use case, it's a web developer. It's a product developer. So I don't think this is super obvious. I wouldn't think that would be the case. There's all sorts. I mean, Bolt's argument is that you should want to use something like a web container where it's like run on your own browser and it's all free and very fast.
21:35I guess the one more critical question, like let's say I do know my infra. What is the point of a cloud for AI, a computer for AI? So basically, why can't I use existing tools like Railway? I'm not sure if Railway wants to go after your customers, right? So why does there need to be an AI-focused virtual machine, a sandbox, or execution environment, whatever you call it? What we offer is sort of orthogonal to cloud. So beforehand, you don't know what kind of code you will be running in your cloud, so you can't really optimize it. So there's no build-deploy step, per se. Everything happens ad hoc during runtime.
22:14So you need to solve problems like you want to install dependencies very fast, you want to be able to pull GitHub repositories very fast. So then how do the workloads we are running can go from five seconds to five hours. So that also changes even the pricing model a lot. How do you make sure that everything makes sense for you and for the user from the pricing point of view, from the unit economics, and also from just the infrastructure point of view where the sandboxes are getting placed inside your cluster. The security model is also different usually because it also comes down to you don't know beforehand what code you will run.
23:01So by default, it's untrusted code. And you need to have complete isolation between these sandboxes to make sure first that if something happens in one sandbox, it doesn't affect other sandboxes. But also you want to know about security inside the sandbox. So as the LMs are getting better, you want to know what's happening inside the sandbox. There's like a fun story from Hugging Face when they were using us and are using us for their Open R1 model. And one of the developers, he shared it online. So I think it's okay to say that. So he lost access to their cluster because the LM decided to change permissions.
23:42If that happens with us, we just like kill the sandbox and get a new one. and it takes like, I don't know, 150 milliseconds. For them, they had to take down the whole cluster, set up everything, didn't even know that this happened. So it's like the model, the compute model, TLDR is the compute model, security model is different from the current cloud providers. And you need to think from the first days about it differently. And then people use, depending on the, because the difference that maybe people don't think about is like you can generate the code with the Python SDK, but you can still run Lua code, or like R code.
24:19It's not always matched to the runtime of the SDK. It's a general machine, so whatever you can run on Linux, you can run inside a sandbox. We had users running, of course, Python code, but C++. We had Fortran, someone running Fortran, which is like... Why? They were using very old API for banking or something like that. But you can also start a server inside a sandbox and then you want it to be accessible from the internet. It's a very general machine and the challenge is how do you make everything fast, for example? How do you make everything secure? At the same time, you need to make it accessible enough and controllable by the LLM and observable for the human.
25:05So you're kind of building for two personas, for the human developer, the AI engineer, and for the LLM who's using the sandbox. Yeah, the composability thing is something I had not thought about before. But later, like you mentioned, it's like, you know, you might need to run Fortran to access one API. And then in the next step, you'll need to take that data and run it in a Python script. And then you're going to expose that through JavaScript to something else. And like, you guys can switch the runtime. Yeah. Halfway. Yeah. And in idle world, you don't want to kill the sandbox, start a new one, or maybe like create a new sandbox template.
25:41you just want to keep using the computer and because we are in cloud we can do it in a way that you can get more ram you can get more cpu you can get less cpu so you are really paying only for what you need and as the llm is doing more and more you can have very elastic sandbox and keep adding features and the goal where we think this is getting going is the llm like decides what it want to do and how it wants to have the sandbox configured. So it basically starts controlling the infrastructure itself and creating sandboxes themselves. While we're talking about the technical details, I just wanted to let you tell people about any other technical details.
26:23Like, what is the box where we get? What kind of Linux? What size of whatever? What details matter here? So you get Ubuntu box. You can customize it and anything Debian based is going to work. you can even add graphical interface so if you want to you can run legit, not headless but legit Ubuntu computer on it. So you can do the sort of take out like an operator type experience? Yeah, we have an SDK called Desktop SDK that does this for you out of the box and it supports VNC so you can also have a human in the loop type of thing and control everything and stream and see what's happening. So by default, you get two CPUs and half a gig of RAM on the free tier.
27:10And you can customize it on our landing page. We say up to eight gigs. But if you tell us what's our use case, you can go to 64 gigs of RAM. We have users using such a BFET sandboxes. You can go to 16 CPUs, I think, if I'm not mistaken. Storage is free. Storage is free. I think there's a lot of things that are free that we probably need to think about a little bit more. You know, with the dev tools, especially infra, I always see this pattern, and I had this like naive idea as well, that the founder says like, oh, this AWS pricing, it's terrible. Like you are GCP, like you have no idea what you are paying for.
27:49There's so many like small add-ins you need to pay for. We will start a new infrastructure company where you're just going to pay$200 per month, and then you just pay for like pure compute. That works until you start scaling and you figure out, okay, actually people are doing weird stuff that I didn't expect, like having lots of traffic, producing. We have a customer that produced petabyte of data. I mean, it's not free to host petabyte of data and that's growing. So then you start introducing, okay, probably they should pay some amount for ingress, egress, for storage. And the pricing gets increasingly complex.
28:30So I just think it's a very interesting phenomenon that you start with this brave idea that everything will be super simple. And we only do code interpreting, so we only need to price for compute, right? Yeah, exactly. And then you wake up in the real world, it's messy. Yeah, so I know about this from my career in cloud. And so the common refrain is that I call this the first principle of technology. everything can be broken down into some combo of compute storage and networking. If you fail to price one of them, you will get abused because... Because you're essentially offering free storage or compute or something, right?
29:09Yeah. For those interested in this idea, there's a fourth one, which is basically the control plane or like the off layer, the IM policies and all that. And that's like, maybe you can call that security as well. It's like the fourth layer that people kind of pay for as its own independent thing. for those also interested HashiCorp has more breakdowns from this like David McJanet which I think is very interesting if you're just in the business of running a cloud infrastructure company you should know these things many hundreds of businesses have run into the exact same problems you should just not repeat them and just learn whatever the best practice is yeah I think like the billing model has been figured out many times so like I don't think it's like the challenge is not figuring it out the challenge is I think is introducing it like sometimes quick enough.
Read the full transcript
29:54You actually need to do changes on your infrastructure to make sure you know about all this data that's happening and moving one way or another. But yeah, I completely agree with you. This problem, we are not the first one that are having it. Do you use one of the usage-based billing providers? Orb. We are talking with Orb right now. Orb in? Meter? Open Meter is another one, I think, or Meter. Yeah, metronome. Sorry, metronome. East meter is the other one. We have been using Stripe usage, which has been a little bit sometimes rougher around the edges. Yeah, this is the thing. We shouldn't spend so much engineering on it.
30:36I want to outsource it because that's not our product. Someone else should be focusing on this full time. And it's actually pretty non-trivial to make sure you have everything right and you really don't want to make mistakes here. Is there anything that you're really looking for that would say like, okay, that's really what we want, that maybe Orb or Metronome haven't really adjusted for AI yet? I don't think this is AI-specific problem. This is infrastructure as you know it. For us, some things that didn't work when we look at some of the providers where like the cut they took from the revenue, for example, they take from you.
31:12So some of the pricings, I don't know what's the latest, but like, I don't remember which one was it, honestly, but I knew that pricing was basically they take a small cut from the revenue each month, which is basically like what Stripe does when you are processing payments. It's just like the value was really high. And then it's a lot about how hard it would be to integrate it. Like how much time we are spending on it. Because we know we don't want to build this in-house. It's more like, is the switch worth it, basically? I'm curious your updated takes. I know you had the Y is in usage-based building a bigger category.
31:47I'm curious now with AI, with token-based pricing, if you have updated thoughts on... So for people who don't have context, at Netlify, we went from relatively flat tier-based pricing to usage-based billing because that's basically how all infrastructure companies should eventually go because you have some whales who use a lot of infrastructure and some who don't use that much and you shouldn't charge the same for both of them. The fun insight was that you would think that at an IS, past company that revenue is like the most important problem to work on and directly impacts the company's revenue and valuation and all that.
32:28No engineers wanted to do it. And I was like, why? Like we actually looked around for a long time. We tried to hire and then we couldn't hire. So we ended up putting one of our most senior engineers on it. And she took a year to ship the whole building projects that was presumably a board level objective, which was like, hey, let's change from this pricing plan to this pricing plan. How hard can that be? Turns out very hard because you have to instrument everything. Even the things that you're like, I don't know if we'll ever use this. Yeah. That's exactly what I meant. That's why storage is free.
33:02The first thing is like, you need to know about everything that's happening inside a cluster. How big are your logs? Yeah. It's crazy hard. The pricing is like actually the, not the figuring out the business model, but the integration of it and implementation is actually a lot of engineering. Yeah, soft limits, hard limits. Do you cut off people once they bust the limits? Probably not because they get pissed at you. And then they also get pissed at you if you don't cut them off because then you send them a big bill. So there's no winning. Yeah, exactly. Also, additional problem that people are asking us, I want to run the agent for five hours and you can do that.
33:38And you might be a very early startup. So our goal isn't to cash you out. our goal is like you can use as much of e2b as you can and just grow but the more longer you run the sandbox is the larger is going to be your bill so i think that's like it doesn't really needs to correlate with how much product market fit you have for example how much users you have because with these agents they can work for a long time even if you are like pretty early stage company yeah if you have few users so sure i mean yeah i think but there's still a question of soft limit, hard limit, that kind of stuff. So just to answer your question on billing, and for those who are interested, we actually had the CTO of Orb send in a talk for our remote track for the New York Summit on what he thinks pricing for agents looks like.
34:25And I think basically you are reselling tokens, right? The base layer is coming from either your open model cloud provider or your closed model lab API, and you resell them. And a lot of them, some people have a lot of markup on them, like certain unnamed AI builders. And some people have negative markup on them. Like they are basically selling you at a discount. You should buy as much as you want because they are using VC monies to subsidize your thing. And I think that seems fine. People often, Simon Willison often asks for like, bring your own key solution where like, you know, I will have my control over like my relationships and my pricing and my credits with my LLM providers.
35:03But I want to use your app. So I'll give you my keys and then you use the keys on behalf of me. that doesn't seem to be as popular as I think as I understand it. I don't know if you've had that request. It's sort of like in crypto use your own hardware wallet always going to be less popular than just going with the thing that's Yeah, so literally Alex from OpenRouter is the only person who's implemented this and it's fine for individual use cases because there's a lot of free tiers for individuals but yeah, I mean I think pricing wise people are trying to move that discussion on like, you know, are you positive margin on your tokens or are you negative margin on your tokens?
35:43And there's some economic reality there, but they're trying to move that to the agent work, which is what is the value of the human labor you're replacing, which is a whole different thing, right? So like, instead of comparing on cost of goods sold, you're comparing on value delivered, and that is much higher. Yeah, I feel like if we could get to a good market on like bid ask of like work being done, and then you can kind of arbitrage how many tokens you need. But I think today the agents are so unreliable and unpredictable that it's hard to price ahead of time because you could price any software engineering task per task.
36:18It's like do a new bun, it's like$500 and then I can arbitrage that. But today there's no certainty of that. And I'm curious. Yeah, you would think like, also this is why I was very excited about Replit when they first launched their marketplace credit thing. I'm not, if Replit can't do it and I don't know if anyone can. Right. The other technical thing we talked about before is forking. You also mentioned it before, and sandbox checkpoints and things like that. Is it something you have today? And then what are people using that for? Like our example was the Cloud Plays Pokemon hackathon. Are there more enterprise use cases where you see people request forking and checkpoints?
36:56Yeah, we don't have this publicly now yet, but that's something we are working on to release somewhat soon. we have persistence which is like the prerequisite to call it. You mount a volume. Yeah, well mount a volume but also memory persistence which is very interesting. So you basically can pause the whole sandbox even when all the code is, you know, with all the context of the code execution and resume to it resume it later and come back to it I don't know, like two weeks later and it's still going to be there. But the continuous session time is limited to 24 hours, right? No, it's limited in the beta in 30 days.
37:34Okay. Do you mean, are you asking? I don't know, I was looking at your pricing pages. Are you asking about like when the sandbox is running? The sandbox can run up to 24 hours, but when it's paused, it can be paused for like a month. I personally think this is like one of the cases where you kind of need to show the developers why it's useful. And this is something I think is going to be very useful as the agents are getting better and LMs are getting better because you will be able to paralyze problem solving essentially. So you will, instead of having a single agent doing one thing, you might have multiple agents trying different paths.
38:09And if you imagine like a tree or a graph, every node is like a snapshot sandbox, like a checkpoint sandbox. And then from that node, you fork the sandbox and go to the next state. Eventually you find the right path, right? It's kind of like a, it's a tree search. So the forking and checkpointing solves the local state problem. It doesn't solve the remote state because that's something you can't really control, but it solves a local state for the agent and can then come back to a state and you don't need to replay the whole session or trying to force the LLM to do the same thing again. You have the chat history and you have even all the steps inside the second box that went to it.
38:55Do you feel like you want to help with the forking and then the re-merging? Because I think people understand the forking, but then it's like, okay, how do I monitor which of the leaves is successful? And then how do I merge that back into thing? Do you think that's something that you want to help people do, like kind of spread out, parallelize, and then find the winner? Or is that something people should do on their own? I think this is sort of like a framework discussion on top of E2B. So this is, we are looking at it. I think eventually we should go higher level and I don't know if framework is the right type of a thing.
39:34We like to in the team think about it as a toolkit. Instead of building opinionated framework how to build agents we give you a wrapper around the E2B that makes it really easy for the LMs to for example merge these states or navigate these three. So I think it will eventually move there. I think it's a good question to ask how it's going to look like. I still think like building a framework is very hard in the AI as things are moving very fast. Curious about frameworks. Do you see any rise in popular frameworks that we should be keeping tabs on? I don't know if this is an unpopular opinion, but people keep telling me LinkedIn isn't popular, but if you look at its stats, it has 20 million downloads per month.
40:21How can you have not popular framework when it has 20 million downloads? And it's growing, I think. I think there's a slight bubble. I don't know if it's like people in SF, but developers thinking blockchain isn't popular or used. There's one framework that's interesting. I think it's called Mastra from one of the founders of... Sam Bhagwat, yeah. He's in the Discord. Yeah, yeah, yeah. So that I like, which looks very promising. I also like it's like TypeScript first, which I thought is the right decision because as we talked about, like being more bullish on the product developers instead of like a Python machine learning type of a developer.
41:04I've seen a lot of more like a toolkits or I don't know how to call them. Like for example, Composio has been one that like gives you all the tools at once. There's been more tools like that, which is interesting approach. I think also browser-based stagehand is super exciting because in my head, it's not a framework it's like it doesn't necessarily like dictate how the agent should behave it just gives it the good tools to navigate a website and it's very elegant like it has I think they added like three methods by act see and something yeah there's three APIs yeah observe yeah observe we talked about it in that episode no it's cool actually I wasn't expecting that many names to come out but these are good names for people to know I would agree with most of them as they're in the conversation of tooling.
41:54A lot of people listen to us for like, oh, what's going on in SF, right? Yeah. Actually, it's a good question because when you ask it, I would expect to be a bigger framework boom nowadays because it would make more sense to build a framework now instead of 23 because things are a little bit more stable. The problem with frameworks is that the idea of frameworks should probably be that things aren't changing underneath your hands, which if you launch your framework in 23, it was. So that's not a fun thing to be at as a developer who's using the framework. I think things are still changing. I mean, things are 100 % still changing, but slightly, I would say like some things are clearer than N23.
42:35I don't think people realize, but like, I'll just say it out here, like chat confessions is dying. So like any framework that was built in that era with like no conception of real time, no conception of omnimodal or multimodal native things, they will probably not age very well. Yeah. Yeah, probably as we, like, what's the next after chat-based interaction with LM? You need chat completions and you also need reasoning with streaming. And the streaming interactions of agents is also not mapped out very well. But, like, all agent frameworks will have to adjust to that, basically. Good question to ask when thinking all about dev tools with LM.
43:13Is my dev tool more relevant as the LMs are getting smarter? And as people need less prompting? I'm, for example, like, really bad at prompting. but like I can get more work done over the years because the elements are getting better so that means like there's less need for prompt management type of thing yeah the talk from RAMP at our New York conference was basically about this like how do you set up your architecture so that you benefit from 10 ,000x improvements in models rather than every time you you know model improves you have to kind of throw out your existing workflows yeah the question we ask very often when thinking about the features and how to position ourselves in the ecosystem, essentially.
43:56We've gone 51 minutes without talking about MCPs, which is tragic. Yeah, how do you do that? You talk about AI infrastructure and no MCP. Right? It's kind of crazy. But since you mentioned the prompts, we just had the episode with the MCP creators and they were, I wouldn't say not frustrated, but maybe they hope more people will use the prompts and resources in MCP servers instead of just the tool costs. I'm curious if you've seen any fun use cases with remote MCPs on E2B or any other things like that. Yeah, MCPs, watching it very closely, but still undecided. Not what it is, but what to do with it.
44:36There's no MCP on your docs, bro. If you go to GitHub, we have an MCP server. Okay. But we've seen people using E2B to host MCPs. but then I would say I don't think you need E2B for that you can just probably you can run it it's not even optimized for it probably there are probably better ways to do that I think people are a little bit too much focused on the protocol side of it even calling it a protocol, I don't know it's just like a server server and client and there's some agreed so I guess in a sense it is protocol, but I see a lot of people comparing to email protocols and such, which seems a little bit far-stretched to me at this current moment.
45:29But maybe I'm missing something. I think it's a super interesting idea. I just haven't had the right insight about it yet. One last thing I wanted to add, because we have a bunch of users that added E2B, our MCP server, to their registries. So I think, at least at the moment, And what's more useful is higher order MCPs. I was talking with Henry from Smithery about it, that you know. And he was telling me exactly this concept. It's unclear who's using the MCP at the moment. Is it a developer? Is it an end user? Or is it another agent? If it's a developer, then it might make sense to have a sandbox, creation of sandbox in the MCP.
46:13If it's an end user, probably it's too low-level primitive, so you want some kind of higher-order MCP. Instead of us offering a sandbox, we would be offering a way to build a co-gen agent with MCP or something like that. Or run a co-gen agent that's using E2B MCP in the background. So I have more these type of unanswered questions in my head about MCPs. I mean, they are all running locally right now, mostly. I don't think there's many remote MCP servers yet. I know that people are pushing for it. It looks like there's more remotely, actually. I don't know. This is what I heard from people managing these registries of MCPs.
46:55Well, they're incentivized to tell you that. Yeah. I fully agree with that confusion about what to do. I do think that every DevTools company needs some kind of MCP strategy, for better or worse, as annoying as it is. I think it does start with having an API though, rather than like a SDK first experience, because then people can just wrap in whatever language that they want. And then you also, for you particularly, you might want to have a distinction in your strategy for MCP clients versus MCP servers, because those might be different things. And particularly, I like what you said about the higher order MCP for the end user who doesn't really care about implementation detail.
47:38I think that is fantastic for E2B, where the agents can just spin up an E2B instance in the background and they don't even know about it. Yeah. In the world, we have like mcp.e2b.dev. First, it needs to have figure out authentication. So the agent should just ask to have, I guess, account created for it, whatever it is, it's going to be. And then it can just like launch a sandbox, do the execution there, and you can come back to it a month later and the state is still there. So that I think makes a ton of sense. Then you can start building higher order things on top of that, which could be very interesting.
48:17I like what you said about having API first versus SDK first approach. I think this is very, very important for the LLMs. And that's how we are building our whole new dashboard, the infrastructure. So everything needs to be API controllable from public APIs for users because eventually the LLMs will want to control and get this data. Okay, so that's a big shift for you to be because you're SDK first. It has been underneath API base all the time, but now we will go more into it. And it was more like in terms of prioritization. So we needed to start with humans to get to LLMs and first build for human developer and then now building for like a LLM developer first.
49:02Yeah. I'll just call out that since we did the MCP episode, they announced their update to the spec that they added an off component to the spec itself. It seems to be just based on OAuth 2.1. And I assume that's the first easiest thing to do, but there's no effective distinction between an agent and an user. So we never had that. I think this touches more like a broader question of you have all the websites that has optimized everything for humans, but now you will have like agents visiting those websites like what are the incentives there you know like probably you as a website owner you want to know it's an agent because you spend so much time and money optimizing everything for humans so i think like the the dynamics in on the internet get might get really weird if you don't have a distinction this is a human this is agent yeah we we tried to record an episode with matthew prince ceo of club player yesterday and then we technical difficulties, but they're building a lot of that.
49:58You can see the stats. They mentioned, for example, used to be Google would be a 2-to-1 crawl to referral ratio. So for every two pages, they would read, they would send you one visitor. He said OpenAI is 250-to-1. So they'll read 250 of your pages and send you one person. And Anthropic was like 6 ,000 to one. So they'll read 6 ,000 pages before they send one person back to your website. So obviously we have Jeremy Howard who's been working on llm.txt to kind of have a separate interface. for that i think today people don't really curate the llm experience uh i think a lot of the llm txt that people are making is like automate it take a website and turn it into llm.txt but like that's not really what you want to do it's like how do you separate completely the two things you know like if you're e-commerce store the llm txt should have your own inventory and like one thing you know shouldn't have a search button i feel like obviously llm.txt is a good movement that improve the legibility of these doc sites for LLMs.
51:00But I feel like it's kind of like a halfway measure. Like, I think we've failed with agents if we are reshaping the human environment for agents. Like, there must be two of everything. What's your human side and what's your agent side? Like, no, like, agents should just be the human side. Like, why are we making any special dispensation for these things? Yeah, I think the monetization is the only thing. Like, too many websites are, like, ads-driven, kind of like... You need to look at it, right? Yeah, I think you need to change. Once we figure that out, I think you can use the same interface. But I think today, all these pop-ups, it's like, fuck man, stop popping.
51:32Bitcoin solves this. Yeah. I'm just kidding. But anyway, yeah. I don't know if you have a take on all this. I have a general rule that usually when something new comes out, people have this tendency to recreate the thing that already exists for the new thing. That's like the internet for agents versus you already have all the infrastructure for the old thing and probably might be easier to teach the new thing to use the old thing. I think it's very common coming from developers that it will be a perfect world when you have nice clear distinction between these two, but actually the world is super messy.
52:09So nothing comes to my mind that they will like end up that you have clear distinction between like two type of sort of entities or we created like a new internet for just mobile phones. you actually have both desktop website version and mobile phone version, right? It's much more always complicated and messy in the real world than having a nice clear distinction, which is something I think developers strive for just from working with code because you want this nice clear distinction, but humans are more complicated than that. So I think everything ends up being sort of mixed and you need to adapt to it.
52:47For sure. Yeah, I think we'll do this for 20 years and then we'll figure out the... Yeah, and probably it will be somewhere in between that you have like a world where agents are using human internet, but the internet changed because you have agents or something like that. This reminds me of some conversation that I think, I don't know who was making this analogy. Like in the mobile era, we had m.yourdomain.com and then we had www.yourdomain.com and now we might have like just lm.yourdomain.com and that's just the LLM experience. MCP. Oh yeah, way better, way better. Yeah, yeah. Consumes everything.
53:21Cool. We're just going to go through all the rest of the use cases. We can refer people to your website, but I'm very just, I always want people to have a good mental map of how, when they should go to E2B and like, or what other people are using E2B so that they don't miss out, right? So it's AI data analysis, data visualization, coding agents, generative UI, code gen evals, and computer use. Do you think that's like the sequence of most popular is data analysis, least popular is computer use? Most experimental is computer use. it still gets a lot of traction when you share the demos of it it's just Manus?
53:56Who else is doing this? I basically don't see anyone else when you see computer use, I imagine a graphical interface Manus, as far as I know isn't running that so, what do you also call it computer use, but without graphical interface because you are using the full computer but more from a code point of view but that's a different discussion I would say computer use is very exciting, but experimental from what I've seen. I think for real computer use, you really want to support more platforms than just Linux. Yeah, Windows, right? And then it might be more licensing battle than technical battle.
54:37Yeah, lawyers always win. We'll probably have Eric from Pig at some point talk about his movements on Windows. So that was going to go into evals. right? And also how that links to RFT. Yeah, just like, can you tell more stories about the OpenR1 projects, how you work with them, and any other academics that are working with E2B that you think could be possible for the research or model training use case, or fine training use case? Yeah, the way HuggingFace who built the OpenR1 project is using us is during the reinforcement learning, reinforcement learning step, where the R1 model, the open R1 model, has a training step where they give it a code problem and the model needs to generate and run code somewhere.
55:28Then you have reward function basically giving you zero or one, telling you if that was a good solution or bad solution, then you improve the model. So kind of like the feedback loop. They are using the E2B sandboxes to run many hundreds of these sandboxes, thousands of sandboxes per training step. So they can achieve like big parallelization. We started very fast. You don't need to use your GPU cluster for that, which is like very expensive to use for this type of workloads. And also, and that goes with the story that I mentioned in the beginning, is that you don't need to worry about like LLM actually like changing permissions in your cluster and then you can't access the cluster because everything is isolated and the secure from each other.
56:13So I think that's a very, very interesting use case because we've had a few other companies like reaching out and started using us this way, building models, foundation models. It's like when we started E2B, that wasn't the use case that we had in mind. But it's like makes total sense. If you, also if you think about like life cycle of an AI agent, like if you, it makes a lot of sense for us to be like from the earliest stage possible. And the earliest stage is probably model training. So this fits very, very, very nicely in that use case. We actually released a case study with Hugging Face that's on our website that people can read.
56:53Have you seen people also use that to evaluate agents they want to use, or is it mostly people doing training? We've seen people using E2B for evals. Yeah, that's what I'm thinking. It should be very easy to run Sweetbench and all these on E2B. Yeah, this is more foreshadowing, but we will be launching in the next couple of months a startup and research program for people and universities and researchers doing exactly these type of things. A different use case, but we, for example, work with LM Arena folks from Berkeley that are using us to compare models in AI app generation and we run the AI-generated app.
57:31He wrote that in the equation, right? Yes. I think that's only possible thanks to Alessio and I'm not joking because first he connected it and then he actually wrote it. Man, the things you have to do to win deals these days. He was already our investor. So I think that also shows it even in a better light because you clearly see a person who's not interested just to win the deal, but actually you know. To increase the value of the investment. How many VCs are actually fixing your buck in your code base? Let's make a YouTube short of this part so I can share it. You should put it on our landing page.
58:09Yeah, exactly. Awesome, man. So let's talk about, since you mentioned VCs, a lot of VCs that passed on you before because you only do code execution. It's kind of like a small market. I would love for you to maybe also paint the picture. So you just mentioned it's cheaper than the GPU cluster than Hugging Face has. Do you want to do GPUs in the future? You mentioned Railway. That was easy to do. Do you want to compete with Railway down the line? Where do you see E2B going? So GPU questions is an interesting one. GPU market is hard. You are competing on the compute internal, I think it's hard. Because eventually you will have competitors and everyone will be pricing it a little lower and no one is making any margin.
58:52So that's like, but it just opens new use cases for you. And for example, even with the data analysis, if you just run, I think it's Pandas code, there's a recent update. if you just run it on GPU it's like twice as much faster if we want to like do really big AI data analysis you need that so I think like GPUs for us and also if you want to like have LLM trained small machine learning models maybe you want to like build full games you really want to offer the full cloud computer but in cloud that's very elastic so I think GPUs are there on the roadmap I wouldn't say it's like the thing that we are immediately working on.
59:36But eventually it makes a lot of sense for us to offer this. And so what was the other question? On the, do you want to host the apps that you are building to? We are very well positioned that the LLM is doing all the development work with us. Then you need to deploy it somewhere. So it's like very natural next step. Eventually, we want the LLMs to deploy these services, apps that they are building and have them manage it. And developer is more like in the backseat, like looking at things if everything is working correctly. If your swarm of agents are working correctly. It also requires a slightly different infrastructure for deploying.
1:00:18But I think there's a big advantage in knowing what developers are building on your platform. And then, because then you can see like what's the ideal, even from technical point of view, you have all the insights that you need to actually effectively deploy it and kind of then cover the full life cycle of building the app. But now it's not built by the human developer, it's built by the AI. So TLDR, yes, probably down the road somewhere. We want to build essentially the new AWS, but for LLMs. So we're just going to move out and just to wrap up. One of the interesting things that I saw you do was you were originally check-based and I'll be very blunt when I first invested in the early round that you did I was like these guys know how to recruit in Czech Republic it's like a competitive advantage and then the next thing I know you're moving to SF your whole team you should show off to our office, that's hilarious actually I think you offered even your place to live for so why move to SF?
1:01:25Do you think that everybody in your similar situation should? Any pros and cons that you're experiencing? I think you can definitely build a DevTool company from Europe. I think it's a lot about question of how easy you want it to be in the earlier days. Especially if you're building like red ocean versus blue ocean waters. if you are building in a field that already exists and you have large competitors you are building something that's 10 times, 100 times better you probably don't need to be NSF eventually you probably will need some kind of US base because for sales and customers but you can very well build this from Europe I know great companies doing that because the knowledge is already among all the developers But I would maybe argue that it might be harder to find people that are comfortable with a fast iteration loop and changing things early on.
1:02:27You know, almost we are pivoting every week, every month. But the main motivation for us was we just wanted to be very close to our users. And it was clear after a few weeks that SF is becoming this AI hub. and what we used to do and we still do it sometimes but slightly less because of like not having that much time and resources but we just met with the customers that had problems and we just like implemented it to be for them like next to them like we made a PR. I call this the Collison installation. By the way I don't know if this is a known thing but do you know how many times they did that?
1:03:10200? Twice? Three times? Oh, long. I asked about like three years ago when I had a chance to ask a question from Patrick. Like how many times you did the installation? And he was like three times, but then after that, you probably don't want to do that because you want to automate things and focus on other stuff. But it's exactly like do the things that don't scale. Three times. Exactly three.
1:03:40but my whole point was that how many such users i can meet in prague in a week versus in san francisco in a week in prague it's gonna be probably one and then i can't meet anyone else for the next half of a year uh because i just don't have users in prague but all of my users our users were here so we could just keep repeating doing that again again i would even argue we did it too much we could have automated slightly faster, but it's very useful feedback that you can get. And you can then start moving much, much, much faster. So that was important. And I think also, eventually you are in the B2B business, even from the early days, and it's good to have good relationship with people.
1:04:26And just like meeting in person is just better than meeting over Zoom. Well, I mean, that's why we do this in person. But yeah, I mean, obviously I run a conference. I'm very sympathetic to people meeting in person, right? But I also want there to be some hope for people who are never going to come to SF that they can still get involved. That's partially why we do this podcast is to get them involved in the community. And I mean, we started a new office in Prague. Yes, you're hiring again in Prague. Yeah, and I think that there's really good talent in Europe, in Czech Republic, for example. I don't know, I can't speak for other countries, but I can imagine it's very similar.
1:05:04Once you have a clear idea of what your product looks like, then you can find a really good expert on a certain part of your infrastructure, on a database, and things like that, and just get top talent for that. The reason we didn't want to do it early on was because we kind of didn't know ourselves what we were building, and we had to figure it out in person with users here. But now we feel much more strong about knowing the roadmap for the company and for the product. And so it's much easier to hire people that have 8-hour difference from us and explain them what they are building. Even remotely, sometimes, if you are not there and you're just communicating through Slack.
1:05:49Just to wrap, what are the roles that you're hiring for? So we are hiring distributed systems engineers. We are hiring platform engineers, AI engineers. We are also hiring account manager and customer success engineer. So kind of all over the place, we see a lot of market pool and momentum being built up. So we want to double down and move even faster because we see all the potential of what we can do. And just want to kind of like you want to pour the gas on the fire on the spark. and that's how I feel where we are now at G2B. Awesome, man. Thank you so much for coming on. Yeah, thank you for having me.
1:06:28It's great.
From the publisher
Vasek Mlejnsky from E2B joins us today to talk about sandboxes for AI agents. In the last 2 years, E2B has grown from a handful of developers building on it to being used by ~50% of the Fortune 500 and generating millions of sandboxes each week for their customers. As the “death of chat completions” approaches, LLMs workflows and agents are relying more and more on tool usage and multi-modality.
The most common use cases for their sandboxes:
- Run data analysis and charting (like Perplexity)
- Execute arbitrary code generated by the model (like Manus does)
- Running evals on code generation (see LMArena Web)
- Doing reinforcement learning for code capabilities (like HuggingFace)
Timestamps:
00:00:00 Introductions
00:00:37 Origin of DevBook -> E2B
00:02:35 Early Experiments with GPT-3.5 and Building AI Agents
00:05:19 Building an Agent Cloud
00:07:27 Challenges of Building with Early LLMs
00:10:35 E2B Use Cases
00:13:52 E2B Growth vs Models Capabilities
00:15:03 The LLM Operating System (LLMOS) Landscape
00:20:12 Breakdown of JavaScript vs Python Usage on E2B
00:21:50 AI VMs vs Traditional Cloud
00:26:28 Technical Specifications of E2B Sandboxes
00:29:43 Usage-based billing infrastructure
00:34:08 Pricing AI on Value Delivered vs Token Usage
00:36:24 Forking, Checkpoints, and Parallel Execution in Sandboxes
00:39:18 Future Plans for Toolkit and Higher-Level Agent Frameworks
00:42:35 Limitations of Chat-Based Interfaces and the Future of Agents
00:44:00 MCPs and Remote Agent Capabilities
00:49:22 LLMs.txt, scrapers, and bad AI bots
00:53:00 Manus and Computer Use on E2B
00:55:03 E2B for RL with Hugging Face
00:56:58 E2B for Agent Evaluation on LMArena
00:58:12 Long-Term Vision: E2B as Full Lifecycle Infrastructure for LLMs
01:00:45 Future Plans for Hosting and Deployment of LLM-Generated Apps
01:01:15 Why E2B Moved to San Francisco
01:05:49 Open Roles and Hiring Plans at E2B




