In short
Dev Interrupted Podcast Episode Summary
Episode Title
Scaling ChatGPT: Inside OpenAI's Rapid Growth and Technical Challenges Guest Host: Ben Lloyd Pearson Guest: Evan Morikawa, Engineering Manager at OpenAI Release Date: [Date Not Specified]
Episode Overview In this episode, the hosts delve into the technical challenges and rapid scaling issues faced by OpenAI during the viral success of ChatGPT. Evan Morikawa shares his insights on the engineering hurdles, misconceptions surrounding generative AI, and the critical role of GPUs in AI computations. The conversation also highlights use cases of GPT-4 and the importance of APIs in facilitating the integration of AI into various applications.
---
Key Discussions
- Scaling Challenges with ChatGPT
- Viral Launch: Unexpected surge in user engagement post-launch, primarily driven by social media.
- GPU Shortage: The difficulty in scaling due to a scarcity of GPUs which are essential for processing.
- Traffic Management: Implemented a "capacity page" strategy to manage user inflow without creating a waitlist.
- Misconceptions About Generative AI
- Black Box Nature: Many users perceive AI systems as omnipotent, but the underlying processes are complex and context-dependent.
- Prompt Engineering: The effectiveness of AI can be improved through well-structured prompts, akin to guiding a human to give better responses.
- Use Cases and Applications
- Generative AI in Development: Tools like GitHub Copilot provide substantial assistance in coding by offering suggestions and educational support.
- Caution in Use Cases: AI should not be relied upon for critical decisions, such as medical advice or legal citations, due to potential inaccuracies.
- Role of GPUs in AI Computation
- Technical Importance of GPUs: GPUs enable the massive parallel computation required for AI tasks, executing quadrillions of operations per second.
- Capacity Considerations: Models often require multiple GPUs, making efficient interconnectivity and memory bandwidth crucial for performance.
- Adapting to Growth
- Nimble Engineering Teams: OpenAI's structure allows for small, agile teams that focus on rapid iteration and integration of research with product development.
- Collaborative Culture: Close collaboration between research and engineering teams enhances product development and responsiveness to user feedback.
- Future Directions and Innovations
- API Ecosystem's Role: OpenAI sees APIs as a means to enable broader innovation and facilitate integration with various applications.
- Expanding Use Cases: Significant opportunities are anticipated in fields like law and education as AI tools become more versatile and widely adopted.
---
Key Takeaways
- Iterative Development: OpenAI's success hinges on a philosophy of rapid iteration and responsiveness to user feedback.
- Safety First: Any product rollout is contingent on robust safety measures and testing to mitigate risks associated with AI technologies.
- Continuous Learning: The need for engineers to understand AI models and their implications is critical for effective problem-solving in AI-driven projects.
Conclusion This episode provides a detailed account of the challenges and operational strategies of OpenAI during a pivotal moment in the AI landscape. Insights from Evan Morikawa underscore the complexities of scaling AI technologies and the importance of fostering a collaborative environment to drive innovation.
---
For more insights from this episode, visit [Dev Interrupted](https://www.devinterrupted.com).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00It took a lot of tweaking to kind of find what the right utilization metrics were and then how to optimize those. and all of those are really critical to getting more out of it. But for us, everything was framed in terms of every improvement that we have represents more users we could let onto the platform. As opposed to say like, oh, it's driving down our margins or making it faster. We kept cost and latency as relatively fixed. And the thing that could get to move was we got to get more users onto the system. What's the impact of AI-generated code? How can you measure it? Join Linear B and ThoughtWorks Global Head of AI Software Delivery as we explore the metrics that measure the impact Gen.AI has on the software delivery process in this first of its kind workshop.
0:48We'll share new data insights from our Gen.AI Impact Report, case studies into how teams are successfully leveraging Gen.AI, impact measurements, including adoption, benefits, and risk metrics, plus a live demo of how you can measure the impact of your Gen AI initiative today. Register at LinearBee.io slash events and join us January 25th, January 30th, or on demand to start measuring the ROI of your Gen AI initiative today. Hey, everyone. Welcome back to Dev Interrupted. I'm Ben Lloyd Pearson, Director of Developer Relations here at Linear Bee. I'm pleased to have Evan Morikawa joining us today.
1:26Welcome to the show, Evan. Awesome. Thank you. Yeah, so full disclosure, Evan and I worked together in a past life at an API company. You may not remember this, but you were actually a big part of why I decided to join that company. I did not know that. Wow. You saw your video on YouTube. Oh, wow. Yeah, yeah. So yeah, we've got a little bit of a history. So it's really great to just have an opportunity to catch up with you, talk about some of the new stuff that you're working on. I know, you know, personally, I was a little bit envious when I saw that we were going to open API because I was like, that sounds so cool.
1:58And here we are talking about how cool your work is. So let's kick it off. I mean, OpenAI, I feel like needs no introduction. I mean, I feel like everyone is talking about it. It's gone viral. At the center of the conversation around AI and LLMs, the release of ChatGPT has kicked off a global phenomenon. And I want to walk through that story, in particular, what it took for you to scale ChatGPT when you had that viral moment. And especially, I know I've heard a little bit But there was a shortage of GPUs that also affected this. I do want to dig into that. And, you know, I just want to learn a little bit about how you've been flexible and nimble while this industry has just like rapidly shifted around you.
2:39So let's just kick it off. Where does this story all start for you? Like, how did you get to where you are right now? Yeah. So, I mean, as you mentioned, I was working at an API company before. And at the time, that's all OpenAI had. In fact, that was the light bulb moment was learning that three years ago or so now, OpenAI was starting this Applied team. And Applied was all about bringing this crazy technology safely to market. And at the time, that was just a single developer-facing API around the very first GPT-3 models. Which is good because up until that, like, it's still the case that I do not have a machine learning background.
3:17OpenAI was much more research lab focused. But then here now, it's starting up this brand new small team doing APIs and products. And I was like, oh, I've been doing APIs and products. Maybe I can help contribute here. I think there was also that feeling, too, that, you know, I feel like I had a reasonable sense of how computers worked by that point in my career, except for this. This still is this kind of magic. Certainly at the beginning was this magic black box. I'm like, OK, I'm going to see what that's all about as well. I also knew of like some of the founders through just the broader network and yeah, reached out and that was the start of it there.
3:56When I joined, Applied was very small. There were only basically like half a dozen engineers working on all of the APIs and systems for that. And that steadily grew as we were trying to work on iterating on these language models. I think the next big, the big really first release or push of any kind was when we tweaked these to write code and worked with GitHub to launch GitHub Copilot. So GitHub Copilot actually initially ran through our servers at launch because it was very difficult to run or deploy these. That was definitely the first time we had any experience running this at any sort of scale.
4:36but still all the way through basically up until chat gpt and still to this day have this very core of an api business that powers all these other ai powered applications that you uh that a lot of people are trying to build on now and then and then along came chat gpt it's actually kind of interesting story because you know we uh when chat gpt launched there it wasn't necessarily sure whether or not it'd be like a scary thing or a totally normal thing. You know, on one hand, the model that was powering it, GPT 3.5, that had been out for several years at that point in various iterations. People could already sign up for free through the developer playground.
5:20And in fact, really noticed a lot of people like playing with the models in that kind of way. So in some senses there wasn't that much different about it you know maybe a new ui also the model had been improved dramatically to make up things less and be more conversational but you know on the flip side this is also the first time we would ever be offering anything without a wait list this would be a free to use application out there and yeah that definitely changes things as well um you know Well, actually, on launch day, I think it launched on a Wednesday. It was like November 30th. And we kind of by design sort of thought this would be a low-key research preview.
6:06Just a blog post, a tweet, and nothing else. And actually, on launch day, nothing crazy happened. Like some people came and used it and never passed like number five on Hacker News. We had all the capacity we needed. Traffic tapered off for like, great, quick little launch, like move on. You know, actually at the time we were preparing internally for the launch of GPT-4, which was coming up next. So this was actually a way for us to experiment with a lot of the recent fine tuning that had gone into the older models to like make them safer and more conversational. It was the next day or rather 4 a.m.
6:45the next morning when our on call starts to get paged because traffic is starting to really rise. There was this graph. We were trying to figure out what was going on and all the traffic was only coming from Japan. And we were very confused. We actually thought we were getting like DDoS or like attacked or something like that. But no, they had just woken up first. It was starting to virally spread through Twitter. and then by the time the morning of the East and the West Coast picked up, it was like very clear that we were getting hammered here. Unfortunately, there, you know, normally you can just like throw more servers at the problem.
7:24But yes, there is like a very finite supply of GPUs here. So there's really not much we can do about it. We did have some contingency in place for this. the idea being that we could throw up a, like, we are at capacity page, this kind of, like, bouncer model, if you will, right? Like, oh, the club is full. When some people leave, we can let more people in. That was actually explicitly done because we also did not want another waitlist. Like, no one likes a waitlist. So we're like, oh, we can try it this way. But unfortunately, that we are at capacity page was up a lot for the first while. Why were we trying to like scramble to get to get more capacity online?
8:06And we'll just like fix the long tail of other stuff, too. GPU capacity was definitely a dominant concern, but we also had all the other scaling problems. Kind of other. Everybody else in engineering has, too. Yeah. Wow. That's a fascinating story. And trust me, I remember those capacity pages quite well. So, yeah. So I want to talk about a little bit before we get into the GPU stuff and some of the scaling issues. about this black box that is AI because it is really how it feels to a lot of people. And many of the productized versions of it, that's really kind of how they perform. And there's no shortage of content on the web that describes how generative AI and LMS are going to do both wonderful and horrible things to all of us.
8:52And as this initial hype wave wears off, I think we're really starting to see concrete use cases that truly bring a lot of value to people. And I'm thinking, And, you know, just from my personal experience, like my grammar checker software giving me more intelligent advice about my writing or, you know, Copilot, we brought that up. That's a great example of, you know, I basically have like autocomplete for my IDE now and it can't do everything, but there's still some things that it just saves me so much time. So, you know, I think it would be really interesting to hear from an engineer that's actually building this stuff, like about the biggest misconceptions, the biggest misunderstandings that you've seen related to all of this.
9:29and specifically if there's like one or two things that you can clarify for the world about AI, like what would those things be? That's a good question. The certainly one misconception is that it is definitely, certainly not as it exists today, this completely omnipotent system here, right? It is, there are quirks about how this thing works. It's actually quite helpful to remember somewhat how these things are trained. They're trained by predicting the next word for all words and phrases on the internet. And that, though, is in itself deceptive, because this is much more than a Mad Lib system or an autocomplete engine, because it has turned out that in order to be able to predict the next word, they kind of need to know a huge amount about society and structure and context and culture and things like that as well, too.
10:18But at the same time, they're also very steerable based on the context that you give it ahead of time. You know, when GPT-3 first came out, this was, in fact, the title of the paper is about few-shot prompting. Few-shot here basically refers to the idea that you just give the model a handful, like three or four examples of what you wanted to do, and that kind of eggs it into the right direction. In some ways, this is not too unfamiliar. I think, you know, if you go on Google, or old Google at least, if you type a question one way, you get Yahoo answers. But if you type it a different way, you get like a scientific paper.
10:52So kind of in the same way, it can steer it in the direction of things. The one thing that's a little weird about this, though, is this has left some very seemingly black box types of prompt engineering into this right now. There were some papers recently that have been getting a lot of press that simply say, if you ask in the prompt, literally, take a deep breath and think step by step. It does much better on a large category of tasks. That's something that feels like kind of wrong on one hand. It feels very humanistic, though, doesn't it? That's actually the great point. On the flip side, though, if you kind of assume that the models are going to approach kind of a human type level of intelligence, it's worth to ask yourself, if you threw a relatively competent human in what you're asking it to do, with as much context as you gave it, how would they perform?
11:50And it's not unreasonable to think that these models kind of mimic that because they are mimicking human speech as we've seen it on the Internet. So, yeah, in fact, if you are actually some of the people who are that could be the best at prompt engineering are like teachers, engineering managers, tutors, people who are used to like asking the right questions and setting the right context for people to like help them arrive at the right conclusions. And if you kind of think about it that way, you get like weirdly better results across there. Yeah, and I think you're actually partially answering my next question.
12:27So, you know, in your opinion, like what are the situations that are ideal for generative AI? And alternatively, like when would you steer someone away from it as a solution? Yeah, yeah, yeah. Some places I think has absolutely been helpful. I've really just started to tap at right now, certainly software engineering, coding, boilerplate type things. that was the first place that we internally really dogfooded any of this was when we ourselves started to use Copilot as an educational tool. I think that is still, and despite how much has been talked about, still an underrated, undertapped ability here.
13:01The idea of a personalized tutor everywhere you go, like TAs, the thing about university, the professor can talk at you all day, but it was the TAs where I like really learned things and those follow-up sessions, because you could ask follow-up questions and you could frame it in a way that makes sense to you. That kind of iterative learning, I mean, this is where I personally use it the most, like the thought of reading any paper without this thing on the side or without being able to like, just being able to like ask it for concrete examples of things, being able to rephrase and reword and take follow-up questions.
13:38That I think is going to be a huge deal. There are some places on the flip side of this, though. Yeah, like it still has. We have not at any way, shape or form solve this like perfectly verifiable problem. This should not be your end state for medical advice on a huge sleuth of topics right now. It should not be the thing that is you are trying to use to cite case law for your own trial, for example. I think a lot of people call this hallucination. That's right. Now, on the flip side, there's actually law. I think it's a really interesting area as well, too, especially some of the side effects of this for like the power of the embedding models that we have.
14:18You know, it is very good at saying what patents are similar to this one, what cases are similar to this one. And in ways that are much more than just do they have similar keywords. But the fact that these models deeply semantically understand what's going on, it can help you find and search like that. Like that's, yeah, those types of search abilities will get dramatically more powerful. Yeah. Yeah. And I know personally, one area that I've really found a lot of benefit is I have to very quickly understand a lot of new technologies as a part of my day and work with things like, you know, I don't, I haven't done a lot of regex in the past, but I find myself doing a lot of it today.
14:56And just getting that like intermediate understanding like immediately without having to like parse through a bunch of resources across the web. I mean, I can't even add up the number of hours that that saved me. So, yeah, it's really great. Regex are probably the best example of that. Yeah, that's been wild. It's actually kind of blown my mind at how quickly, you know, because I mean, Regex is like, it's not complicated, but it can be time consuming if you don't work in it frequently. So then, you know, as an engineering manager, like what expectations are you setting with your team in regards to the use of generative AI?
15:29So, you know, it sounds like you were an early dog fooder of Copilot. You know, is that, are you, as an organization, are you like systematically adopting tools like that? And, you know, what are the like changes that you've seen? Yeah. So getting this really, we definitely want to more and more have this help us be productive. I mean, the, we actually have a research team called the AI scientists team, which is very much long-term about being able to make this work. At the same time, though, there's a pretty wide gap between what works just straight off the bat from the prompt in chat GPT and like an actual tool you'll use day to day.
16:10Like, yeah, some people were like did a quick plugin to like hack in before VS Code. But I mean, it still takes a lot of product work to make to make to put the whole experience together and make it work really nicely. I think this kind of immediate generation of like developer productivity tools, it will take a fairly large investment to like really put it naturally into a workflow. This is actually why I'm very optimistic about the coexistence of both a tool like ChatGPT and the API ecosystem. Like, ChatGPT is about, we think it can be useful in lots of different places in a much more kind of generic sense.
16:49But there are also a huge number of industries where, like, being in flow matters a huge amount. Developer tools, law, medical systems, like all these other places. You'd also need and want a lot of these, like, integrated applications as well, too. But, yeah, like, we absolutely believe this will make us and everybody else substantially more productive over time, for sure. Cool. So let's pivot a little bit to talk about what I think is going to be the most unique aspect of this discussion. So, you know, in recent years, there's been something that has brought gamers, crypto enthusiasts and LLM experts together.
17:25And that is the frustration over this shortage of GPUs. Right. I tried to upgrade my PC a couple of years ago. So, you know, before we get into that shortage, I think it would be beneficial to just step back for a moment. So can you describe like the technical role that GPUs play in the OpenAI tech stack? And, you know, why are they so important and how do you all use them? Yeah, absolutely. So at the end of the day, when you ask ChatGPT a question, it's taking your text and it's doing a huge amount of matrix multiplication to kind of predict the next word, right? That's what all these hundreds of billions of model weights are for.
18:01And at the end of the day, we're basically doing one math operation, multiplying and adding a lot, which like we're talking like quadrillions and quadrillions of operations a second here to do this. So the ability for a GPU to be able to, GPUs are just many orders of magnitude faster here. For a sense of scale, the latest GPU that will be running, this NVIDIA H100 that everyone's been trying to get, that can do about two quadrillion floating point operations per second. Your laptop CPU probably can do on the order of a couple hundred billion. So we're talking like thousands of times difference in speed here.
18:44So yes, you can run these things on CPUs, but the performance difference we're talking about is 100x. So especially for models this size, it's really important to do that. The other thing that's significant is that the models are so large, they don't fit on just one GPU. We need to put them on multiple different GPUs. That's actually where things jump dramatically in complexity. Much like the rest of this world knows that, yeah, your simple single-threaded application makes sense. But once you run it massively concurrently on a globally distributed system, that's where the hard problems come from.
19:24Similarly here, once you run these on multiple GPUs, things get a lot more difficult. Now you really care about how fast your memory bandwidth is. You really care about how fast your interconnect, your network bandwidth is between GPUs, between boxes. And it's gotten to the point now where every single one of those metrics can become a bottleneck at various points in the development cycle. So we really care about all of them and we really maximize them. And any time there's a newer, faster interconnect, that usually almost directly translates to improved speed for ChatGPT. that directly translates, if you can make it run twice as fast, that's twice as many users as can access it on the same hardware.
20:10Or that's twice as large of a model as you can run on the same hardware today. So these really make a huge difference in what is capable going forward. Nice. So I think that describes pretty well the impact that, you know, a shortage of GPUs would have on your company. So beyond a page that says, hey, we're really busy right now, come back later. But what other strategies did you all take to adapt to the sudden influx and the lack of this hardware resource that you need? I have to do a little bit of everything. Certainly one was trying to find more GPUs where we could. We were working very closely with Microsoft, who subsequently was working closely with NVIDIA on this to build out capacity here.
20:53But also it was about making the most of the resources that we had. So optimizations are hugely important here. And this is the long tail of sort of your classic, put it through a flame graph, figure out the parts that are slow, optimize those. But for us, those optimizations are all across the stack. It's from low-level CUDA kernel compiler optimizations to sort of more business logic. How are we batching requests together? How are we maximally utilizing these things? It took a lot of exploration. Like we discovered that we started with a very basic GPU utilization metric, you know, from whatever the NVIDIA box spits out.
21:35We found that was actually misleading because we were not, like it could be doing more math when the same time it was on or we're actually running out of memory instead. So it took a lot of tweaking to kind of find what the right utilization metrics were and then how to optimize those. and all of those are really critical to getting more out of it. But for us, everything was framed in terms of every improvement that we have represents more users we could let onto the platform. As opposed to say like, oh, it's driving down our margins or making it faster. We kept cost and latency as relatively fixed.
22:13And the thing that could get to move was we got to get more users onto the system. Gotcha, gotcha. And, you know, I've always kind of felt that that GPUs were chosen for LLMs mostly out of convenience because it was the hardware that was available at the time that was most closely adapted to the needs of that community. So are you looking at other hardware options out there? Like, are there things that are coming up in the market that you think have potential to replace GPUs specifically for generative AI? Yeah, well, so even though we all call them GPUs, like, these are not the graphics processing units of your desktop PCs anymore.
22:49In fact, most of the ones targeted at data centers can't even do graphics. They would have awful frame rates for your machines. And especially with Google calling theirs TPUs, this is why it's sort of like AI accelerators, more the generic trend now, but I still call them GPUs. I mean, at this point, they are hyper specialized to do this exact one type of matrix operation. another concrete example here that they've been optimized only for AI things is doing lower precision math most people when they have a floating point number you get 32 bits to preserve it we do math with you can do math with 16 bits so with 8 bits which means you can just like do more at the same time and now they have dedicated circuits to do that on the upcoming GPUs that are coming out so in a lot of ways they are super specialized the other thing though is the software stack So, NVIDIA has their CUDA stack, their software stack, the kind of compiler layer on top of there has been hugely specialized to that.
23:59It's currently very difficult for people to use AMD or Intel or the other manufacturers out there. Actually, there is a product, OpenAI Triton, which is explicitly designed to try and better abstract that. And that's potentially a huge deal because the ability to use other hardware much more easily is definitely a big thing, will be a big thing for this market. But right now, yeah, NVIDIA has an incredible hold on this. It's reflected in their share price right now. A lot of it is because they own the hardware, the software stack, and a lot of the interconnects. For example, we use InfiniBand, which is a like ultra high bandwidth interconnect.
24:42That company that developed it, Mellanox, is also owned by NVIDIA. So they like really have the stack top to bottom here. Yeah, they definitely got in early because I remember playing around with LLM tools years ago. And like NVIDIA was the only option in the market. There really wasn't anything else to look at. And I remember CUDA in particular, that was around the time that it was really taken off. Yep, yep. Yeah. Actually, I should note, though, it's been very, very difficult for the chip manufacturers to even get the right chips, though. Another example, despite NVIDIA's dominance here, their upcoming chip, this H100, is kind of widely known, well, within this specific subset of the industry here, that it doesn't have enough memory bandwidth relative to how much compute they added into it.
25:32So it's getting increasingly difficult to utilize this. But the reason that happened is because they didn't know about how large the models were going to get. They didn't know. It was very difficult to predict this on the scale of like chip development cycles. ShadGPT is less than a year old. And one year in semiconductor manufacturing design is nothing. So, yeah, it's very difficult for anybody to predict to do this. So I want to transition a little bit into talking about how you approach scaling the engineering function at OpenAI as this was going on. So, you know, the rapid success, no secret at this point, your leadership has been very open about sharing the challenges of this rapid, sudden virality.
26:16And, you know, it's one thing to deal with sustained growth over a long period. But when you deal with like this sudden surge, it's an entirely different beast. because, I mean, not only do you have to deal with potentially much higher peaks, but you don't know how much of that is going to stick around for the long term, right? Yep. So, you know, walk me through how that played out for your engineering organization. Yeah. So for a sense, staying nimble has been a huge piece of this. One thing that is like very much helped in by design is doing everything we can to try and treat everything like a tiny early stage startup.
26:51Yeah, this originally was true in terms of raw, like headcount here. Yes, there are a huge number of people that contributed to the research and the model training, but at the end of the day, the kind of product engineering design and like parts that is applied, it was only several dozen people when ChatGPT launched. So it was still like a much smaller group. Even still, we intentionally set up ChatGPT as a more vertically integrated sort of separate product team within Applied. If you think of Applied and the API as this three-year-old startup, ChatGPT looks, feels, and acts like a 10-month-old startup.
Read the full transcript
27:31And concretely, that was in the form of we intentionally started on a separate repo, separate clusters, different controls, taking on a little bit of that kind of tech debt and duplication at the start to really optimize for iteration there. whereas gradually the API started to optimize a little bit more for like stability and SLAs and stuff like that too. Now this is of course changing. ChatDBT also has like huge stability and SLA concerns. We are kind of working to build out more broad platform teams as well. But nonetheless, this idea of like keeping things really product focused, fast iterating was important.
28:11The other kind of key piece about this was having the research teams deeply embedded here. So while I talk about applied, because that's a group I'm in, in reality, ChatGPT effort heavily involved a huge chunk of researchers from various research teams. They were the ones who were actually constantly tweaking and updating and fine tuning the models based on end user feedback as well here too. So keeping these as very vertically integrated, like teams with both product engineering design and research was also super important. Yeah. So what would you say is like the biggest improvement that has come out of this from an engineering perspective, from, you know, just dealing with all of this scale?
28:53The biggest improvement that's come out, actually our ability to work together as like a single research product group. There was an early fear that it would be like the worst case scenario for us would be the type of place where research trains a model, throws it over the wall, go productize it. And it was like this one-way street. And we spent a lot of effort making sure that that was not how we developed anything. But, you know, that was all like abstract. You like are actually in the trenches developing and like really working on a product here. And of course, the reality is about like, just like, it's very messy to begin with.
29:33You just kind of like have to tweak it as it goes along. But now that there's a much stronger focus around sort of these clear products we have with this clear API product that we need to build, We have this clear consumer app that we're focusing for. I think that has helped a lot really integrate the research and product and engineering and design and kind of this one push. And I imagine that probably gives your engineers an opportunity to learn more about how this stuff is created, right? Yes, it has definitely been necessary. So it has been the case, actually, that everybody in applied and engineering did not need to have a machine learning background to do this.
30:09I do not have a PhD in machine learning, but that's fine for now. Certainly the interest to pick up and learn a lot of things along the way is important. But yeah, a lot of our, most of our problems are product problems. They're engineering problems. They're distributed systems problems. They're kind of classic like that. But at the same time, it has been really important for everybody to at least get a reasonable understanding of how all these models fit together. because a lot of the engineering considerations are deeply tied with the way these are structured, the hardware that we're using, the way things are deployed.
30:47Those all definitely matter. Yeah, this might be a tough question to answer, but if you could go back in time to Evan a year ago and tell him, hey, your product is going to go viral someday, but are there any changes that you would have made in your approach to respond or to anticipate that? You know, maybe not because I I'm inherently a bit of a skeptic when it comes to things like I would not have believed my thing would go viral. Like that's not like also because like I believe in not like prematurely optimizing the system, like not really like we have this like deeply iterative model like baked into here.
31:30So I actually would have been afraid that we would have spent this huge amount of time, like, making sure the infrastructure was load tested up the wazoo, that making sure the product is perfect before it even got out there. So you're saying just swing for the fences and deal with what happens after the fact? Yes, no. So, I mean, one thing that we do now, I should note, though, that iterating a product quickly does not, especially here, it's very important that that does not compromise the kind of safety mission that we have to begin with as well. So I should note that the, like, safety, not being happy enough with our safety systems, with the red teaming results that we're getting, that is the primary thing that will delay launches.
32:21That is the non-negotiable before we can ship something. Yeah. So that is the place that we can, should, and would, like, spend even more time iterating on. But here again, a lot of the philosophy around this is that we kind of see the safety systems of the red teaming layers coming in layers. Like we have a very active network of experienced red teamers who will go in and test stuff. We have very controlled rollouts through various stages to try and like catch everything. But it's still the case that there's no way to catch all of it until you actually get it out into the world. Yeah. I mean, did you test, take a deep breath?
33:00and absolutely not. Yeah. Yeah, like having the research field on this is significant. Having a lot of people thinking and working and doing this is a big part of what it takes to make these things move forward. So I got one more subject I want to talk about and it kind of brings us back full circle to how this conversation started and that's APIs. So, you know, we mentioned that the part of generative API that really excites me is seeing it pop up in all the tools that I use every day. you know, at Linear B, you know, we've trained ChatGP to use multiple aspects of our platform. So we can do things like ask it questions.
33:37It can write configuration files for us that are highly specialized. And, you know, that's like, to me, that's like the real, like, fascinating use cases. Yeah. But, you know, obviously for that to happen, there's got to be a really strong API for developers to build functionality on top of it. And, you know, I think since we both worked at API companies, we understand that, you know, it really comes down to making sure the API is performant and that it's built for real world use cases rather than theoretical situations. So what role, generally speaking, do you think APIs or the APIs for your products are going to play in the success of OpenAI and ChatGPT?
34:12Yeah, there is absolutely no way, despite how I think we have a very talented team, there is no way our one team can out-innovate the vast sum of all of the creativity and startup energy and like company focuses that it has on making this work right now. As I mentioned earlier, like the ability to have this, to have AI deeply integrated with everything you're doing somewhat seamlessly and transparently, that's where a lot of real power is. Yeah, and it's really only possible through a type of API environment too. Also developers and APIs are one of the best places for us to sort of try out new ideas first.
34:55That's actually a great point. I wanted to ask, like what kind of feedback are you getting from developers in the field? That's one of the most important pieces about it. Consumer apps, it's very difficult to get feedback on unless you start doing really aggregate stuff. But yeah, if you sit down and you like talk with the developers building on the API, we really learn where the actual problems are in the system. Like it's the gap between a cool demo and something that's useful. is massive. And that is true here. That is true in every industry. It is especially true here. The more hype that there is, the further that gap becomes.
35:29And that gap can only be closed when you're really working with companies and developers, figuring out the hard way and learning all the quirks of it. It is definitely the case that the API ecosystem has discovered far more quirks of the models than our own research teams have. So what are some of the engineering challenges that have been unique to the API versus the the more general purpose tools. We have lots of similar challenges to other API companies, you know, like database connection limits, networking shenanigans. The GPU constraints, though, I would say those are quite different. We've needed to jump immediately to tons of clusters all over the world.
36:10That was mostly done because we were chasing GPUs wherever we could have. So we are like a, we suddenly found ourselves multi-cluster, multi-region. On the flip side, though, we've also spent a fairly large amount of effort making it such that there really aren't that many special unique things about our deployment stack. We've been using stock Azure Kubernetes service. We use a lot. We use like Datadog. We use a lot of just like off-the-shelf tools. I think that's actually been really important. That's really helped our development team stay small. It's meant that new hires come in, kind of know what they're doing.
36:46the more we can do to try and treat things as just another service that takes text and spits it out the other side. You know, a cube service is a cube service. People know how to deal with that. At least it's a known unknown. I'm trying to like minimize the unknown unknowns here. But yeah, at the same time, the scaling characteristics of this are very strange. I was mentioning that when, yeah, talking about these like utilization metrics were hard to figure out initially there too. And the scaling challenges of this, I think, are also going to get pretty nuts. Like, the ambitions for the scale of capacity that we need to ramp up to, it's both usage going up, the models getting bigger, the models having all these different types of modalities for them.
37:35Yeah, that's all just getting started right now. Yeah, like, we found ourselves suddenly, when Dali came out, we're like, oh, we're suddenly not dealing with text anymore. Now it's all everything that is image processing and image generation. Now we're in audio as well with all speech in and speech out. So, yeah, it's going to become a lot more complex. Awesome. They're all very fascinating. I'm really happy we got to learn about all this from inside the organization. And I have a couple of just quick questions because they're things that I think that our audience is really going to be interested in.
38:10The first is, what is the most interesting or useful adaptation you've seen so far of one of your products? So I think this is still emerging yet, but there's a lot of work of people trying to build longer form agents right now on the system. That's been a huge focus of a lot of startup activity. I'm very excited to see where a lot of that goes right now. As I mentioned, a lot of the law applications, I think, have a huge potential to feel like that industry should feel very different. I would love to have a lawyer in my pocket. Just saying. Yeah. Then the other one, of course, is like the education side of things.
38:49What's this? We have a real effort to figure out how to use these tools. To me, it is kind of analogous to, I don't know, math classes had to do something different when my calculator showed up everywhere. But we will, figuring that out and sort of like getting in a world where people start to use these as tools and can like help them be just dramatically more impactful things. I think that's a huge deal. I actually really liked Spotify's feature they launched recently, which was using text to speech. So they release these podcasts in other languages with the voice of the original podcasters. Wow.
39:31So yes, you can like listen to Lex Freeman in Spanish, but it is clearly his voice. And if you think about how much work or money it used to take to dub something or to hire somebody to do that, like, that's insane. The other partnership I think is really cool is Be My Eyes. They've been using the GPT-4 with vision capabilities to help visually impaired people. Take a photo of your closet. Like, what should I be wearing today? Yeah, that's nuts. You now have a device in your pocket that can deeply, meaningfully, and semantically describe what you're looking at. And if you're a visually impaired person, that's a huge deal.
40:12Yeah. Awesome. Well, it's been really great learning the inside perspective from you on OpenAI and all the engineering work that you're doing over there. If people want to learn more about you or the work that you're doing, where's the best place to send them? Yeah. So I'm E0M on Twitter. And actually, I actually point, including a lot of our new hires, to OpenAI's blog. which both is a mix of like the product and research releases too, but kind of that whole arc gives a pretty good view of the kind of the state of what we're doing, but also say to the kind of the industry too. Awesome. And I know personally, I've been silently watching your LinkedIn too and seeing a lot of your posts about all the cool stuff that's happening.
40:53So yeah, well, it was really great having you here today. I'm glad we got the chance to catch up and talk about what you're doing. So thanks for showing up. Likewise. Thank you.
From the publisher
What can you learn from the scaling issues OpenAI experienced when Chat-GPT went viral?
On this week’s episode, guest host Ben Lloyd Pearson is joined by Evan Morikawa, Engineering Manager at OpenAI. Join us for a first-hand look at the engineering challenges that came with Chat-GPT’s viral success, and the difficulties associated with scaling in response to the sudden platform popularity.
They also discuss misconceptions around generative AI, OpenAI’s reliance on GPUs to carry out their complex computations, the key role of APIs in their success, and some fascinating use cases they’ve seen implementing GPT-4.
Show Notes:
OFFERS
- Start Free Trial: Get started with LinearB's AI productivity platform for free.
- Book a Demo: Learn how you can ship faster, improve DevEx, and lead with confidence in the AI era.
LEARN ABOUT LINEARB
- AI Code Reviews: Automate reviews to catch bugs, security risks, and performance issues before they hit production.
- AI & Productivity Insights: Go beyond DORA with AI-powered recommendations and dashboards to measure and improve performance.
- AI-Powered Workflow Automations: Use AI-generated PR descriptions, smart routing, and other automations to reduce developer toil.
- MCP Server: Interact with your engineering data using natural language to build custom reports and get answers on the fly.
