In short
The episode traces the evolution of servers and cloud computing from the late 1990s dot-com era to today, then connects it to what Oxide is building next: a hardware+software “clean sheet” data center computer designed for modern, API-driven infrastructure.
Guest backgrounds
Brian Cantrell is a distinguished engineer who worked at Sun Microsystems during the dot-com boom and bust, later built a small AWS competitor called Joyent, and is now co-founder at Oxide.
Key claims
Innovation often comes from “desperation,” so the dot-com bust produced more technically significant work than the boom. The shift from proprietary, tightly bundled hardware/software to open source (Linux growth; Solaris open-sourcing) and x86 commodity performance enabled mainstream server building. Cloud’s real breakthrough was API-driven, elastic infrastructure (not just elasticity). Kubernetes later enabled “cloud neutrality” by reducing lock-in to specific cloud APIs. Hyperscalers learned that correctness matters (e.g., memory corruption), so they moved from “velcroing” cheap parts to engineered systems and internal tooling.
Notable examples
Sun innovations during 2001–2005: ZFS, DTrace, and service management facility; dot-com bust timing (telco buildout collapse); AWS price-cut dynamics at re:Invent; Kubernetes and CNCF; hyperscaler server design lessons; Oxide’s “blind-mate” power/networking and building its own switch using Intel Tofino.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOWelcome to Brian Cantrell
1:14 to 1:24
Brian Cantrell discusses his career and early experiences in the tech industry.
“This podcast episode is presented by Statsig, the unified platform for flags, analytics, experiments, and more.”
The 1990s Tech Landscape
1:24 to 2:36
Exploring the vibe and developments in technology during the 1990s.
“I'd love to jump back in time a lot, back in the 1990s, because you're someone who's been around the block.”
The Rise of Sun Microsystems
2:36 to 5:13
Brian shares how Sun was positioned during the dot-com boom.
“So at Sun, those next couple of years, I mean, I got very lucky, really, because Sun was in the right place at the right time with the right technology, which, you know, sometimes you only appreciate in hindsight.”
Experiences During the Dot-Com Boom
5:13 to 7:36
A deep dive into the frenetic energy and challenges during the tech boom.
“So that gives you an idea of kind of how it was being used.”
The Transition to Bust
7:36 to 10:07
Brian recounts the abrupt shift from the boom to the bust of the dot-com era.
“When you're in frothy times, that boom will go on longer than you think possible.”
Lessons from the Dot-Com Era
12:43 to 13:08
Understanding the valuable insights gained from the dot-com boom and bust.
“how those believing the internet will be the future back in 2001 were right.”
Impact on Silicon Valley
13:08 to 14:00
Discussing the aftermath of the bust on the tech community in Silicon Valley.
“in November of 2000 in particular, there were zero orders from telecoms at Sun.”
Silicon Valley's Resilience During the Bust
14:00 to 15:50
Learn how the tech community in Silicon Valley adapted and innovated during economic downturns.
“So I would say that like lots of people left and you had like the statistic of, you know, the U-Hauls were 10 to 1 out of the Bay Area.”
Innovations Born from Adversity
15:50 to 17:40
Discover key technological advancements that emerged from the tech industry during the post-bust period.
“So all of those things happened from 2001 to, say, 2005.”
The Rise of Open Source and Linux
17:40 to 19:40
Understand the transition from proprietary systems to open source and the impact on server technologies.
“But like, nope, apparently in high tech, we've got to be like on or off.”
Show all 39 chapters
The Shift to Cloud Computing
19:40 to 21:20
Learn how cloud computing transformed the tech landscape and altered server management.
“And we, you know, because we ran the operating system on that was in Solaris on both Spark and x86, we could see how fast these x86 machines were.”
Oracle's Acquisition of Sun and Its Impact
21:20 to 23:00
Explore the implications of Oracle's acquisition of Sun and the subsequent changes in the company.
“It's like that every single company, no matter if you were a website, you had your own server room.”
AWS Dominance in Cloud Markets
23:00 to 26:50
Analyze the factors behind AWS's early dominance in the cloud computing space and its business strategies.
“I'm like, no, no, you're misunderstanding.”
Kubernetes and the Era of Multi-Cloud
26:50 to 28:00
Examine how Kubernetes changed the landscape for cloud computing and introduced multi-cloud strategies.
“they would buy our software and they would run it on there and get a cloud.”
The Rise of Kubernetes and Cloud Neutrality
28:00 to 29:22
Learn about the emergence of Kubernetes and its impact on cloud service flexibility.
“because they could never be API compatible.”
Google's Decision to Open Source Kubernetes
29:22 to 31:08
Explore the motivations behind Google's choice to open source Kubernetes and its implications.
“like why would Google, like what is the business reason?”
Building Custom Hardware for Big Tech
31:08 to 34:26
Discuss how companies like Google and Facebook built their own hardware and the tools that support it.
“They, so it's kind of funny because for all of these folks, they took a somewhat similar path.”
The Shift to Building Internal Infrastructure
34:34 to 36:08
Understand why major companies shifted from buying servers to building their own infrastructure.
“And now let's get back to the conversation about the history of computing and what might be coming next.”
The Economics of On-Premises Solutions
36:08 to 37:52
Examine the financial motivations for companies to own rather than rent cloud infrastructure.
“one of the things that we had seen is that, and we felt, we earnestly believed that one cloud computing is the future of all computing, not a deep thought.”
Designing Hardware for Modern Data Centers
37:52 to 42:00
Learn about the design considerations for building efficient data center hardware.
“be folks that were born on the public cloud that would outgrow the economics of the public cloud and want to go on-prem.”
The Importance of Building Their Own Switch
42:00 to 43:14
Learn why creating their own networking switch was critical for operational efficiency.
“So we also did, in addition to doing our own compute slide, we did our own switch.”
Programmable Networking and Industry Challenges
43:14 to 44:26
Discover the challenges of networking components and the significance of programmability.
“And so you're saying that like buying – because a switch to me sounds like a somewhat simple component.”
The Complexity of Building a Computer
44:26 to 45:46
Understand the intricate process of computer design beyond just assembling parts.
“We were not going to get a bunch of the things that we needed in building that switch.”
The Challenge of Innovative Engineering
45:46 to 48:12
Explore the challenges and philosophy behind innovative hardware engineering at Oxide.
“These boards, by the way, ultimately this is all analog.”
Transparent Compensation at Oxide
48:12 to 51:16
Learn how Oxide's transparent compensation model attracted top talent.
“I would say that in computer design in particular, the high-speed designs are so hard.”
Developing Custom Software for Hardware
51:16 to 56:00
Discover the journey of creating custom software, including an operating system for their hardware.
“Yeah, because your composition was both the same and you also put the number specifically.”
The Complexity of Distributed System Updates
56:00 to 1:01:56
Learn about the challenges of updating distributed systems at Oxide and the importance of robust design.
“but it's, and it's also got, you've got your API, you've got your CLI and your provisioning instances.”
Integrating AI Tools in Software Development
1:01:56 to 1:10:01
Discover how AI tools are being used in software engineering and the challenges they cannot solve.
“And so the way Dave and team did this is, you know, with that foundation and then very slowly lighting up different aspects of the system and making it more and more automatic over time.”
The Role of AI in Startups
1:10:01 to 1:13:26
Discussion on the integration of AI tools in startups and their limitations.
“No, that's actually like, that's something to go check.”
Perceptions of Programming Jobs
1:13:27 to 1:16:37
Exploration of how LLMs affect programming roles and perceptions in the industry.
“And I would say we're seeing a wide variety of experimentation.”
Diversity in Team Composition
1:16:38 to 1:20:59
Insights on team dynamics and the value of diverse experiences within a company.
“And I saw you have those machines in there.”
Remote Work in Hardware Development
1:21:00 to 1:24:00
Discussion on remote work challenges and strategies in hardware engineering.
“Larry Ellison makes every hiring decision at Oracle.”
Oxide's Growth and Culture
1:24:00 to 1:25:16
Discussion about Oxide's compensation culture and maintaining company values during growth.
“So we are all, in fact, we've got a bunch of folks out there this week for Benchmark Electronics in Rochester.”
Hiring Discipline and Company Culture
1:25:16 to 1:27:26
Exploration of Oxide's disciplined hiring process and its impact on company culture.
“And it is very important to me that we continue to have absolute discipline in the way we hire.”
Impact of AI on Engineering
1:27:26 to 1:28:52
Insights on how AI tools are revolutionizing engineering and the fears surrounding job displacement.
“us to do, but we need to do it in a way that protects and preserves what got us here.”
Opportunities in a Changing Landscape
1:28:52 to 1:32:45
Discussion on the importance of new company formation and self-improvement for engineers.
“one of the things I am a little bit worried about is a little bit of despair from younger software engineers in particular who are like, what's the point?”
Learning and Using AI as a Tool
1:32:45 to 1:33:52
Advice on leveraging AI tools for learning and self-improvement in engineering.
“There is lots that you don't understand.”
Recommended Reads for Engineers
1:33:52 to 1:37:44
Brian shares book recommendations that are essential for engineers to read.
“or two books that you would recommend to folks and why oh so many good books you know my my My, uh, my, I've got a, I've got a 21 year old, an 18 year old and a 13 year old.”
Reflections on Oxide and AI Tools
1:38:00 to 1:38:59
Explore insights on Oxide's unique approach and the limitations of AI in hardware engineering.
“We went from, from the nineties all the way to the future.”
Transcript
Automatic transcript. May contain errors.0:00Can you tell us about the dot-com boom? We did much more technically interesting work in the bust than we did in the boom. There's a degree to which innovation requires some level of desperation, that good economic times are kind of hard to summon that desperation. How have AI tools changed how you're working at Oxite? Certainly we're using Cloud Code a bunch and people are doing that, but for a lot of the work that we're doing, it is helpful as maybe a polishing tool, but less as at the epicenter of its creation.
0:26The Pragmatic Engineer Host:Can you tell me what it actually means to design or build a computer? Oh, it's very involved. Yeah, it's very involved. So first of all, how have servers and cloud infrastructure evolved since the late 1990s? And what is next? Brian Cantrell was a distinguished engineer at Sun Microsystems during the dot-com boom and dot-com bust, built a small competitor to AWS called Joyent, and is now the co-founder at Oxide. Today, we go into the history of servers and the cloud from the late 1990s to today. The challenges of building hardware like the Oxide computer from scratch. how the Oxide team uses AI, and why they find it practically useless for hardware during challenges, why Oxide builds everything as open source, and how they manage to work remotely as a hardware startup, and many more.
1:07The Pragmatic Engineer Host:If you'd like to understand more about how the cloud works and learn how nimble hardware plus software startup operates, this episode is for you. This podcast episode is presented by Statsig, the unified platform for flags, analytics, experiments, and more. Check out the show notes below to learn more about them and our other season sponsor. So, Brian, welcome to the podcast. Oh, it's great to be with you. Thanks for having me. I'd love to jump back in time a lot, back in the 1990s, because you're someone who's been around the block. And back then, you worked at some interesting companies, including at Sun.
1:36The Pragmatic Engineer Host:And if you could give us listeners and viewers a sense of what was it like in the 90s in terms of software, servers, what was the vibe like? Yeah, it was an interesting inflection point because I was interviewing in 1995. I started in 1996. So I would say that the internet, I mean, HTTP had been developed in like 93, 94. We had kind of the first web browsers, but it was still very, very, very new. And the internet was just kind of primed for takeoff. Java had been, Java had come out in maybe 1995. Java had kind of taken off immediately. So there was a lot of really exciting energy, but it was nowhere near what would become a couple years, Even a couple years later, it became very frothy, of course.
2:24And it was exciting. It was very clear to me. I went to school, actually, in the East Coast. But just coming out here to Silicon Valley, the energy was extraordinary and really knew that I wanted to come out here for my career. So at Sun, those next couple of years, I mean, I got very lucky, really, because Sun was in the right place at the right time with the right technology, which, you know, sometimes you only appreciate in hindsight. because it was so explosive. And if you wanted to build a website as part of that dot-com boom, you were buying Sun servers, you were buying Cisco switches.
3:00The Pragmatic Engineer Host:Now, why was this the case? Because, again, just taking myself back, just being a bit naive, I would assume that, let's say, I'm in the 1995, I want to build a website. Could I not just use the PC and split up a server? Did it not work like that, or how did it work? I mean, a PC, like maybe, but you didn't really have an operating system, right? Because Linux is very, very new. Linux is not. I'll back down. Oh, yeah, definitely. Linux is, you know, would be like Haiku today, which is an operating system you haven't heard of for a reason. It's kind of like a hobbyist operating system. You know what I mean?
3:33You'd be like, what? No, you wouldn't. And then you kind of had the BSDs. The free BSD was certainly out there. Also still very much under the shadow, though, of this lawsuit from AT &T. So the Unisys, there's not really open source operating system options. There was the, actually, this is kind of funny because, so where was the GNU option? It was going to be the H.E.R.D. operating system. So H.E.R.D. was kind of like the Duke Nukem Forever of its time. It was the operating system that was constantly coming kind of next year and next year and next year. And it was going to be microkernel based.
4:07And so you know that it's kind of amazing, but you really couldn't do it on PCs because of the lack of system software. And actually part of my attraction to Sun was I had used Solaris on Spark, but I knew Solaris existed on x86, but I never used it. So I was excited to use Solaris on x86.
4:26The Pragmatic Engineer Host:And so what did Sun build? You mentioned Solaris. That was the operating system, right? Solaris was the operating system. We built servers. So we built Spark-based servers. We built desktop machines. So we, Sun was a computer company. It was a systems company. So we built desktop machines, built some ill-advised laptops, so basically desktop machines, workstations. But then at that time in the 90s, what was really exploding were everything from those kind of workgroup style servers up to really getting bigger and bigger servers up to very large machines. Machines that are as physically the same size as what Oxide makes today.
4:58And I remember vividly in what would have been like 97, 98, maybe, Greg Papadopoulos, then the CTO of Sun, giving it to the entire company saying, here are the top three applications for Sun Microsystems, databases, databases, and databases. So that gives you an idea of kind of how it was being used. And this is, again, as that kind of in that knee up of that dot-com build out where if you, again, if you wanted to really build a web presence, you were going to use Java. You were going to do it on Solaris. You were going to do it on Sun servers.
5:34The Pragmatic Engineer Host:And you were going to, and it was kind of, it was a wild time for sure. And can you tell us about the dot-com boom? Because, you know, right now I know AI is pretty exciting and it feels like we're in a special time. But what was it like, especially working on something? It sounds like. It was the epicenter of it. And you know what was funny is I did, it was frenetic in a way that was not always positive. So one of the things that is just a point of fact and one can take from what one will, I did, we did much more technically interesting work in the bust than we did in the boom. Because I think that when you're in boom times, you know, everyone kind of like secretly believes that this is because of me.
6:15Like it is because of the thing that I am working on. If I, you know, I once had, you know, one of the early technologists behind Java once told me with a straight face, every server that Sun sells, they sell because of Java. And I'm like, you know what? You know what's most amazing? I, you believe that is actually the more interesting fact that, I mean, it is like obviously false, especially with, you know, databases, databases, databases being the top three applications. But that kind of reflects the zeitgeist of the time, that everyone believes that this is, you know, if I work on the microprocessor, it's because of the microprocessor is perfect.
6:49If I work on the operating system, it's because, oh, this is the operating system that people are buying the machine for. And it's like that doesn't really lend itself to really, to real innovation, I think. I think there's a degree to which like innovation requires some level of desperation that good economic times are – it's kind of hard to summon that desperation sometimes. So I think that during the boom, it was – and it was just – it was frothy and it felt like there was a period of time where I'm like this obviously can't go on forever. And, you know, The Economist is having these very like gloomy covers about how this is all going to end and it's going to be an apocalypse, which I believed.
7:24And then I just stopped believing it. I'm like, well, maybe The Economist is right. It just went on longer. And one of my early life lessons from the boom and bust is these things go on longer than you think possible. But when they switch. In terms of the growth? In terms of the boom. When you're in frothy times, that boom will go on longer than you think possible. And when it switches, it will collapse faster than you can fathom.
7:48The Pragmatic Engineer Host:In the boom, do I understand correctly, that customers were just like wanting to buy your servers. They were flying off the shelves. All these companies were. Everything. And on a day-to-day work, what did it mean for you? So I'll tell you, like, day-to-day, it meant, first of all, it meant that traffic was terrible. That the, you know, there is, you couldn't get housing. You couldn't get, you know, everything was in short supply. You couldn't, customers are, you know, they are buying, we had a customer that, you know, was going to buy 19 ,000 servers, which is obviously a very big number. And these were these massive big servers, right?
8:20Yeah, well, in that case, those were actually one-use servers to build out a broadband initiative. That actually was a company called Enron. You know, I remember vividly we were at a dinner here in the city at a restaurant called Aqua, which is a very kind of fancy restaurant, long since out of business. And I don't think Aqua survived the bust. And we were at Aqua with a bank who was a customer of Sun's. And they were spending a galactic amount of money every year with Sun. And we were at a dinner, and I just remember – I mean, it was the kind of like 19th century Gilded Age kind of dinner. People are ordering, you know, nine courses.
9:00What I remember is at the end of that having Chateau de Quim, which is a Sauternes. So I don't know very much – I don't know very little about wine. I know nothing about Sauternes. What I did know is there was someone who knew wine, and it's like we are going to all drink the 1952 Chateau de Quim Sauternes.
9:16The Pragmatic Engineer Host:Oh, wow. Which is – and I remember being like – I'm like – I'm not much of a drinker, but I was like too drunk at that point to really appreciate it. So I have had this Sauternes that, you know, that enophiles kind of live their life to drink. And I'm sad to inform you that there's one less bottle of this precious vintage because it was poured down the gullet of a 20-something dot commer who really had – And I just remember being back in my apartment, being literally drunk on Chateau de Quim, thinking about in Petrero Hill. And I remember thinking to myself, this can't last. This is not sustainable.
9:53And I swear the dot-com boom turned to a bust like that night. That is September of 2000. So the pets.com had kind of busted out and a bunch of NASDAQ had busted out early in 2000. The traffic got lighter early in 2000. Anyone who was here would be like, the absolute spookiest thing is it went from like gridlock to like COVID-like traffic in the span of like a month. Without COVID happening. Without COVID happening. With only the NASDAQ collapsing. And you're like, okay, that's very odd. And then 2000 kind of muddled along. And then that dinner was in September of 2000. And what really stopped was the telco buildout.
10:35So there was a lot of telco buildup because people are like, the internet is the future.
10:39The Pragmatic Engineer Host:And telco buildup meaning the towers, the server. The servers, the infrastructure for, and then all of the concomitant, the fiber, like JDS Uniface was a huge company. You had these companies that were, you know, Global Crossing and MCI WorldCom and all these companies were explosive. And everyone believed that the internet is the future. And this is like an important thing. And they were right, right? They were right. Brian just said how an important lesson with the dot-com boom was that people who believe the interim will be the future, they were right. Today, we're in a similar stage with AI.
11:11The Pragmatic Engineer Host:It's pretty likely that AI will be part of the software stack in the future, even if timing is harder to predict. The latest shift is how AI agents are becoming a lot more commonly used for development. And this is a great time to talk about our seasoned sponsor, Linear, and how they think about collaborating with agents. Linear has taken an interesting approach here. Instead of building one proprietary AI assistant and locking you into it, they built an open API and SDK that lets any agent plug into your issue tracker. That means you don't need to wait for a linear to build the features that you need.
11:40The Pragmatic Engineer Host:You can connect the best coding agents on the market like Cursor, GitHub, Copilot, OpenAI Codex, and Devin. Or you can build your own agent for your team-specific workflow. It's a fundamentally different approach from most issue tracking and project management tools on the market. You get optionality. And the experience is surprisingly natural. You assign an issue to an agent the same way you assign it to a teammate. Or you can simply mention the agent in an issued thread. Cursor then can pick up a bug, understand the context from the issue, open a PR. Codes can explore a fix while you're focused on something else.
12:10The Pragmatic Engineer Host:Centric and root cause analysis when something breaks. It's pretty powerful what you can get these agents to do. And here's what I like. You, the human, stays the accountable owner. The agent works for you, not instead of you. You review the work, you decide when it's good, and when it ships. If agents are going to be a part of the tool set of building software, and it feels to me they increasingly are, you'll want a system that's actually designed for them. Linear is a system like this. To learn more, head to linear.app slash agents. And with this, let's get back to the point where Brian was saying how those believing the internet will be the future back in 2001 were right.
12:47This is the other thing. It's like they're right. And so like a very famous impact creator from the dot-com boom is Webvan, right? Webvan was delivering groceries, which many people today are going to get the groceries delivered, right? Right, right, and the Instacart. It's like they weren't wrong, but their timing was off, and they lost track of the underlying economics completely. And so when it busted out, so in the fall of 2000, in November of 2000 in particular, there were zero orders from telecoms at Sun. Like it went to zero. Wow. And you're kind of used to kind of ups and downs, but that's like just like off a cliff.
13:26And from that point, you know, going 2000 and then 2001. And it was then very, very grim. I would say that the thing that happened through the bust and layoff after layoff after layoff, and because companies had kind of built themselves and geared themselves around these fat times lasting forever, and now they were gone. And expectations, as frothy as expectations were during the boom, they were that much negative in the bust. People were like, everything is, it's the end of days.
13:55The Pragmatic Engineer Host:And were you a software engineer back then? Yeah, a software engineer. Yep. And then, so as a software engineer, like both you and also thinking about your colleagues back at the time or friends, how did it impact you? Were you kind of just chugging along? So I would say that like lots of people left and you had like the statistic of, you know, the U-Hauls were 10 to 1 out of the Bay Area. So you, they moved away. They moved away. And the thing that I noticed is that the people that had moved out to Silicon Valley because they really had an interest in the technology, all were there, all stayed.
14:29And we're not adversely affected, honestly. I mean, yes, every one of us, if you had equity in your company, which, of course, you all did, like you tried not to overthink it, right? You just tried to like – you tried to remind yourself like I never had it to begin with. So like it's hard to – but it's definitely gone. Sun lost 98 % of its value. Um, so it's like definitely gone. And, you know, there was some thinking, and I think it also like a boom can get you to care about things that you actually don't care about. And a boom can get you to, because in a boom, everyone is so financially driven that it's hard not to become financially driven, but it's like, that's actually not why I got into this.
15:07And so during the bust, I'm, you know, definitely able to put, you know, put a meal on my table and a roof over my head. but the it was really a reminder about what's important and again because we did do better technical work in the bus than we did in the boom and I think it's because in the bus it's like okay now we really we have to focus we have fewer resources that the fewer resources actually force more creativity so you know all of the things that we did certainly speaking at Sun and system software so So ZFS and D-Trace and the service management facility, all these things that were really revolutionary for the operating system all happened in the same kind of post-bust period of time.
15:51So all of those things happened from 2001 to, say, 2005. And so what were these specific innovations? So I'd gone to work at Sun to work with Jeff Bonwick. And as long as I'd known Jeff from the mid-'90s, Jeff had wanted to rethink file systems. And now finally in the early 2000s, he and Matt Ahrens were able to really go take a clean sheet of paper from the file system, and that's ZFS. I had a chip on my shoulder about the way we understand debug systems, by the way we observe systems. So I, along with two other colleagues, did D-Trace, which allowed us to dynamically instrument systems. and you can kind of go down the line and there were a bunch of things like this where we, and I don't know that all of this is related to the bust, it's just that the timing lined up such that it was all happening during the bust and what we ended up with was a whole bunch of interesting technology coming together actually in a single version of the operating system and then very, I mean, fortunate for us and I do think this is a bit of a consequence of the bust because Sun was definitely open to new approaches, we open sourced all the operating system.
16:59So that happened in 2005. And that was very important to give these kind of technologies eternal life.
17:03The Pragmatic Engineer Host:But I think, you know, we can never predict the future. But to me, it is pretty positive in the sense that even in the bus, hearing the stories that innovation did not stop. Sure, you know, sounds like it was probably harder to get jobs and there might have been fewer of them. But, you know, industry kept innovating. And what you said, I didn't expect to hear that it was a bit easier to innovate. It's just less manic. We were able to focus more. And so not that now, I mean, not that one should necessarily pine for a bust because busts are brutal, but there is a clarity that you get, too. So, I mean, ideally, you would like to have just like, can we just be like normal economically?
17:41But like, nope, apparently in high tech, we've got to be like on or off.
17:44The Pragmatic Engineer Host:So bust aside, in the early 2000s, leading up to this internet boom, the way to, you know, most companies went about buying Sun servers with Solaris installed. Everything was hardware and software came together. It was beautiful. It worked well. Again, I heard from folks who did it. What happened then? Because when I got into SEC in 2000, I did not hear about Solaris. And that was not how it did. No, no, no. Right, that's right. What was the shift? So the shift was, first of all, open source, right? So then, so, you know, we said in the mid-90s, Linux was kind of still very much the hobby project.
18:15Not so by the 2000s, right? It grew up. It grew up, absolutely. And it grew up because you had a bunch of companies that really backed up the truck. and, you know, the things that at first, IBM and SGI, data general, some other companies, those companies were very important because they decided to contribute their technologies. Like XFS, right? XFS, many people still use today on Linux. That's from SGI. XFS was at SGI on IREX. That was happening in kind of the late 90s. And then in the 2000s, I mean, Google was always built on Linux, right? And so you had kind of the companies that became that next boom were all built on open source and indeed needed to be built on open source.
18:54They economically relied on open source to be able to build what they build. So then it became much more practical to certainly run Linux and I think the other BSDs or we open sourced Solaris. So there were a lot of options that were now available. So that shifted. I think the other thing that shifted is that, I mean, Spark bluntly lost to x86. And Sun for - And Spark is a Harvard architecture. Spark is a microprocessor. Yeah. And there was – because there was a time in the 90s when if you wanted the fastest microprocessor, it was a RISC microprocessor. It was from – it was a Spark microprocessor or it was MIPS or it was Alpha.
19:37And x86 was a commodity but was – and obviously available with a personal computer but was not faster than those RISC microprocessors. That shifted. That shifted in the late 90s. And we, you know, because we ran the operating system on that was in Solaris on both Spark and x86, we could see how fast these x86 machines were. And could see, frankly, how, like, you know, you talk to the microelectronics folks. They really did not. They kind of dismissed x86 and dismissed Intel. And you shouldn't do that. And in particular, Intel was very focused and architected the way around what was called the memory wall.
20:15and they were able to, in part because they used speculative execution, they were able to actually make these microprocessors that became much faster than the RISC microprocessors. So by the time, say, you are in 2004, 2005, if you want a leading edge microprocessor, it's X86. So that was a big and important shift. So by the time you're coming up, it's like, okay, yeah, if I want this, I'll just like, I don't know, get like a Dell box or a Supermicro box and then I'll put Linux on it or maybe FreeBSD and away I go. Then the next kind of big and important shift that happened started in 2006, you could argue, with S3.
20:55But then especially in those next kind of 7, 8, 9 with the introduction of EC2 and now you have the cloud that starts to come into play. And now people are like, well, why would I even screw around with the server at all? I mean, it was so great to be able to just spin up infrastructure.
21:13The Pragmatic Engineer Host:Yeah, I remember one of my early companies, mid-2000s, we had a server room. We had server administrators. The server room was always hot. And this was a small company, mind you. This was not a big one. Every company needed to do that. It's kind of amazing to think. It's like that every single company, no matter if you were a website, you had your own server room. And if you were a dev, you wanted to be friends with the server admin because when you wanted to deploy your stuff, they could do stuff for you. That's it for you. That's it totally. And so I think that cloud computing was really important.
Read the full transcript
21:45This is not a deep thought, that elastic infrastructure was really important, but the ability to have API-driven infrastructure. And so for me personally, so I was at Sun, and then in 2006, I started a storage group inside of Sun, which was great, really successful group, but so successful that it actually attracted Oracle as a customer for the first time in a long time. This is like a little bit of residual shame that I have that like, did I attract the marine apex predator that ate the company?
22:17The Pragmatic Engineer Host:Because Oracle later acquired Sun, right? Oracle bought Sun and that closed in early 2010. I left shortly thereafter because I could see what Oracle was. Well, I never heard the story of your potential role here. Right. So I, yeah, and Oracle, and I gave some, maybe a year later, I gave a talk in 2011 with some rather unvarnished opinions about Oracle and Larry Ellison. In particular, I caution people about anthropomorphizing Larry Ellison. You have to treat Larry Ellison as a machine, like a lawnmower. You stick your hand in the lawnmower, it'll chop it off. All right, so I'm giving this talk in 2011.
22:57Again, this is after I've left what was then Oracle. And, you know, like I was just saying things that I felt were obvious, but people, you know, the audience is kind of gasping and, you know, it's like, and people are coming up to you after the talk, like, do you think there's going to be like, there's going to be retribution from Oracle? I'm like, no, no, you're misunderstanding. Like there's no, the lawnmower's not angry at you. It's a machine. It doesn't have the mirror neurons to be, I would almost like it would almost show me that I'm wrong for Oracle to resent what I'm saying about them.
23:27Anyway, so but all the videos for that conference go up and my video does not go up. Oh, right. OK. And so my colleagues were like, this is an Oracle conspiracy. I'm like, this is not an Oracle conspiracy, which it wasn't. It wasn't orchestrated by Oracle. But what I did, what I underestimated was the fear of the conference organizers. So they themselves were terrified of offending Oracle.
23:48The Pragmatic Engineer Host:Yes, even though it probably would have been fine. No, the talk did finally go up. Before the talk starts, there is a disclaimer. The views in this talk do not represent the views of the USNX Association. And you're like, all right, I get it. Like, I've never seen this disclaimer before, but fine. Then during the talk, you know, the format of the talk is you've got a slide, and then you've got like a little blank strip, and then you've got this talking head in the little right corner. So there's like kind of this dead space above the speaker. They took this disclaimer, and they re-justified it, And they put it above my head the entire time I'm speaking.
24:21So if you – and maybe in this regard they were prescient because to this day, if Ellison is mentioned on Hacker News or Oracle is mentioned on Hacker News, someone will immediately cite minute 33 of this talk, which is when I go on this kind of Oracle – again, I don't view it as a rant. I view it as just like me describing what is obviously true that we all know. But anyway, I had left. I left Oracle after they bought Sony.
24:48The Pragmatic Engineer Host:So we're now around like 2000 or so. Cloud has taken off. x86 architecture is everywhere. Linux is now winning both for small time servers, but also on the cloud. And then what happens? This was an interesting time when Google started to figure out that, hey, they could do something interesting on their cloud, right? Yeah, that's right. So this is still a little bit before that. So this is in kind of from, I would say from 2010 to about 2014 is when, is a period of relentless execution from AWS. AWS is executing so extremely well. There are not really other public cloud options. There's like kind of Azure's kind of drifting out there.
25:25The Pragmatic Engineer Host:I think people forget that, you know, like GCP on paper has been around from 2009, but up to like 2014, it was like, it was almost like a joke. It was a joke. I would say before it was like, it existed, but it was a joke. And in particular, at every single reInvent, Amazon would announce a new price cut. And if you were a competitor to AWS, you were like dreading reInvent because here comes another price cut. If you are a partner of AWS, you're dreading reInvent because here comes the announcement of a new service that competes with what you're making. I think people who have not been around have forgotten, but it really has happened because it's not been the norm the last, like, let's say, five, ten years or so.
26:01Well, and in particular, they did a couple things that were just like, man, you got to tip your hat to just. I mean, Jeff Bezos is the apex predator of capitalism. Like Larry Ellison may be the lawnmower, but Bezos is ultimately the apex predator because the thing that was so impressive is they were able to give people the idea that this was a terrible business. So in particular, they did not break out their financials. So everyone's like, oh, my God, what an awful business. Like they're cutting the price every year. Like you do not want to – like this is a classic red ocean. It's bloody. You don't want to compete.
26:33And so we were at Joint. We were actually competing head-to-head with AWS.
26:37The Pragmatic Engineer Host:So you were offering a public cloud. So we have a public cloud and then unlike AWS, taking the software that we'd use to run the public cloud and making it available for people that wanted to run a cloud on-prem on their own hardware. So people that would buy Dell or HP or Supermicro, they would buy our software and they would run it on there and get a cloud. So we ran a public cloud and we knew what the economics of a public cloud were, namely pretty good. Margins were good. And so what we knew that Amazon wasn't volunteering, but what we knew is that AWS S3 was underwriting a war on big box retail.
27:14S3 was paying for your prime shipping. It was a genius move.
27:18The Pragmatic Engineer Host:Also some insider information that you had because you did your own thing. Well, we don't know that the margins are very good. And then, of course, I mean, you will be unsurprised to learn that several of Joyant's most prominent customers were retailers. Retailers, this was not lost. Retailers were like, gee, I wonder what's happening. Retailers are like, if you think I'm going to take my dollars and spend them on AWS, so AWS, can I, so Amazon can go to war with me? Like, no thank you. There was a period of time when it felt like in order to be in the cloud, you have to implement every AWS API.
27:50So there was this idea that you had to be API compatible with EC2. There was a company called Eucalyptus that tried to do this. It was just a disaster. and part of the reason it was thought that GCP and Azure could never compete with AWS because they could never be API compatible. And so I am convinced that the, because what changes? What changes in like 2015? What starts in 2015? Kubernetes. And I think that part of that initial attraction to Kubernetes is that people wanted to get some optionality around their cloud. And they felt locked into AWS. They're like, I'm not using all this stuff. I'm not using Elastic Beanstalk.
28:25I'm not using Greengrass. I'm not using kind of these more, I'm not using Redshift. What I actually want is this kind of basic infrastructure and Kubernetes now gives me this layer upon which I can deploy and get some sort of true cloud neutrality. So multi-cloud didn't really exist, I would say, before Kubernetes. And I think a lot of that, especially early momentum behind Kubernetes is around this idea of like, I need to get some optionality in here. I want to actually be able to go to GCP. So I think, you know, I don't, I think it's giving Google slightly too much credit, but only slightly too much credit to say that is Masterstroke.
29:00The Pragmatic Engineer Host:On the podcast, I had Kat Cosgrove, who's released a project manager on Kubernetes. And, you know, she's been in the project for a long time. And I asked her, she's not, she was never a Google employee, but I asked her, why do you think Google open source Kubernetes, which, you know, they have Borg, which is amazing. And they kind of built on as a better version for external. And they just released it just like that. They put a lot of work in it. And to me, it didn't really compute, like why would Google, like what is the business reason? And she told me that she thought, again, speculation from the outside, that she thought that they probably thought that it would help Google Cloud.
29:33The Pragmatic Engineer Host:That's right. To have a container which is now portable and now you can give the promise that if you run this on Azure, especially AWS, you could come over. So it kind of makes sense. Is this your thinking? Yeah, absolutely. But I think that is definitely the argument that Kubernetes proponents would make inside of Google. In terms of like why they did it, nobody prevented it. You know what I mean? It's like they kind of open sourced it. Google was a pretty cool place, but in the sense that it was very bottoms up, as I understand back then still. And then I think part of their, you know, it was Craig McCaukey who really pushed for the CNCF, the formation of the CNCF around Kubernetes to give it kind of a foundation home.
30:11I do remember one conversation with Craig and I were talking early as he's contemplating the CNCF. And he's like, wait, I think this is going to allow Kubernetes to get the marketing dollars that it needs. So I'm like, don't you work for the most profitable company on earth? Like, do you really? Isn't it just like gushing cash over there? And you can't get like, you know, a couple million bucks for marketing for this thing. But no, apparently you can't. So I think that the argument that people are making internally was about we should be encouraging cloud neutrality because we are the ones that have something to win.
30:42And they're right. And they did. And GCP is now not an afterthought. GCP is very important. It's a very big business. And I think that they've got – is Kubernetes to thank solely? No, but I think it's played an important role for sure.
30:54The Pragmatic Engineer Host:And where are we today in terms of the hardware and the software stack running, specifically thinking of these big clouds, what's happening inside the likes of meta, these giants? As I understand, you know, they're no longer just like, you know, ordering servers from Dell or wherever. Never were. Never were. Never were. Never were. What did they do? They, so it's kind of funny because for all of these folks, they took a somewhat similar path. They never were because in Google's earliest days, they were assembling machines from fries you know rip fries fries being a local electronics shop that has long since disappeared but they were kind of famously velcroing machines together and finding so so they bought like the processor the that's right the different car networking switch whatever and they had this idea that like it doesn't matter what junk we run on because you know our our software is going to run as a distributed system it actually doesn't matter we don't need ecc protected memory because it doesn't matter if your dims fail and so i think they learned well it does matter a little bit if If your DIMMs have rampant data corruption, like DIMMs failing, that's actually not a problem.
31:54DIMMs, your memory returning the wrong thing, like that is a problem. You can actually like, you turn that like next thing you know, like your software inserts that into a row into a database. And like, yeah, now you got it. That is correctness is a problem. Yeah, correctness is a problem. It's like, okay, overshot the mark. So by the time they're like, okay, we're not going to Velcro machines together. But by that point in time, you know, the business was established enough that they actually did – they built the machines that were fit for scale. So they have a great book that was written in the kind of the mid-2000s, The Warehouse Size Computer, where they talk about all the things they did at DC bus bar, really thinking about power across the entire DC.
32:34So they kind of – they went from being kind of too cheap for kind of Dell or even Supermicro to then being much better engineered than those systems ever were. So they were never really meaningful customers. And ditto for Facebook, Meta. They were never really meaningful. I mean they kicked them out very early and did their own stuff.
32:54The Pragmatic Engineer Host:Brian just talked about how Facebook built their own servers because off-the-shelf solutions didn't work at their scale. And what's interesting is that companies like Meta and Google didn't just build better hardware. they also built incredible internal tools. Tools for safe deployments, feature flagging, experimentation, debugging, analytics, the whole stack that lets teams shift fast and with confidence. Most companies never get access to this level of infrastructure. You either build it yourself, which takes years and large engineering teams, or you make with scattered tools that don't talk to each other.
33:23The Pragmatic Engineer Host:That's exactly where StatSec comes in. StatSec is our presenting partner for the season and they give every engineering team access to the kind of tooling that only the biggest tech companies used to have internally. At its core, Static is a toolkit for safer deployments and experimentation. You ship a new feature to 10 % of users behind the feature gate. You validate that it behaves correctly, watch the metrics, and expand to remaining 90 % only when you're confident. And if something goes wrong, you can turn it off instantly, long before it affects everyone. And safe deployments require visibility.
33:54The Pragmatic Engineer Host:Static includes analytics, both product analytics and infrastructure analytics, so you can actually see what your code is doing in production, errors, performance changes, funnels, user behavior. Because you cannot ship safely if you can't see what's happening. Companies like Microsoft and Notion run hundreds of experiments per quarter where Statsig, velocity that used to require entire platform teams to build and maintain. This used to be infrastructure available to maybe 10 or 15 tech giants. Now startups and mid-sized teams use Statsig to ship quickly without breaking things. If you want to give your engineering team world-class tooling from day one, go to Statsig.com slash pragmatic.
34:28The Pragmatic Engineer Host:There's a generous freeze here, a$50 ,000 starter program, and affordable enterprise plans. And now let's get back to the conversation about the history of computing and what might be coming next. And this was independent. So like both Google and Meta both came to the conclusion of like we should just build our own stuff. And Microsoft and Amazon all came to the independent conclusion because the scale at which they needed to run was not at all the scale at which Supermicro and Dell and HP were geared. What they were geared to do was to run the servers in your server room where you needed to know the devs, right?
35:02Where it's like, I'm going to have a little rack. It's going to have six servers. Then maybe it's got 12 servers. Okay, maybe we grow to 24 servers. That's what they were designed to do. If you're like, no, I want to buy servers by the thousands because I've got a public cloud business. Like if you want to buy servers by the thousands, there is no product from those companies for you. And in very, very basic ways. Well, like the DC bus bar at every juncture, they've been designed to be a personal computer that you happen to be slapping many personal computers together, but they're not designed to actually run infrastructure at scale.
35:32So, and that was happening inside of effectively all the hyperscalers and joint. Meanwhile, it was bought by Samsung in 2016 joint was bought by Samsung because their cloud bill was off the charts and.
35:43The Pragmatic Engineer Host:They bought, they bought you to bring it on, bring it in a house. And, and there was not a product they could go buy. So they went to go buy a company. Wow. So you're like, wow. And then it's like, wow, that's a big AWS bill. It's like, yes, very big AWS bill. But then that was not a product that our company that was available for, you know, the next, what does the next Samsung do? It's like, well, that's one less company available to buy. So when we were contemplating the next thing in 2019, one of the things that we had seen is that, and we felt, we earnestly believed that one cloud computing is the future of all computing, not a deep thought.
36:17That elastic infrastructure, API-driven infrastructure, that is modernity, one. Two, you shouldn't be able to only rent that. You should be able to buy that, own it, run it in your own data center. Why would you want to do that? Well, you might want to do that for risk management, for security, or for economics. Because it, you know, if you're at a certain scale, you'd rather own it than rent it.
36:39The Pragmatic Engineer Host:And I think, you know, before Oxide or like in 2019 or even in like, you know, 2020, 2021, if you were like a mid-sized company, you know, like not big enough to build out your own custom cloud and build everything that the hyperscalers did you could like buy some off the shelf like hp or dell like a bunch of them i think that's what base camp did i think they posted that they they bought a bunch of bunch of these things they rented a space in a in a one of these shared or i think two different locations they put in their boxes with all the memory and then you know they kind of set it up and put it together so i guess those were the two options right yeah those are the two options and i think that you know base camp ended up being a real poster child for the economic advantage.
37:18Because, I mean, DHH, obviously outspoken, and the economic advantage was really, really, really clear. They're also at a scale which is like not the scale that we're targeting, right? The scale we're looking at is a much larger scale. And so the economic argument is actually even more compelling when you're at that larger scale. I love it when, you know, the VCs that passed on us because they felt there was no market, then would send me like the DHH blog post. It's like, why are you sending this to me? I should be sending this to you. Like, I know this. We just knew the economics of it. And we knew, couldn't predict exactly what the trends would look like, but believed that there would be folks that were born on the public cloud that would outgrow the economics of the public cloud and want to go on-prem.
38:00The Pragmatic Engineer Host:Economics aside, what does it take to build one of these things? And I saw one of these things. We'll put in a picture of it. It's like a proper, like, you know, like nine feet tall rack it's it's big it's it feels like you're putting like i don't know like 16 or 32 of those of those like the dell things they get in terms of size just to get a sense of yeah we would 32 compute sluts in there that's right and and what does it take what did it take to actually build what did you need to design in terms of hardware and then software yeah so well and we knew this too that going into the company we knew we were taking a clean sheet of paper right and so we were deliberately like no we're going to start with a problem we're not going to build it out of Dell HP Supermicro.
38:39We are going to start with a problem, and how do you best solve the problem? And as it turns out, like, there were a whole bunch, there's a lot of technical debt that had been accrued by this kind of PC ecosystem. So, I mean, you know, Kyle, where do you start? Just on the environmentals, like on power, right? The fact that you've got AC power in each of these Dell HP Supermicro.
38:57The Pragmatic Engineer Host:Yeah, so if you, like, put 16, you have, like, 16 separate AC. Times two. Oh, times two, because you have two power supplies per 1U2U chassis, two power supplies. By the way, there are two fans sitting on those power supplies, and those fans are actually what wear out. If you go to the, like, in terms of, like, the whirring fans, it's not just coming from the computer, it's coming from the power supplies, because those power supplies are dense, they're packed with stuff, so they've got to overcome a huge amount of static pressure. So, like, that's not the way anyone does it at scale. What people do at scale is you've got DC bus bar, you've got a power shelf that is much more efficient, and that rectifies from AC to DC, and then you run DC up and down and then you blind mate into that.
39:37So we knew we were going to do that.
39:38The Pragmatic Engineer Host:So that's the law of electronics engineering right there. Yeah, yeah, yeah. That was, yeah, power engineering for sure. And we knew we were going to do that. We also knew that by taking a clean sheet of paper that we would have opportunity made available to us that we weren't necessarily thinking of. And that manifested pretty early. So we blind mate into power, which is to say that when you feed a sled in, that power connector, you don't see it. It's at the back. You lock the sled in, blind mate's into power. And we had assumed that we were going to do what Facebook and Google and others had done, Amazon had done, and had networking out the front in the cold aisle.
40:11But as we were, you know, taking a clean sheet of paper, talking to some connectivity vendors, they asked us, like, why are you, wait a minute, you guys are, like, taking a clean sheet of paper. Why are you putting cabling in the front? Like, why wouldn't you also blind mate in the network and the network connection? And we were like, can you do that? They're like, oh, you can definitely do that. It's like, well, why don't the hyperscalers do that? It's like, oh, they would all tell you that if they could start over today, they would blind me the networking. And they're just too afraid to do it at this point, which is like, I mean, that was like catnip for us.
40:40You know, like they're too afraid to do it. Like, OK, we got it. And one of the very early holy God, we're going to bet the company decisions was blind mating networking. Because if blind mating networking doesn't work, you've got nothing.
40:53The Pragmatic Engineer Host:You don't have a problem. And so what is the difference in blind mating networking? It means there is no cabling in the system at all. So when you've got a sled, you are blind mating into a cabled backplane. So it's cabled in the factory. So the operator. So when the box comes in, that's why I didn't see any cables. It's inside. It runs inside. It runs down the back. And so. Versus when I look at the pictures of a data center of, let's say, Google, you see they're very neatly organized. It's like, I love organization. So it's like beautiful. But it's cables everywhere. And you can see. So you don't have that.
41:25We don't have that. And in particular, so because there's no cabling, there's also no miscabling. Right. So every computer is not actually on just one network. It actually needs to be on three. It's on a power detect, a presence detect network. It is on a service processor network. And then it's on that high-speed network that you really care about, like the actual network. In any facility, you need another network for power, environmentals, and so on. It's very easy to have miscabling. That's got to go to a different router. There's a bunch of just complexity that we eliminate because we do. And then part of that decision came out of an arguably earlier bet the company decision, which was we did our own switch.
42:04So we also did, in addition to doing our own compute slide, we did our own switch.
42:07The Pragmatic Engineer Host:And last time you told me about this in our deep dive, we did a little bit that, like at first you said we did our own switch. And I was like, yeah, okay, cool. You did your own switch. And then you told me that actually like that is a second computer to build. Can you tell me why? It's funny because when you went through Sand Hill, initially raising money for the company, nobody asked us about it. It's the Sand Hill Road where all the VCs are. Yeah, exactly. And we were definitely, so people would be like, I've got a technical question for you. And you're like, oh God, here it comes. It's the switch question.
42:34But then we'd know some other random ass questions like, all right, that's not a very good question. But nobody was asking us about the switch. And we were concerned about the switch because we'd already come to the conclusion in order to make this thing really work, we had to do our own switch. And the reason you have to do our own switch, if we didn't do our own switch, it would be a third-party integration nightmare. And we wouldn't be able to actually solve the problem that we're trying to solve, which is when this thing shows up in your data center, We want this thing to come out of the crate.
42:59We want you to wheel it up. We want you to put in power and networking and go. We do not want you to have to cable anything. It should be the level of operator involvement should be really minimal. So we'd already come to the conclusion that in order to make this thing operable and manageable, we need to do our own switch.
43:15The Pragmatic Engineer Host:And so you're saying that like buying – because a switch to me sounds like a somewhat simple component. And you're going to tell me why it's not? Oh, yeah. It's definitely not. No, but that answer is very important. If you want to go build your own switch, I encourage you to have that attitude as long as you possibly can, because otherwise you won't go do it. So what is your switch? What is a switch being obviously the networking switch? What is your networking switch do or or that made it so important for you to build it as opposed to like going to one of the many suppliers and saying, you know, let's get here.
43:43Not many suppliers. Oh, if you actually go to the actual switching silicon is coming from like it was like one and a half providers. Oh, it's all Broadcom. And so what you're actually talking about is Broadcom silicon. What we discovered is this actually interesting piece of actually Intel Silicon from a company they had bought called Barefoot. And we found Intel Tofino, which allowed us to have true programmable networking. So we use Intel Tofino. Intel later killed Tofino. So complicated relationship with Intel over this. We fortunately have procured enough Tofino to be able to – we bought ourselves the time we need to kind of design our next-gen switch.
44:19But that programmability was very, very important for us and that we were not going to get from – Broadcom is a very proprietary company. We were not going to get a bunch of the things that we needed in building that switch. We were not going to get out of Broadcom. So it ended up being very important. We were concerned. I mean, again, another one of these kind of bet the company decisions. Very, very concerned about having our own switch, integrating our own switch. And what we found is that was a win in so many dimensions, so many dimensions that we did not anticipate. And as now, you can't imagine the company without having done a switch.
44:53The Pragmatic Engineer Host:I guess sometimes do stuff and you might get some wins. Absolutely. Well, I think also like whenever you're deliberating something big like that, the fact that it is big kind of forces you to really deliberate it. And then once you commit to it, to taking that big risk, you often see unexpected dividends. It's like, well, as long as we're going to do this, as long as we are taking a clean sheet of paper, as long as we're doing our own switch, we can blind mate the networking. If we were not doing our own switch, we really couldn't blind mate the networking. We really needed to be able to own both sides of that in order to be able to do our own switch.
45:23The Pragmatic Engineer Host:A lot of us, you know, listeners, viewers are software engineers, so we don't know as much about hardware. Obviously, we know how the things work. But can you tell me a bit on what it actually means to design or build a computer? Because, you know, I'll give you the novice approach, which is obviously going to be wrong. But the novice approach is like, oh, here's a processor. Here's a few chips. Here's a main board. I'll just put it on there and I'm done. But when I was in your lab in Oxide, you told me that one of the first engineers turned out to be a radio frequency engineer. engineering you told me how this is great because of the all the fda approvals and all these things and i was like okay this is way more involved than i ever imagined yeah it's very involved how do you build so it's all i mean it would be a lot easier if we're all slower right the problem is it's very fast it's high speed so the connection to memory via now ddr5 double data memory five is ridiculously obviously high throughput is very, from a signal integrity perspective, really complicated.
46:23These boards, by the way, ultimately this is all analog. We think of it as digital and it is digital, but digital is like a lie that, that double E's allow us to tell ourselves. It is actually like you are talking about signals that are racing through a, a substrate and the, and with a PCIe or DDR5, the, all of the, so those signals are very complicated to lay out. That's complicated. the actual like how does the computer start like this computer is like it's like a it's like a triple seven right or you know i 747 used to be my favorite jet to kind of pick on but now the 747s have retired so i got to pick something out and i'm not going to pick another boeing aircraft i don't think i don't know in a380 i guess right i should pick an airbus but you think about like the okay an airbus doesn't just like come by itself like it needs an airport it needs like a runway.
47:13It needs all the infrastructure to feed it. Well, so too for a microprocessor. It doesn't, like just the power sequencing for those things is very complicated. It needs another surround that manages the power distribution network, that actually manages its power on sequencing, that manages all of its environmentals, that manage its connection to memory, to IO. So So it's just fractally complicated to the point that people often just take reference designs and iterate on them. They don't actually really innovate on this stuff
47:44The Pragmatic Engineer Host:because it takes so long. And you told me, this was really interesting last time, that as I understand, reference design means, correct me if I'm wrong, that you're an electronics engineer or hardware engineer and you want to build a new hardware and you take an existing reference that has been tested, measure it out, like it doesn't create accidental, like all sorts of radio frequency things, and then you implement that. but you told me that this is not what you did. You also told me that it's pretty hard to find electronics engineers who are used to not doing reference design, but who are brave enough to like...
48:15Who are brave, yes. I would say that in computer design in particular, the high-speed designs are so hard. People got very accustomed to taking the reference designs and it was harder to find folks that were willing to take a clean sheet of paper. And we ultimately found them. I mean, and we've got a double E team that is extraordinary.
48:38The Pragmatic Engineer Host:And double E is electronic engineering, right? Yeah, and absolutely fearless. And in part because like they're actually, but they didn't spend their careers at Dell and HPE. Like they're coming, no, they're like coming from like GE Medical where they worked on CT systems. Wow. How did that happen? How did they come to Oxide? It's not, but it feels like such a different field. I would have assumed naively that, you know, if you're building a computer, you'll try to get electronics engineers who have built computers. You would think. And that was probably our thought as well. And then we discovered that we were not getting along with those engineers very well.
49:15We didn't hire them because we were – but we were just, like, finding, like, there's a lot of friction because there wasn't a real first principles approach from those folks. And this is where you get to – especially you get to talk to folks that, like, have been at Dell for a generation. and like for any design, they're used to calling what's called the FAE, which is the field applications engineer for the voltage regulator. It's like, well, the FAE gives me the design. It's like, all right, well, how do you know that it's the right design? Well, no, he, so it's like, all right, so like let's go hire that person then, let's forget you.
49:47And we were really just, we were struggling. I was struggling to get outside of my own personal network to find the right engineers and we were kind of brainstorming, like how can we get people to see the company who wouldn't otherwise see it?
50:04The Pragmatic Engineer Host:And specifically for hardware engineers, right, we're talking about. Yeah, and just in general, but especially for double E's. Especially, yeah. Yeah, for double E's, it was feeling especially acute. One of the thing, you were kind of brainstorming as a team and one of our engineers said, you know, the values are very important to us at Oxide, which they are. And I relay Oxide's values and our principles to people outside of oxide. And they're like, that's just bullshit. And I explained that like, you know, normally I would agree with you, but it's when I get to the compensation, people, their heads turn because our compensation is transparent and uniform.
50:40And people are like, wait, what? And oh my God, I can write a blog entry on it. Like, yeah, that'd be great. I'm like, okay. And so up to that point, we had not talked about it at all. We had not talked about it publicly at all. I just came up with the idea that like compensation is just private. It's just not something you talk about with people, you know?
50:57The Pragmatic Engineer Host:And you go to a level of FYI or some of the forums. You're like anonymously asking. People are anonymously sharing. That's how you get information. That's how you get information. And so I kind of had this idea that it was that it just is not something to you. And so we wrote this blog entry in March of 2021. And it sent our hiring nonlinear. And it wasn't that people were like, oh my God, I want to work for a company where everyone's paid the same. Like that is like, that's like. Yeah, because your composition was both the same and you also put the number specifically. I think it was something like$200 ,000 back then.
51:28The Pragmatic Engineer Host:Yeah, it was a little bit less back then, but it's now a bit more than that. Yeah, now we just got another raise. So now I've lost track. It was$207 ,000, but now it's more than that. I actually don't know because the one thing is when compensation is uniform, like you don't keep total track of like, oh, like literally people were like, wait a minute. Like I got there's an error in my paycheck. I just got paid more. People are like, no, no, we got a raise. Like, when was that? Like, no, it was at the last all hands. Like, oh, you know, I did have to go to the bathroom. Like at the end of the last hour, all hands, I didn't listen to the recording.
51:54I guess I missed my raise.
51:56The Pragmatic Engineer Host:Like, yeah, you got to pay attention around here. But it was more that what drew attention was that people, engineers in particular, but just in general, people drawn to a company that would be so nuts as to do that. And ultimately, like that engineer that made the suggestion was absolutely right. It was the compensation that convinced people that we take our values really seriously, that we're a really principled company. Which is you're paying everyone the same base salary. That's right. Exactly the same. Yeah. They're making the same as you, the electronics engineer, software engineer, the whatever other role you might have.
52:32That's right. And I don't know if, you should just go ahead and say it if you want to, but many people are like, would you pay support engineers the same amount? It's like, why do people always like pick on support? They would ask.
52:45The Pragmatic Engineer Host:Exactly. Answer to that is yes. And the answer to that is, if you do that, you find supportive support engineers. And so we have got, I think we've got the best support engineers in the business. I think we've got really, really phenomenal folks in support. I heard a small company called Gumroad do this, where they paid their support staff really high, again, about the same as software engineers. And then they got support staff who were software engineers, and they could fix the code or like write tools for themselves. And you get people for whom, because, you know, there's a certain thrill in support that because you've got someone with a problem, it's technical, you get to come up, you get to be technical, you get to solve a hard problem.
53:29And then immediately you get such gratitude, you know, and like that's a rush. And if there are people that are really drawn to that, like I love helping other people. I love that feeling that I get when I resolve a problem for someone, that immediacy. So one of the things that we've heard repeatedly from several of our support engineers is my heart was always in support, but my career path was forcing me into a different career path. And I love the fact that I can get back to where my heart is.
53:59The Pragmatic Engineer Host:Yeah, that's nice because now like it. Yeah, you're not going to make more by doing something that you're not as into. I love that. So going back to where we were, which is like you build the hardware, you build this like really complicated piece and you went through electronics engineering, putting it together. Let's put a software because that's super exciting. What does it take to build software for this? Did you start from scratch? Let's talk from the low level. Did you start from scratch from operating system? Did you have to or could you use? Yeah, and there's kind of different answers at different levels of the stack.
54:31So on our service processor, we did start from scratch. We did our own de novo operating system in Rust, appropriately called Hubris, because we had the hubris to do it. The debugger, by the way, for Hubris is called Humility. Feels like appropriate for a debugger. So that was de novo.
54:46The Pragmatic Engineer Host:And this is open source, right? Open source, yeah. Open source. The entire stack is open source. Everything we've done is open source. We can go on GitHub and check it out. Go on GitHub and check it out. And yeah, I mean, we've got God's own revenue model because like you're like, well, what if somebody like can download it, run it on a different computer? It's like knock yourself out because, you know, we think the best way to run this is on the machines that we make. And those are not free. An Oxide machine is not, you know, that's not freely downloadable. But all open source. So that was for the service processor.
55:13For the host CPU, we really had it kind of at a quandary. Like what are we going to do in the host CPU? And with that to say, like on the actual, like what was then AMD Milan, now AMD Turin Silicon, we knew that we wanted to do in the product, we would do our own hypervisor and our own control plane. It was very, this is not something that you run ESX on.
55:33The Pragmatic Engineer Host:The control plane, is that controlling multiple? Like the whole, like you have a bunch of processors and memory and all that and control plane controls all that. You plug this thing on, you power it on, you put in networking. What you get is a console that looks a lot like a, what would like look like AWS of AWS looked better. I mean, it's a console. I mean, look, not disparaging AWS, but like we know that like design is not really the strong suit. We agree with that. Yeah, exactly. So it looks gorgeous, of course, but it's, and it's also got, you've got your API, you've got your CLI and your provisioning instances.
56:06Where are those instances provisioned? It's the control plane that makes those decisions. You are attaching virtual storage, those instances. Where does that storage live? It's the control plane that makes that decision. So just like with AWS, you don't need to know that stuff. That's just happening. You're using Terraform to spin up your cluster. You're running Kubernetes on it. You're knocking yourself out. So we are delivering all of the software from that lowest layer, that service processor, the operating system that's running on the host CPU, and then that distributed system. very importantly, that distributed system, which we called Omicron before the Omicron variant of COVID, which was feeling very like ill-timed for a very brief period of time.
56:48It was feeling ill-timed. And now I feel like the Omicron variant of COVID is just like, it's just forgotten. And now it's a good name again.
56:53The Pragmatic Engineer Host:So it's like, you know, we just, it was a really short live. It was a short live. Yeah. So we, so, so we, so we, you know, we, we, we lived longer than the Omicron variant of COVID. And that is our control plane. And that is a very sophisticated body of software. In addition to, because it's not enough just like provision in an instance, right? And you need to do that robustly. You need to do that via API, COI, and so on. But then you have all of the software that does that and keeps track of your instance and so on. It's very important that you can actually update that software. That whole distributed system, you need to be able to update to a new version of the software.
57:31And this gets really thorny, right? Because in a public cloud, you do that with a runbook, right? I mean, even the, you know, we don't feature it prominently, but even in GCP and AWS, yes, there's a lot of automation, but there's also humans involved. And there are humans that are taking the responsibility for actually updating software, for sure. Yeah, I mean, again -
57:55The Pragmatic Engineer Host:For the most part. Yeah, I mean, there's a lot of automation involved, But in particular, if something goes wrong in an update, you know, you've got DevOps that can hop in and figure out what's going on and get it rectified. We are shipping a distributed system across an air gap in an oxide rack that's potentially running in a secure facility. We cannot be there if it goes wrong. So we need... Especially because a lot of your customers are buying it because they want to do it themselves, right? So in many ways, the thorniest software problem for us... We had actually several thorniest problems.
58:27I couldn't pick between them because they're all thorny for different reasons. One of the very, very thorny problems was how do we ship a distributed system that we can then update? And one of the things we did that was important was like, okay... because it's very easy to paint a roadmap that is very complicated for update. You'll never ship anything. So what we needed to ship in that first product that we shipped when you were in Emeryville two years ago, we needed the minimum viable update. We needed an update where the software could be updated, even if it was painful. So what we did is we had this thing called Mupdate, which is the minimum update.
59:01And Mupdate, in particular, required the control plane to be parked. So we're going to take this rack that's running Instas, take it offline we're going to update it and then bring it back online and that was robust it was great and we got that working that's great that it's great and that you can update it that's actually not what you want in a cloud right you're like i sorry i'm like using this thing 24 7 like i actually i i want to these instances need to remain up while i update it but that gave us the platform to go build that update functionality into the software extraordinarily sophisticated and really an extraordinary body of work.
59:39And actually just recently we had at our internal meetup, the engineer who led the charge on that, Dave Pacheco, gave a presentation on looking back of two years of update. And I got to tell you, I think this is one of the best single talks
59:53The Pragmatic Engineer Host:on software you'll ever see. And we will link this, but can you give me just a short overview of like why this update is so difficult? Because like some listeners will be used to just building applications, for example, on the iPhone. and an update there it what it means obviously i know this is way more complicated but an update is there's a new binary version and it replaces the old binary version now of course you know you're saying this is an operating system update or or you know like with the car and of course you might think like well you know you could just replace the old version with the new version and there's some downtime but where is the complexity that actually like puts all this thorn because i'm sensing this is like i'm missing something something very obvious so because it's a distributed system like when you've got an app on an iphone it's not a distributed system oh and distributed system meaning that you've got a bunch of different nodes or components that are going to speak to one another and it's like and those might need updating as well definitely updating oh they all need up yeah the whole thing is the update you got to be able to update all of the software in the rack oh this is not just operating updating the operating system this is updating absolutely everything.
1:00:57The Pragmatic Engineer Host:So you might need to update some parts or all parts. You update the service processor, the root of trust, the drive firmware, the host operating system, and then all of the components that speak to one another. Okay. And then it's like, okay, so I mean, this is challenge is fractally complicated. I mean, one of the very basic ways it's complicated is like, so when we're updating, we are moving the system from one version to another version. In between, it's going to kind kind of be in both versions. Like, what does that mean to have the system that's operable while you've got some new components and some old components?
1:01:29What if you change your database schema from one version to the next version, which we definitely have. Like, you have to have a method of doing that. And for every one of these components, how is it updatable? We've got to reason about the system when it's in this hybrid state. And then it needs to be done in a way that's very, very robust. So first and foremost, we had to develop the foundation that allowed us to do this absolutely robustly. And so the way Dave and team did this is, you know, with that foundation and then very slowly lighting up different aspects of the system and making it more and more automatic over time.
1:02:08And, you know, first started running that on what we call our dog food rack and did our first automatic update on the dog food rack. Like, it was a really great feeling for that team because this has been a very long software road. And it has been one that has been very deliberate. And ultimately, like, and, you know, full credit to Dave and team, took us about the amount of time that we thought it would, which is kind of very rare for software. Because I think software is so practically complicated. But that's only because they've been very carefully managing scope versus schedule making. And because quality has got to be the constraint.
1:02:44And Dave's talk goes into that in detail in a way that I think is just extraordinary.
1:02:49The Pragmatic Engineer Host:So I'd like to talk about the topic that is, you know, a lot of people's mind is AI, specifically and AI tools. Yeah. How have AI tools changed how you're working at Oxide? Specifically, think about software engineering, maybe even hardware. Are you using these tools? Are you experimenting with them? For sure. When we've been early on in terms of using them. Yeah, I mean, use them for different – and people are using them in different ways. I mean, no part of the oxide stack is vibe-coded. I think that that is safe to say. But we are using it, and we're using it to – and again, different people are using it in different ways.
1:03:25We are using it to do things that are tedious. We're using it to do – generate test cases. You know, generate the – I use it for – because I think the thing that is just like unmatched at is just document comprehension. We've got a very writing-intensive culture. We've got a lot of documents. It is great. You've always had that. Yeah, I always had that. And if you've got a writing intensive culture, like you're LLM ready not to generate those documents, but to consume them. And to, you know, one of the things that I've always wanted to do, and it's still like now is possible. I haven't quite found the time to do it.
1:04:00Early on, I wanted to make an RFD glossary. So RFD are a request for discussion. We've got a lot of technical terms. I want to make a glossary. I tried to do that for like three hours. This is like in 2020. I'm like, this spreads to the horizon. This is so, just making a glossary is so complicated. A glossary is something that an LLM can just turn out. And so there are lots of things that we're doing to use LLMs in particular. It's clearly a very real, very, very big shift in lots of different aspects of software engineering. I think that, you know, but of course there are people that are being kind of productive about it.
1:04:37I am definitely not a doomer. There are a lot of doomers that are out there. And, you know, I tried to give this talk about building the oxide itself, the oxide rack. And in particular, the problems that we had along the way that an LLM was never going to be of any assistance on. And so, and I, the title of the talk was Intelligence is Not Enough. And one of the prominent doomers actually did a reaction video to my talk. It's like the only time I've ever had someone. And my daughter, who was then like 11, was just like, thought it was hilarious that someone had held their own time in such low regard that they would spend it recording a reaction video to my talk.
1:05:20And so she was like, I want to watch this. I'm like, oh, God, I do not want to watch this again. Ultimately, the thing that was really frustrating is this person obviously disagrees with what I was saying. but then when i was giving these very concrete examples of here are the specific technical problems that required more than intelligence to resolve that an lm was not going to be able to resolve he literally fast forwarded through those parts he's like we just don't need this this is like this is just you're like bro this is the talk like you you can't do this like you're fast forwarding over to the actual like meat of the talk can you give an example of like a problem
1:05:55The Pragmatic Engineer Host:which you felt was this like even you know if we fast forward to like the arbitrary future yeah yeah yeah so yeah super simple i mean like the i mean we've had many many scary problems but um we had a uh the cpu when we did our first bring up of our first machine and then what does it bring up mean a bring up means taking a board and powering it up and trying to get it to work for the first time i think you mentioned that the term smoke tests comes from electronics engineers Oh, I mean, a smoke test, I think it was smoke test more from aerodynamic injuries, but yes, I mean, aeronautical injuries, but yes, I mean, you're definitely like, smoke is definitely a possibility.
1:06:36That's a very bad, you do not want smoke. That is bad. No smoke, please, and bring up.
1:06:39The Pragmatic Engineer Host:So the bring up. But we are doing bring up, and we are unable to get the CPU out of reset. And after 1.25 seconds, the CPU resets itself. What's going on? Is the power network bad? We're doing all, and like, when you have something like that happen, it's like, well, what's happening? It's like, I mean, it's just not working. I mean, like, what do you tell your LOM to be like, like, it's not working? I mean, and it can maybe give you some suggestions, but in this case, it wouldn't. So we are going deep in this, understanding like, maybe the power network is like marginal. No, no, no, we resolve that.
1:07:17No, no, we've got a man, actually, we're working with AMD at the time. And Andy's like, no, these power numbers are amazing.
1:07:23The Pragmatic Engineer Host:Like, your margin is very good. you're measuring it out you're like eliminating that one eliminating that when you're going to eliminate eliminate eliminating and uh couldn't get we and this was weeks and you're like we are we don't have a company like we're wow we are absolutely dead and i feel like this is the kind of thing that desperate you know you get desperate you're like we're gonna try kind of anything and what we uh the engineer was working on this um actually looked at the protocol between the cpu uh and the voltage regulators there's a protocol that it goes back and forth says hey I need this voltage and this is a voltage chain.
1:07:56And one of the things he notices is that there is no acknowledgement packet from the regulator. So the CPU asks for a voltage to be set to a certain level. And he's noticing that there's no acknowledgement packet back from the regulator. Which should come. Which should come. And the test, they've got something called SDLE, which is this great test goober that you take the CPU off, you put on the SDLE, and it will measure the power for you. Well, the SDLE didn't care whether it got an acknowledgement packet or not. The CPU definitely did. And the CPU, so the CPU says, I want you to go to 0.9 volts.
1:08:30It never gets an acknowledgement back. And meanwhile, sitting at 0.9 volts, and it's just like, well, I never got an acknowledgement, so we're going to reset and I'll do it again. And that was due to a firmware bug on the Rennesos controller. And so we got a firmware update from Rennesos and done. And I mean, to be fair, the Rennesos FAU is great. It was like, well, you guys should reach out a lot sooner like yeah i know we really wanted to make sure that we got like everything uh and and that's the kind of problem and there were many many problems like this where it's not merely intelligence it's not building a a board is not an iq test it's more i mean you need to be intelligent to do it but intelligence is not enough you need these other kind of characteristics and i feel we also need a team in this case right absolutely a team 100 100 you need to like you're
1:09:16The Pragmatic Engineer Host:you're gonna solve these problems with you know you had that engineer who just like thought of measuring this out. Right. Well, and she who was desperate, you know, because we were all getting desperate. And, you know, we, and again, we've had many of these over the history of the company. And you're right. You absolutely need a team. You need a team. And you see also the value when you have a team. People have different ways of approaching a problem. That diversity is really important because you need, and actually sometimes, this has happened more than once to the company where somebody kind of like is just kind of like walking through the problem and like someone's like hey i'm just joining you know that's about a remote company anyone joins you know they're joining the google meet yeah i'm just joining because you know i think that i'm following along and you get someone will be like just making like hey i got a dumb question are those virtual addresses like those look like similar virtual addresses you get something where someone's making and you need someone to kind of like come and make that observation that is maybe less grounded in it.
1:10:17And people are like, oh, wait a minute. No, that's actually like, that's something to go check. And so you need that different kind of approach that is really a team kind of uniquely summons.
1:10:29The Pragmatic Engineer Host:And, you know, I think you might've alluded to it, but on the previous podcast, Armin Roner sure mentioned to me, he's the creator of Flask. He's been around the block for quite a while and he's now doing a startup. And he said that right now it's just him and his co-founder and he's got an army of, AI interns right now, there's prototyping him. But he told me, I'd like to start to hire people soon because people bring energy and you need energy for a company to live and thrive. And I'm kind of sensing the same thing. Oh, for sure. No, for sure. And I, you know, just listen to this great piece with Richard Sutton, who was the inventor of reinforcement learning.
1:11:06And I think rightfully, and I agree with him, it's like, you guys are conflating an LLM with artificial intelligence. It doesn't have goals. This is really important. So like a prompt is not a goal and guessing the next word is not a goal. And, but like us together as a startup and like wanting to make it together, not wanting to die here together, that's a goal. And that, so we can use that creativity. Maybe we, you know, use an LLM certainly as a tool to help us achieve our goal.
1:11:37The Pragmatic Engineer Host:But I do think that that's a very important distinction. And can you tell me like what kind of tools you use and what are the areas that you find it helpful? I understand you're experimenting with stuff and, you know, this is all work in progress, but where are areas that, you know, and you mentioned like summarizing was one example of glossaries. Oh, yeah. I mean, I use LLMs as an editor all the time. I find it to be a really, I mean, actually it was funny. I had a blog entry that went on Hacker News and someone's like, oh, this is LLM written. I'm like, actually, it is LLM edited, but the only thing that I did based on the LLM is I deleted an entire paragraph.
1:12:11So there's a paragraph that, like, wasn't working, and the LLM was like, this paragraph's not working. And I'm like, you know what? I'm just going to delete the paragraph. So it's like, I don't know. You want to say that's LLM edited? Because, like, every word there is written by me, but there were some words that there was written by me that an LLM suggested I deleted, which I deleted. So, I mean, I use it in writing for sure. I mean, I also like to use – I mean, this is like a stupid reason, stupid thing. But when you're writing Rust, and we write a lot of Rust, especially when you're new to Rust, you wonder like the way I just phrased this, is this like idiomatic?
1:12:45Is there a better way to do this? That's a great little problem. I've got this small little snippet of code. Is this an idiomatic way of doing this? Is there a better way of doing this? And that's a great thing for an alum to make a suggestion or not or tell you like, no, that's an idiomatic way of doing it. maybe I would make this small adjustment. So I find it really valuable. I find LLMs to be more valuable in the small than in the large. So like, again, this kind of, I might, you know, hats off to people who want to spend their lives acting as a middle management for robots, but like, that's not necessarily for me.
1:13:17Certainly in Oxide, I mean, our belief is that people take responsibility for their own work. So if you want to have an LLM help you out on that, that's fine. But ultimately, like if there's a bug in this, like you can't blame the LLM. The LLM broke my code is like not it not interesting that that's lms don't have accountability and so one thing
1:13:34The Pragmatic Engineer Host:that is starting to spread across i think a lot of engineering is engineers using lms either inside your id with autocomplete or or and also kicking off now agents now there's more advanced ones with like cloud code and codecs where it can actually run command prompts and run your tests are you seeing engineers use some of these tools and there's a little bit of back and forth as well you know like it's very clear that when you're doing kind of more boilerplate things that are so-called on distribution which is they've learned like reactor type script it can spit out a bunch of stuff but you strike me as someone who's doing a lot more nuanced things yeah i mean you're writing a bunch you're writing a bunch of c code in the operating system kernel it's like it is it is less valuable yeah but so what are you seeing across the team in terms i think you I encourage people to experiment.
1:14:23And I would say we're seeing a wide variety of experimentation. Certainly we've got, we're using Cloud Code a bunch and people are doing that. And, but I would say, you know, broadly speaking for a lot of the work that we're doing, it is helpful as like maybe a polishing tool, but less as a kind of at the epicenter of its creation. It's not true of everything. There's some software for where we're really.
1:14:46The Pragmatic Engineer Host:But that's also nice to hear because I'm kind of asking you more to putting on your CTO hat, who's also very like, you know, you're very hands on and you know what's going on with the industry because a lot of non hands on executives are kind of looking their finger and thinking, oh, we must be 10 or 20 or 30 percent more productive. But what I'm hearing is like things are kind of the same as before. Right. Yeah. I mean, my big belief is it's a tool. It's a powerful tool. I mean, I will say the thing I, you know, occasionally people are like, well, I don't want to use it at all. And I'm like, you should.
1:15:15So like you should try. Yeah, yeah. Like, let me get you off of that position. And let me, you know, we had Simon Wilson on our podcast. Simon's delightful. And, you know, one of the lines that he has that I really love is people should run these LLMs on their own laptop where they run slowly and poorly so they can see the bad output that they generate so they can understand what some of the limitations are. So I definitely, I love that. I do think that people should use them enough to know where they are valuable. It's a very important tool in the toolbox. You want to be aware of it. But it's definitely reductive to think it's the only tool in the toolbox because it isn't.
1:15:49The Pragmatic Engineer Host:Now, you're in such an interesting company because, like, you know, you don't not just do software, but you do a lot of hardware. Yeah. Have you found any use? No. No. No. Zero. I mean, okay, zero is a bit reductive. I have found it to be useful when, for example, you know, you've got a waveform of an I2C transaction. It actually, amazingly, you can send that to an LLM and have it, like, interpret this. Like, hey, am I seeing I2C kind of compliant behavior? and it can help you out on that a little bit but it's like absolutely at the edges okay so that's a 0.01 also like i think people don't realize like there are already tools for that like that's what eda is you spend a lot of money on like we're not laying this stuff out like by hand with graph paper like this is like you've got you know when you do layout for a board there are a bunch of rules that are automatically checked for si you know we we've got a we do a bunch of simulation work.
1:16:43Like we're not doing that by hand. We're using software. Yeah.
1:16:46The Pragmatic Engineer Host:And I saw you have those machines in there. Like I saw that. I think it's a bit reassuring to hear because I think it's very clear, like maybe we don't realize software engineers, but programming is such a great use case for LLMs. It's a simple grammar. You can validate it. And I think it's sometimes nice to just, you know, like touch sand of like an area that is very, very different. But it's cool that you're checking and, you know, you're seeing if it changes over time. I guess you always keep checking. Yeah, for sure. And I think that it is frustrating to me because programming is such a good use case for certain kinds of programs.
1:17:21So as a result, you end up with certain kinds of programmers who just, in part because of their own self-centric view of the universe, believe that, oh, this is just going to replace every job. And it's like, no, not even close, not even close. And you need to spend more time. You need to get outside a little bit more.
1:17:36The Pragmatic Engineer Host:Yeah. So speaking of getting outside and meeting different people, what I noticed when i went to oxide is just like it was great we had double easa as you say software engineers people used to work on virtual wheel alley at oculus all in the same room can you tell me about how big is the team what's the composition yeah so we were on you know we've i think you know we've got some more offers going out tonight so i think we've got on the order it will be at like 85 i should probably keep better i should keep better mental track track of it but we've got like 85 plus minus. And we, you know, we've been very blessed by, uh, we've really put a beacon out there.
1:18:12We've got a lot of people rooting for the company. We've got a lot of people. And as a result, we got a lot of people who want to work for the company. So, um, you know, we, as we talked about last time, um, we really put a lot on folks to describe, you know, the work they've done, what's important to them, why they want to work for oxide. I mean, a lot of my LLM use is I will look at someone's materials. You can imagine we've started to see materials that are heavily LLM authored potential applicants to Oxide. Please do not do this. We get people who human author their entire materials, and then they get to the last question, why do you want to work for Oxide?
1:18:45Why do you want to work in this role? And they have an LLM spit that out. And you're like, do you think you want to work here? Let's leave aside whether this is right or wrong or cheating or not. It's like, fine, I guess. But I don't think you want to work here. You're not going to get a job here because I don't think you actually want to work here. Put it in your own words. But that process really has allowed us to attract people who themselves are attracted to the company and attracted to the culture, the problem, the team. And it's just extraordinary. I mean, I just feel so lucky to be with such an unbelievable group of people across more and more and more and more disciplines.
1:19:26I mean, the great thing about our approach is it brings people in who are, you know, God, it's like I love this approach. for we talked about support engineering are we i people who are like god i love this approach like finally qa can stand on its own two feet i feel that that qa has been kind of subjugated by by these other disciplines now qa is kind of really thought to be as important as anything else in the company and it is because at some like at some like monetary perspective it is as important as
1:19:55The Pragmatic Engineer Host:anything else um yeah but but i remember like when i worked at microsoft back like 15 years ago or or so the QAs were just on a lower pay grade, you know, like the senior QA was at the same as like, I think software engineer two or something, which just kind of implied. Yep. You're less important. You're less important. You're just less important. And so like, if you tell the world that we think it's as important, you know, you get, you get people who are extraordinary at QA, you get the best, the best. And so that has been really exciting. And now we've got people coming. I mean, I do love how many different companies.
1:20:30My belief is that like every company has something to teach us, that there is something positive you can take from every company. Now, there are some companies just like, oh, you're really scraping the bottom of the barrel.
1:20:44The Pragmatic Engineer Host:Maybe not Enron, although they did buy some. Yeah, yeah, that's right. There are like even Oracle you can find. There are maybe a bit of a challenge. Let's not do that one. But you know what? And at the time, I thought this was a negative, but now I'm like, I see it. Larry Ellison makes every hiring decision at Oracle. So what's positive about that? Exactly. But you'd be like, I really think that the kind of the founder mode, the Paul Graham essay on founder mode, is talking about founders that lost track of their own hiring. So I think, no, I don't like the way Ellison does it. I think that you want to have, you want to trust a team to make a decision.
1:21:20But ultimately I believe that the, that the CEO of a company bears responsibility on every single hire. And I think should be looking at every single hire coming into your company. And that is a, to me, that is a very important check on these kinds of companies that, that, so that is, there you go.
1:21:36The Pragmatic Engineer Host:Something that I've, something that positive, I get that I take to Oracle and it's telling that your immediate reaction is like, wait, what's positive about that? Yeah. Yeah, I'm not sure you undid that talk on Oracle. Yeah, fair enough, exactly. Yeah, and they're from some companies more than others, but I think that there are, and so I love having all of these different experiences present at Oxide, because I do think that there's so much to learn. And we're trying, you know, you want to take all the positive things. Because I also think that every company, including, you know, people, I had actually one of the questions I love that I got once, is like, what do you not want to emulate from Sun?
1:22:12I'm like, oh, thank God. because like people think of Oxide as kind of the second coming of Sun Microsystems. And like, there are lots of things I loved about Sun. There are lots of things I did not love about Sun that I did not want to emulate. And so I think for any, also any company, there are things we want to leave behind. And, you know, I think when you've got a big diverse team, you get to go do that.
1:22:31The Pragmatic Engineer Host:And one thing that really surprised me last time I was at your office is, it turns out that most people are not in the office and they work remote. And I would understand for software, but how do you make that work for hardware development where physically you do need to, you know, be at the hardware? Sometimes I understand you need to measure stuff. I saw a lot of like, you know, units. Sometimes you need to go to like check while manufacturing. How does that part work? Yeah, so I mean, a lot in people's basements. So, you know, fortunately we're making, you know, this is the advantage of making a server and not making like, you know, a tractor or like, you know, we're not making like a, you know, I don't know, like a wind turbine or something.
1:23:08You know, this is something that people can actually model in their basements. So that helps. But then a lot of even hardware engineering is using these software tools, using EDA tools, you're using SolidWorks, you're using Altium. You're kind of putting this thing together. When you're doing layout, for example, which is a very important task when you're laying out a board, all of that is that can be done anywhere. That's all just software.
1:23:30The Pragmatic Engineer Host:Okay. And so there are things that are where that physicality is very important. And then when you're doing bring up, you actually need to be at your manufacturer when you do that. So like that is also not an office. You would need to travel anyway. Yeah, you need to travel anyway. And anyone coming to the electronics industry is like, okay, I'm interested in oxide, but please tell me I never have to go spend any time in Taipei or Beijing. Because you go out there for, you know, or Shenzhen or wherever, and you're out there for two weeks in a windowless office trying to get this thing brought up.
1:23:57And we, all of our assembly is done here in the United States of Minnesota. So we are all, in fact, we've got a bunch of folks out there this week for Benchmark Electronics in Rochester.
1:24:06The Pragmatic Engineer Host:So this is wonderful. And one thing that you told me is one of the things that's on top of your mind right now. As oxide is growing, you still have this culture of the same compensation for remote. So, like, it's kind of been the same since the start. What will be the challenge in maintaining it? Because, again, you worked at large companies. You've seen how it goes. It can get tricky. What are the things that you're seeing and what are the things that you're trying to do to, you know, keep this kind of startup vibe even as you might be just bigger? Yeah. So I think that the thing that is, that is top of mind right now for me is, and especially because, you know, we raised a big Series B, which is great.
1:24:42I think much more importantly, we're seeing a lot of customer traction, which is great. So we've seen really exclusive. Wonderful. It's paying off. Right. And we kind of knew that was going to happen in the abstract, but it's fun to actually see it happen and fun to actually see the customers that are, you know, like, you know, I bought one rack and I mentioned it, but now I want to buy a lot more racks. I love what I'm seeing and I want, you know, that's great. Very, very, very exciting stuff. That means we're growing the company a bunch. And one of the things that's very important to me because I've seen this happen so many times is companies take their eye off the ball when it comes to hiring in particular.
1:25:15order. And it is very important to me that we continue to have absolute discipline in the way we hire. And we're doing that. And fortunately, you know, the nice thing about our hiring process is every single Oxide employee has gone through it. So it's like, I'm not having to persuade anyone about the importance of our process because everybody has gone through it. And that, you know, the thing that we've got overwhelmingly in our favor is because we've used our values as a lens for that hiring, Oxide's culture is important to every single person at Oxide. That's what it takes to really preserve that.
1:25:49And it doesn't mean that it won't change at all, but the bones aren't changing. Like what will change is it will be bigger and it will be, I think, you know, and I love the fact that, you know, even at like 85, we're already so big that, you know, Steve and I know everybody at the company, but very few other people know everybody at the company. So when we get everyone together, it's like the best party you've ever been to. Because in college, I used to throw the best parties in college. And the reason I threw the best parties in college, it's not because of me, it was because of the roommates that I had.
1:26:22So I was a computer science student who played ultimate. My roommate was an engineer who was on the water polo team. My other roommate was a history student who was in the course. That's six different demographics that don't normally overlap. And then very importantly, we made sure that the women's swim team was always invited the women's swim team they were like the foundation of a waterfall yeah exactly waterfall you always check their calendar to make sure they can make what and people loved the parties we had why because they would meet people that they never met before who were really interesting and they and what i love about oxide is we've got this when when we get the whole team together people get the all these delightful surprises so people take me aside and be like god you know rye is awesome i'm like yeah i know i know i know you know too now that's great, but like, you know, or, you know, whomever it is, it's, it's just, it's really exhilarating.
1:27:12And I think that also serves to reinforce how important what we've got is, as I tell the team, like we have lightning in a bottle and we cannot take it for granted. And that means that every single one of us need, we need to rise to the moment. We need to do what our customers need us to do, but we need to do it in a way that protects and preserves what got us here.
1:27:31The Pragmatic Engineer Host:So thinking a little bit ahead, let's assume that, you know, these AI tools will just get better eventually they'll be able to you know help more even on on your kind of low level things you've been in the industry for quite a while you've seen a lot of shifts what are what do you think are some of the things both in software engineering or in hardware engineering or just engineering that will probably not change even if we predict uh these things being like more capable yeah i think what we i mean i think that that it's certainly a revolution i think it's going to allow us all to do more. I do think that we are going to hit a point where people understand that this is a tool where, because there's a little bit where we're still have this tension of like, oh, is this going to be AGI?
1:28:17Is this going to replace all jobs? And this is like nonsense as far as I'm concerned. And it's distracting kind of nonsense. And we actually need to get back to putting the tools in the toolbox of the human that's building it. Now, these tools have become much more powerful. And I think that's going to be, I think it's extraordinary. I think it's important. I think that also we'll be, you know, we've got a lot of experiments right now. We humanity that I'm not sure are going to make economic sense. So, you know, we'll, we'll, we'll, we'll be figuring that out as well. But I think that, you know, one of the things I am a little bit worried about is a little bit of despair from younger software engineers in particular who are like, what's the point?
1:29:01The Pragmatic Engineer Host:Like an AI can do all this. Well, and there's also the news, even for more experienced software engineers in the mainstream media, there's this news that company X is laying off half their workforce because of AI. And by the way, when we look closer, it's not because of AI, but it is coming across and it does give not younger people a lot of anxiety, tons, even like mid-level folks or even some more experienced. Like it does give a sense of, I think it's the first time in computer history that most of us to remember that there is this thing that could threaten my job. And I think we've just never had to deal with this.
1:29:34The Pragmatic Engineer Host:I think, you know, there are other industries that might have been a bit more used to it. Yeah, I would say that we, I mean, there have been busts before. The dot-com bust was a bust. Like a lot of jobs did disappear, right? So I think that we, but the busts really come in what feels to be a broader and more permanent way. I mean, my view is like, this is an opportunity for, I mean, I think one of the things we should be societally really encouraging is new company formation. Because now, I mean, just like you're talking to Armin about how, you know, just a small group, you know, just Armin and his co-founder were able to do so much together, right?
1:30:08We should be really encouraging that. And what are some of the gaps that we can all go fill? Because ultimately, like, we all need to find a livelihood. We need to find meaning. And the way we do that as engineers is we build useful things. And so we're like, we can now build many more useful things. What would we go build? What would, if you could build anything, what would you go build? And that's kind of the question that people need to ask themselves. It's scarier. It's scarier than like, go to this school, get concentrated in this. And then mama Google will hire you and take care of you and feed you breakfast.
1:30:40It's like, no, that's not going to like, that's not what's going to happen. And it feels a lot scarier because it feels like there's at some level, like less security, less job security. But yeah, that's true. You know, and that's scarier, but there's also a lot more opportunity.
1:30:57The Pragmatic Engineer Host:And for a college student or someone in school or with a little experience who says like, look, my goal would be one day in like five years time to be as good that I could get a job at a place like Oxide. It doesn't need to be Oxide, but again, a place that has a high bar there, they often hire experienced people, but I want to get there. And yeah, there's all this AI stuff is happening. happening, what would you advise them in terms of what to focus on, what areas to study, what things to do, or how to think about like, you know, like they have the goal is there. What advice would you have them part with?
1:31:32Yeah. So I think that they need that you need to have a different mindset. And that mindset needs to be not around how do I create as much as possible, but rather how do I get better? How am I getting better every day? And I think LLMs are a great tool to get better. How can I learn about something new, go deeper, go into something that I wouldn't go into before, get over that kind of that fear. And one needs to, especially if you're in school now, you want to work in a place like Oxide. It's like, you kind of have to view it as like, all right, like you want to play major league baseball. That's great.
1:32:09Like you're a great high school player, you want to play major league baseball, it's really hard. Got to get better every single day. And you're going to need to be really focused on getting better. And you need to be like really realistic about like what I need to go do to get better. And it's hard, but, and it's chancy because you might not get there, but you could get there. And the, and you're certainly not going to get there if you don't focus on that kind of self-improvement. So I, I really think that, that it, there is a shift in mindset that, that needs to happen or that one needs to have.
1:32:39I would put that way. You really got to have a mindset towards getting better, understanding more. What do you not understand? There is lots that you don't understand. I mean, I think one of the challenges of modernity is that we delude ourselves into thinking that we understand it all. You don't. I don't. Like one of the things that I've learned, I've joked at Oxide that like I keep waiting for the day that I know how computers work. And it like – like it wasn't today. It definitely wasn't yesterday. It's not like it's going to be wrong.
1:33:06The Pragmatic Engineer Host:You understand how they work. But I mean that earnestly in that the amount of complexity that I definitely, I mean, I knew but also didn't know. It's like every day I feel I'm still learning new facets and not just like a computer but actually delivering a computer to people. Like there's so much to learn out there. And now with the way you've got to view LLMs is not like this thing is coming from my job. you got to view it as like no i've got now this like private coach tutor what have you that i can ask any question to it's not going to i gotta like fact check its answers for sure but now you've got the opportunity to and you've got it is easier to get into this domain than it ever has been and that is that's great and it's powerful but it can also be scary and as closing what's a book or two books that you would recommend to folks and why oh so many good books you know my my My, uh, my, I've got a, I've got a 21 year old, an 18 year old and a 13 year old.
1:34:08And when the 18 year old was in his, he's now a freshman in college. He's a high school senior. He got this assignment, great assignment from his, his English teacher. Name, we'd go to someone that you, that, that, that you know, and ask them for three books that they would recommend that you read. And I'm going to assign you one of those three books to read and you're going to read it. And then you're going to talk with them about that book. I'm like, oh, I love this assignment. So he's like, dad, I'm coming to you. And I'm like, oh, you have, thank you. So I know of course my wife was like, why didn't you come to me?
1:34:36Like, hey, look, I'm, you know, I, sorry. You know, look, it was great. So yeah, I'll give you those three books that I gave to him. And I think that each of these is really terrific. First is Solve a New Machine by Tracy Kidder. So this one won the Pulitzer Prize in 1980 or 1981, but about the building of a new computer at Data General. And it's a, it's extraordinarily well-written. and even if folks are like, well, I'm not like, what do I have to do with a computer company in the late 70s and early 80s? Any engineer will see something of themselves in that book. It is just masterfully told. Tom West, who's the, is kind of a complicated figure, but that is, soul is still, I mean, it's literature for us.
1:35:22So I would absolutely, Soul of a New Machine, every engineer should read Soul of a New Machine by Tracy Kidder. For me personally, very influential was Skunk Works by Ben Rich. So about the history of Skunk Works, Clarence Kelly Thompson was kind of the originator of Skunk Works at Lockheed Martin. Extraordinary story about what engineers can do when they kind of task themselves on the impossible. It's such a good book. It's such a good book. Amazing book. And then the other one is Steve Jobs and the Next Big Thing by Randall Strauss. so um steve jobs is kind of like lionized by the industry but people forget about a very important chapter of his life namely next and i believe we are it was just an anniversary maybe it was the 30th anniversary must have been of the or maybe the 40th anniversary jesus of the the announcement of the next machine so the um steve jobs left apple was fired from apple started a computer company called Next, a really interesting company in a lot of ways, was at Next for a very long time.
1:36:28It was a 13-year journey before Next was bought by Apple. Next is bought by Apple. Steve Jobs returns to Apple when they buy Next. This book, Steve Jobs and the Next Big Thing, is written before Apple buys Next. And it is at Steve Jobs' lowest moment. It is not here to praise him. It is to bury him. And it is very interesting about all the missteps at Next. And the thing that we cannot know, because Jobs obviously died, but I believe, having read the book, which gets basically, Next gets essentially no treatment in the Isaacson biography. Next is like six pages of glory. It's like, that's not what it was.
1:37:06But Randall Strauss's book is masterful. And And in particular, I believe that Jobs' failures at Next were essential for the resurrection of Apple. And because you look at the way he handled himself coming back to Apple was very different from the Jobs that got fired from Apple. And I think that like when people look at Jobs, like they don't really take him apart. And I think you should because I think he's a really interesting guy. He's enigmatic. He's someone who's like – he did things that I think are really fascinating and also things that I really strongly disagree with. So just to be clear, I'm not like, but I think that he's, he's an indisputably an important figure and that book is by far the best book.
1:37:47So Steve Jobs.
1:37:48The Pragmatic Engineer Host:I'm adding that. I actually want to read that. Oh, it's extraordinary. It's very good. Well, Brian, this was such a fun discussion. Oh, my, my pleasure. I mean, we knew this was going to be long and wide ranging, so hopefully it delivered, but, uh, I, I really appreciate that. We went from, from the nineties all the way to the future. There you go. Awesome. Well, thank you so much for having me, Gary. It was terrific. I've got to say, Oxide is one of my favorite companies, and I say this as someone who has zero affiliation with them. It's just so rare to find a startup that built both hardware and software and are world-class in doing both of these, and are so open about talking exactly how they do it all.
1:38:24The Pragmatic Engineer Host:Honestly, the only downside I can think about Oxide is how their server racks are built for pretty large companies and are definitely out of reach for hobbystuffs. In this episode, I really appreciated how much of a straight shooter Brian was, especially about the impact of AI tools. Yes, everyone at Oxide uses them and they do find use cases for coding and working with documents, but it's eye-opening how it gives them basically zero help with hardware engineering. This is a good reminder that LMS might be the single best fit for coding-related tasks, and as devs, we should know that these tools might be more specialized than many people think.
1:38:56The Pragmatic Engineer Host:I hope you enjoyed the stories in this episode as much as I did. If you'd like to learn more about Oxide, I did a two-part deep dive about the company, and you can read it linked in the show notes below. If you've enjoyed this podcast, please do subscribe on your favorite podcast platform and on YouTube. This helps the podcast a lot. A special thank you if you also leave a rating on the show. Thanks, and I'll see you in the next one in the next year.
From the publisher
Brought to You By:
• Statsig — The unified platform for flags, analytics, experiments, and more.
• Linear — The system for modern product development.
—
How have servers and the cloud evolved in the last 30 years, and what might be next? Bryan Cantrill was a distinguished engineer at Sun Microsystems during both the Dotcom Boom and the Dotcom Bust. Today, he is the co-founder and CTO of Oxide Computer, where he works on modern server infrastructure.
In this episode of The Pragmatic Engineer, Bryan joins me to break down how modern computing infrastructure evolved. We discuss why the Dotcom Bust produced deeper innovation than the Boom, how constraints shape better systems, and what the rise of the cloud changed and did not change about building reliable infrastructure.
Our conversation covers early web infrastructure at Sun, the emergence of AWS, Kubernetes and cloud neutrality, and the tradeoffs between renting cloud space and building your own. We also touch on the complexity of server-side software updates, experimenting with AI, the limits of large language models, and how engineering organizations scale without losing their values.
If you want a systems-level perspective on computing that connects past cycles to today’s engineering decisions, this episode offers a rare long-range view.
—
Timestamps
(00:00) Intro
(01:26) Computer science in the 1990s
(03:01) Sun and Cisco’s web dominance
(05:41) The Dotcom Boom
(10:26) From Boom to Bust
(15:32) The innovations of the Bust
(17:50) The open source shift
(22:00) Oracle moves into Sun’s orbit
(24:54) AWS dominance (2010–2014)
(28:15) How Kubernetes and cloud neutrality
(30:58) Custom infrastructure
(36:10) Renting the cloud vs. buying hardware
(45:28) Designing a computer from first principles
(50:02) Why everyone is paid the same salary at Oxide
(54:14) Oxide’s software stack
(58:33) The evolution of software updates
(1:02:55) How Oxide uses AI
(1:06:05) The limitations of LLMs
(1:11:44) AI use and experimentation at Oxide
(1:17:45) Oxide’s diverse teams
(1:22:44) Remote work at Oxide
(1:24:11) Scaling company values
(1:27:36) AI’s impact on the future of engineering
(1:31:04) Bryan’s advice for junior engineers
(1:34:01) Book recommendations
—
The Pragmatic Engineer deepdives relevant for this episode:
• Startups on hard mode: Oxide. Part 1: Hardware
• Startups on hard mode: Oxide, Part 2: Software & Culture
• Three cloud providers, three outages: three different responses
• Inside Uber’s move to the Cloud
• Inside Agoda’s private Cloud
—
Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email podcast@pragmaticengineer.com.
Get full access to The Pragmatic Engineer at newsletter.pragmaticengineer.com/subscribe




