The Co-Creator of Kubernetes On Convincing Google, Building It, and Scaling for LLMs

23 Mar 2026 · 1 h 7 min · 29 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Notes on "The Co-Creator of Kubernetes On Convincing Google, Building It, and Scaling for LLMs" - The Peterman Pod

Episode Summary In this episode of The Peterman Pod, Brendan Burns, co-creator of Kubernetes and current technical fellow at Microsoft, shares insights into his experience building Kubernetes at Google. He discusses the challenges of gaining buy-in from leadership, the development process, the importance of open source, and how Kubernetes is evolving to handle new workloads like AI training.

---

Key Topics Covered

  1. Convincing Google Leadership
  2. Articulating Value: The hardest aspect of building Kubernetes was convincing Google leadership of its importance. Brendan emphasized the need to articulate the strategic significance of Kubernetes to gain approval.
  3. Influence on Tech Landscape: Google wanted to influence technology rather than just releasing white papers like MapReduce and Hadoop. Kubernetes was seen as a pathway to help Google maintain a leadership position in the cloud space.
  1. Development of Kubernetes
  2. Minimum Viable Product (MVP): The initial MVP was built in under a week. It showcased core functionalities like load balancing and health checking.
  3. Team Dynamics: Brendan collaborated with a product manager and a design engineer, leveraging their strengths to create an effective prototype.
  4. Use of Existing Tools: The MVP was built by integrating various open-source technologies, which helped expedite development.
  1. Open Source Philosophy
  2. Rallying the Community: Early buy-in from tech partners was crucial. Brendan posited that they were more willing to contribute to Kubernetes because they didn’t see it as their core business value.
  3. Democratic Governance: The establishment of the Cloud Native Computing Foundation (CNCF) and governance rules were imperative to ensure Kubernetes remained an industry standard without being dominated by Google.
  1. Scaling Kubernetes
  2. Current Challenges: Kubernetes is now facing new workloads, particularly for AI training. Kubernetes had to adapt to handle larger clusters and optimize its core systems.
  3. Storage Layer as the Bottleneck: The scalability of Kubernetes heavily relies on the efficiency of the storage layer, primarily etcd. Any significant scale-up would necessitate enhancements in this area.
  1. Reflections on Career and Education
  2. PhD Perspective: Brendan reflects on his PhD in robotics, stating that it provided him with valuable skills in writing, presentation, and critical thinking.
  3. Advice for Aspiring Engineers: He encourages individuals to pursue areas they are passionate about rather than chasing trends. The focus should be on continuous learning and personal interest.
  1. The Inevitable Trajectory of Software
  2. Software’s Lifecycle: Brendan believes that all software eventually faces obsolescence. He emphasized that while Kubernetes may not disappear soon, it will likely evolve or be replaced by simpler, more effective solutions in the future.
  1. Book Recommendations
  2. Key Reads:
  3. Design Patterns: A seminal book for understanding software engineering concepts.
  4. Leadership on the Line: Insightful for those in leadership roles.
  5. Five Dysfunctions of a Team: Offers perspectives on team dynamics and management.

---

Key Takeaways

  • Importance of Open Source: Open-source projects that invite community participation have a better chance of success and longevity.
  • Career Development: Engineers should focus on learning through passion and interest rather than strictly adhering to trends.
  • Scalability Challenges: As Kubernetes scales, the efficiency of its storage layer will dictate its performance and capability to handle emerging workloads like AI.
  • Software Dynamics: The nature of software is transient; being adaptable and open to change is crucial for long-term success.

---

Final Thoughts Brendan Burns' experience offers a wealth of knowledge for software engineers and leaders alike. His insights into the development, implementation, and community engagement surrounding Kubernetes not only highlight the project's importance in the tech landscape but also provide valuable lessons for those navigating their careers in the industry.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Origins of Kubernetes

0:46 to 4:34

Brendan discusses the early motivations and challenges of building Kubernetes.

“I don't fully understand the business motivation.”

The Case for Open Source

4:35 to 8:06

Brendan explains the strategic advantages of making Kubernetes open source.

“to ignore you and then they're going to go build their own.”

Building the Initial MVP

8:07 to 11:30

Brendan shares details about the creation of the initial Kubernetes demo.

“It was like, hey, if these eight people go off and, you know, do something and it turns out to be stupid, like, we'll just kill it off.”

Advice for Aspiring Innovators

11:31 to 14:00

Brendan offers insights on balancing existing work with new ideas and projects.

“I think a lot of software engineers, when they hear this kind of story, they think, oh, I have my existing responsibilities and I can't necessarily go off and build this thing, even though I think it's a great idea.”

Balancing Work and Passion Projects

14:00 to 18:02

Explore the trade-offs between personal time and professional commitments.

“And a little bit of family, like, and sleeping, right?”

The Importance of Side Projects

18:02 to 20:00

Learn why maintaining side projects can enhance your career and creativity.

“And so like, it's all about saying like, I'm not going to ask permission.”

Navigating Risk and Innovation

20:00 to 22:44

Understand the risks involved in pursuing innovative ideas and how to manage them.

“And I think you just have to be comfortable with that.”

Convincing Leadership for Open Source

22:44 to 25:36

Discover strategies for persuading leadership to embrace open-source initiatives.

“you can still have a very successful career, but it's a little bit harder for it to like, be directly attributable to you, right?”

Building the Initial MVP for Kubernetes

25:36 to 28:00

Learn about the decision-making process behind the MVP development for Kubernetes.

“So how'd you decide this is the minimum set of features that we need for this to, before we launch it?”

Challenges of Loose Coupling in Kubernetes

28:00 to 28:50

Explore the difficulties in debugging a loosely coupled system in Kubernetes.

“But when things go wrong, it's really hard to figure out why it went wrong.”
Show all 29 chapters

The Complexity of Troubleshooting

28:50 to 30:20

Understand the complexities involved in reproducing and troubleshooting issues in distributed systems.

“But then you're like, okay, but what happened?”

Using Etcd for Consensus in Kubernetes

30:20 to 31:20

Learn how Etcd facilitates consensus and leader election in Kubernetes.

“that's controlling everything right like there's some leader or maybe some nodes that are looking at one of them to kind of coordinate everything.”

Design Principles: API Server and State Management

31:20 to 32:50

Discover the role of the API server in managing state and its implications for system stability.

“Well, probably somebody understands it, but not a lot of people understand it.”

Declarative vs. Imperative Programming in Kubernetes

32:50 to 35:10

Delve into the advantages and challenges of a declarative design approach in Kubernetes.

“So you don't have to worry about corruptions.”

The Benefits of a Declarative Approach

35:10 to 37:40

Analyze how a declarative approach enhances clarity and self-healing in Kubernetes.

“So you kind of just say, I want this to be true.”

Scaling Kubernetes and Gaining Buy-In

37:40 to 38:50

Examine the strategies for gaining support from other companies for Kubernetes.

“Like there's there's not a lot of downside.”

Maintaining Independence for Kubernetes

38:50 to 42:00

Learn about the importance of governance and independence in Kubernetes' evolution.

“And it sounds like a a large part of this was getting buy-in from other companies and other people.”

Establishing Governance for Kubernetes

42:00 to 43:35

Learn about the importance of governance rules in the Kubernetes community.

“Because it's hard to partner if somebody else has trademarks on the Kubernetes logo or whatever.”

Creating the Bootstrap Committee

43:35 to 45:26

Discover how the Bootstrap Committee was formed to lead Kubernetes governance.

“But also, I think we're critical to a success.”

Challenges in Open Source Contributions

45:26 to 46:46

Understand the barriers to community contributions in open source projects.

“So we didn't have to be like, oh, you know, we grabbed this side and not that side.”

Legal Concerns in Open Source

46:46 to 48:56

Examine the legal worries that companies face when contributing to open source.

“It's really hard to get people to contribute.”

Scaling Kubernetes for AI Workloads

48:56 to 50:18

Learn how Kubernetes has evolved to handle AI workloads effectively.

“if I wrote the JPEG, open source JPEG library that ended up in the smoke detector that caused the house, you know, like you can sort of imagine the chain of logic that gets you there.”

Handling Increased Load on Kubernetes

50:18 to 55:01

Explore the potential limits of Kubernetes as demand scales up.

“there's these new workloads coming in for AI.”

Reflections on Pursuing a PhD

55:01 to 56:02

Gain insights into the value of pursuing a PhD from a personal perspective.

“But I think like anything else, when there's motivation, people go and figure it out, as long as there's not something inherent in the design.”

Career Reflections on Education and Experience

56:02 to 58:26

Learn about the speaker's journey through education and how it shaped their career.

“One story is that at one point in my career, I ran into a guy, same company, this guy who I'd went to undergrad with.”

Navigating Career Choices and Learning

58:26 to 1:00:22

Discover insights on choosing what to learn and overcoming career fears.

“I actually kind of don't care what you learn.”

The Future of Kubernetes and Software

1:00:22 to 1:03:28

Explore thoughts on the longevity of Kubernetes and the evolution of technology.

“I mean, what you're describing, it reminds me of, I don't know if you've seen that Steve Jobs commencement speech, but he literally says exactly that.”

Influential Books and Personal Development

1:03:28 to 1:05:15

Hear about key books that influenced the speaker's career and advice for others.

“Plugs are still the same shape-ish, stuff like that.”

Final Thoughts and Reflections

1:05:15 to 1:06:23

Reflect on lessons learned and the importance of note-taking in one's journey.

“Last question for you is, If you could go back to yourself when you just graduated college and give yourself some advice, what would you say?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00There's going to be an open source one. Do you want it to be ours or do you want it to be someone else's?

0:05Brendan Burns:This is Brendan Burns, the co-creator of Kubernetes, and I asked him about the stories behind building it. The hardest part actually of the project was actually articulating that. How long did that initial MVP take to build? I don't know, a little under a week maybe? He had interesting career advice grounded in his career story. I believe you can hide order 10 % of your effort from your management. Is there any upper bound where Kubernetes just cannot handle that load? Every time you change an order of magnitude, the problem moves. Here's the full episode.

0:43Brendan Burns:I mean, let's start with Kubernetes because that's super interesting. I don't fully understand the business motivation. Like, let's say I was your director or something like that. And you came to me with this and you said, hey, let's do this for everyone. I don't fully understand what would be in that strategy doc or that, what would you say in that meeting that would say, here's the impact for Google if we invest in building this for the industry? Yeah, it's funny because like the hardest part actually of the project, I would say in those early days was actually articulating that. And I think it was really clear in our heads, but like figuring out how to convince people was tricky.

1:26and you know I think there was there were a variety of different ways that we articulated why it was important one of them was related to the MapReduce white paper right so MapReduce at the time especially like Hadoop and big data were were a big deal I think that you know other things have kind of replaced them at this point but like MapReduce was a big deal and the big data revolution or whatever they called it. And, you know, Google had written the original white paper. But Hadoop was an open source project that Google had nothing to do with and got no credit for. And they just read the white paper and they re-implemented it.

2:11And it's not the same. It's similar, but it's not the same, right? Because anytime you re-implement something, it's not the same. And so part of the argument was like, look, like we have this cloud and we want to like be influencing the technological landscape. If all we do is kick out white papers, we're not like, if it doesn't run, if it's not something that people can run, we're not going to be in the driver's seat. And so that was one of the arguments. I think, you know, the other couple of arguments were like, why containers? Why not? Why? why people are using VMs, why containers? And a lot of that was talking about like, look, the demands of writing software.

2:52I mean, we know internally, we know from doing this internally that the demands of writing reliable software necessitate having systems that are sort of like autopilots for your application. And we know that this is something that as software becomes more and more critical to more and more businesses, this is something that they're just going to have to have. And so that's sort of the why containers part. And then I think that the third part was the why open source. And in some states, that's the most interesting conversation because people are like, wow, this would be, you've convinced me. You've convinced me we should build it.

3:30We should make it available to the world. It's something that they're going to find useful. But wouldn't it be so much better if they could only use it on our platform? And you're like, yeah, that's absolutely the case, but you can't win if you make it only on your platform.

3:47Brendan Burns:Why is that? Well, because there's other platforms out there, right? And so if you make it an exclusive, then the people who for other reasons are on other clouds or on on premise, they're shut out. And so they're going to just go build an alternative, right? Like the open, it's sort of like the Linux, you know, the reason Linux won, right? It's because Linux could go anywhere, right? The whole reason that open tech and open ecosystems win is because you, the majority of people are going to be not on your platform. Like if you're, if you're not the leader and we, you know, GCP was not the leader, then the majority of people are not going to be on your platform.

4:32And so if you make it such that the majority of people can't use your thing, they're just going to ignore you and then they're going to go build their own. Right. Whereas if you go and build it and you build it for everybody, but you make sure that it's awesome on your platform, then you have a chance of attracting people, attracting more people to your platform than otherwise. And then, I mean, also in some ways, it's just like the aesthetic of the time also. Right. It's like if everybody's using Linux and everybody's using Docker and everybody's using these programming languages that are all open, like if everything else is open source, you don't want to be the thing that's not right and and there's only a few places where like where that hasn't been true historically in technology where you could be different and still succeed and you have to be so differentiated that and i don't think that we were different that differentiated you have to be so differentiated such that people are like oh actually i want that thing so bad i don't care

5:25Brendan Burns:that it's closed so at the time just thinking what was the competitive landscape i guess if i if i remember it was, I mean, AWS was dominant. They were there first and doing very well. GCP at that time, probably an up and coming company, or I guess offering. And so my understanding then that the idea is let's pull market share away by giving kind of open source distributing Kubernetes to more and more developers. And then they'll be more open to kind of migrating or using GCP because you'll make a they'll pay attention right they'll pay attention and also like you change the dialogue too right like taillight chasing is hard right like if if someone else has built vms and everybody's using vms and all you're doing is saying like well we're building the same thing that they have only maybe it's a little bit better because we do something else over here or you know whatever like that's a hard market strategy to articulate but if you create a brand new playing field where you are the thought leader um suddenly like people are listening to you even if they're not using it on your platform they're listening to you and that gives you way more voice you get you get a lot more voice in the market it changes the narrative it changes who people are listening to um and so that control of of the story is an important aspect of of of how i think how you break through that dynamic or you try to break through that dynamic obviously like, you know, they're still in third at this point.

6:55So like, didn't work out. But I mean, I wouldn't say it didn't work out. I think it worked out in general, but it's still hard, like overcoming those kinds of market dynamics is hard. And, you know, I think the other thing that happened is everybody in the cloud consolidated around it. And so now Kubernetes is just sort of a utility everywhere.

7:12Brendan Burns:You mentioned that, I guess, that perception benefit of being the dominant offering, which, I mean, if you look at what happened, in hindsight, it makes a lot of sense. I mean, that is what happened, and it's a wonderful benefit. But I guess when you were looking forward and you were talking to leadership, were you cognizant of those benefits and saying, we need to do this because we're going to kind of become the dominant offering, and it's going to have all of these optics benefits? I mean, I think we absolutely wanted to make sure that we were front and center in terms of thought leadership.

7:48And we definitely articulated that. Right. Absolutely. Right. Like being being in a thought leadership position is value is valuable.

7:56Brendan Burns:it's interesting because um i feel like it's hard to quantify i mean if i was in that meeting and trying to make a call yeah how much is that worth yeah well i think you have to also realize though at the time like it was pretty cheap right it was like eight or nine engineers and we kind of and this is in some sense i mean it's both a blessing and a curse like part of the reason why we articulated and argued for having such a distinct brand where the kubernetes brand was separated from the Google brand was that it kind of gave us freedom to fail also. It was like, hey, if these eight people go off and, you know, do something and it turns out to be stupid, like, we'll just kill it off.

8:37And it won't have, like, it won't, you know, it won't, it won't damage the broader perception of the cloud. And so I think there's that benefit of the open source part of it too. It helped with adoption, right? Like, it helped, it helped us. And especially as we went to like the Linux foundation and things like that and truly established like an independent entity, it helped ensure that, you know, people like Red Hat or Azure or AWS could take a bet on Kubernetes and feel confident in that bet. But it also was an insurance policy against failure. And also, to be honest, like it was it simplified a lot of things, too, because, you know, we were competing against startups, right?

9:13Like at the time, Docker is a startup. They can be way more agile than a big company. And so by virtue of sort of like being a separate entity, we could be a little bit more agile also.

9:25Brendan Burns:I mean, the earliest conception of this project was you and two others kind of hacking something together. Yeah, it was a demo. I mean, it was sort of a demo almost. It was like, look at what we can do if we just smash a bunch of existing open source tech together. What did that demo do? I mean, it was basically like a basic cube control. it was like hey here's like here's a container i built i mean at the time you had to explain docker to people you were like hey here's docker i used it to like build this container image um and and then you could run it and deploy it and see that it had gotten distributed across a bunch of machines and that you could you know load balance to it because you you'd hit a single endpoint and it would go i'm replica one and then you'd hit reload and go i'm replica three and you So it showed that it was replicated and then basic health checking.

10:15So if you killed it, it would come back. And a V1 to V2 upgrade. That was about it.

10:20Brendan Burns:How long did that initial MVP take to build? I wrote it in, I don't know, a little under a week maybe. Something like that. I mean, I don't work on the weekends, so maybe five days, four days, five days. And did you drop all of your existing work? Cause I'm imagining you're, you had existing project work and this is kind of extra credit stuff that you were working on. Yeah. Well, I mean, I wouldn't say I, I dropped it, but like in, in a timescale of that timescale, like you can kind of like slack on it a little bit, you know, like, like you could be sick for a week. I mean, I guess the thing is like, you could be sick for a week, you know, and I'm not saying that that's not what we did, that that's not what I did.

11:03But there was enough flexibility in the system that you could hack it together. And I mean, believe me, it was hacked together, right? Like every possible shortcut to take. And I think one of the things I've been good at historically is integrating other open source projects together, seeing how you can take stuff off the shelf and put it together. And so, you know, a lot of the nuts and bolts were pieces that we could take from other open source projects and kind of combine together with glue and glue code to give the feel of it. So that helps too.

11:43Brendan Burns:I think a lot of software engineers, when they hear this kind of story, they think, oh, I have my existing responsibilities and I can't necessarily go off and build this thing, even though I think it's a great idea. Do you have any advice for someone who has that opinion? Yeah. Well, I mean, I think that what I would say there's two things to I have two answers to that. One is advice that I've always given to every single person that's ever worked for me, which is I believe you can hide order 10 percent of your effort from your management. right so like you know there's there you have slack you have the ability to slack no matter what right um and you know as you get a bigger and bigger org actually the percent that what you can do with that 10 actually increases um and a lot of really of really influential good ideas that i've had have come out of that i mean it's another i mean it's kind of a flip way of saying I want to empower people locally to make local decisions that they think are optimal for the business without having to consult up the chain, without having to ask permission, right?

12:53And you tell people while, when you tell them that you're like, by the way, you're also going to make a bunch of bad decisions and you're going to waste a bunch of time. And so like, you have to be comfortable with this idea of like, I'm going to try some ideas. Some of them are going to fail. Some of them are going to succeed. When I look back retrospectively, the ones that have failed were effectively a waste of time. You know, and it might be the difference between an exceeded expectations and a met expectations, right? Like you don't want to drop below meets expectations, but like it might be the difference between an exceeded expectations and a meets expectations.

13:25And you have to be comfortable with the notion that you're going to bet five times and the payout from one of them hitting is going to be way better than the grinding to get that exceeds every single time. And, you know, I mean, I think there's equally valid paths where you don't do that, right? And I think you have to be the right kind of person who's willing to take that kind of chance. And that's not everybody, and that's okay. and I think the other side of it that I always remind everybody also though is like you know people say things like that to me sometimes and I'm like so do you play Call of Duty?

14:03Do you watch Netflix? Do you watch YouTube? You know because like I'm pretty sure there's probably 10, 15, 20 hours in your week at least when you're doing something that's not work right and I can tell you in that time period I wasn't doing anything except for this and work, right? And a little bit of family, like, and sleeping, right? And so, like, sometimes it's also about saying, like, well, what are you willing to give up, you know, to have the space to do that? And, you know, I'm not a big, like, I'm not a big, like, work all night kind of person, but, like, it does mean, like, maybe not going to watch YouTube for a while, maybe not going to watch, you know, sporting events for a while.

14:44Brendan Burns:That makes sense. And I mean, on that second point, in this case i mean it was such a the returns of this project were exponential obscene i mean if you put in two times the the time for a year but you you get 20 times the the impact so it just kind of makes sense in terms of investment of time well and also i think i mean i found personally that like the it was addictive right i mean i think i benefited from two things one is i really like to write code right like i i enjoy it as an activity um and uh and so like if i'm choosing between netflix and coding i'm actually pretty happy just coding you know so that's a benefit um for me um and then for me anyway like once people start using it and they're excited about the project and they're putting issues on GitHub and all this kind of stuff, like, I'm just addicted to it.

15:43Like, I just want to close that issue. I want to help that person out. I want to like, I'm going to, I'm going to go till I'm falling asleep on my laptop. And, and that's just because that's, I enjoy that. Like, that's why I'm in the industry. So even in the moment, I wasn't, I mean, I was definitely not thinking about like, oh, here's this payout that I'm going to get for the rest of my career. I was definitely in the like wow I just want to keep this thing going I want to keep this rush going

16:13Brendan Burns:because it sounds like it took a while for you guys to get buy in from leadership I think there was a solid six months of going from a very hacky prototype to something that like legitimately we thought somebody could take and use and laying down the right kind of groundwork for that you know There's a lot of little details that you have to get right along the way. And it's always nice also because, you know, a lot of the people who we brought in in the early days had built similar systems before. And so they were having this opportunity, this kind of clean. It's rare in your life as an engineer to get a clean room opportunity to rebuild something that you have ideas about how it could be better.

17:01It's like getting a second chance, you know. And so that was also, I think, really attractive to people because it's suddenly like, oh, wow, this is a clean room. We don't have any users right now. So we don't have to be fixing bugs because some big company who pays us a lot of money is asking for something or whatever. We've got this clean room time. And we've all spent a lot of time thinking about what the system could look like. And so now we get to go and just build the thing that we imagined.

17:33Brendan Burns:You mentioned that first point on hiding maybe 10 % of your bandwidth from management. And I mean, that's super interesting to me. What does that look like in practice? Well, I mean, I think that what it means is that you should always have sort of like a side project that you think is relevant. Right? Like, you should always have something that nobody told you to do, but that you think is important that you're working on. i see so it's kind of like i remember google i don't know if they still do this but at the time there was a 20 time yeah it's similar to that kind of idea yeah yeah exactly but okay so when you say when you say hide you don't mean how you say manage expectations so that your manager is also okay with you working on another project oh no i actually do mean hide right like don't ask permission right like sooner or later you'll show it to them but like it's pretty like you need a solid, I don't know, a couple months or whatever to get to like a place where it's something you could show to somebody.

18:34Right. And so like, it's all about saying like, I'm not going to ask permission. I'm going to go build something that I think is important and useful. And then obviously when it comes time to launch it or whatever, or put it out there, like then you do have to ask permission. Right. And, and so then, yeah, you say like, Hey, I built this thing, but, but you've had that time to like get it from because i i mean i don't know i feel like it's hard to articulate the value of something like in a doc or in a powerpoint it's way more effective if it's like a running thing that you can like somebody can interact with right so getting that time to basically build it up into something that's real and could ship because also like in some sense your manager is always assessing like, well, should you spend your time on that?

19:25Or should you spend your time on this? Right. And by building it, you kind of like force their hand because it's no longer like, should you spend time building this? Or should I spend time building this? It's like, I've already built this. Like, do you want to ship it? And that's in some ways a much easier decision. Like, I don't know about easier, but like, it's like, it's not an either or it suddenly becomes just sort of about like, is your idea good? Right. Because the work is done.

19:51Brendan Burns:I mean, in the happy path, I feel like it's a great idea. You launch this thing, has impact. It's great. But what about in the case that you work on this thing and you didn't tell anyone about it and then no one cares when you launch it or it's not as good as we thought? yeah and and then like that's the flip side right like you have to be comfortable i mean as i said like you have to be comfortable with that idea that um you know you're gonna waste some time uh and maybe there'll be waste time in the sense of like wow i could have been like watching netflix or you know whatever uh and i instead i wrote a bunch of code that nobody liked um it could be wasting time in the sense of like wow i could have you know it could have gotten promoted i could have because I thought this was the great idea that was going to get me over the hump, but it wasn't.

20:42And I think you just have to be comfortable with that. Like, you know, it's taking, I mean, it is taking a risk. It's not unlike in some sense, like doing a startup or something like that. It's taking a risk. So, you know, I think, and you can't assume, I think sometimes people are like, go into any of these sorts of things and they're like, oh, I will have the idea and it will be amazing. And then it will hockey stick. And like, I think if that's your mindset, when you go in, you're probably setting yourself up first in disappointment. You know, you have to go in with that mindset of like, I think this is good.

21:13I'm going to try, but I'm okay if I fail. Right? Like, I know that I'm making an explicit choice here.

21:19Brendan Burns:I also imagine at some point, at the highest levels of engineering ladders, you need to take that risk to get promoted to higher levels. For instance, like if you're a staff engineer or senior staff engineer and you ask your manager, how do I get promoted? They'll often tell you you need to figure out what that project is because I don't I can't just hand this to you because it's starting to become more ambiguous. Oh, yeah, for sure. Yeah, I could totally see that this becomes a necessity at some point if you want to kind of grow. And I mean, this this project also got you promoted as well at Google, right?

21:57Brendan Burns:From from staff at some point. Yeah, no, I mean, certainly my career absolutely benefited from the success there. Absolutely. Right. But, and I think you're right that like, there's also this aspect of, at a certain point, like, you're just expected to be the person who knows enough to come up with the really good ideas. And that's just the expectation. Like, it's no longer about like, can you execute on the ideas that other people give you? You know, and that's a big part of it as well. And I think there's also like, I mean, I was, the other thing I tell people sometimes when they're thinking about getting promoted is, if you create the idea yourself entirely, it's like blindingly obvious that it was you who did it.

22:39right? If you succeed and have impact in something that is a bigger project or someone else's idea, you can still have a very successful career, but it's a little bit harder for it to like, be directly attributable to you, right? And, but again, I mean, I think that, again, it's a roll of the dice, right? And so like, at some level, like, there are probably your people out there who've tried over and over and over again and just have never had the right idea, right? Or just had the right idea at the wrong time. I mean, I think one of the other things that is interesting about innovation that is disruptive is that it's a combination of being the person who has the idea and being in a time in which the idea can take off, right?

23:26And so you could have the idea, but it could just be the wrong time. And it won't do the same, and it won't go in the same direction.

23:33Brendan Burns:That point on promotions, I think, I mean, if you create the scope, not only is it obvious that the credit should go to you, but it also feels kind of permissionless. Like you don't need to wait for, I guess, management or someone to give you the opportunity. You can kind of create it. And so you have a lot more control in that process. One thing on the business strategy, I guess before we leave that for Kubernetes, is at that time, Borg to me felt like a competitive advantage for Google, like some secret infrastructure sauce. I would have thought people would be kind of worried about giving away any part of that to the industry.

24:15Brendan Burns:So what was the thinking there? How did you convince people that, hey, it's okay? Yeah, I mean, I think that there was a little bit of that. and I think to be a little bit make it into a little bit of a joke or whatever one of the things I said to people was it's not like you men in black flash people as they leave Google it's not like everybody comes to Google and it just stays there forever and in fact as we talk to people at Facebook, as we talk to people at Twitter we talk to people at other scale out tech companies they were all building this stuff it wasn't really a secret and there was also I mean, like Mesos at the time was, you know, not the same, but similar.

Read the full transcript

24:57And like, you could just see that there was going to be an open solution. And so in some sense, part of the argument was like, look, there's going to be an open source solution. Do you want it to be one that we can influence or not? It's not like, do you want there to not to be an open source one or not? It's there's going to be an open source one. Do you want it to be ours or do you want it to be someone else's? And it's just reframing the choice, right? and making it clear that people understand that that is the choice, right? That you don't get to choose the proprietary option because it's just not viable.

25:28Brendan Burns:When you were building that MVP for the orchestrator, well, how'd you decide? Because there were no customers or anything like that. So how'd you decide this is the minimum set of features that we need for this to, before we launch it? Sure. Well, I mean, I think absolutely, you know, we benefit from the fact that there were three of us working on it, right? So Craig was a great product manager. and joe was a great engineer and and and fantastic at api design um and and i could write code fast basically i think is sort of the you don't think if i if i had to sort of stereotype all of us uh you know that's the like craig was the product business guy and joe is the like i know how to design really good at design kind of person and and i was basically like i can hack prototypes like there's no tomorrow.

26:18And I think we reflected a lot about our own experiences. And also, we had seen the pain of people deploying into traditional VM infrastructure. And so we had that kind of knowledge of the pain that people are going through. And then, you know, at the time, there were people like Netflix who were talking about immutable infrastructure, and they were kind of advancing some similar concepts. And so there was also kind of like a broader movement happening that we were taking part in. And so, you know, obviously there were literally no customers, but in some sense there were customers. They just weren't customers yet.

27:03So it's not like creating something brand new, I guess, in some sense. Like I feel, I really don't feel like Kubernetes was something that was brand new. I feel like it was a coalescing of a lot of ideas that were kind of circulating in the industry at the time. And it just became an anchor point and a really good expression of those ideas.

27:23Brendan Burns:So when you talk about, I guess, you wrote a lot of code quickly. Did you write most of the code, I guess, for this initial MVP? Yeah, I don't know what the number is, but high 80s percentage, maybe more of the original code. and I think I'm still number I mean I haven't contributed significantly to Kubernetes in a while and I'm still I think number five on the overall contributor you know overall contributor list on the GitHub commits and I was number I was number one for a long time I mean after writing that much code for Kubernetes which part of this system was the hardest to build I think I'm going to say like because i don't think any of the specific code was that hard um i think that the core the hard part was the decision that we made early on that it was going to be a really loosely coupled system and so it's very um which is great for resiliency like we made this decision around very loose coupling a lot of independent actors taking actions there's all these control loops running all over the place, which is really good for resiliency.

28:36But when things go wrong, it's really hard to figure out why it went wrong. Because you've got 15 different processes that are all having to work together to achieve an outcome. And so you can see that the outcome wasn't achieved. But then you're like, okay, but what happened? And now I have to sift through a bunch of different logs and a bunch of different operations of executables and sort of reconstruct in time what happened and and especially early on we didn't have very good we didn't have very much consistency around logging we didn't have very good consistency about event like events and things like that um and so i think the hardest part i mean it's the hardest part of anything i guess but like when something went wrong figuring out like why it

29:21Brendan Burns:went wrong because your logs are all distributed everywhere and yeah and everything's out of time sync and i mean and hopefully you logged the right things but a lot of time early on like you didn't log it like you know and and and and because it's a interaction effect it's hard to reproduce um oftentimes it's oftentimes it's hard to reproduce the problem like if the problem reproduces easily then it's pretty easy to fix even if you don't have the logs because you just go and add the logs and then you do it again and you see what happens um but for problems that are transient uh because of it's just you know race condition between two or three different things happening you can go in and add the extra logs but then you have to figure out how to make it happen right um and that's that can be pretty tricky um and also i mean i'll say the other thing is like we were all learning go so maybe the other thing was like there's some gotchas in golang and we were all kind of like learning all the gotchas on the fly i would have thought there's something that's controlling everything right like there's some leader or maybe some nodes that are looking at one of them to kind of coordinate everything.

30:30Brendan Burns:Like leader election, if the leader goes down, I would have thought would be some really challenging problem. Yeah. I mean, well, I think the reason it wasn't that big a deal was because we relied pretty exclusively on etcd to do that for us. Oh, so etcd is another framework or open source? Yeah, etcd is an open source. I mean, it's kind of really part of Kubernetes now. But at the time, it's a raft-based consensus system, key value store. And so it was a pre-existing piece of software that CoreOS had written that implemented the raft protocol. Because at the time, like, because Paxos was the original for this, and Paxos is really hard to implement because the algorithm is really complicated.

31:24And nobody understands it. Well, probably somebody understands it, but not a lot of people understand it. So then, but it's provable, right? And then right around that time frame, people had come up with Raft, which is a provably correct consensus algorithm, but it's way easier to implement. And etcd implemented the Raft protocol and then gave you this consensus reliable store so that you could do it had you had multiple replicas and it would do it would do the consensus there. And it doesn't do leader election exactly, but it allows but it gives you enough primitives that doing leader election is relatively straightforward.

32:10forward. And, and, and I guess I would also say that, I mean, two things about that. One is we decided to force all of the access through an API server. So no, nobody had, and this was actually something I pushed really hard initially. But we, I think the general agreement, but I definitely pushed for it, was that nobody got to use store. Everything had to be remote store. Like nobody got to write stuff to disk themselves. Every, piece of the system had to use the API server and had to use etcd behind the API server as the way that it did any sort of persistence.

32:47Brendan Burns:What's the main benefit of that? The main benefit of that is that everybody gets to restart all the time and they just come up and they work. So you don't have to worry about corruptions. You don't have to worry about schema changes. You don't have to worry about any of the... Everybody was effectively stateless except for the database. There's the etcd consensus algorithm database. And that's the only place where there was state. And so as a result, the whole system was just a lot easier to make stable. The downside of it is it leads to this like loosely coupled, like loose coupling, right? Where it's a bunch of independent loops mediating everything through this storage layer.

33:32And which made the debugging part harder, right? So those are the sort of the trade-offs is like, if you have a complete log of like, I'm in, you know, if you think of it as a state machine, like it's much easier to understand where you are and where you got to if you're in a state machine. But state machines are a nightmare to make reliable. And so like, they're easy to debug, but they're hard to make stable. Whereas the system we built was like designed to be stable, but hard to debug. The trouble with the state machine is a state machine says the world looks like this. And unless you get it exactly right, sometimes the world looks like something you didn't imagine.

34:14And at that point, you're kind of screwed because like your state machine doesn't know what to do. Right. Whereas because we have these control loops that were based on a desired state and a current state and trying to drive the current state to the desired state, like no matter where you woke up and found yourself, you kind of knew where you were supposed to drive to and that and that's the stable part that's the stable part of it is that like it didn't really matter what where the system got itself it was always trying to drive towards uh the the desired state you know inspired honestly a lot by like control theory from robotics actually like oh like uh pid yeah the same same idea right like you know you could imagine if you tried to write a pid controller to balance a beam with a bunch of if else loops it like it just doesn't work very well.

35:05Brendan Burns:And this kind of reminds me because I was reading like in some of the design, like a large, I guess it's a feature of Kubernetes is that it's declarative rather than imperative. So you kind of just say, I want this to be true. I don't care how you get there. Instead of saying, just do X, Y, and Z. I'm curious like the pros and cons of that design decision. Well, yeah, I mean, and that was something that was happening broadly in the industry. Like that's a part of the whole like infrastructure as code movement that was happening at the time. So we're not the only ones who said that, but we definitely embraced it.

35:40You know, I mean, I think the benefit is, you know, you have clarity about the way you want the world to work. Right. It's not like if you execute a bunch of instructions, start this, run that, do this. like you've done a bunch of stuff in pursuit of some objective, but you never wrote down what the objective was, right? There's no record of like what you were trying to achieve. I'm trying to create a website. Well, you didn't write that down. You just took a bunch of steps, right? With a declarative approach, you actually write it down. You say like, I'm trying to create a reliable website and here's what a reliable website looks like.

36:19Hey system, could you take the steps to get there right and so you have that record um and it obviously makes it easier for things to be self-healing because if you've written it down now i know where i'm supposed to go back like if i get perturbed from that state if something fails or something restarts well i know where i'm supposed to go back to um and similarly like it has side benefits of like once I write it down, well, I can apply code review to it. I can apply unit tests to it. There's a lot of the mechanics of how we do software development that apply once you write down that declaration.

37:00So a lot of those are the benefits. I think the downside is probably just complexity,

37:05Brendan Burns:right? In comparison to going click-a-click through a wizard or whatever, you know, in a GUI, learning and, you know, everybody complains about the YAML and I have to learn all this stuff. And, you know, like it does introduce a learning curve. Now, I think fortunately at this point, there's enough education material out there that it's not, and Gen AI too, for that matter, that it's not that bad a learning curve. But certainly in comparison to what people have done before, that's probably the biggest downside. But I don't know. I think that the upsides are like up here and the downsides are like way down.

37:42Like there's there's not a lot of downside. I see.

37:45Brendan Burns:Yeah, I could also see that being helpful for I guess if you want to optimize anything under the hood, because you're just making a promise to people that this is going to happen. But if you want to do it in a more efficient way or something like that, then I guess it just gives you all the power to do so. Yeah, well, I mean, and that does make things like machines failing a lot easier, right? Because people don't say run this on this machine. They just say, I want three of this to be somewhere. And if a machine fails, well, just move somewhere else, right? Because the application isn't tied. Because I can't, in some ways, I don't know your intent, right?

38:20If you log into a machine, and you start a process on that machine, is it because you wanted a process? Or is it because you wanted a process on that machine? I don't know. You didn't tell me. And so if that machine fails, what should I do? Well, I don't know. But if I know that you said, hey, I just want three replicas, well, then I know it doesn't matter that it's on that machine. It could be on a different machine. You'd be just as happy.

38:46Brendan Burns:I want to shift a little bit to kind of when Kubernetes was scaling. And it sounds like a a large part of this was getting buy-in from other companies and other people. And so, you know, how, how did you get the buy-in from, I know OpenShift was an important part of, yeah, Red Hat, and other companies that joined on, like how did you sell those companies that Kubernetes is what you want to use? Well, I think for a lot of them, especially in the early days, it was kind of that quote around undifferentiated heavy lifting, right? Like they had some other objective, Like OpenShift was trying to build a platform as a service.

39:27Or, you know, they were, you know, a lot of our early users who were also contributors, you know, they were trying to build some sort of reliable web service or something like that. Right. And so it was like, well, we're going to have to build this thing anyway. Why don't we all build it together? And we don't really care because we don't think that's our value. Like we don't think our value is tied up in that layer. So we'll go contribute to your thing because we're going to get more value out of the collective. than out of trying to do it ourselves. And so for a lot of the early partners, that was a big part of the argument was like, hey, we'll let you in.

40:03And part of that is making sure that they understand that they're going to be equal partners. Because it's one thing to take a dependency on something, but then you're kind of like taking dependency on someone else's roadmap. And so it was really important also to say, hey, like, you can come take a dependency on us, but also we'll give you a seat at the design table. So when you need new features, you can, you know, contribute those features. And here's what we're trying to achieve. And it matches up with your roadmap and, you know, that kind of stuff. So I think that's how we approached it. And then over time, you know, people became more and more interested in being part of it because there was a growing ecosystem.

40:46So when you look at like networking providers or storage providers, you know, as their users were starting to become Kubernetes users, they were motivated to make sure that their networking system worked well with Kubernetes or their storage system worked well with Kubernetes and things like that. So that was sort of a secondary layer of partner discussions that we had.

41:06Brendan Burns:Right. And that's downstream of becoming the dominant player. And I guess that's validating the open source strategy, which is you become dominant, everyone's kind of got to integrate with you and all of that. How did you prevent Google from dominating in the roadmap or I guess controlling what Kubernetes would be, given that it started at Google, was funded largely by Google? Yeah, I mean, I think that was really important. And I think it was a critical part of gaining adoption. and becoming the industry standard was giving it independence. And I think there's two pieces to that. The first is getting it to foundation.

41:47So the creation, it was only a year in that we created the Cloud Native Compute Foundation, that we donated all of Kubernetes to the Cloud Native Compute Foundation. And so getting the project, the logos, all of the legal stuff, trademarks, all of that stuff into an independent software foundation with the Linux Foundation was critical, right? Because it's hard to partner if somebody else has trademarks on the Kubernetes logo or whatever. And then I think the other piece that was important that came a little bit later was writing down the governance rules. So, you know, for the first time, for a few years, Kubernetes didn't have any really governance rules written down.

42:37It was a mistake, I would say, right? Like we didn't realize how, it was really, it was something we should have done earlier, but we didn't. And so we sat down in 2016 to write the governance rules. and I think all of us were aligned on this idea that we didn't want any one company to be able to take control of the community and we really built the community and the rules of governance to be democratic I think that's an aesthetic from Craig and Joe and I we never set out to be a benevolent dictator for lifestyle project we always set out to be a distributed ownership, democratic kind of project.

43:24And we codified a lot of that into the governance docs that continue to this day. So I think both those things together really helped make sure that it was an industry standard and not any one particular company standard. But also, I think we're critical to a success. I think they're duels of each other, right? You can't have one without the other.

43:45Brendan Burns:people wouldn't have come on if they didn't see that it was governed well and open. Yeah. I mean, because obviously, like, you're thinking about adopting it, or you're thinking about, you know, putting it in your service, like, the thing you're worried about is, like, whose roadmap am I betting on? When you said governance, does that, I mean, is that literally like, like, when I think of government, like, there's a constitution somewhere? That's literally what we wrote. Did you write that yourself? Or is that something that, like, lawyers do? No, no, we wrote it ourselves. Yeah. In the span of about, it was a couple of three, a couple of three fairly intense meetings amongst like six or seven of us.

44:25We got together and just kind of talked it through and looked at a bunch of other communities and kind of like what had worked, what had not worked, what were we worried about, what were we trying to achieve. some of it was codifying stuff that already existed. So we had some loose organization stuff that already existed in sort of a de facto way, but didn't exist in an explicit way. Some of it was, you know, we created the steering committee that had never existed before. Right. And we just basically, and we were lucky, I think that we were able to gather. So the people that came together, We called it the bootstrap committee.

45:13You know, we were lucky in the sense that we had enough people who kind of were not, who the entire community would look at as being leaders. And they weren't, we weren't fighting with each other. You know, we weren't infighting. So we were all kind of aligned. And we kind of got everybody. So we didn't have to be like, oh, you know, we grabbed this side and not that side. We were able, and it was like seven people, I think, seven or eight people, we were able to pull together a group of people that really represented everybody and that everybody kind of all respected each other and respected each other as leaders in the space.

45:52And a lot of credit there, I think, goes to, I mean, everybody who is involved deserves a lot of credit. But Sarah Novotny, who is our community leader at the time, deserves, I think, a ton of credit for bringing that thing together.

46:07Brendan Burns:When you look back on Kubernetes, with an open source project, there's obviously the read aspect, which is everyone can use and duplicate this code and execute it. But there's also, I guess, the writing part, which is people making contributions. What percent of the contributions actually come from the community and what percent is actually just the main stakeholder companies just putting in their code? Yeah, I don't have the specific numbers for Kubernetes, but my experience in open source says it's like 80, 90 percent the core contributors and like less than 10 percent to other people. It's hard, I think, in general.

46:48It's really hard to get people to contribute. Part of it is companies, honestly. Right. Like, you know, companies like Microsoft, we make a commitment to contributing to open source. And so, you know, we, at a leadership level, we've decided that this is something that we want to invest in. And so we're willing to have teams of people who specialize in working in upstream open source projects. But for a lot of users of Kubernetes, you know, they're a retailer or they're a banking industry or they're like, it's tech isn't their core thing that they're doing. Tech is a means to an end to deliver an app for their user.

47:25and in that world it's pretty hard to justify well i'm going to take 10 of my people and i'm just going to do upstream open source contributions right um and especially if the leadership is like not a technical leadership and so they didn't necessarily grow up in those communities and if you grow up in finance it's hard to explain like what's the value of contributing to the i mean the value of taking the open source is very clear right it's free um but the value of contributing back, it's harder to explain. Or legally, we ran into people, even today, it's getting better, I think. But even today, I've run into people who say, we would really love to contribute.

48:04Our engineering leadership is aligned that we would want to contribute. But our legal team is worried that if we contribute, we'll be liable if we introduce bugs. Would someone sue them? Yeah, I think that's what they're worried about. It doesn't hold water legally. And I think the linux foundation can give you plenty of like case law and things like that to show why it doesn't hold water um but sometimes that's enough to block uh someone from contributing right i never

48:37Brendan Burns:thought someone would get sued for adding a bug i mean everyone adds bugs on accident well but i mean but on the other hand like if you write a proprietary piece of software and you sell it to somebody and it has a bug and it causes your house to burn down, like you can imagine you're going to sue the people, right? So like, it is sort of a legit worry at some level of like, if I wrote the JPEG, open source JPEG library that ended up in the smoke detector that caused the house, you know, like you can sort of imagine the chain of logic that gets you there. I don't think it's true. I don't think it would hold up.

49:11I think a lot of the licenses, you know, a lot of the open source licenses include indemnification language that basically says, if you use this, you're using it under your own risk and like, you can't sue us if it burns down your house. But I think that that's a worry for, I mean, well, I've heard, I don't think I've heard from people that their companies do have that worry. And again, in some sense, it's because they're like, well, what's the value? If I see this potential risk and I don't necessarily see the value and i mean and like also like again if it's not a core thing if you're not a tech company you know is that developer really capable of like arguing with legal arguing with legal about what you can and can't do right probably not right like they're probably just going to fade away right um

49:58Brendan Burns:so you know there's that aspect too i remember this is many years ago i read this blog post that open ai put out before i think open ai was kind of huge and it says here's how we scaled Kubernetes to 7 ,500 nodes or something crazy. I forgot exactly how. Yeah, yeah. And so I want to know, there's these new workloads coming in for AI. There's training, which is this huge, I guess, all at once workload. And then there's inference, which is latency sensitive and you kind of need it to come out instantly. How has Kubernetes adjusted over the years to handle these kinds of workloads? Yeah, I mean, I remember when, you know, like we couldn't really handle more than about 100 nodes.

50:49So it's definitely been a lot of optimization in the core systems. And there's places where the APIs were pretty noisy and we needed to reduce the noise level or we needed to extract components into another component it so that you could scale etcd in particular actually is one of the can be the main bottleneck to that kind of scale and so figuring out how to run etcd really well is a core part of figuring out how to run kubernetes really well at scale um i mean i don't think it's that different than like learning how to run a database or anything else like that at scale like large scale is just weird and you just have to you know run it see where it breaks figure out how to fix it rinse and repeat, you know?

51:39And I do think what's interesting is that while, you know, AI training as an example is, is a really large, large cluster, large scale kind of thing. You know, I think by virtue of being in the cloud, a lot of our users actually have much, much smaller clusters, but lots of them, right? So hundreds or thousands of clusters where each cluster itself is a little bit smaller. And I think that's not something we anticipated because we came from a world of like physical data centers where, you know, you only want one because like you don't want to have to set it up a bunch of times, right? You just want to set one up for the entire data center, call it a day.

52:17But because of the cloud, because AKS, you know, you press a button pops up in two minutes, right? Like it's really easy to get yourself a cluster. So people create lots of clusters. And, and so I think we've also invested a lot in the communities community and in Azure as well on managing lots of clusters. How do I manage clusters at scale? I think one of the jokes we spent a lot of time talking about containers is replacing Snowflake servers, not Snowflake the company, but specially handcrafted servers. And now we just have a bunch of Snowflake clusters. So the VMs all look the same, but the clusters are all weird.

52:59So we have to provide people with tools to make sure that the monitoring software is the same on all of them and that the you know all of the kubernetes versions are the same and like you know all this admin users are the same and like all this kind of stuff right um so that's another aspect of scale out that i think we didn't anticipate that we had to go and and build which is number of clusters as

53:20Brendan Burns:opposed to size of cluster i always hear in the news that i mean this the anticipated scale is even higher than today's unprecedented scale. And I see people are purchasing GPUs like crazy. I'm curious, is there any upper bound where Kubernetes just cannot handle that load? Like let's say you 10x it from where it is today. Is it going to break down at some point and you need something more custom? Well, I mean, I think it all comes down, it all comes back to the storage layer Because everything, again, because there was this design decision that everything routes around the storage layer. Everything else is basically horizontally scalable.

54:08So you have more API requests coming in from more nodes. Well, you just need more API servers. You know, you want to do scheduling faster. Well, you need to just have more schedulers. So everything else, more or less, you can just horizontally scale out. it's the, it's, it's the storage layer that is the, is the, the bottleneck. Um, and that's where the work comes. And so you want to say, go 10 X up. Well, you're going to have to probably figure out, um, if you can make at CD scale that way, or if you need to replace at CD with something else that has the same characteristics, but can operate at scale.

54:44Um, so, so I don't think there's anything like inherent in the design that would prevent it. Um, but obviously that, you know, there's a famous quote that that every time you change an order of magnitude the problem moves um and so i think that's really true is every time you increase by an order of magnitude what you thought was the main problem is going to become easy and then like the problem moves somewhere else so you were network constrained now your cpu constrained you were memory constrained

55:10Brendan Burns:now your network constraint yeah that'll be cool to watch how because it seems like everyone wants to yeah i think it's definitely it's definitely the case that people continue to try and push the limits of scale.

55:23But I think like anything else, when there's motivation, people go and figure it out, as long as there's not something inherent in the design.

55:31Brendan Burns:The last part of this conversation, I just wanted to reflect over your career a little bit, maybe ask you a few questions about things. You mentioned that you had a PhD in robotics, and I hear a lot of people say they don't recommend PhDs. Some people do. I'm curious what your take on on getting a PhD is? Yeah, that's probably like, like, if I had to have a top 10 questions or top five questions that people ask me, that's definitely in the top five, top 10 questions. And I guess I'll answer it with two different stories. One story is that at one point in my career, I ran into a guy, same company, this guy who I'd went to undergrad with.

56:13And he'd gone off and done startups and done the tech industry thing and ended up in the same company that I was working at. And I'd gone off and done my PhD and come back into the industry. And we were at the exact same level. We graduated the exact same year, same degree, and we were at the exact same level in this company. And so I guess that's one way of me saying like, eh, you know, like it probably doesn't matter. It probably doesn't matter one way or the other. but I'll also turn it around and say A, I had a lot of fun I had a lot of fun doing a PhD in robotics so that's worth it to me and then two, I think I learned a lot about how to I think from the PhD and my PhD advisor I learned a lot about how to write and put my ideas forth in both written and presentation form that you don't necessarily learn in the industry.

57:10And I think that benefited me.

57:12Brendan Burns:I think I benefited my ability to argue. We talked about that six-month period where we were arguing for why we should be allowed to open source this thing. I think the skills I learned in terms of presentation and in terms of writing benefited me during that time. And have continued to benefit me. and then I think when I went out as a professor for a couple years teaching CS101 and having to explain stuff to students who didn't really know anything about computers I think really helped me organize the initial parts of the Kubernetes project so that somebody could learn about Kubernetes because people were coming in and being like what is a container?

57:54What is orchestration? How do I do this? There's a lot of just teaching that you had to do and I think that experience as a professor thinking about how do I teach students something really helped me do a good job with teaching Kubernetes to people. And so I think those things were really beneficial. And so I guess there's the two different arguments, which is one is like, it doesn't matter. The other is I learned a ton of stuff that I think was pretty useful to my career. And I had some fun.

58:24Brendan Burns:Earlier, you said top questions that people ask you and I'm kind of curious what's the top question I think the one the other one that I get a lot um is uh how do I know what I should learn like a lot of especially when I talk to the interns or the first couple years you know a lot of the questions are right revolving around like AI seems really hot right now but I'm really interested in systems like should I like go learn AI because it's hot or should I learn systems because I think systems are interesting. I actually kind of don't care what you learn. I care that you're learning. And so the most important thing is to find something that you're excited about and energized about because, you know, you'll do that instead of watching YouTube.

59:10You know, so if you're not excited about AI, like, well, you're probably not going to do a very good job learning AI, which means that you're kind of going to waste your time. But if you're really excited about systems you'll probably put a lot of passion and energy into it and you know we still need systems engineers like so um you know i think that's the that's kind of a pretty popular question i think there's a lot of i i sense anyway a lot of like fear of making the wrong decision and um i always tell everybody like there was no plan i've never had a plan for my career like never ever ever like i've always just chased after things that i thought were useful and were fun and interesting um and uh and you know obviously like that can work out badly for people i'm sure it's good to have a plan probably for some people but like i also want to make sure people understand that like when you look back sometimes the things you think were mistakes or dead ends like actually were critical things that taught you stuff um and so like worrying about did i choose the wrong thing?

1:00:17Am I going to choose the wrong thing? As long as you're learning, you're probably doing okay.

1:00:22Brendan Burns:I mean, what you're describing, it reminds me of, I don't know if you've seen that Steve Jobs commencement speech, but he literally says exactly that. You wrote down somewhere, I don't fully remember where you wrote this down, but I have this in my notes. It says, the inevitable trajectory of software is death. And I just can't imagine Kubernetes dying, but How do you see that happening? And what do you think about that if it did? I mean, I definitely stick by that statement. Although I think that the sentence before I said that was, you really should never fall in love with your software because the inevitable trajectory of software is death.

1:01:07Which means don't stick with it past when it's dying. right um like you should always be willing to throw away stuff just don't stick with it just because you wrote it you should always be willing to throw it away but like obviously i think if you look historically across the industry it's it's true too right like like um and and quite frankly like even within kubernetes like the source code that i wrote has been rewritten a number of times um over the 10 plus years history of the project um so what does it look like i mean i think it looks like uh something coming along that achieves similar things but easier with more uh you know with less complexity and and more utility um and and i think that you know i can imagine what that looks like like i think some of these natural language stuff if you could actually really get it to be an interface that worked 100 of the time like obviously it's way easier to come in and say I would like a reliable web service than it is to say yamely yamely yamely yamely yamely you know I think sometimes I think it's sort of two different trajectories like sometimes the trajectory is it goes away sometimes it just becomes so hidden that nobody sees it right like underneath Linux there's I mean excuse me underneath Kubernetes there's Linux and underneath Linux there's a processor but you know people don't pay much attention to that and there's a lot of attention now on AI and underneath a lot of the AI is Linux but it could be that people focus so much on the AI that they forget about the Kubernetes part.

1:02:41And I think that's happening already, honestly. Like, I feel like if I look at the volume of changes and things like that, like, I think it's, you know, I think we've sort of plateaued in terms of like the amount of change that's driving through the system. Stuff needed to support AI is kind of like the exception to that category. But, you know, I'd be shocked. I mean, I guess Let's put it this way. Let me take the long view and say, in 100 years, is Kubernetes still going to be running? I'd be pretty surprised. Right? No way. It's hard to imagine, right, that that would be true. I mean, I don't know.

1:03:18We haven't had computing systems for long enough to maybe know for certain. And there are things that we still use. I mean, there is some stuff that we still use from back then. Plugs are still the same shape-ish, stuff like that. so maybe but even like something like x86 like if you'd asked me six years ago and said is the x86 processor going away i'd say like well maybe i mean obviously on mobile it did but like in the server maybe not um but now two things have happened one is all the processes on gpu now and two like arm 64 is now a pretty important platform on the server for energy usage and other reasons.

1:04:00Right. And so it's pretty dangerous to predict the future because like it has a tendency to show up sooner than you, you, you predict. So, or, or longer than you predict too, right. Self-driving cars. Like I've heard, I've heard self-driving cars were five years away for like the last 15 years. Yeah.

1:04:17Brendan Burns:Me too. I don't know if you read books for career sake, but if you do, is there a book that impacted your career the most? Well, I mean, I would say like early on the book that impacted my career the most was a book, was The Gang of Four book. Was it Software Engineering Designs and Patterns or whatever? Like it's a software engineering book. I see it. It's Design Patterns, Elements of Reusable Object-Oriented Software. Yeah, there you go. So that like early on, that was a very influential book. It's like a late 90s or mid 90s kind of book. There's a much more recent book called Leadership on the Line that as I've become sort of a large org leader, I really like that book.

1:04:58And then there's this, what's this book? It's called, I think, Five Dysfunctions of Teams, I think. That's a really good book too from a how teams operate perspective.

1:05:07Brendan Burns:If I'm understanding, if you're an engineer, check out that first book. If you're a manager or a leader, check out the second two books. Yeah, that's probably about right. Yeah, I think that's right. And it's an evolution over time, right? So maybe you'll do both. Last question for you is, If you could go back to yourself when you just graduated college and give yourself some advice, what would you say? Keep better notes. I feel like there's a great MBA thesis or a great book in the whole Kubernetes journey and beyond. I just don't have enough notes to do that, to write it down. We went through a lot.

1:05:49like a lot of different stuff happened and I remember some of it and I don't remember a lot of it and it would I would have been nicer if I'd kept better notes I feel like right well well you

1:05:58Brendan Burns:got all the code there maybe an LLM can parse it or something you know it's not so much about the notes part it's not such a code part it's like all of the like the stuff you were talking about like all the partner discussions and all of the like interpersonal stuff and you know all that kind of stuff and like I remember a lot of it but I don't remember all of it cool well thank you so much for your time, Brenner. I really appreciate it. Yeah, for sure. Thank you. Thank you for listening to the podcast. It's a passion project of mine that I really enjoyed building. Another passion project that I've been working on kind of in secret is building an ergonomic keyboard that I wish existed.

1:06:34Brendan Burns:And I finally have a prototype, so I'd love to show you what we've built. It's ultra low profile and ergonomic, and I couldn't find anything like it on the market. So that's why we built it. I'll put a link to the keyboard in the description. You can take a look and learn more about the project there, we could definitely use your support. Also, if you have any feedback for me about the show, I'd love to hear it. Comments on YouTube have led to guests coming on like Ilya Gregorik and David Fowler. I wasn't aware of them until someone dropped a comment. Also, feedback in the comments helped me learn to reduce the number of cliffhangers in the intros.

1:07:08Brendan Burns:So your comments definitely make a difference. Please keep letting me know what you'd like to see more of in the show, and I'll see you in the next episode.

From the publisher

This is a conversation with Brendan Burns, co-creator of Kubernetes and current technical fellow at Microsoft working on Azure. We discussed what it was like building it at Google, how he got buy-in, and what he learned along the way.


🔸 My keyboard project: https://read.compose.llc/p/our-keyboard-design-reveal


𝗣𝗼𝗱𝗰𝗮𝘀𝘁 𝗹𝗶𝗻𝗸𝘀:

• YouTube: https://youtu.be/FKijpCEH9D8

• Apple: https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835

• Transcript: https://www.developing.dev/p/the-creator-of-kubernetes-on-building


𝗧𝗶𝗺𝗲𝘀𝘁𝗮𝗺𝗽𝘀:

00:00:00 - Intro

00:00:37 - How he convinced Google leaders

00:09:26 - Building the MVP

00:11:43 - How he made time for Kubernetes

00:25:28 - Technical details on building Kubernetes

00:38:46 - Rallying the open source community

00:50:01 - Scaling Kubernetes up for AI training workloads

00:55:31 - Reflections on getting a PhD

01:00:22 - The inevitable trajectory of software is death

01:04:16 - Top book recommendations

01:05:22 - Advice for his younger self

01:06:21 - Outro


𝗪𝗵𝗲𝗿𝗲 𝘁𝗼 𝗳𝗶𝗻𝗱 𝗕𝗿𝗲𝗻𝗱𝗮𝗻:


• LinkedIn: https://www.linkedin.com/in/brendan-burns-487aa590/

• Twitter/X: https://x.com/brendandburns

• Github: https://github.com/brendandburns


𝗪𝗵𝗲𝗿𝗲 𝘁𝗼 𝗳𝗶𝗻𝗱 𝗥𝘆𝗮𝗻:


• Newsletter: https://www.developing.dev/

• X/Twitter: https://x.com/ryanlpeterman

• LinkedIn: https://www.linkedin.com/in/ryanlpeterman/

• Threads: https://www.threads.com/@ryanlpeterman

• Instagram: https://www.instagram.com/ryanlpeterman

• TikTok: https://www.tiktok.com/@ryanlpeterman

More from The Peterman Pod

All 60 episodes
The Co-Creator of Kubernetes On Convincing Google, Building It, and Scaling for LLMsThe Peterman Pod · 1 h 7 min
Listen in VO