Cilium, eBPF, and Modern Kubernetes Networking with Bill Mulligan

26 Mar 2026 · 58 min · 24 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Episode topic: How eBPF enables programmable Linux kernel networking, and how Cilium (the Kubernetes CNI) uses it for scalable, secure, observable networking; includes network policies, “Cilium service mesh,” and observability via Hubble.

Guests and backgrounds

Bill Mulligan, Cilium maintainer and team member at Isovalent (company behind Cilium). Gregor Vann, security-focused technologist and former CTO across cybersecurity/cyber insurance/software engineering; based in Singapore.

Key claims

  • eBPF safely “reprograms” the kernel via verification, avoiding slow kernel upstream cycles.
  • Cilium replaces kube-proxy and iptables with eBPF for faster, hash-map-based lookups (O(1) vs linear rule processing).
  • Cilium shifts networking/security from IP-based to identity/label-based connectivity to reduce churn in ephemeral Kubernetes.
  • Hubble surfaces eBPF-observed traffic for flow logs, UI service maps, and policy-drop explanations.

Notable examples

  • Trendall (Turkey e-commerce) reportedly saw ~40% cluster throughput improvement by replacing kube-proxy.
  • Bloomberg “data sandbox studio” uses network policies to isolate tenants and prevent data exfiltration/egress.
  • ESnet (US national labs) uses Hubble for IPv6-only Kubernetes debugging, reducing multi-day work to ~30 seconds.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Bill Mulligan's Journey to Cilium

1:06 to 2:14

Bill shares his diverse background and how he came to work with Cilium.

“works under the hood, why Cilium has become one of the most widely adopted Kubernetes networking projects, and how the future of cloud-native infrastructure is being reshaped by programmable kernels.”

Understanding CNCF and Cilium

2:14 to 3:59

Discussion about the Cloud Native Computing Foundation and its projects, including Cilium.

“So like, maybe just walk us through kind of all of that.”

Cilium's Role in Cloud-Native Networking

3:59 to 5:42

Cilium's significance as a container networking interface in Kubernetes environments.

“And really kind of the CNCF was created with Kubernetes as the core project.”

Introduction to eBPF Technology

5:42 to 8:04

Explains eBPF, its history, and its fundamental role in modern Linux networking.

“I think different than a lot of open source projects, it actually is open source from the first commit.”

How eBPF Reprograms the Linux Kernel

8:04 to 11:16

Detailed explanation of how eBPF allows safe modifications to the Linux kernel.

“people upstream, decide on the right path forwards, and then it gets into the kernel.”

The Evolution of BPF to eBPF

11:16 to 13:11

Discussion on the transition from BPF to eBPF and its expanded capabilities.

“And they're a very efficient way to reprogram what's happening.”

Cilium's Networking Capabilities and Kubernetes

13:11 to 14:01

Explores how Cilium leverages eBPF for Kubernetes networking.

“So I think you've set the stage pretty well in terms of what does eBPF allow within Linux and like Linux kernel.”

Cilium's Role in Cloud Native Networking

14:01 to 18:03

Understanding how Cilium enhances Kubernetes networking through eBPF and identity-based policies.

“It's next year a 35 year old technology.”

Transitioning from IP to Identity

18:04 to 19:50

Exploring the shift from IP-based networking to identity-based models in cloud environments.

“So we shouldn't be able to route traffic to it.”

Deep Dive into Cilium Features

22:49 to 25:59

Detailed examination of Cilium's features, including network policies and service mesh.

“And I think that'd be also helpful maybe when we get there to just touch on like what it even is a service mesh, because I think some listeners may not be familiar with that.”
Show all 24 chapters

Understanding Service Mesh Concept

26:03 to 28:00

Discussing the service mesh concept and its relevance in modern networking.

“So network policy is a really important thing to be able to secure your Kubernetes clusters.”

Challenges in Modern Networking

28:00 to 28:48

Explore the complexities of networking, observability, and security in microservices architecture.

“And if you're trying to do it at just one specific layer, you might be missing a lot of the context from all the other layers.”

Rethinking Service Mesh

28:49 to 30:13

Discuss the limitations of service mesh and why it must be considered within the full networking context.

“You can't be like, oh, I'm just going to look at only layer three.”

Cilium's Approach to Networking

30:14 to 32:00

Learn how Cilium integrates both layer three and layer seven networking for improved efficiency.

“It was originally just a CNI doing a lot of this layer three, layer four, like routing.”

Understanding the Service Mesh Landscape

32:01 to 33:57

Examine the confusion surrounding service mesh terminology and its practical implications.

“But kind of how the Cilium service mesh came around is like, we also were like, people started asking us about service mesh.”

Introduction to Hubble

33:58 to 35:30

Discover how Hubble enhances observability within Cilium's networking architecture.

“We're always looking for a term that just solves all our problems.”

The Power of eBPF in Networking

35:31 to 37:43

Understand how eBPF transforms networking capabilities and enables new debugging tools.

“So it's basically saying, here are all the packets going through.”

How Cilium Operates Under the Hood

37:44 to 41:28

Dive into the architecture of Cilium and its data path for efficient networking.

“And there's also another project under Cilium called Peru, like Packet, Where Are You?”

Understanding Cilium's Architecture and Setup

42:00 to 44:22

Learn about the architecture of Cilium, its components, and installation methods.

“and they work in concert with each other.”

Migrating to Cilium: Strategies and Benefits

44:22 to 47:28

Explore the migration strategies to Cilium and the benefits of doing so.

“And our security team says we need to have these layer seven network policies.”

The Growth and Community of Cilium

47:28 to 50:18

Discover the community and growth metrics of the Cilium project.

“And at that point, you can uninstall the old CNI.”

Future Developments and Innovations in Cilium

50:18 to 54:24

Learn about upcoming features and innovations in Cilium, including IPv6 support.

“Or do you also just have kind of diehard Cilium fans that just work on this?”

Getting Started with Cilium: Resources and Guidance

54:24 to 56:00

Find out where to access resources and guidance for learning Cilium.

“And then the second part is connecting to the outside world.”

Exploring Cilium Resources for Developers

56:00 to 58:11

Learn about the best resources to get acquainted with Cilium, including hands-on labs and documentation.

“So I'm extremely biased because I'm a maintainer of the website.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Modern cloud-native systems are built on highly dynamic distributed infrastructure where containers spin up and down constantly, services communicate across clusters, and traditional networking assumptions break down. Linux networking was designed decades ago around static IPs and linear rule processing, which makes it increasingly difficult to achieve scale in Kubernetes environments. At the same time, modifying the Linux kernel to keep up with these demands is slow, risky, and impractical for most organizations. The Extended Berkeley Packet Filter, or eBPF, is a Linux kernel technology that allows sandboxed programs to run safely inside the kernel without modifying kernel source code or loading kernel modules.

0:45Cilium is an open-source, cloud-native networking platform that's built on eBPF and provides, secures, and observes connectivity between workloads in Kubernetes and other distributed environments. Bill Mulligan is a maintainer in the Cilium ecosystem and a member of the team at Isovalent, the company behind Cilium. He joins the show with Gregor Vann to discuss how eBPF works under the hood, why Cilium has become one of the most widely adopted Kubernetes networking projects, and how the future of cloud-native infrastructure is being reshaped by programmable kernels. Gregor Vand is a security-focused technologist, having previously been a CTO across cybersecurity, cyber insurance, and general software engineering companies.

1:31He is based in Singapore and can be found via his profile at van.hk or on LinkedIn.

1:50Hello and welcome to Software Engineering Daily. My guest today is Bill Mulligan. Hey, thanks for having me. Yeah, great to have you here today, Bill. We're going to be talking all about Cilium and the technology ePBF. Before we get there, as we like to do, it would be great just to hear a bit about you and what was your journey to joining Cilium? And I believe like the company you work for is sort of a wrapper around Cilium, for example. So like, maybe just walk us through kind of all of that. Yeah, definitely. So I like to say it's a little bit of an accident how I've ended up here, just a series of circumstances kind of going on.

2:28So I actually originally got my undergrad in biochemistry, so very, very far away from technology. Got my master's in social science and then ended up at the first startup that I was working at. And they were doing, back in 2018, an AI platform on top of Kubernetes. And at the time, nobody was doing AI and nobody was doing Kubernetes. So obviously, I went out of business pretty quickly, moved on to the next startup, and I worked for the CNCF, the Cloud Native Computing Foundation, kind of looking at the global cloud native community before ending up at iServellant as kind of like this promising startup in the cloud native space.

3:04and I was excited about going to Isavellant because Selim at that time was like really starting to emerge onto the scene as kind of a new and exciting way to do networking in the Kubernetes and cloud native world. So like this company seems pretty interesting. They just emerged from stealth and I was like, let's see where this rocket ship goes. And it's been kind of a wild ride since then. Awesome. For those that are not totally familiar, like what exactly is CNCF and could you just sort of then describe like what isovalent and psyllium what is the relationship between the technology and the company and that kind of thing yeah definitely so CNCF is or the cloud native computing foundation is a sub-foundation of the Linux foundation so Linux hosts obviously the Linux kernel but I think it's 900 other projects and CNCF is the largest sub-foundation under that and CNCF itself hosts just over 200 projects now Cilium being one of the projects.

4:00And really kind of the CNCF was created with Kubernetes as the core project. And then kind of all of the cloud-native projects have been brought in around that. And Cilium being one of those. And what Cilium does in the cloud-native world is, so Kubernetes is a way to orchestrate containers and other types of workloads now too. But it actually doesn't come with any networking, right? And in the world of Kubernetes, it's all distributed systems. and the most important part of a distributed system is the network, right? Because everything's got to talk to each other. So Zillium is the CNI or the container networking interface that plugs into Kubernetes and basically says this packet needs to go here.

4:40This is how traffic is getting into our cluster. This is how we egress traffic out of it and a lot of other things. So essentially you can think about Zillium at a very high level as networking for the cloud native world at the very beginning. It's expanded a lot beyond that, which I guess we'll dive into a lot more after this. And then Isabelent is a company that originated Cilium. Then they gave it to the CNCF. So it's a neutral governance under there. And we have a lot of contributors from different companies around the ecosystem. And Isabelent is the company that's creating commercial products around the different projects that are in the Cilium ecosystem.

5:16Yeah, I saw the annual report came out just a few hours ago, actually, for Cilium effectively. And it was said on December 16th, 2015, Thomas Graff pushed the first commit for Cilium. So we're literally almost to the day 10 years on from that first commit. Yeah, exactly. So it's a decade in the making, right? Decade in the making, overnight success, as people like to say. And it's kind of wild. I think different than a lot of open source projects, it actually is open source from the first commit. If you go look at the first commit, it's, I think, like 200 lines of code, the license, and like dot get ignore file like that that's it awesome full of win source proper yeah yeah awesome so i think we maybe should sort of get sort of super base level and just understand what is so psyllium is what could be described as e bpf and what is that i think that's kind of where we need to start i imagine some of our audience know this already and it's great that you're here equally i think a lot of our audience have maybe no clue what this is yeah so ebpf is also a technology that's not 10 years old birthday's a little bit earlier it's about 11 years old now and it's a linux kernel technology and psyllium was founded from the ground up based on ebpf as a technology and what ebpf allows you to do is to reprogram the linux kernel and the comparison a lot of people like to make is like eBPF is to the kernel what JavaScript is to the browser.

6:52And so if you think back, like before we had JavaScript, you kind of had like static web pages, right? You can kind of consume information off of it, but you couldn't actually kind of do anything with the web page. JavaScript comes along and suddenly you could add interactive elements. As you start to interact with the web page, it changes what it's doing. And that's exactly what eBPF is doing for the Linux kernel. And if you're not familiar with how like the Linux kernel development cycle works and how the Linux kernel kind of works as a whole, I'll kind of jump into that. And so the way the Linux kernel works, it's not just, you know, you get this distribution and you download it and you're like using it.

7:24The way it works is you need to upstream things into the kernel. And Linux is the largest open source project in the world. And kind of the development cycles are a little bit longer because it's deployed on literally billions of devices. Every single Android phone is running some portion of Linux. So they need to be very careful about what actually goes into the upstream kernel. So the development cycles, and if you know anything about like the Linux kernel mailing list, there's a lot of technical discussions that go on on the mailing list. So to be able to get something upstream is a long process that can take years or maybe even for some more controversial things, a couple of years.

7:58So not just like, okay, like we need this new feature in Linux kernel, like let's just ship it. It's like, okay, well, you need to have the discussions, work with the people upstream, decide on the right path forwards, and then it gets into the kernel. And then you have it in the kernel, right? So that's like the latest one that like Linus just goes out and produces, but that's not actually what you're running in production. If you look at most people, what kernel they're running, it's actually two years old, three years old, five years old. I mean, people don't run the latest kernel. They wait for it to actually kind of like bake.

8:31They wait for like an LTS, like a long-term stable release, or they get something from their vendor. So if you look at the actual like kernel version you're running, it's most likely a couple years out of date, right? So if you have one year development cycle, a couple years till it actually gets into you, maybe it's five years from like, okay, I have this idea to when I can actually receive this future, right? So for most people, all intents and purposes, like the Linux kernel is pretty static, right? Like you're not going to change it. eBPF came along and completely changed this programming model.

9:02What it allows you to do is you're saying you want a new feature. And the way that applications usually interact with the kernel is they interact by making like system calls into the kernel or other ways to interact with it. And the application says like, okay, can you read this file? Can you open this networking socket? Can you do these different things? And it makes a call into the kernel. The kernel does that thing and sends things back to user space. But what eBPF allows you to do is to modify how the kernel is actually running. So you take a program that you write, you insert it into the kernel.

9:33And now when the application makes a call from user space into kernel space, instead of working how the kernel normally does, your program now runs when that specific hook is called. So if somebody's, say, malicious process is trying to read a file, you can be like, we don't want this process to read this file, so it's blocked. If you want to, in the case of Cilium, be like, we want to do networking faster or more programmatic, we can be like, okay, this packet is coming here. We want to actually just reroute it directly without it going through the whole networking stack. And so what eBPF is allowing you to do is to actually add functionality on the fly into the Linux kernel.

10:14And if you've gotten to this point, you're kind of like, OK, well, Linux kernel, pretty important. I know if I crash it, that's really bad. All the systems are over. And so that's kind of what makes eBPF really powerful is because it's not just extending the Linux kernel, but it's doing it in a safe way. because if you just like throw random code into the kernel, you're very likely to crash it. If you're familiar with like the CrowdStrike incident where they took out half the world's IT, right, that was because there was a bug in the kernel and they crashed all the kernels around the world. It's obviously something you want to avoid.

10:46And so what EVPF allows you to do is it lets you insert programs into the kernel in a safe way. And the way that it does that is for each of the programs that you're adding to your kernel, it goes through a verification step. And what this verification step is basically checks that the program is safe to run in the kernel. So it's not going to crash the kernel. It's not going to call memory out of bounds. It's not going to do these things that will essentially harm the kernel. And so these programs are safe to run in the kernel. They won't crash the system. And they're a very efficient way to reprogram what's happening.

11:19And so that's kind of like the foundational technology, right? EVPF got merged into the kernel about 11 years ago. So the Cilium team was like, okay, this way of programming the kernel is going to let us rebuild everything in the kernel better. Like, what can we rebuild first? And the team at the time was working on, they came out of like the OpenV switch team. They were doing a lot of networking stuff. And they're like, okay, we're going to start rebuilding networking in the Linux kernel better, faster. I'm ready for the cloud native world. And so that's the birth of Cilium. Got it. and i think it's actually helpful what does e bpf stand for its extended berkeley packet filter i guess packet filter on the end there is kind of the thing i mean i believe there was there was in theory a retro actively named classic berkeley packet filter and eppf is the sort of advanced version of that is that correct yeah so bpf or berkeley packet filter was kind of like the original thing from like decades ago it's like okay can we put like a packet filter into the kernel.

12:20And that was the thing. And originally, kind of like this story was, Alexi, who was one of the co creators of EVPF came to Linux kernel and says, like, I want to be able to insert byte code into this. And they're like, this huge patch that come into the kernel, like, no, no, we don't want to do that. And so Alexi went back to work with like Daniel and a couple other people. And they're like, okay, like, how can we actually get to this into the kernel? And they're like, well, there's actually this like packet filter already in the kernel, what if we just improve that to be able to get what we wanted into it.

12:47And so piece by piece, they started improving the original BPF, like the packet filter, and then also extending it. And so what they're able to do was to improve the existing subsystem and then change it into a more generalized version. So like the CBPF, like the classic BPF is like the original kind of like Berkeley packet filter. And now EVPF is the extended version. We don't really call it extended Berkeley packet filter because it does so much more than just networking now and saying that it's a packet filter isn't really true it does things across observability security profiling scheduling interacting with devices it's basically like a generalized way to reprogram the linux kernel so we just kind of use it as like the standalone ebpf term but that's kind of history and why if you hear somebody it's a little bit confusing because there's bpf there's cbpf and ebpf and if you talk with some of the kernel people they use eBPF and BPF kind of interchangeably just because of that history.

13:46But technically now it's called eBPF. Gotcha. Okay. So I think you've set the stage pretty well in terms of what does eBPF allow within Linux and like Linux kernel. So it's kind of become this framework for enabling things that can run in the kernel, almost like a sort of set of, I guess, rules that mean that if you want to build things that touch that they're not going to break some of the big things and i think yeah the crowd strike example was a good one like that's that's what happens when this isn't done properly that's obviously windows so it's not linux but yes that was a problem so where did then psyllium come from in that sense and i guess psyllium is handling the networking side of what can be done with this capability there must be other products companies that then deal with other bits that are now possible with epf being there but yeah what does psyllium enable and kubernetes is obviously a big part of that so let's kind of go there so if you rewind back again about 10 years ago we had like the whole containerization movement that was like really exploding at the time and kubernetes had also just been released as you know open source not from the first commit but as like this actual standalone project and so So we kind of have this new cloud native world.

15:05And if you think about like the transition into the cloud native world, like what you're kind of seeing is a lot more ephemeral dynamic environments, containers are coming and going and kind of like the traditional world of how Linux was built, right? Linux is not a 10 year old technology. It's next year a 35 year old technology. And so like kind of like the programming model for the Linux kernel is very different from what you need in the cloud native world, right? A lot of Linux networking is built based on IPs, based on IP tables. think of like, okay, we have this list of IPs that we trust, a list of IPs that we don't trust.

15:38But if you think about the cloud native world, you're spinning up containers up and down all the time. And so how can we have this decades old technology and bring it into the cloud native world? And that's really the challenge that Cilium was set out to do is how can we do cloud native identity-based networking? And so some of the original challenges that Cilium was looking at, Like one is like IP tables is the way that we do a lot of networking. So as you're being like, okay, we need to send a packet from this IP to this next IP. And the way that you do that is you have like a list of IP tables and you go through them linearly.

16:15But if you're having, I don't know, thousands, tens of thousands, a million containers in a cluster, going through a linear list of rules is not very efficient. So one of the first things I do, and one of the reasons that a lot of people like Cilium is because it replaced IP tables and kubeproxy, which is the proxy that routes most of the traffic in Kubernetes. We replaced that with what we call kubeproxy replacement, and that's all written in BPF. And so it replaces IP tables with eBPF. And rather going through things linearly and having to do things kind of like in an ON order, what eBPF allows us to do is have everything in the hash map and it's able to look things up in 01.

16:57So a lot faster and a lot more scalable. So if you look at the difference when you have like 10 services in the cluster, it's not that big, right? Because you can run through a list of 10 services pretty quickly. But if you have 10 ,000, the difference between reading through the rules linearly and being able to just look them up in the hash map is very significant. So that's one of the things. Being able to write things in a modern way for the modern technology and the modern way of doing things is making networking a lot more efficient. And there's a lot of stories now about being able to replace Kube Proxy.

17:29like Trendall, like a e-commerce company in Turkey, by replacing QProxy, they increase cluster throughput by 40%. So big performance gains there. Then the next thing about Cilium is, right, so you have a lot of containers, you have these IPs, but the IPs aren't fixed because they're coming and going all the time. And a big thing in networking is not just, yes, one is like connecting things, but it's also making sure that things that aren't supposed to be connected don't let's say connected, this whole like network security part. And if you're doing that based on IPs, you're going to have to be updating all these rules a lot and being like, okay, like this IP is like no longer being used.

18:06So we shouldn't be able to route traffic to it. And so being able to understand like network routing and like network security, if you're doing that with like fixed IPs as containers are coming and going, you're going to have to be updating these rules a lot. So it's going to cause a lot of churn in the cluster, a lot of overhead. And Syllium was like, okay, as we're moving from a world of IPs towards identity, like, you know, the classic DevOps analogy from pets to the cattle, we're not looking at like individuals anymore, we're looking at like groups or sets of people or things with like labels.

18:35And so Cilium switched the whole networking model from like this IP based model to this identity based model. And so rather than saying like, IPX can talk to IPY, we can say front end talks to back end. So then as we kind of rotate the containers behind these labels, it doesn't actually matter. And as you spin up a new container, you can be like, okay, this is a backend label. And so it can now automatically talk to all the front end labels. And so if you think about the cloud native world, it's like, how can we switch to this identity based model? And by being able to give things identity, it makes things a lot easier because you can swap things out on the backend and the identity is still the exact same.

19:13And it reduces a lot of the churn in the cluster too. So what Selim was trying to do at the beginning is like, you know, we have this new modern cloud native world, things are a lot more dynamic, ephemeral, the current networking technology that we have isn't going to be able to keep pace with what we need to do in this world. So how can we rethink networking for the modern world with eBPF? And the way eBPF allows us to do that, we can take out IP tables, we can route things like very efficiently, we can move from IP towards identity for both our networking and our security model and allows us to bring networking into the modern cloud native world.

19:50Yeah, that's a really good way to explain it. I guess there are some analogies with just like cloud IAM, like identity management, like how the cloud providers effectively added this, what is unfortunately now an incredibly complex thing. You've got layers on IAM now to try and make it easier to administer because people end up creating the wrong identity profiles and all sorts of things but yeah so like is that kind of a good analogy yeah exactly so if you like you think about it when you join a new company they give you like an example would be like or in my personal life i log into google and it gives me access to a lot of different services based on the identity that i have i'm like i say i'm bill mulligan i give this like identity to google and google goes out and says like this is bill mulligan to all these different services you could do the same thing in Kubernetes.

20:37You can be like, okay, this new pod is now front-end. It's front-end to all these other services, or it's the back-end that all the front-end just want to talk to, essentially. And if you think about when you're adding a new developer to your team, you don't want to give them access to GitHub and your cloud resources and your developer environment and to all the other services they need. The way that you probably do it is you probably give them one identity to something like Okta, and then Okta provides the identity out to all the other services that they need access to. Yeah, that makes sense.

21:10In mobile application security, good enough is a risk. GuardSquare uses advanced, multi-layered code hardening techniques and automated runtime application self-protection and mobile application security testing, combined with real-time threat monitoring to deliver the highest level of mobile app security. Discover how GuardSquare brings all these together to provide mobile app security for your Android and iOS apps without compromise at www.guardsquare.com. If you're an engineering leader, you know this cycle. Your team's focused on building product, but someone in ops needs a dashboard. Marketing needs an admin panel.

21:52Finance needs a custom workflow. The requests pile up. You can't get to them all, so people start building their own solutions. shadow IT spreads, and eventually, you're the one stuck cleaning up tools that were built with duct tape and good intentions. Retool breaks that cycle. Their AI AppGen platform gives teams a governed place to build the tools they need so everything stays secure and under your control. Someone could type, build me a customer admin panel that manages accounts from Postgres, and they'd get a real, production-ready app with proper permissions built in. Your teams get unblocked, and you don't inherit a pile of technical debt down the road.

22:30So if you're tired of being the cleanup crew for shadow IT, head to retool.com slash SEDaily and see how other engineering teams are democratizing app building without creating chaos. Because honestly, we could all use a better way to handle internal tools. Sometimes you just need Retool. So maybe if we just look at what are the features, feature set, if you like of Cilium, you know, we've got things like we've touched on it there, but like network policies, I think that'd be interesting to kind of dive into a little bit more service mesh as well. And I think that'd be also helpful maybe when we get there to just touch on like what it even is a service mesh, because I think some listeners may not be familiar with that.

23:13And then we could maybe just sort of get onto some of the more advanced features as well. But yeah, so let's maybe just start with like network policies that seems to be kind of the core of Cilium, like Maybe just dive into that a little bit more. I talk to a lot of users out in the Cilium community. And the main three reasons that they choose Cilium, because there's a lot of different networking solutions in Kubernetes, is one is network policy. Another one is kubeproxy replacement, getting the performance and scalability benefits, encryption of network traffic, and observability with Hubble.

23:45But starting with network policy. So this is going back, distributed system. We want to make sure things can talk to each other, but we also want to make sure things can't talk to each other. And so in Kubernetes, there's Kubernetes network policies, and these are layer three, layer four network policies. So you're looking at things like IPs, like this IP can or can't talk to that other IP. So Solium implements Kubernetes network policies for layer three and layer four. But the additional thing that a lot of people look at is also, it's not just, you know, these low level. we also want to look at like layer seven network policies.

24:22So Cilium has layer seven network policies too. And we call them like Cilium network policies. So things like allow traffic from star.google.com or don't allow traffic from this domain, right? So being able to look at the actual domain with layer seven network policies is super helpful for a lot of people. You can also do things like cluster-wide network policies, looking at which namespaces cannot talk to each other or can talk to each other. So one interesting one, that use case that we had was like Bloomberg, obviously a lot of financial data, and they were coming out with a new product that was a essentially like data sandbox studio.

25:04So customer logs in, they're able to access the financial data, they're able to do different types of work with it. So they had a Jupyter notebook and they They could write different programs against the data, get what they wanted to, and then see the data. But the important thing is Bloomberg's financial data. They want to make sure data is not being exfiltrated out of this data sandbox studio. They want to make sure they had multiple tenants. And so each of the tenants can't talk to the other tenants. One person can't see what the other person is doing with the data. And a lot of that you can do with network policy.

25:40So you basically namespace each of the tenants within the namespace and make sure that they can, with network policy, talk across the different namespaces in the cluster. And you can also write network policies basically saying that data can't egress out of the Kubernetes cluster at all, too. And so with that, they're able to create a new product for their customers while still keeping their sensitive financial data secure. So network policy is a really important thing to be able to secure your Kubernetes clusters. That's a really good example. Funnily enough, I'm working on something similar, so I'm going to just ask a question on that basis.

26:14So why does Cilium make that easier than if not using Cilium? If you're just using Kubernetes, there's different CNIs that you could use in Kubernetes. And some of them, it's not a requirement that they implement network policies. So some CNIs don't have any network policies. Then you can write Kubernetes network policies, but they won't be enforced. It's really not effective. Some of them just do the Kubernetes network policies. Then you only get the layer three, layer four. So you can't write more complex or advanced use cases around network policy. And then the other one is like other things like if you're doing like multi-cluster network policy.

Read the full transcript

26:52So if you're running multiple Kubernetes clusters, this is another thing a lot of people turn towards Cilium for because it simplifies that. Cilium allows you to look at network policy, not in just one Kubernetes clusters, but across multiple Kubernetes clusters. Kubernetes gives you basic network policy and Cilium allows you to do much more advanced use cases around network policy. Yeah, makes sense. So how about, is there anything more around network policies or do you want to go to service mesh? We can go to service mesh. So yeah, talk to us about service mesh. Again, I think what is a service mesh first and foremost and then how is Cilium helping to that end?

27:26Yeah, so anybody that knows me might know that I'm trying to kill the word service mesh and the category service mesh. If you go look at it, there's an article that I wrote that I think explains a lot of my opinion that it's called like the future of service mesh is networking. And so service mesh is a little bit newer than Kubernetes. And it's once again, like, okay, new cloud native world, like there's kind of a lot of things that we need to rethink. And a lot of this is like service routing. So service mesh is a term that tried to mean a lot of different things. So like if we have like microservices, we have a lot of new challenges of how do we do the networking between them how do we do the observability how do we do the security between all of them like microservices running all over the place and service mesh tried to solve this with a lot of like layer seven networking stuff so right it's like networking observability and security a lot around like layer seven stuff and this is why i kind of have a problem with like the category like service mesh is like if you look at all those like those are all like fundamentally like networking things.

28:31And if you're trying to do it at just one specific layer, you might be missing a lot of the context from all the other layers. So I think we had a big arc where service mesh was very hot and a lot of people were trying to implement it. But I think we're kind of getting into a phase where people are understanding that you can't separate out the different layers of the networking stack. You can't be like, oh, I'm just going to look at only layer three. You actually, if you want to have the full context for your application, you need to look at all the layers and you need to look them like holistically right because like an example that i've heard is people are running psyllium as a cni and then psyllium has a service mesh which i'll get to in a second but you can also run a different service mesh on top right so they're running psyllium and they're running a different service mesh and they're like okay well psyllium does a lot of smart things in ebpf it doesn't use ip tables it just reroutes traffic and do things like just route packets from socket to socket within the same linux host you do a lot of things that will make it much more efficient, scalable, performance, save you CPU cycles, but it doesn't mean all the other parts of the networking stack know what's happening.

29:37So an example would be like, Cilium can route things directly from one socket to the next, and it doesn't go through the whole Linux kernel networking stack. So the service mesh is looking at the end of the Linux kernel networking stack and it's like, okay, I'll do this like layer seven processing. Once it comes out of that, what like actually never goes through the networking stack. So you never see the packet. And so you're like, okay, like all this traffic is disappearing or we don't route it or we don't see it. And it's because it doesn't go through the traditional networking stack as the service mesh was expecting.

30:07And so service mesh, I think, can't be its own standing run-alone category. You need to think of it in the context of your whole networking stack. And so this is why Cilium came along. It was originally just a CNI doing a lot of this layer three, layer four, like routing. But then we're like, okay, well, networking is not just like a couple layer standalone category. It's actually like you need to have the context of the full stack. So what we did is we came up with Cilium service mesh. And this is where some of like the layer seven network policies started to come in and also doing other things like traffic routing in layer seven, some of the observability stuff, but integrated within the rest of the CNI and the rest of the networking story.

30:47And I think I also have a problem with the term service mesh because it's kind of nebulous. It's like, where does service mesh live? Where does networking start? Where do they overlap? Where they're all just overlapping concerns. What I think of today as Cilium service mesh is Cilium's gateway API implementation in Kubernetes. And some gamma. I think, sorry, I have a lot of problems with service mesh. Yeah, there's a lot of emotion comes up with service mesh. Yeah, it's good, it's good. But I think it's also really funny. So I look at the analytics for the website, and one of the top, I think, three pages is the Cilium service mesh page.

31:26So it's what people are interested in. But I'm like, so what are you actually interested in? Is it like the routing? Is it like the observability? Is it like the security? It's like people are trying to solve, like service mesh isn't a problem. What they're trying to solve is like, okay, we need to do like layer seven routing and like host tasks or something, or we need to do layer seven network security. Like that's the actual problem you're trying to solve. Like service meshes is like this kind of like nebulous term that somebody told me I needed it. Yeah. I mean, I think I can probably give then my perspective on that because yeah, working on a specific problem, you know, I'm not an engineer by day anymore, but working pretty sort of hand in hand with pretty advanced engineers.

32:06and when i was looking at what we're needing to achieve and did a sort of bit of running around sort of understanding okay what are the bits to the network stack that we need to think about here service mesh just kept popping up so i'm like okay guys do we have a service mesh uh it was kind of my first question just to kind of maybe get a sense of do we have a concept of this thing or not and then you know i got my answer and then we move on from there so i guess it is it's just sort of a catch-all term to help to that end and maybe that's why people are looking up so much on the website because it's sort of like yeah but that's just an anecdote i guess but yeah yeah i think you're exactly right i think it's kind of like as we saw this transition to the cloud native world there is kind of like a lot of like quote-unquote new problems that were like as we have a bunch of microservices like there's like new networking problems and new like network security problems and like service mesh was the label to solve a lot of those problems it's like okay, we want a service mesh to solve the specific problems that we have.

33:08But kind of how the Cilium service mesh came around is like, we also were like, people started asking us about service mesh. And we like looked into it and we're like, okay, what does the service mesh actually look like? And we're like, okay, with what we have so far, we actually have like 80 % of the service mesh because it's a lot of like networking, network security, network observability parts. The only part we're missing is like a bit of like the layer seven stuff. And so we're not actually building like a whole service mesh from scratch. We're actually just adding that last 20 % is what people are looking for.

33:37And so that's how we originally came out with what was the Cilium service mesh. Got it. So let's maybe move on from service mesh. That's been, I think, super helpful. And I'm sure there's a lot of people that will sort of look at the term service mesh differently now. Sorry if I'm destroying people's hopes and dreams. One thing to solve it all. Yeah, exactly. That's all we're looking for. We're always looking for a term that just solves all our problems. Yeah. psyllium also does observability to my understanding through i guess a sort of arm of the product called hubble could you maybe talk to us a bit about about hubble yeah definitely so this is like everything else in kind of the psyllium ecosystem is based on ebpf and so what hubble does is like okay since our ebpf programs are in the kernel routing all the traffic and we can see all of that what if we just like took that information and surfaced it to the user.

34:31And to be honest with you, I think Hubble is the favorite feature of basically every user I talk to. The quote from the ESNet Energy Science Network, which is all the national laboratories in the US, they're doing crazy stuff like IPv6 only Kubernetes cluster, and it's like, Hubble's a godsend. It lets me what used to take multiple days of engineering time, I can now solve it in 30 seconds. And so, going back, if we think about it, distributed computing packets are flying everywhere. We need to be able to understand it. It's not just like, okay, let me debug my one application. I can follow it through the whole program.

35:10It's like, okay, applications are making calls out to different programs. It's going over the network. We don't know where the packets are going. We don't know where the information is going, where things being dropped. This is where Hubble came along. It's like, okay, so if we're actually routing all the packets already with eBPF, why don't we actually observe them? So Hubble just piggybacks on top of the CIMM-CNI and basically takes the information, from the CNI and surfaces it to the user in a couple of different ways. So one is network flow logs. So it's basically saying, here are all the packets going through.

35:39This is where they're going. And then the other one is the Hubble UI. And this allows you to create a service map of everything that's going on in your cluster. And you can see where things are being routed, where different things are connecting. And also, I think more importantly, where traffic is being dropped. Because to be honest, when most people run into networking, Like everybody likes to not have to think about networking. The only time they do is when things are going wrong. And with Hubble, they're able to very easily visualize either through the UI or through the flow logs. They're able to see, OK, where is our traffic being dropped?

36:13Right. Because that's probably like most people are concerned. The security team is concerned about like, OK, where are things going that they shouldn't be? But like most developers on a day to day basis, they're saying, why isn't my traffic reaching the destination? And Hubble is a great way to understand that. So it can give you the reasons for like the policy drops. You can be like, okay, our security team wrote all these new network policies and deployed them into the cluster. And now all of my traffic's being blocked because they didn't want this type of traffic going in the cluster anymore.

36:42And then you can go have that conversation or deploy a new service. Why isn't any traffic reaching it? Well, okay, like it's in a new namespace and we don't have the network policy. We have a default deny all network policy in our cluster and we forgot to write one. Okay, we should allow traffic to this new namespace. So people love Hubble because it allows you to have insight into where all the network packets are going in your cluster in a very easy way. How would this be done, I guess, without Hubble? You just have to kind of roll your own? So Linux, right, like decades old technology, there's a lot of networking tools to be able to understand where things are going, like TCP dump.

37:20But that's right if you're using the networking stack. And the problem with BPF is like there's kind of like this sometimes people say EVPF magic that people sprinkle on. It's like this black magic happening in the kernel and the packet just disappears because it's routed from one socket to the other and it doesn't go through all the traditional tooling. Right. So as you're switching how you're doing things, you also need to come up with new tools. So, yeah, Hubble is one of them. And there's also another project under Cilium called Peru, like Packet, Where Are You? and that's another great debugging tool that allows you to essentially pull more information out of the kernel.

37:57One of the reasons that people love eBPF is because you can hook anywhere in the kernel. You can pull any system information that you want. You can modify any system information and you can essentially pull out this firehose of data. The limitations of some of the previous tooling is it's either one, it doesn't pull out the information you need or it's designed in a certain way. It's like, this is what it does. but eBPF allows you to look for anything that you want to. You want to pull out new type of information from the kernel, well, you can write a program to be able to do that. And so what Peru and Hubble allow you to do is to surface exactly the information that you want rather than relying on tooling that may not even work.

38:34Yeah, I think that's a helpful call out given, as you say, eBPF could be seen as a bit of magic. Then, unfortunately, with the magic often, yeah. I mean, my sort of long ago version of that was like when you start using Ruby on Rails, for example, and then it's like yeah but we need like tools to actually then understand what's going on because it's just a whole layer of magic passing data between the back end and the front end and i need to sort of understand what's going on there for example exactly so you can think about it new tooling for the new world so we've kind of looked at what psyllium does and why developers might want to look at it or why they're using it already let's maybe just spend a little time going a bit deeper on how psyllium actually works sort of under the hood so we've got the idea of like what is a psyllium data path for example and then we can maybe sort of jump into just the actual component architecture after that so i believe there's sort of a concept of bpf hooks maybe we could kind of start there and just sort of how does like psyllium actually work i guess so kind of like the basic architecture for psyllium is there's a psyllium operator running into your cluster and this runs kind of like the life cycle for all the psyllium agents and the way the psyllium agents work it's like daemon set so it's one psyllium agent running on each of the nodes in your kubernetes clusters and what the psyllium agent does is it installs all the ebpf programs onto that specific node.

40:02And so what the agent does is it gets information and basically writes all the BPF programs and installs them actually into the cluster. And the interesting thing about this architecture is it in some ways like simplifies a lot of the networking and upgrade lifecycle because the data plane, which is the BPF programs running in the kernel is actually separated from the actual lifecycle of like the control plane, which is the agent and the operator running in the cluster. And so one thing I didn't touch in before with eBPF programs is that you install them into the cluster and they start running and you can de-install them.

40:46And so what that allows you to do, and there doesn't have to be any communication between like kernel space and user space between like the control plane and data plane. So the agent is installed on the node. It installs all the BPF programs. the BPF programs start routing all the networking packets, or they're with Hubble, and they're observing all the networking packets. And with that, you're able to essentially, if the agent goes down, it doesn't matter because all the programs are pre-installed, and they're still going to be routing the packets. The only thing that won't happen is you're not able to update any of the data path because the agent's not there anymore.

41:19And so what you can do is you can update the agent, new agent is there, and it can then modify the BPF programs. It's a little bit different with Envoy. So that's with all the BPF stuff. Some of the Layer 7 stuff, which I guess I didn't touch on yet, is done with Envoy, which is a very popular service proxy. It's what a lot of the other service meshes are based on. And Envoy is also run as a daemon set. So one Envoy in every single node in the cluster. And that works with the Cilium agent and the BPF programs to do some of the layer seven based routing. And so the BPF programs and the Envoy together kind of create the whole networking story and they work in concert with each other.

42:03And then layer seven is a little bit different because this is something that they're working on right now. It's like some layer seven connections are like more long live. So if you restart the Envoy on the node, then it like resets some of the connections, but now they're working on like hot restart of Envoy. So that story's changing a little bit different, but I guess what's different about like the ceiling architecture is like The agent is separate from the BPF programs that are running in the kernel. So the data path can keep on running even as things are changing on the control plane too.

42:32So then we've got the Cilium agent. I guess there's also the Cilium operator as well? Yeah, the operator manages the lifecycle of all the agents in the cluster. Gotcha. And then the CNI plugin and then identity management, which we've obviously touched on a little bit as well. so in terms of getting up and running with psyllium maybe we could take sort of two brief examples one is like total greenfield which is obviously hopefully the easier one and then one is you've already got like a kubernetes cluster kind of running something medium advanced so maybe we take those kind of two cases like what are we talking in terms of getting up and running so you have a brand new kubernetes cluster there's a couple different ways to install psyllium so the first and the easiest one is you're on a cloud provider doing something like GKE or AKS and Cilium's already the default on that cluster.

43:24So the cloud provider sets up your cluster, you get Cilium off and running to the races, it's already in there. I think that's one of the cool thing is like Cilium's already the default CNI for a lot of managed Kubernetes clusters. The next one is you're setting up your own Kubernetes cluster. And this is like super common use case for like on-prem or other things. Cilium has like a couple of different tools. I think most often people install Cilium with the Helm chart. So install that with Helm into your cluster. You can also install it with the Cilium CLI, but this is, I guess, maybe not as recommended because then it's what things you pass into the CLI and trying to remember that versus like a Helm chart.

44:01So you're going to probably see most people install Cilium with the Cilium Helm chart. And then the migration story is something that we see commonly because when I look at the way like people users set up the kubernetes cluster it's like what i was saying before nobody wants to care about the network until they have to care about it and they don't usually turn to psyllium or they might not always use psyllium until they come across one of the problems that they're having with like that psyllium helps solves like i was saying before so something like performance and scalability the network policy or encryption the observability aspect or like multi-cluster networking because like if you're running on something like open shift open shift has their own cni that they install into their communities clusters or aws has the aws vpc cni so there's already one installed in there or maybe for instance a lot of tutorials start with like flannel and so you already have a kubernetes cluster running with a different networking cni in there and you run into one of these challenges that you're like okay well I need to have better performance and scalability in my cluster.

45:08And our security team says we need to have these layer seven network policies. And our application developers are having a hard time debugging the cluster because they don't have any insight into where the traffic's actually going. And we're thinking about spinning up some more Kubernetes clusters. And you're like, OK, well, Cilium solves a lot of these problems. So I think we should migrate to Cilium as our CNI. And the question becomes, OK, how do we do that? And so the migration path that we see most commonly that people do is the cool thing about Cilium is it's not like kind of like a big bang where you're like, okay, like we need to like switch this over.

45:46It actually allows a lot of incremental things. And so one common path is I see people doing is like, okay, we need better observability in our cluster. And so what you're able to do is to do CNI chaining. So you essentially have like your first CNI say, like, fine. And you're like, okay, I need better observability because it doesn't have any observability. And so you're able to essentially install Cilium on top of Flannel. Flannel still does all the network routing, but Cilium is able to see all that networking and being able to surface that information through Hubble. So you now have the network observability without having to change any of your data plane.

46:20And you also now have the added benefit of having Cilium as a CNI in there. Or people also install Cilium for network policy. And so we don't want to change our data plane. Actually, there's a lot of companies that actually wrote their own in-house data planes. So for example, Alibaba wrote their own CNI, but they wanted to add network policy. So they installed Cilium on top to do the network policy part. And so now you have two CNIs in there. One's doing the networking, one's doing the network observability or network policy. And what you're able to do is Cilium has this flag called Cilium node config.

46:55and you're able to specify on each node in your cluster which CNI you want to do the network routing. And so at the time you're able to essentially say, okay, I want all of my original CNI to do the network routing. But what you can then do is you're able to drain traffic off of a node and say like, okay, we want this new node to be, as you're adding new nodes to the cluster, we want this new node to be using Cilium as a CNI for traffic routing. And so you can basically roll over your whole cluster node by node. And each node that's coming online now is able to have Cilium as the CNI until eventually you're basically switching over all your traffic to the new CNI.

47:36Cilium's your CNI. And at that point, you can uninstall the old CNI. Wow. Yeah, that's really cool. Yeah, I like the CNI chaining. That's super cool in terms of... Yeah. Yeah, because as you say, migration is probably the more likely case for a lot of people maybe listening today that aren't using already. but unfortunately that's also often the reason it doesn't get adopted as quickly because it's often challenging yeah nobody likes to mess around with the network yes quite yeah if it's working just like leave it right but yeah if you start to run into some of these problems you're looking at a migration story and a lot of people want to do live migrations so if you go on the ceiling website there's actually quite a few stories like the most recent one was from db schenker which is like the German national rail like logistics thing and they needed to migrate over to Cilium and they did the Cilium node config and were able to do like a live migration of Cilium.

48:27Exactly that's infrastructure sort of at some of the most important levels so that's pretty cool. So kind of looking at as we touched on right at the beginning Cilium is a very open source project and again looking at the so just some of the stats from the report that you guys just put out there's a number yeah a hundred a thousand individual contributors you cross that line basically on the project in october 2025 so that's a huge number that's a huge number of contributors so like this this is surely one of like the i guess largest open source projects out there yeah i think this is like kind of interesting for me because it's also if you're looking at like other like vanity metrics like github stars and things like that it's like okay how do you measure the size, success of the project?

49:12And a lot of the people that I'm talking to, like Cilium is run by the platform team, and it's four engineers supporting 200 developers. And so if you look at a lot of other projects, they're going to be a lot bigger, but it's just because the number of people actually interacting with that is a lot bigger, like a front-end framework or something like that. It's going to have 200 developers for one SRE supporting those 200 people. And so it's kind of wild to me that Cilium is kind of growing into like such a large project, right? If you're looking at, there's a lot more developers that can write like HTML than can write BPF code, right?

49:47So like a thousand contributors is actually quite a lot in some ways. In terms of like stats, so depending on how you look at it, Cilium is in the top three projects in the Cloud Native Computing Foundation. So 200 some projects are the largest ones in the Cloud Native space. Kubernetes, obviously number one, because it's also the second largest open source project in the world. Number two and number three is Cilium. So yeah, one of the fastest moving projects in the whole CNCF ecosystem. Yeah, I think I said in the report, now the second largest. Cilium is now the second largest. That's pretty awesome.

50:23So yeah, I mean, I guess sort of on that note, sort of in terms of community and contributions, is it usually someone who's kind of, I guess, using already through their company, I guess, that sort of ends up then jumping in and making a pull request for something that they've seen? Or do you also just have kind of diehard Cilium fans that just work on this? Yeah, I would say the most common thing, it's kind of like what I was saying before, like the number of people that can write eBPF code in the world is not that large. The Cilium agent is written in Go, so it's a little bit different there.

50:54I think it's a bit more approachable. But yeah, Cilium is like a pretty deep networking technology or like technology as a whole. And so, yeah, the most common use case is we're running Cilium in our production cluster. We're running into this issue or bug. It's open source. So we're contributing this fix because nobody else in the community is working on it and we really need this. So that's actually how some of the earliest maintainers of the project came along. So some of the earliest ones were actually from Palantir and Datadog because they were running Cilium in production. They needed things solved.

51:28Easiest way to do that is to upstream the changes into the project. They got more and more involved and became maintainers of the project. And then in the exact same way, two of the other big companies, like I said, Google and Microsoft use Cilium as the data plane in their managed Kubernetes cluster. So they need to get involved to be able to upstream their changes. So it's people running Cilium in production that are trying to solve the issues that they have. And that's really how they get involved. Looking ahead, when it's sort of this open source, like roadmap, it can be a bit of an ephemeral term, but like, what are you, you know, 2025 looks like it was a pretty big year.

52:04Like, what do you think is kind of on the horizon through 2026? I think there's a couple of different things. So one, like in the Kubernetes world is really starting to see this transition to IPv6. I was just at Cilium Con in Atlanta, right next to KubeCon. And there is actually two talks from ESnet and also from TikTok talking about using IPv6 only Kubernetes clusters. I think maybe this is finally the year of IPv6, but I think we're really starting to get there. And a lot of the work done this year by the project was around how can we basically bring IPv6 feature parity up to IPv4. So I think there's a lot of work being done around there because we're actually starting to really see IPv6 only clusters going into production.

52:49And not only that, going into production at scale. If you look at the scale of the national labs infrastructure in the US, it's pretty big. Also, TikTok also has quite big, very different use cases. So that's one that's in Kubernetes itself. The next one that I didn't touch on that much is around, I guess, everybody's kind of aware of the whole VMware thing. And so people are looking to migrate off of that. I think that plays into Cilium a couple different ways. One is how can we bring VMs into Kubernetes? and another part is how can we connect Kubernetes or like the cloud native world with the rest of our IT estate, like sitting in virtual machines outside of that.

53:32If we're migrating and modernizing, like how do we still connect it to what we already have? So there's kind of like two pieces there. So one is how do we run like VMs in Kubernetes? And Qubect, I think a lot, is what a lot of people are using. And one thing that I'm really excited about that Cilium's coming up with is this thing called NetKit. and this was originally developed to solve the container networking overhead. So the problem with containers is it's a process running on the host in its own network namespace, and by traversing into the network namespace, it creates essentially some overhead.

54:07And what NetKit allows you to do is to basically take something off of the NIC and put it into the container with essentially no overhead, so eliminating the networking overhead of containers. right since kubbert is kind of like a vm running in the container in kubernetes what we're doing now is like can we use netkit to get the packet into the vm directly from the nick so eliminating the overhead of not just the container but of the vm running inside the container running inside of the host so it's once again how can we reprogram the networking stack to make it faster more efficient more scalable as we add more layers of abstraction how can we still kind of make those abstractions, things that we can still get the performance that we need to out of our clusters.

54:49So that's on running VMs inside of it. And then the second part is connecting to the outside world. So we're actually doing a lot of work at Isolvion, like getting Cilium to connect to like the outside world. So VMs running outside of your cluster or VMs that you're trying to migrate into your Kubernetes cluster in a seamless way. You know, it's once again, how do we make that transition that migration story as easy as possible and doing that in a very smart way with eBPF. Awesome. Very cool. So NetKit, is that that'll be available in 2026? No. So that's actually out now on a part of Cilium. So this is another thing.

55:29It was made by Daniel Borkman, who's also one of the co-creators of eBPF. That was kind of like the next project he's working on was NetKit. And that was a Linux kernel feature. So it's actually part of the Linux kernel, Cilium uses NetKit to be able to eliminate that container networking overhead. And then we're also implementing it to work with Kubevert too. So the whole part with four containers is already, it's in the kernel and it's in Cilium if you're running like the right versions. And then like the part with VMs is kind of for next year. Awesome. So just as we wrap up, where's best for a developer or just someone who's kind of interested in Cilium and maybe thinking about, I don't want to say throwing over the fence to the developers, but throwing it into the mix, like where's the best to go and just sort of get acquainted with Cilium?

56:17Yeah. So I'm extremely biased because I'm a maintainer of the website. So I would say go to Cilium.io first. And I think there's a lot of like helpful resources for the different things that I've talked about. So if you're interested in like, okay, like I'm interested in like this QProxy replacement or other things like that, we have pages for all the different features. So like, I want to learn more about Silliam as a CNI, about kubeproxy replacement, about BGP or cluster mesh or host firewall. I want to learn more about Hubble. We can do that. We also have like created different pages for different industries, kind of talking to the challenges that a cloud provider or consulting company or financial services company.

56:55So it's not just like, here are the features, it's here are these features and how do they apply to your actual industry. And then the last part that we have is like these outcome pages, because once again, it's like you have these features but like companies aren't buying a service mesh they're buying a specific outcome like how do we do layer seven routing or how do we do zero trust networking or how do we do network automation how do we consolidate our networking tools and so like actually looking at those outcomes and so it's kind of like moving up the stack from here's the future to here's like the business value that we're getting and how do we do it for the industries but if you actually want to get hands-on what i recommend is going to like the getting started and going to the labs.

57:35And there's a lot of hands-on labs around Cilium. And so what these will actually allow you to do is to, you don't even have to set up your own Kubernetes cluster. It's set up for you. Cilium is installed. And it'll walk you through the different features. So basically one is like, okay, how do we install Cilium? And it'll walk you through that. The next one is, okay, like I need to do network policy. How does Cilium network policy work? And there's a really famous Star Wars demo. It's like, how do we blow up the Death Star? How do we protect the Death Star? So that's like a fun lab. And it's different, like hands-on labs, you're actually in a Kubernetes cluster and walking you through these different features, how they work and how do you actually apply them to your cluster.

58:10Yeah, so those are all great sources. Obviously, there's always GitHub. We have a Slack channel if you want to jump in there too. But I would recommend if you want to get hands-on with Celium, go to the labs. I know I like to actually be in a cluster and be able to do things or just read through some of the stuff. And then there's also the documentation pages there too. Awesome. Sounds pretty fully featured on that front. So yeah, cool. Well, Bill, awesome to have you on today. Thanks a lot for coming on. As we were talking about before recording, you've been doing a lot of traveling, so you've managed to find a slot on that schedule for this.

58:44So that's really appreciated. And no doubt we'll be following along and maybe catch up again in a couple of years or something. Yeah, that'd be great. All right. Thanks a lot. Yeah, thanks for having me.

59:02Thank you.

From the publisher

Modern cloud-native systems are built on highly dynamic, distributed infrastructure where containers spin up and down constantly, services communicate across clusters, and traditional networking assumptions break down. Linux networking was designed decades ago around static IPs and linear rule processing, which makes it increasingly difficult to achieve scale in Kubernetes environments. At the same time,

The post Cilium, eBPF, and Modern Kubernetes Networking with Bill Mulligan appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
Cilium, eBPF, and Modern Kubernetes Networking with Bill MulliganSoftware Engineering Daily · 58 min
Listen in VO