ByteDance’s Container Networking Stack with Chen Tang

1 Jul 2025 · 48 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: ByteDance’s Container Networking Stack with Chen Tang

Podcast Overview

  • Title: Software Engineering Daily
  • Episode Title: ByteDance’s Container Networking Stack with Chen Tang
  • Description: This episode features Chen Tang, an engineer at ByteDance, who discusses the company's innovative container networking solutions, including the use of eBPF technology.

---

Key Themes & Discussions

  1. Introduction to ByteDance
  2. ByteDance is a global tech company known for platforms like TikTok.
  3. Operates at a massive scale with over a million servers running containerized applications.
  4. Challenges include maintaining performance and stability across their vast data centers.
  1. Introduction to eBPF
  2. What is eBPF?
  3. eBPF stands for Extended Berkeley Packet Filter.
  4. A technology allowing dynamic and safe reprogramming of the Linux kernel.
  5. Functions as a way to run sandboxed code in the kernel without the need to load kernel modules, improving performance and safety.
  1. Use of eBPF at ByteDance
  2. Chen Tang's role involves networking in ByteDance's data centers, utilizing eBPF for enhanced connectivity and stability across containerized applications.
  3. eBPF is particularly advantageous in container environments due to its lightweight nature and ability to efficiently manage packet processing.
  1. Networking Challenges with Containers
  2. Each container has a unique network namespace, necessitating efficient communication between containers and the host network.
  3. eBPF captures packets and redirects them to the appropriate container, bypassing the need for traditional virtual switches or heavier methods like IP tables.
  1. Benefits of eBPF Over Traditional Networking Solutions
  2. Performance: eBPF reduces the overhead associated with traditional networking stacks, leading to better performance.
  3. Simplicity: eBPF allows for more straightforward management of networking rules compared to complex virtual switches and IP tables.
  1. Challenges and Developments in eBPF
  2. Early versions of eBPF were limited in what functions they exposed and how complex programs could be, requiring ongoing development to enhance capabilities.
  3. Development focuses include improving verification mechanisms and expanding functionality for more complex networking needs.
  1. Integration with Hardware Offloading
  2. Future developments may involve combining eBPF with SmartNICs for hardware offloading, which can dramatically enhance packet processing performance.
  3. The challenge lies in translating eBPF logic to hardware rules, maintaining a balance between flexibility and efficiency.
  1. Strategic Differences Between ByteDance and Other Major Tech Companies
  2. ByteDance's approach to cloud-native technologies differs due to its relatively recent establishment compared to older companies like AWS or Meta.
  3. While legacy companies may be constrained by their existing infrastructure, ByteDance has the advantage of adopting newer technologies from the start, such as Kubernetes and eBPF.

---

Key Takeaways

  • eBPF as a Game Changer: The introduction of eBPF has transformed how ByteDance handles networking in its containerized environments, emphasizing performance and stability.
  • Scalability Challenges: As ByteDance scales its operations beyond conventional limits (e.g., 100,000 machines), it faces unique challenges that require innovative solutions like custom service discovery frameworks.
  • Future of Networking: The evolution of networking technologies, including the interplay between kernel-level operations and user-space applications, raises important questions about the future of system design and resource management.

---

Conclusion This episode provides deep insights into the innovative networking strategies employed by ByteDance and the transformative role of eBPF technology within the context of modern cloud infrastructure. The discussions illuminate both the current state and future potential of networking in highly scalable environments, making it crucial listening for software engineers and system architects.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00ByteDance is a global technology company operating a wide range of content platforms around the world, and is best known for creating TikTok. The company operates at a massive scale, which naturally presents challenges in ensuring performance and stability across its data centers. It has over a million servers running containerized applications, and this required the company to find a networking solution that could handle high throughput while maintaining stability. eBPF is a technology for dynamically and safely reprogramming the Linux kernel. ByteDance leveraged EBPF to successfully implement a decentralized networking solution that improved efficiency, scalability, and performance.

0:43Chen Tang is an engineer at ByteDance, where he worked on redesigning the company's container networking stack using EBPF. In this episode, Chen joins the show with Kevin Ball to talk about EBPF, the problems it solves, and how it was used at ByteDance. Kevin Ball, or KBall, is the Vice President of Engineering at Mento and an independent coach for engineers and engineering leaders. He co-founded and served as CTO for two companies, founded the San Diego JavaScript Meetup, and organizes the AI in Action discussion group through Latent Space. Check out the show notes to follow KBall on Twitter or LinkedIn, or visit his website, kball.llc.

1:35Chen, welcome to the show. Hi, Kevin. I'm excited to get to talk to you. Yeah, me too. Let's maybe start with, you can introduce yourself. Just tell me a little bit about your background and how you got involved doing cloud native stuff and networking. Okay. My name is Chen, and I'm currently a software engineer in Python. So my job is focusing on networking parts in our data center to keep every service in our data centers, especially deployed in containers, rounding, and to make sure their connectivity and stability. And in this field, we use a lot of different technologies. We are using kernel technology and hardware technologies.

2:19But in the kernel part, we use eBPF. It's a customized program you wrote and you can somehow put it inside the kernel without loading a kernel module. I think for us, the eBPF technology has been developed rapidly in the recent decades and now it's very popular. You can use it in not just networking, but basically you can do everything with eBPF. And you can just write your program and you want them to run inside the kernel. and you don't have to be afraid that your code might jeopardize the entire system. I think that's worth digging into, especially for older developers like me. I remember to get anything into the kernel when I started, it was this months or years long process where you would go back and forth on email and all these different things.

3:10But eBPF, as I understand it, lets you essentially run sandboxed code similar to how you might run JavaScript in a browser. Is that a fair analogy? I'm not quite familiar with JavaScript, but yes, I think you are right. You run virtual machines inside the kernel, and you have the environment prepared for you, and you just wrap your code, your program, and you get them run inside the system. That is really cool. So before we dive into some of the specific ways that you're using it, maybe we could talk a little bit more about eBPF. So what does the programming environment look like for these virtual machines?

3:49First, BPF means Berkeley Packet Filter. It is a technology developed by the UC Berkeley, and it was first to filter packet inside the kernel. And then when people find that this mechanism to do something inside the kernel, it can be used in a different way. So the people just expand this entire technology and we call it the eBPF. The e means extend, extend packet, Berkeley Packet Filter. But we can do more than just filtering packets. And we can run different code inside the kernel. You can filter the packets. You can monitor the entire systems. And you can trigger your custom functions when somebody loads a file.

4:34You can do basically a lot of things. And the interesting thing is that when we talk about kernel, we know it's a big system and it's fragile. If you do something wrong and you probably blow up the entire systems, but EVPF provides you with a mechanism. They check every line of your code. They make sure your code can run safely inside the kernel. So this will help provide developers with powerful tools for them to, they don't need to care about the check, the verifier, everything. They just run their code inside the kernel. If the verifier checks your code is not safe, that might blow up the system and they will tell you you cannot load it inside the kernel.

5:14But once you pass the check and you don't need to be afraid of that. So basically, this is what eBPF is about. So to make sure that I understand, it essentially provides a set of APIs or hooks into the kernel that are more stable than kernel internals. So you can hook into networking stack, you can hook into system calls, things like that. And on top of that, it does a static analysis verification up front to make sure that this is going to be safe to run. Yes, exactly. That makes a ton of sense. So I think this is super cool because there's a lot of times when you need to get down into the kernel.

5:57and I know writing application code, crossing the kernel interface, doing a system call, it's expensive. Yes. What was the motivator for you to start working in eBPF? Yes, let's go back to the networking part. Because why we need eBPF in networking? Because it's simple, because we come back to the container, cloud native things. We run multiple containers on the server and in each container, they own a unique network namespace means the network environment of the container is isolated from the host. So each container, they have its own IP address. They probably have its own network interface card.

6:35And that's caused a problem. So if your application inside the container and you want to send a packet to the outside world, and there is a wall between the container and the physical nick is the host, and you need something to connect between the container and the host interface. And that's we use eBPF. And the mechanism is simple. Let's assume we run a customer eBPF program in our kernel stack. When we receive a packet and the eBPF program will capture the packet when we analysis the packet and we see, okay, this packet belongs to this container and we send them to the container. That's simple.

7:19That's why we need eBPF in container networking. just for a very simple target. Got it. So to make sure that I understand when you're in this container environment, inside the container, it doesn't know it's in a container. It wants to treat its network like a network stack. But outside, you can look at the kernel layer that's running all of these containers. You can say, hey, this packet actually is going to another machine or a container inside of my same virtual machine. There's no need to go through an expensive interrupt stack to go through an actual physical device. I know in software, I can just route it straight to that other container.

7:54Exactly. It's a lightweight virtual machine. You can think container is just a lightweight virtual machine. Yeah, that's really interesting. So what were you doing before and how did you kind of realize you needed to do something like that? Because let's go back to the network virtualized things. Before we run eBPF for container networking, what we use, We still need network virtualization for virtual machines, right? We have virtual switch. I'm not sure you're familiar with virtual switch. It's kind of like a switch running in a software road switch on your node. And when you receive a packet, that virtual switch will capture the package and determine where to send it out.

8:41So basically, these virtual machines, they need dedicated cores to run them because to achieve the maximum performance. So it's heavy. Virtual Switch is heavy and the virtual machine is heavy. And this is for the, well, for most cloud provider because virtual machine is a technology probably like 20 years ago. And container is something new. And the feature of container is lightweight. We don't want those heavy things, but we want to achieve the same purpose. We want isolation. We want security. And we want resource designation to make sure the container is running in a safe environment. And we need two cores, probably four gigabytes of RAMs.

9:26And we make sure the connectivity between the container and outside work. But we don't want that virtual switch. And what we need to do next, and let's go inside the kernel to see what we have. what the kernel can provide. And the kernel provides us with eBPF. But at the very beginning, the eBPF is still not functioning as what we have today. And there is still a lot of work to do. So there's many develops. We get together to think how to make eBPF more powerful for container networking. So it's been developed for five years. And now I think it's mature technology for container networking. So basically you have a very many popular projects, open source projects, aiming to use eBPF to serve for container networking in each aspect.

10:14For example, not just connectivity, but like security, and you bring a lot of cloud-native concepts inside eBPF. That's interesting. So actually, let's maybe walk through that journey a little bit. When you started working on this and you said, hey, we have these heavy virtual switches. We want to get rid of those. And you looked at eBPF. You said it was not ready. What was missing and what were the things that you needed to add or develop to get this to work for you? Yeah. So the thing is you use eBPX, but fundamentally you are leveraging the ability from the kernel. And eBPX, they help you to expose some of the key functions from the kernel for the user.

10:59But at the very beginning, that function exposed is not enough. So you will have these problems. The performance is not good enough, and you can write a very simple program. If you write the program more complicated, like I said before, we have a verifier. And the verifier will reject the program because it's too complicated for you to recognize if it's safe or not. So you will encounter many problems. So the people keep working on this, keep optimizing the entire system. So basically, this is what all the eBPF programs have been doing in these five years. Yeah. To make sure I understand, they were, one, extending the surface area of what you could do, what the hooks were, things around that.

11:48And two, improving the static analysis verifier to be able to handle more complex cases and still prove them to be safe. Exactly. Got it. Okay. So you mentioned pulling a lot of different cloud-native tooling into eBPF. What are some of the different ways in which you are using this in your stack today? Yeah. So basically what we are doing now, focusing on two parts. The first part is networking. We have packet direction, and we have to enforce different networking security policies. It works as a data plan. We receive policies from the remote controller, and the eBPF help you to implement those policies.

12:30And the second different part is the kernel tracing. Like we said before, the eBPF, it provides a lot of hook points inside the kernel, and we call them trace points. And each trace point, you can, like, when the trace point is triggered, and you can run your custom eBPF program to help you to understand what is going on inside the kernel. Like, for example, when the functions is called and you assume some bugs has happened and you want to know what is going on inside these functions, and you can bury your eBPF program inside this hook point and when the event is triggered, eBPF will help you to print everything you want, especially the context at this moment, and to deeply understand what is going on inside the kernel.

13:16Because most of the time, it's a black box. But with the help of eBPF program, you can really understand the kernel. It's almost like you can insert your observability stack down into the kernel and say, oh, we think there's something going on with this function. Print me the context on entry. Print me the context on exit. How much overhead is there in this tracing? The kernel tracing, actually, the cost is very high. So we don't use them often. We use them only when there's a bug is reported. We need to analyze the kernel. But for packets filtering, because there are different kinds of trace points inside the kernel, some trace points, the cost is low and some is high, especially when you want to bury your customized trace points inside the kernel, the cost is very high.

14:02But for most, especially in the networking part, all the trace points, that is not a trace point. The hook point, the trigger of those hook points, the cost is low. So we use it for networking. We direct tons of millions of packets and the costs can be limited. Let's see, the cost is acceptable. Yeah. Now, what does it take to roll out? Like, say you want to turn on tracing. Can you do that live on a running container? Do you need to restart the server? What does that lifecycle look like? Yes, and this is the attracting part. You can run the eBPF, the kernel tracing in the live container. You don't need to restart anything.

14:47You just inject your code inside the running system, and you get everything you need. And once it's done and you just cancel it, you can remove the program, and the system will become back to normal. That's wild to me that you can inject observability code down into your kernel on a live system with knowledge of safety, run it as long as you need it and just pull it out. Yes, exactly. That's why eBPF becomes so popular today. That's cool. So diving in maybe a little bit, you mentioned also there's a set of open source projects that have been building and bringing these cloud native technologies into eBPF.

15:27Are you working on any of those or are those areas that you're connected to? Yes. For now, I've been working on some of the open source projects, but I'm mainly a user because I work for a company. So my purpose is first to serve what the company wants me to do. And I will use a lot of open source tools. And when I found there's things I can modify, and yes, I will like to, it's kind of my way to, how to say this is what community do, right? Yes, absolutely. Which projects are you using? Since I first worked for the company, so my top priority is help the company to get things done, to get the job done.

16:09And during this process, I use a lot of open source projects and I get help from the communities. And especially there's an eBPF library called eBPF Go. Now it is the most popular eBPF library for Golan developer. And I use this library to help me to load eBPF program inside the kernel. I think this is the tools I use the most, the open source project. That brings up kind of an interesting area, which is what does the development environment for eBPF look like? Because you're compiling down, as I understand it, to essentially a bytecode that is what gets analyzed and loaded. So what is the environment that eBPF Go, for example, exposes to you?

16:56Okay, so for most developers like me, we wrote the C program. We want the C program to run inside the kernel. And once the C program is done, first we need to compile them with bytecode. and we load them into the kernel. And before we load them into the kernel, we need a verifier to make sure all the code can be run safe inside the kernel, and to run a syscall to the kernel to load the program to the bytecode inside the kernel. And after that, once the program is inside the kernel, it's still not functioning. You just somehow, this kernel helps you to store the program. And if you want them to run, you have to attach the program to a hook point.

17:39And this is what those libraries help you. You develop your program in a C code, but the library helps you to do the load, to do the verify, to do the attachment. Everything relates it like, yeah. Got it. So the core eBPF program is developed in C, but all of the different sort of attach, detach, load, all of this stuff is what you're using eBPF Go to do. Exactly. Awesome. Awesome. I'd love to go back a little bit to the use you used in the network. What was the impact of moving from this big virtual switch approach to the eBPF networking stack? Okay, so that's a very interesting story because it is related to the company's development.

18:23So let's take a look at the companies like Meta or Amazon. Those companies, when they run their data centers, it's probably like 20 years ago. So by the time the technologies they had for virtualization is virtual machines and the virtual switches. So this is what they used at the very beginning. And now we have to know one thing. If you migrate from one technology to the others, especially you have to upgrade the entire data center, the cost is huge. So probably they will keep using virtual machines. They keep using virtual switches for their data centers. But I work for Bythons. I don't know if you know ByteDance, but I think you know TikTok is one of the apps developed by our companies.

19:11ByteDance is founded in 2012. So it's only 10 years ago. So when ByteDance is growing and we need to build our data centers, cloud native technology emerged. So at this time, we have a better choice to manage our data center. That's why we use cloud-native technologies with Kubernetes to run everything as a container in our data centers. So I think this is luck. I don't know, because at this time we have better options and we still have the chance to make a choice. And we choose Kubernetes and container. That's really an interesting point of how coming later, you can often jump over the learnings that the earlier companies had to do.

19:59Maybe we can expand a little bit now into some of these other cloud native areas. What other places do you feel like going cloud native first has made a difference for ByteDance relative to, say, Amazon AWS or something like it? Okay, so now most companies, when they embrace cloud native, first let's make a statement. For companies like Meta or Amazon, they run their data centers in a different way. So when cloud-native technologies emerge, of course, everyone wants to use this, but we use them in a different purpose. For Amazon or Meta, they use cloud-native technology as a service, especially for Amazon.

20:43They are a cloud provider. They use cloud-native technologies mainly as a service for their customers. They don't use themselves. They provide them. But for us, we use ourselves. So the difference is since we have the Kubernetes running inside our clusters, and I think the biggest difference is the problem of scalability. For Amazon, they use Cloud Native as a service for their customers. And their customers, they mainly have a small scale of cluster because they cannot afford to build their own data center. That's why they want to buy machines from Amazon. So the scale of their clusters, let's assume 1 ,000 machines, I think that's a lot.

21:27But for us, we have over a million. So the scalability is the major problem. And we found that the cloud-native technology seems it's very powerful, but it has a fatal problem is the scalability. If you run a Kubernetes on a cluster with 1 ,000 machines, that's enough. and you can have a powerful cluster management tool with basically everything you want. But if we have 100 ,000 machines and you have the Kubernetes, it becomes a bottleneck of performance. So by the time, if you still want to use Kubernetes, you have to do a lot of modifications. Some of the concepts from cloud native ecosystem is powerful, but it's just not so efficient.

22:16we have to optimize them. So I think this is a major difference from how Python uses cloud-native technologies from the others. Yeah, there's a scale factor there that is really, blows my mind. You mentioned finding progressive bottlenecks in Kubernetes and cloud-native abstractions as you scaled up to 100 ,000 and beyond. Where were those? and what have you replaced or improved? Yes. For example, there is a very important concept in container networking. We call it service. Because once you have a client container, you want to talk to your server container. You cannot just ask for the IP address of your destination and you send the message to the destination IP address.

23:04In Kubernetes, they build a concept called service. You ask for the service and the service will return to you a target and they say you can connect to this destiny. But the service, the total cost for you to run a service is huge because a service, you have to store all those backend containers. You have to choose from one of the real server. If you have 100 servers, that's okay. But if you have 100 ,000 and the cost to run this service concept is huge, it's not acceptable. So we have to remove them. So we use a different framework. is like a service discovery. We develop our own service discovery framework to help the client container to discover their target.

23:51So this one difference. And also because for us, because we run our own data centers and the cost things is in recent years, all those IT companies, they're not running so well. You see there's layoff everywhere. so we can see the cost becomes a big problem for all those companies. They want to save money. They want to buy less servers because the server is the biggest cost for our company. So let's see if you have 1 % of improvement for the total cost of the machine and we have over a million servers, the 1 % would be a lot. So this is what drives us to seek for new solutions to optimize the entire system.

24:40We know Kubernetes is powerful. The cloud native concept is very useful, but sometimes it's just not what we want the most at this very moment. Yeah. So let me make sure I understand the service example. So in Kubernetes, you have the service abstraction, and it kind of assumes a global view. is going to index all of the different servers or containers that might respond for this service type. And when a client asks, it looks up in its index, who's free, sends them the target. That kind of assumes a bounded set that is small enough to perform quickly in that type of index. And when you scaled up large enough, you needed to say, oh, we actually need a much more contained version of service discovery that's not going to be trying to index 100 ,000 machines.

25:35Exactly. That makes a lot of sense, I think, and points towards a potentially even just class of problems where Kubernetes is providing this nice abstraction where it's going to do all the management for you of all these different things, package it up in a single location. And that's going to run into scalability concerns as you go far enough. Maybe you bump out of a cache or a memory or something like that. Were there other areas that you found needing to kind of almost scope down the abstraction exposed to be able to deal with that level of scale? Actually, I'm not expert in this field because we have a different team that manages the resource orchestration.

26:21And they are the teams to build the Kubernetes systems. I'm focusing on the networking part. What I know about the service things reaches my knowledge boundary. Okay. Looking within networking then, again, we talked a bit about using ePPF to make a much sort of more efficient networking stack within the physical machine. So you connect to other areas of the networking stack that being cloud native from the beginning has really made a difference. Yes, we have IP tables. That is old technologies being first used when cloud native. Actually, at the very beginning, when container networking becomes a problem to be solved, people use IP tables at first.

27:11But IP tables, it's slow. and then we use eBPF to replace IP tables. So now eBPF becomes the majority and IP tables just emerged from at the very beginning and then it gets replaced because of its bad performance. So I understood the eBPF approach to network device, replacing the virtual switch. Instead, you just drop into this kernel hook that looks in different places. What does it look like for IP tables? Iptables is also a function inside the kernel stack, but Iptables works in a chain format. For example, if you can write a chain of rules, you have the match and the actions, the packet match and hit these rules and will proceed with the action, it is a chain.

28:02And you can pass the packet to another chain. So the entire Iptables system is just a bunch of chain, it's a chain system. And we can say it is not efficient. And especially once the rules expand and the cost will not be acceptable. But the eBPF is a single program. You write your own program and one program will be enough. So that's why people replace IP tables and choose to use eBPF. And when it comes to the virtual switch things, I remember you also asked how eBPF replaced the virtual switch. And no, So EBPF doesn't replace virtual switch. They just work in different scenarios for different purposes.

28:45Got it. Let's maybe go in on there because I misunderstood. Which scenarios are you using the EBPF networking direction versus the virtual switch? So why we need a virtual switch and EBPF? Because they serve for a different purpose. EBPF focusing on container networking. Why eBPF can be used in container networking? Because it's a lightweight, isolated environment. There is still only one system, one kernel. Even containers run in an isolated environment, but actually the kernel is the same as the host one. So the eBPF just play the tricks on the host kernel help you to redirect the packet from the physical nick to the container.

29:34But in the virtual machine, virtual machine is a full isolated environment. They use a different kernel than the host one. So it's impossible for the host kernel to talk to the guest machine. So you need a different mechanism to redirect the packet to the guest machine. That's why we use virtual switch. So basically, these two technologies serve for a different purpose. To make sure that I understand it, the way that I'm seeing it right now, it's almost like there's layering. So if you have containers within a single machine or virtual machine, you can route between those containers purely with EVPF.

Read the full transcript

30:12As soon as you start to go outside of that machine, you need to go over an actual physical NIC or a virtual switch if you're going to a virtual machine. Yes. Okay, that makes sense. I'm curious then, from a performance gains standpoint, how much of the traffic that you're directing stays within that single machine and is able to leverage eBPF? And how much still ends up going past? What types of performance gains do you end up with on an aggregate basis? You cannot make a full comparison between eBPF and the virtual switch because the difference is the performance varies between how you use them.

30:59So the main advantage of eBPF is easy management. If you want to run virtual machines and you have to do a lot of preparation for that and you have to reserve a lot of resources simply for the virtual machine, it cannot allocate those resources to guest hosts. But for eBPF, you do have to consider that because all eBPF programs runs inside the same one kernel. So all these things make sure it's much easier to manage a container than the virtual machine. So I think this is the biggest advantage of using eBPF. But when it comes to the performance comparing between eBPF and virtual machine, it's hard to see which one is better.

31:46And I think in a scenario, let's say the machine is full loaded. I think the virtual machine can have a better performance than ABPF. That is the fact. Interesting. What are the scenarios in which you still want to have virtual machines? It feels like containers is the cloud native way to do it. Okay, so when we want to use a virtual machine, let's go back to AWS as a cloud provider. You run the virtual machine not just for your own services, you run the virtual machine for your guests. The people buy your virtual machines, they run their own applications. So what they want is full isolation and security.

32:25That is top priority. They don't want their information to be accessed by, even they run their service on AWS, but they don't want their data to be accessed by the cloud providers. So AWS, they have to build a full isolated environment for their customer. And that's why they still choose using virtual machine. But in ByDance, the entire data center is run. We run our own service on the data center. We don't have this requirement for security and isolation. So a lightweight method will be enough. And what we want to achieve is the easy management. That makes sense. Okay. What do you see as the next frontiers in this space?

33:15What are you working on for eBPF or within the networking stack that you think is taking this to the next level? Yes. I think what we are considering at this moment is still the cost. We see the eBPF brings a lot of advantages, easy management, but still the cost of the kernel stack is still inevitable. Because if you look into the kernel stack, we have multiple interrupts and memory copies. When we receive a packet from the NIC, we have the first copied packet from the NIC to the kernel stack. And when we go through the kernel stack, we have to copy the packet from the kernel to the user application.

34:00And this cost is inevitable. And it becomes when we want to optimize the entire system, there is no way for us to ignore this cost. And this is the bottleneck for eBPF technology. Even though it's popular, it's because it's easier to be used. But if we want to save more resources, we have to optimize eBPF. At this moment, for us, that's why we have several solutions. First is NetKit. we mentioned before, it helps us to reduce one interrupt when packet is transmitted between the container and the host. NetKit using, let's say, a special mechanism helps us to reduce one interrupt, but that is not enough.

34:46So we're asking help from the hardware. And now what we are doing next is to combine eBPF with hardware offloading, because we know the difference between eBPF is that it's powerful. We can write our own custom program in the kernel, but the cost is higher. But once we leverage the ability from, especially now the SmartNIC, we have a hardware interface from different vendors. They help us to offloading packet processing ability from the kernel to the hardware. But the problem is difficult to use. You cannot write your own program inside the hardware. You can just inject the rules or policies, like what I mentioned, like IP tables.

35:33So then somehow we need to translate the eBPF program into hardware rules and to load those rules inside the hardware to using SmartNIC help us to process the package. And this is what we are doing at this moment to, let's see, to achieve the best performance. So kind of curious there, is there a standard for how those rules are defined for the hardware offloading? Could you create a compiler essentially that takes your eBPF rules into them? It's predefined. Yes. Let's say if we write an eBPF program, everything can be defined by the code. So we can write whatever we want. But let's see a typical hardware offloading rules.

36:18It's just like IP table rules. you have a match, you match the header of the packet, and you have an action. Match action, match action. You have multiple rules with different priority, and each rule, they have a match field and action. Then you need to write a different program to translate the eBPF program to those rules that behave the same. So this is the difficult part. And what we are doing now is we combine these two technologies because the rules cannot be predefined. If you predefined everything you want into rules, first we don't think it's possible. It's just too difficult. And what we are doing now is we have a separation of the slow path and the fast path.

37:12The fast path is hardware offloading. And once the packet is done, miss the rules and the packet will go back to the kernel again and the eBPF program will process the packet. And when the eBPF program decides, okay, this packet will allow them to be processed and we will inject a rule. The match of the rules is the header of this packet, the destination, the source, IP address, and the destined port and the source port. And the action is, as we say, redirect to container A. And we inject this rule to the hardware. And the hardware will recognize the following packets. And this is what we are doing.

37:57That's interesting. So to make sure that I understand, you're in some ways treating the hardware as a almost caching layer of rules, where a packet comes in, starting from a blank state. A packet comes in, you don't have a rule for it. It goes to the networking stack. Your eBPF code picks it up, analyzes it, says, okay, here's where this needs to go. And furthermore, here's the rule that the hardware can use to do that fast next time. It loads that up into the hardware, which then for subsequent packets with similar patterns or following that rule, knows what to do. Exactly. What's the sort of lifespan of those rules?

38:43Are they durable or is there like a timeout or how does that work? Every time a packet matches the rules, we have a counter. And we see when this rule is being idle for about 30 seconds and we recycle them because there's no way for the rules to delete themselves. We have to delete the rules. But since if, let's assume the root is in the hardware and all the packages, the kernel cannot capture the packet anymore because the kernel has been bypassed. So how could we know, okay, the session is over. We need to delete them. There's no way for us to know that. So what we're doing is that we're running another program on the host.

39:26They periodically to fetch all the roots from the hardware and to analyze them. if this rule has been idle for like 30 seconds or one minute, we have a timeout setting and we just recycle the rule to make sure there's no rules leak inside the hardware. That makes sense. That's cool. And you're able to get all the data you need from the hardware itself or does there need to be some sort of communication between the eBPF that sets them and the program that's clearing them? Well, for now, we use the user, we run another agent to fetch all those data from the hardware. And since there is no way for eBPF to communicate to the hardware, since this part is still missing, and maybe somehow in the future, we can find a way for eBPF program to talk to hardware directly, but now there's no way for us to do that.

40:19How does it set the rules then, if it can't talk directly? We first let the eBPF program to talk to our agent. There is a channel for the eBPF program to communicate with the host application. And the agent will analyze the message from eBPF and the agent will translate the message to our hardware rules. And the agent will help us to inject the rules to the hardware. This is what we're currently doing. But we see there is a two-way communication from the host to the user space and from user space to the hardware. And the user space application will periodically fetch data from the hardware. And yet it's not, it does look a bit ugly, but there's no better ways for us.

41:03Now, I think, let's see in the future. I hope we can find a way for the EPTF program to talk to hardware. Honestly, that kind of makes sense though, because the agent needs to be somewhat durable. It needs to run every 30 seconds, check things, keep track of it. Whereas EPPF, as I understand it, is event-driven, right? It's always happening. Just a thing comes in, we do it. So to make sure that I understand the whole thing, that you have this durable agent that is responsible for keeping track of what rules are currently on the hardware and translating when there's an EPPF rule that triggers, move it into a hardware rule.

41:44And so network packet comes in, misses the hardware, cached logic, goes into eBPF, eBPF applies its logic, puts a rule on, and sends a message to your agent. Your agent then says, ah, here's a new rule. Let me push that up into hardware. Yes, it's a bit complicated, right? It is, but I think it's quite clever. That's cool. What else are you doing in this space? And besides this, since we have rules offloading in the hardware and now we can leverage the RDMA technique. I'm not sure if you're familiar with RDMA. You essentially map user space memory to the NIC and it can send it directly? Yes, exactly.

42:29And this will, because RDMA is a technology developed by our NIC manufacturers called Menalox, now it's a part of NVIDIA. Since we're using eBPF, help us to leverage the ability of hardware offloading. And now on top of this, we can use RDMA. So let's talk real quickly about what that looks like. So you send, eBPF figures out where it needs to send the network packet. Does it know the memory that is mapped for the NIC to be able to go? Or does that go through your agent? Or like, what is the flow? No, the eBPF doesn't know anything about this. eBPF only knows about the packet redirect. We need to pass this packet or drop it.

43:15And once we translate the eBPIF policy into a hardware rule and we push them into the NIC, and the subsequent packet can be goes directly through the hardware and bypass the kernel. And at this moment, we use our RDMA because RDMA provides a different mechanism of packet transmitting help you to especially is mainly bypassing the kernel and since the hardware know where to route the packets and you can use rdma to talk directly to the destination and there's no need you don't need to involve kernel anymore so and it's been many years since i did anything with rdma if i understand you need to give it not just a network address but you need to actually give it the location in memory.

44:05Is that correct? Yes, this is what RDMA is doing. But let's see, if we have two physical machines running NIC support RDMA, you can use RDMA directly because the NIC is working. They know exactly where the destination is located. But in container networking, the NIC doesn't know at the very beginning. So this is the trick. You need EDPF to help the NIC to identify the destination. and once you provide the connectivity for the NIC, you can leverage RDMA. Got it, got it, got it. Yeah, because previously RDMA is being used in host machines, all the applications that deployed on the host. They've never been deployed on the container because once you have your application on the container, there's no way for the NIC to locate the destination.

44:57They don't know where the destination container is. But since with the help of eBPF, yes, it works. That's clever. So when you do that sort of fallback, eBPF locates it, sends this information also to your agent, which can then load all of that back up into the hardware, and you can bypass the container boundary and do RDMA straight to a container. Yes. That is super cool. Well, I think we've covered a lot. And I actually really like the example with the hardware because it connects not just this kernel piece, but shows how this can be a bridge between container technology and essentially the old world.

45:37Anything that was living outside of the container world didn't understand it. You can kind of intercept with eBPF, pass that data off to an agent that understands both sides and make this sort of bridging connection. Yes, and you also mentioned to me about one thing I want to share, that is the relations between the old world and new world. Because we see the Linux kernel is very big and it's been developed for decades. And for us, for most people, it's kind of old. And nowadays, we have more technologies emerge from different, from open source communities or from hardware manufacturers. and the people will say that the kernel is big and there's no way for us to change the kernel by our will, but we can find a way to bypass them.

46:32So we see, but we actually are on a crossroads. What we will choose next, if we embrace the technology that bypassing kernel or that we embrace the technology that we still stick with the kernel, actually we are facing a choice. And I think it's a very interesting topic And actually, I don't have the answer yet because we're still wondering if they can coexist in the future or they have to be enemies. And I think it's a very interesting question. And yeah, I'm just bringing it up. And actually, I don't have answers for that. It is a really interesting question. I feel like, and it's been around for years, how much can you do in user space?

47:16How much can you do without having to jump down into the kernel and face the performance costs that go going back and forth over a system call? It does seem like eBPF is a nice middle ground where you can write your own custom code, dynamically load it without having to restart or do anything with your kernel. And yet it runs inside the kernel and has access to all the privileged information. Yes.

47:46Thank you.

From the publisher

ByteDance is a global technology company operating a wide range of content platforms around the world, and is best known for creating TikTok. The company operates at a massive scale, which naturally presents challenges in ensuring performance and stability across its data centers. It has over a million servers running containerized applications, and this required

The post ByteDance’s Container Networking Stack with Chen Tang appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
ByteDance’s Container Networking Stack with Chen TangSoftware Engineering Daily · 48 min
Listen in VO