AI for Network Management with Shirley Wu - #710

19 Nov 2024 · 54 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: AI for Network Management with Shirley Wu - #710

Episode Overview In this episode of *The TWIML AI Podcast*, host Sam Charrington interviews Shirley Wu, Senior Director of Software Engineering at Juniper Networks. They discuss the role of machine learning (ML) and artificial intelligence (AI) in transforming network management, examining various use cases that enhance network quality, performance, and efficiency.

Key Themes

  • AI and ML Applications in Networking: Exploring how these technologies are being integrated into network management processes.
  • Data Science Integration Challenges: Discussing complexities and trade-offs when merging traditional networking methods with AI solutions.
  • Feature Engineering: The importance of data in feature engineering for network-related applications.
  • Future Directions: Insights on future advancements in proactive network management.

Key Takeaways

Introduction to Shirley Wu

  • Background in software engineering and data analysis before the advent of big data technologies.
  • Early career experiences in cybersecurity and AI-related technologies in networking at Juniper.

Importance of AI/ML in Networking

  • Networking is a mature domain, making it suitable for AI integration due to established heuristics.
  • Juniper’s approach focuses on cloud-based solutions to capture and analyze data from networking devices.

Use Cases for AI in Networking

  1. Cable Degradation Diagnosis: Using ML to detect and predict cable performance issues.
  2. Proactive Monitoring: Identifying potential network coverage gaps before they affect users.
  3. Real-time Fault Detection: Implementing anomaly detection models for ongoing performance assessment.

Challenges in AI for Networking

  • Feature Engineering: The necessity to work closely with domain experts to define problems and ensure data adequacy.
  • Model Complexity vs. Cost: Maintaining a balance between model efficacy and operational costs, especially given the scale of deployments.
  • Multiple Model Deployment: Utilizing a range of smaller, specific models rather than relying on a single large model.

Technical Insights

  • Reinforcement Learning and Simulation: Juniper uses reinforcement learning to optimize network configurations dynamically.
  • Multi-Feature Time Series Models: The complexity of analyzing multiple features in time series data for network performance.
  • Anomaly Detection Models: Challenges in deploying effective models across a large number of sites without high costs.

Future of Network Management at Juniper

  • Proactive Network Testing: Aiming to identify issues before they impact users by using data from deployed devices.
  • End-User Self-Service Capabilities: Development of features allowing individual users to diagnose their network issues, potentially easing the burden on IT administrators.

Conclusion Shirley Wu emphasizes the importance of integrating AI and ML into networking to improve the user experience while also addressing the complexities involved in such implementations. The conversation highlights the ongoing evolution of network management solutions and their future potential.

---

Additional Notes

  • The episode stresses the significance of data collection and analysis in improving network performance.
  • Future advancements in Juniper’s solutions point toward smarter, more responsive network management tools.

For more details and the complete show notes, visit [TWIML AI Podcast - Episode #710](https://twimlai.com/go/710).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Recently, there is a pre-trained time series anomaly model, right? That's also exactly the need. But there's a key difference that existing pre-trained anomaly model time series is for single feature. What we are talking about is multi-feature. So, for example, you want to look at the throughput for your given site. You definitely should look at the number of clients and the total traffic volume we are going through and how many destinations they are going through and how many different applications they are routing the traffic through. So this is going to be multiple features.

0:47All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Shirley Wu. Shirley is Senior Director for Marvis at Juniper Networks, where she leads the data science team. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Shirley, welcome to the podcast. Hi, Sam. It's a great pleasure to be on your podcast. I'm looking forward to jumping into our conversation. We're going to be talking about the many ways that you're applying data science and machine learning to the challenge of ensuring high-quality network connections for your customers.

1:27before we jump into that topic i'd love to have you share a little bit about your background yeah so yeah and so i started my career as a software engineer focusing on the data analysis you know trying to find the insights from customer data that is beyond before the big data hadoop and the cloud um at that time there are a lot of consistent challenges regarding the data volume and the latency issue. Then big data, Hadoop, technologies and the tools and the cloud computing came to the world and quickly took over the whole industry for the data analysis. So the capability to process large amount of data really accelerated and the development and deployment of AI-related applications.

2:22So the primary application we're doing the business insights are definitely not enough anymore. We quickly moved to the capability to predict and forecast for customers' needs. So at that moment, I took the leap of my career to get into the startup. I joined the early stage startup in cybersecurity, trying to use AI data to solve the cybersecurity-related problem. So we are, you know, one of the first industry, trying to profile user behavior to identify security threat. So that company eventually was acquired by Aruba Networks. Then that's our first early stage startup. Then I came to Myst. Myst was a Wi-Fi company, but it's different from the Wi-Fi networking company.

3:14It is also a cloud company. So I joined the minister to start to develop AI-related technologies. So the company spent quite a few times, quite a lot of rounds of development to cloudify all the networking Wi-Fi access points and build the cloud infrastructures to be able to capture the data. Then we have data in the cloud. That's pretty much the prime time to develop AI ML solutions. So I was kind of architected with a group of very talented data scientists and engineers to implement Marvice self-driving network solutions. And currently, we continue expanding Marvice to the end-to-end networking, starting with MIST with Wi-Fi-only deployments.

4:05Now we are developing into the Wi-Fi, wired, and the van, trying to utilize the data and AI to help the customer to have a better networking experience, have a better internet experiences. When I think about networking systems, and I spent my early career at AT &T working on data networks and configuring routers and all that stuff. So I have some degree of familiarity with all that. You know, I think a lot about the type of data that, you know, those systems tend to generate. And it's mostly log files and like time series data. And to a larger degree, you know, while that space is evolving, you know, rapidly, some of the most exciting changes in AI haven't necessarily.

5:01you know included time series it's been about language and images so i'd love to have you share a little bit about how you think about the evolution of mlnai from a time series perspective and also more broadly you know networking as a application area for mlnai yeah so your question is right spot on why networking is really the domain suitable for AI and ML. First is networking is a mature domain, right? Not much of the significant, you know, revolutionized changes for the networking domain perspective. So because of a heuristic and the domain level, it's maturity, so that's capable for AI and ML.

5:55So this is the first layer. Second is the MIST. We started to develop the network gears, purely like an IoT device. So purely, we have a visibility about the stats and events from the cloud. So as the device started to power it down, and they auto connect to the cloud, starting to report its stats and events. So in the cloud, we have a large amount of data purely can clearly identify current state of that device. So that gives us a capability to do the AI and ML. Now, a third piece is about revolutionize from AI and ML perspective. Like you said, currently a lot of revolution development in the AI ML is focused on the image and the large language model.

6:53Time series, from people's perspective, it's mature already. It's a down problem, right? But for networking, it's another domain. If you talk about time series to forecast a particular stock market or for the advertising, for financial marketing, that pretty much is mature. But for networking, it's a special domain. A lot of the, you know, you go to market, you can find a lot of the people that have a good experience with data science and ML in advertising financial services, not necessarily networking domain. Networking domain is a very special domain. So it's really hard to find that type of talents, which they know the expert about networking and also about AI and ML.

7:44You know, when you think about time series as a problem class, a lot of the things that you're trying to do are very similar in terms of prediction. Are the problems fundamentally different in networking? Problem is fundamentally different. For example, for Wi-Fi, how the Wi-Fi mobile client to join a network, there are multiple phases it goes through. So this we have to sit together with our domain experts to understand multiple phases and each one of handshake between your Wi-Fi client with your access point, the acceptable latency, which layer they're going to retry. So this is a type of use case perspective is purely well-defined while we work in a specific use case with our domain experts.

8:35So this is one point. Second is, like you said, if we can really interpret the domain-specific problem with the data, then the problem is pretty much done. A lot of times the data scientist team, we're working together with the domain expert. first to identify the true problems they discover at the customer's site. Then we study the data to make sure our data be able to interpret, have all the descriptions about that particular problem. So in the data science perspective, they call it the feature engineering, right? So when we do the feature engineering, a lot of times we kind of really identify that, hey, maybe we even don't have the right data.

9:15So our device did not send that particular stats or events to the cloud. Then we had to go back to work on the AP firmware side to able to capture those stats, send it to the cloud. Then we started to redo the feature engineering. Then the third piece was that when we identify this type of issues, for example, for retail customer, because their Wi-Fi configuration is different compared to college campuses, which also could be different compared to the business offices. So we take that particular domain, the deployments as a problem. We also try to generalize for college campus and for the campus deployments, for enterprise campus deployments.

10:04So in this case, we have a generic solution, not just for a specific one customer, which is not going to work for customer B, right? So it's the three layers of the problems that we are talking about. Yeah, it's kind of lay out the problem. You've got these Wi-Fi access points that are spread throughout whatever facility and you've got these clients that are trying to connect and you ultimately want to have the clients have the best experience possible. One aspect of the problem that's kind of offline is there also an online component where you're deploying models to either clients or access points that somehow aid them in making the right connections to one another that facilitates a better experience?

10:52Yeah, there's multiple fronts, multiple layers, actually. AI ML, it's not just one solution. It's every front. So, for example, let's say I can give you multiple examples. So first to start is what is the feature you are talking about, the auto radio resource management. So when a mobile client is trying to join an AP, you need to make sure the client can hear the AP on the particular channel. You can hear the Wi-Fi signal strength should be optimal, right? So you deploy the access point everywhere in the parking lot, inside the roof, on the ceiling of your room, and in the dorms. So every environment that also, you know, maybe the other radio signals, you know, if you're close into the airport, if another, you know, if the shopping mall in the retail store, the next door is another access point maybe from the different vendors, right?

11:46So all this very dynamic environment, we need to make sure our access point have a right signal and on the right channel so that as a client for this particular store, they are able to hear you clearly. So in this case, we have the models deployed inside the cloud and they are adopted for each environment. And based on, you know, we will call the reinforcement learning. So we select a particular channel for every site. And based on the feedback information from the stats and the events, we continue to fine-tune the selected channels and the power strengths. So this kind of more like a global model, but being fine-tuned for each site.

12:37And also based on the feedback loop, based on the stats and the events data, we continue to fine-tune it for each environment. So it's kind of a wonderful use case, right? Of course, when we start this journey, you first need to do the offline analysis to define that type of model. Then we deploy it to the cloud. Then utilize each side data to fine-tune the model. So this is kind of a wonderful use case. So this use case actually compared to traditional, you know, networking company for Wi-Fi, they typically use a controller. Controller is a piece of equipment. they deploy it at the customer's site.

13:13Then you need the network administrator to manually configure that controller, right? So in this case, our solution is purely utilized cloud, utilized the data for each access point. We totally eliminate that extra components. So this is kind of enable our customer, for example, one of the largest retail customers, they have a very small team. they can support the Wi-Fi deployments globally all their retail stores. So our solution enables them to do that. Another, of course, we call it the AI for operation. So we use AI now to support the customer's operation. What does it mean? Is that you first get your networking configured, deployed, everything's great.

14:02Then our SC work away, so there's our networking running. But, you know, typically network hardware companies, they do is if there's a problem, your IT administrator just pick up the phone working with the support to get the issue resolved, right? For us, actually, after the first initial site going alive, the data and the events continue streaming to the cloud. We monitor your network operations. If there's something that's not right, we are supposed to be able to notify you ahead of time. So, for example, we are monitoring the site, number of clients, you know, the counts. This is just, you know, if you're the college campus, you can see daily trending, right, up and down.

14:50You go into the college room and go to the classroom, then they come home and back to the dorm. You know, the line is up and down, up and down. We can build that line to build the anomaly detection model for every site. We can notify you of a normal sudden drop of a client count that definitely indicates something is not right for your given site, right? So this is kind of another, back to your question, this is a time series detection anomaly model. But here's a challenge for us is that how you want to make sure your time series anomaly detection model is accurate Think about our scale we're running.

15:32We have close to 100 ,000 different sites deployed across the world. And if we have every site, we have a anomaly model, which is a neural network-based model. So which means we're going to have 1 ,000 neural network-based models. This is too expensive, right? One model is fine, but this scale is too expensive, too costly to maintain. So cost also is another factor. Efficacy is definitely the number one issue, but cost is also second to the most important factor to make a decision. Can you elaborate on that? What's the relationship between a single model and the cost? So, for example, let's give an example.

16:24Say, hey, for the retail store, you have a daily pattern. You know, the customer going to your store to visit. Daily pattern probably is a weekend should be a higher volume versus a weekday, right? And the weekday during the normal business hour, probably not many people go to retails, go to your store. Probably going to be even time. But versus for the - The university is going to be almost opposite, right? Yes, yes. And the training is different. So in that case, you can't have a neural network-based model to learn your data pattern for the retail store versus for the university. if the pattern is totally opposite or totally not in line with each other.

17:07So in this case, the best efficacy training was that for each, you know, the nine, each trending nine, we created their own neural network model, right, to the forecasting. Same thing, you know, like say, hey, for a stock market, you want to predict a particular stock market, stocks trending, you train a neural network. But if you want to follow multiple stocks, they're trending. You probably need to train multiple neural networks, right? So we talked a little bit about time series as being somewhat of a, you know, solve problem, quote unquote. But that isn't to say that people aren't trying to create transformer-based models for time series.

17:50I'm curious if you, you know, to what degree you've explored or found any interesting solutions there. Yeah, it sounds like that could be interesting for your problem where you've got some base level model that you can fine tune as opposed to creating many distinct models. Yes. So that's exactly what you know here. Recently, there is a pre-trained time series anomaly model, right? That's also exactly the need. But there's a key difference that the existing pre-trained anomaly model time series is for single feature. what we are talking about, multi-feature. So this is actually, that's another piece.

18:34A lot of times, recent development advancement is only for single feature. Got it. So you've got a log of stock prices, and you're trying to predict future stock prices, or you've got, in your case, a log of RSSI. Would that be one of your... Exactly, yes. I forget what that stands for, but it's a measure of noise. Signal strength, yeah. And you're trying to predict the future RSSI so you can figure out which channel to join or something like that. But I'm guessing that you've got this log with a lot of different features and you're trying to make multiple predictions based on as opposed to a single one.

19:17Yeah, yeah. So, for example, you want to look at the throughput for your given site. You definitely should look at the number of clients and the total traffic volume we're going through and how many destinations they are going through and how many different applications they are routing the traffic through. So this is going to be multiple features in the real world problem. So this is also the reason why the pre-trained time serial anomaly model is very promising, but unfortunately it cannot work for us as a single model. I'm relating it a little bit to my personal experience. I've got a small network here with multiple access points.

19:59And I remember when I first installed this latest version, I would go into the management software and look at what client was connecting where and to try to address performance issues. And I noticed that the decisions didn't really make sense to me. I would be in one room and the laptop would connect to a distant access point. And I'm just thinking about if it was that complex for me. And I was never able to resolve it. I just had to wait for the software to get updated and it kind of resolved itself. But if it was that complex for me, like at the scale of a campus, the complexity multiplies, obviously.

20:48And so it seems like an obvious place to add some intelligence, I guess is the point that I was getting at. So it's another thing that is, for example, we wanted to transform a networking. Just networking previously, just like I said, you need to really domain expert network administrators. They do a lot of troubleshooting, doing a lot of cool stuff to find the data, to correlate, to find the root cause. Because that domain is mature enough so that we can write the software to embed those domain experts into the software to help some of the no-handling fruits. So the network administrators, they can focus on more complex problems like security-related, like a large deployments configuration.

21:41So we are also transforming a lot of network administrators from just a manually cool shell guy and understand the network gears to be API engineers. Because we export all our data to make it available through APIs to the customer. So they can utilize the data to develop use cases for their support, their own specific environments. So it's a key difference here as well to empower our customers. You mentioned in your earlier explanation the use of reinforcement learning. That is sometimes associated with using simulation. Is simulation a technique that you use often there? Actually, we did not really utilize simulation.

22:33We just test our own. So our office is a test lab. So we are all part of our test, you know, our latest algorithm to reinforcement learning. Reinforcement learning is a permit you develop that, you know, the awarding function, right? What is a positive or negative award? So this is the one that we continue in the cloud. We always do this type of AP testing. So first, put a smaller, friendly, different type of deployments into this latest algorithm. So after we have a confidence, then we deploy it to the MIST. We call it the MIST universe, which means globally to all the cloud environment. And so MIST is kind of an umbrella term for all of the Juniper AI-related capabilities?

23:25So actually, MIST was a startup. you know, actually this week will be hit our 10 years anniversary. We started as a Wi-Fi, only focused on the Wi-Fi. But the next is Wi-Fi is different. It's a Wi-Fi access point. It's totally cloud, cloud-driven access point. So about five years ago, Juniper acquired the MIST. So this acquisition is a benefit for both sides. You know, for MIST, we used to use Wi-Fi after we parted the Juniper family. We got the switching, wired, and also get the VAN gateways. For Juniper perspective, Juniper does not have a Wi-Fi, never been a Wi-Fi company. So now Juniper got a Wi-Fi.

24:10And also Juniper is not a cloud company. So now it's a cloud. And of course, Juniper continues to utilize the most cloud technology and the AI ML to try to expand into not just for enterprise, which is wired in the VAN, and also to data center. So this is kind of currently the five years after acquisition. Then we are going to be soon embarking another journey to be part of the HPE family. But the acquisition is the same goal. We wanted to continue to utilize the MIST's AI ML solution for networking to be exposed to the much larger networking deployments. I came across one of the initiatives under that umbrella called Marvis.

24:56Can you talk a little bit about Marvis and what that is aiming to accomplish? So now, you know, one hand, we have a lot of bank-hand jobs running, right? Conscious numbers to process user data, to generate alerts, events. Now, second is how we want to make sure, how we want to change the user interact with our application with their data. So we utilize Marvis. Marvis is, think about it, actually Marvis where it came from is from Jarvis. Ms. Jarvis? Yes. And it's a bot. They can understand a lot of languages. And we replace the first letter is a mist, right? That's why it's called Marvis. So first, you know, Marvis is how users can interact with their networking, their network issues.

25:47So we first started with troubleshooting. so you can always go to the UI, just troubleshooting for a particular client, particular device, or particular issues. That's early days, you just commanded to do the troubleshooting. Now we evolved into a chatbot. So you can just go to the UI, you can just go to the chatbot asking questions about why I cannot connect the Wi-Fi, why my AP is continuous rebooting, why my AP was disconnected from cloud, right? Why my AP, can you, how can I do the AP firmware upgrade? So we added the chatbot support. But the chatbot support actually was developed even before this whole large language model took the world.

26:39So now we are sort of another, you know, enhancement to our current chatbot to integrate with large language model. and you know recently you had a great conversation with Nantium CEO, Nantium Nantium, so actually we'll integrate that Nantium graph into our chatbot so and also you know you could go to Marvis to ask questions about a specific networking domain or Marvis or Juniper related product questions instead of we give you the documentation to grade three, we can summarize the specific answers and provide a real precise response to you. So this is kind of, you know, just the way, first we develop our own RAC solution to be able to do this type of summarization, focus on our public, you know, documented in the release notes and documentation pages so that we can provide a precise Juniper-related product features to our customers.

27:51So you don't have to read through all the technical documentations and the wiki page for yourself. This is one part. Yeah. Second part is that we want to utilize large language model to better understand what customers asking and also develop our own text to sql text to yes type of solutions because all our data inside the sql database or elastor or the non-sql database we so before in order to expose those type of data we have to create a lot of different apis while we just use a large language model help us to expose all the data together with the customer, right? Got it. So based on whatever the customer's intent is, you turn that into a query against the database?

28:41Yes. That's another use case. Because as of now, we are not, we're committed to not sending the customer data to large language model. But we actually utilize large language models to create a lot of embeddings for our internal metadata. So for example, how we store the data in our database. We create this data catalog, or we actually create an embedding and also provide the RAC solution so that we can match the customer, what they're asking with our data much better. So a couple of follow-up questions there. First, you mentioned working with LangChain and LangGraph. I'd love to hear a little bit of your experience and motivation there.

29:32I talked to, there seems to be bifurcation. Some folks appreciate the kind of standardization, if you want to call it that, that the LangChain offers. Other folks like to go straight to the metal and kind of do it themselves. and skip an intermediate framework that might hide some of the either complexity or features that they need to know? What drove your decision to use that? So currently, our integration with NAND graph, we did not go fully. So for example, a lot of large language more related features, we build our own smaller model. So it's kind of multi-model type of a scenario in order to, because like I said, first, networking is a very special domain.

30:26And a lot of, you know, networking jargons that our fine-tuned model probably going to do better. But on the other hand, some of other questions that the large language model can help us to close the gaps. So we actually running this as a large language model, our own fine-tuned model is running in parallel. So because of this, we does not go full in into the name graph to handle all the dialogue and also the assistant. So we kind of build those tools a little by little. Like I said, we feel don't have a confidence that we want to all in with the name graph. Like exactly what you said, you have some customer some of our audiences, they want to go all in with Niagara.

31:16Some of them just, we are probably the second layer. We wanted to have our own control and do our evaluation. Then we kind of integrate with Niagara piece by piece. Got it, got it, got it. Yeah, I hear that quite a bit. And then the other question is with regard to RAG. You've mentioned that a couple of times. I have talked to many organizations that, you know, start down this journey with RAG and primarily expose it via a chat bot, but find that, you know, kind of going beyond that chat bot and integrating more tightly into existing, you know, workflows and user interfaces without requiring that chat dialogue experience is a source of even more value from that.

32:12that from the RAG as a technology? I'm wondering if that's something that you've seen, and if so, what are examples of those workflows that you're looking to enable? So, for example, for us, when customers try to open up a support ticket, before they open the support ticket, when we look at their questions they're asking, the problems they're dealing with. So if it's questions about our product and the features, then we just use Rack to generate response for them. Just another layer to say, hey, maybe this is what you want. You don't have to create the support tickets. So this is kind of a lot of integration with customer support.

Read the full transcript

32:58Good user experience as well, right? So we can expand that, not just for the chat bot. We can integrate with customer support. Another example, actually this is a good question, is a large language model, how are you going to prevent hallucination? For networking, you know, beyond troubleshooting and customer support, the good user experience, we are still hesitating or has not really prioritized this feature. Like say, hey, can we just utilize, you know, Rack solutions to generate a config for my network? Well, even though LOM, let's say we can get to the efficacy to 90 % or 99%, but that 1 % of failure is not acceptable, right?

33:52So because of this, we have some experiments. We're probably going to be just more like the recommendation perspective. Still need a human to do the final validation before we can release it. So this is kind of the challenge. It depends how much the efficacy is susceptible. For us, 0.1 % of a failure rate is not acceptable. Interesting. So you talked about a couple of examples so far. You talked about kind of channel selection. You talked a little bit about like proactive monitoring and fault detection. Are there other examples that you can share with us about where you've integrated, your integration of ML and AI have produced interesting results for your users?

34:48Yeah, so anomaly detection is a general, right? We monitor customer. But beyond that, for example, it's okay. After customer deployed the site in the beginning, it's perfect. as number of users is going on and the usage pattern is different, then we kind of realize that your current deployment may not be optimal. So for example, there's suddenly a coverage hole showed up on the AP access point. So in this case, we are, yeah, so how are you going to detect that? So we continue to monitor your coverage usage and the capacity and also the RSI, the transmission failure rate, then we'll report to you around this corner of the office.

35:36There is a coverage hole. You probably need to adjust your AP access point and also increase AP deployment on this side of the office. So another very good use case is we are monitoring your user experience. That's the case, user experience. And the second point is after you get the network configured, then you continue making the change to your network. What if accidentally all this configuration, you're missing some configuration between a switch? How you do it now? So that's the same. We always compare in a switch, every interface port configuration for VLAN. And then if there's any missing, it will report to you.

36:18What you're trying to identify there is like a network engineer is in a console and like FatFinger is a configuration and deletes a VLAN or something like that. and all of a sudden like a segment of the campus disappears or connectivity to an application disappears. Okay, interesting. Yeah. Then other pieces are, hey, you get everything's running, but suddenly the cable somehow deteriorate in the performance. How do you know? There's certain error counter that will give us indication. So we kind of based on the broken cable build up, you know, for copper cable versus, you know, fiber cable build a different ML model.

36:59and based on the switch, the device reported the counters errors. Then we are able to identify the cable had broken, so you need to do a replacement. Otherwise, the user experience of traffic will be getting impacted. For that particular one, why is an ML model needed? It seems like either you would have the cables really broken, like you have no electrical connectivity and the link is just down, or maybe it's not all the way broken and you just have a really high packet loss rate or something like that. And there's some threshold that you would set to know what's the complexity that requires ML?

37:43So a packet drop could be a lot of reason. One of the reasons could be the cable broken, right? The packet drop is specific for the SD-WAN. could be the ISP-related service problem, ISP server problem. So not necessarily cable. Cable is really the... So yeah, that's the thing. For network key, if you look at the data, there's a lot of issues you can find. The package drop always happens. If you build a normally model, you probably could find, hey, suddenly one day the package drop is, you know, it's suddenly increased. But what is the cost? Right? If you just want to look for problems, a lot of problems, but you truly identify the root cause, the reason, contribute to this package drop to be a cable issue, that's a hard part.

38:33Right? So that's the reason why for cable issue, we specifically look at the device reported error counters, such as error counters. So that's also this type of problem is very vendor-driven. So for our device, Juniper device, we have a certain confidence about the stats, the reporting, the way they're reporting the data. And our model is pretty much tailored to our device. So for other vendors' device, how they report those type of issues, that may not going to work. That's also the challenge for this networking domain as well. Meaning you've got a specific way that you're collecting that error data?

39:24Yeah, so how the device report that error, that is how frequently they report that error. So those are going to impact our detection logic and the model we're building. Pretty much the feature engineering that effort, because this is more like a buildable supervisor learning based on the label data. We build a supervisor learning, right? But also time series or not time series for this? This is not necessarily time series. Is it that a single instance of this error or are the errors like class, a single class of error? It's going to be class of error. And you look at the class of errors and also look at the duration of the error.

40:14So you kind of featurize those two things and based on that, it will flag a cable issue? Yes, yes. So this is based on, like I said, based on label data. So we have a few of the cable issue for copper and fiber, different type of cable. Interesting, interesting. Yeah. So all this stuff is kind of AI ML related. It's very specific to networking domain problems to help the AI for operations. So after your network deployed and we are doing the monitoring for you instead of, you know, from customers. that maybe someday someone's complaining, why are my performances going bad? Then the network administrator, then doing some digging, they realize, oh, it must be cable issue.

41:02Oh, there must be one of the VLANs we forgot, right? So this fundamentally enables our customers that they be able to have large deployments. They can focus on much larger, bigger scale problems. Got it, got it. Do you track, in the case of a larger scale customer that's using all of the features that are available, like how many ML models are coming into play and helping them manage their network? Like is it tens or hundreds? Or like where are you with that? You mean the numbers of the models we deploy in the production? I guess the thought is that you've got some ML solutions where it does one thing, it's a big thing, and it's kind of central to the solution.

41:56In this case, what's kind of interesting is that there are a lot of little things happening behind the scenes that are kind of silently made better by machine learning. And I'm just wondering how many models are required to do that. I'm assuming that it's a lot of different smaller models as opposed to one giant model. We kind of talked about that earlier. Yes, exactly. That's why I was asking. Exactly. We have a lot of smaller models to do this little magic, not just one large model. Actually, we have, you know, large model is a large language model. That's for, you know, that's actually we don't have that.

42:40We utilize a third-party hosted large language model. Internally, we have quite a lot of neural network-based model, like I told you, shared with you about the time serial anomaly detection. But that one actually is, we are at the stage, we need to change the way we did the time series anomaly detection model. Maybe the next episode, I'll share more about what we're doing. Yeah. Then on top of that, we have a lot of smaller models, like cable detection and misconfigured port, even for radio resource management. We have very specialized small model. The reason is, like I shared with you, is actually three factors to decide for this.

43:29One is efficacy. Second is latency. Small mode has a small latency. Third is debuggability and also the cost as well. So we do gather the problem question, why you make this decision? How are we going to debug? So, of course, one is a supporting evidence. We have all the data. We can show you the time series of supporting evidence. But a lot of times, even looking at supporting evidence, we kind of puzzled why the model behaving the way it was. So that's also the debug capability for us to fine tune our smaller model is much faster, quicker to get the new solution deployed in the production. I'm imagining that for many of these features, like there's a traditional way of doing things like maybe, you know, the traditional way to do faulty cable was some threshold on an error rate or something like that.

44:26And then, you know, there's a machine learning based way to do it. And I guess I want to get a sense from you, like, you know, how you think about balancing, you know, the complexity versus the data requirements and like when to make the transition and that kind of thing. Is it just a product management thing? If there's some gap that you're not able to get in terms of capability and this is a way to get there, then you kind of go in a direction? Or are there other ways that you think about making that transition? Yeah. So actually, you brought up a good point. It's actually a lot of our product management, we are just working with them about the particular use case.

45:20How are we going to solve this problem? First, is data in the cloud? Do we have the right data? Then after we have data in the cloud, so data science team actually, we're working not just with the product management. A lot of times we're working with the customer support team. So then to study the use case. So like you said, maybe heuristic is good enough to solve this problem. why we need to spend the effort to train a model, to utilize AI MLs. But anyway, the key point is, first is the data. Second is, do we have the heuristic? Do we have a solution? It doesn't matter if it's a human, you know, magically sifts through all the data to find the problem or write a piece of software to solve the problem.

46:12Fundamentally, we need the data to solve the problem to make our customers' internet experiences better. So that's a key. So for the data science teams, actually, a lot of times we are not really doing training a fancy model every day, but actually a lot of times we are doing the feature engineering. Feature engineering, you probably heard a lot from data science team ML. A lot of times they are doing the feature engineering, trying to study the data. Does the data really contain the information we can solve the problem. Do we really need a neural network or just regression or heuristic baseline that can solve the problem?

46:56So we talked a little bit about some of the ways that you are incorporating LLMs and Gen AI into your solutions. Are there any future directions that you can talk about? Where do you see all of this heading for Juniper? Yeah, so it's two things I can share with you with Juniper. One is not related to JNAI, but it's actually currently, you know, Juniper, we utilize the data and the events. Actually, we are detect the problems which has already happened, which already impact the customer experience. So how can we take it to the next level? we can proactively identify the problem before user experiences this issue.

47:45So this is kind of, you know, we utilize all the devices we deployed in the globally. We are able to inject, you know, from cloud to push down a little piece of software to proactively test the customer's network. So that way we can proactively identify the issues. For example, every night you did the configuration change for the network, we can start to automate the tests on your networking. So before people show up at 8 o 'clock in the morning in the office, we pretty much report to you your office is ready for people to work or not. If there's issues, you should already send alerts to the network administrators to address this issue.

48:30So this is called the Mavis Digital Twin, the Minis. So that's why we are already running in production, but we are trying to progressively to increase the capability of the minis to the test scope. Were these minis running? Are they running on like our devices? Meaning from one access point to another that you're testing across as opposed to like from a client, like a user laptop? So it's inside access points. We're going to have minis inside the switches, minis inside the SDVS. All the Juniper devices were capable to support the minis. Yeah. So this is kind of really proactive identifying the issues before any user experiences networking problems.

49:22Yeah. So this is the next big direction the company is moving to. Another, you know, for GNAI is that right now our, you know, application is a little counterintuitive. For example, our application first is tailored for network administrators. That typically if your network is running perfectly, you don't need to log into our UI. You don't need to look at the author graph charts, right? So they want to move to the next level. But then think about if you are college students, you cannot join to your college classroom. You cannot join the Wi-Fi. What are you going to do? You call your college administrator, complain about your problem.

50:10So then we wanted to say, hey, let each individual Wi-Fi user to be able to do a self-service, self-identify the problem. So we're going to build another feature so that as each individual user, you can troubleshoot. Then we'll tell you specifically what is wrong. Maybe your guest Wi-Fi password was changed. You need to contact the IT administrator to find the correct password or your cert expired. So this type of problem can give you much more information versus you call your network administrator. So I'll tell you your device has a security patch, it didn't apply. That's the reason why you cannot to your Wi-Fi, cannot connect onto your network, right?

50:59So this type of problem that the need to each individual Wi-Fi user to solve their problem instead of flooding the IT administrators. So that's another one we are trying to utilize Marvice to not just serve the IT administrator, also serve the individual end user. And I'm imagining that the logical next step there is some of these problems are going to be the result of network issues as opposed to user configuration issues. And then you're just kind of collecting user experience data that you can use to either inform the network administrators and you already go figure out what the root cause is or at some point just fix it if it's fixable.

51:42Exactly. That's when we wanted to collect the user input data. Do they really have a capacity issue? Do they have a really slow experience? So for Marvis and we try now, we are doing a very good job to identify why you cannot connect a network. We solve that problem is pretty good. Now we are moved to the second is a slowness. Slowness is a hard problem. First is the different applications have a different tolerance for the networking capacity. Second is the slowness also depends on what type of applications we're using. Third is a very subjective for each individual's personal perceptions, right?

52:27So the user feedback can give us a lot of input data. Very cool. Well, Shirley, thanks so much for jumping on and sharing a bit about some of the ways that you're using data science, ML and AI to help us get better networking experiences. Yeah, it's a great pleasure to chat with you, Sam. So have a good day. Thanks so much.

53:03Thank you.

From the publisher

Today, we're joined by Shirley Wu, senior director of software engineering at Juniper Networks to discuss how machine learning and artificial intelligence are transforming network management. We explore various use cases where AI and ML are applied to enhance the quality, performance, and efficiency of networks across Juniper’s customers, including diagnosing cable degradation, proactive monitoring for coverage gaps, and real-time fault detection. We also dig into the complexities of integrating data science into networking, the trade-offs between traditional methods and ML-based solutions, the role of feature engineering and data in networking, the applicability of large language models, and Juniper’s approach to using smaller, specialized ML models to optimize speed, latency, and cost. Finally, Shirley shares some future directions for Juniper Mist such as proactive network testing and end-user self-service.

The complete show notes for this episode can be found at https://twimlai.com/go/710.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
AI for Network Management with Shirley Wu - #710The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 54 min
Listen in VO