In short
AWS Podcast Episode Summary
Episode Details
- Podcast Title: AWS Podcast
- Episode Title: #732: How to gain Multi-Cluster Visibility across Kubernetes Clusters with the EKS Dashboard
- Release Date: August 4th, 2025
- Hosts: Simon Elisha, Hawn Nguyen-Loughren
- Guest: Shriram Ranganathan, Product Manager at Amazon EKS
Episode Description This episode focuses on the Amazon EKS Dashboard, which addresses challenges in managing Kubernetes clusters across multiple AWS accounts and regions. It emphasizes centralized visibility into cluster health, versions, and costs, aiding governance, operations, and infrastructure optimization.
Key Topics Discussed
Challenges in Managing Kubernetes Clusters
- Visibility Issues: With many clusters spanning various accounts and regions, organizations often lose oversight of:
- Kubernetes versions in use
- Cluster distribution across accounts
- Security and compliance statuses
- Operational Complexities: As organizations grow, they encounter issues related to:
- Managing numerous add-ons with different versions
- Ensuring clusters meet governance and security standards
- Example Scenarios:
- Identifying outdated Kubernetes versions.
- Understanding the exposure to vulnerabilities across clusters with varying add-on versions.
Introduction of the EKS Dashboard
- Centralized View: The EKS Dashboard provides a "single pane of glass" view to manage multiple clusters efficiently.
- Operational Planning: It assists in upgrade planning by visualizing which clusters need attention based on internal guidelines.
Key Features of the EKS Dashboard
- Version Distribution: Visual representation of clusters by Kubernetes version to prioritize updates.
- Support Policies: Information on whether clusters are on standard or extended support and the financial implications of upgrading or delaying upgrades.
- Security Compliance: Ability to identify clusters with public access to API servers to enforce internal security policies.
- Node Group Management: Insights on managed node groups, including the distribution of different Amazon Machine Image (AMI) families across clusters.
- Add-on Management: View and manage add-ons running on clusters, facilitating quick identification of vulnerabilities.
Cluster Health Metrics
- The dashboard provides metrics on low-severity health issues, such as IP address shortages and insufficient replicas in add-ons.
Cost Forecasting
- Users can forecast management fees associated with clusters, especially those on extended support, allowing better financial planning.
Upgrade Management
- Dashboard facilitates planning cluster upgrades by providing insights on potential issues that may arise during upgrades.
Reporting and Export Capabilities
- Users can export data to CSV for deeper analysis or reporting, with various filtering options available for customized views.
Integration with AWS Organizations
- The EKS Dashboard integrates with AWS Organizations, allowing controlled access to the dashboard through a management account and delegated administrator accounts.
Roadmap and Future Enhancements
- Ongoing improvements to dashboard widgets and insights.
- Upcoming feature for "scope delegated administrator" to restrict visibility based on organizational units.
Resources Mentioned
- Whats New Post: [EKS Dashboard Announcement](https://aws.amazon.com/about-aws/whats-new/2025/05/eks-dashboard-multi-cluster-view-kubernetes-infrastructure-aws-regions-organizations/)
- EKS Dashboard User Guide: [User Guide](https://docs.aws.amazon.com/eks/latest/userguide/cluster-dashboard.html)
- Deep Dive Blog: [Deep Dive on EKS Dashboard](https://aws.amazon.com/blogs/containers/deep-dive-amazon-eks-dashboard-for-visibility-into-multi-cluster-operations-and-governance/)
- YouTube Video: [EKS Dashboard Overview](https://www.youtube.com/watch?v=N1I1H2wA1Sk)
Conclusion The episode provides a comprehensive overview of the EKS Dashboard's capabilities, emphasizing its importance in managing Kubernetes clusters effectively across AWS environments. Listeners are encouraged to utilize the resources provided and to continue providing feedback for future improvements.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00This is episode 732 of the AWS podcast released on August 4th, 2025. Hello, everyone. Welcome to another episode of the AWS podcast. In today's episode, we are going to dive deep into how you can centralize visibility of your Kubernetes clusters across AWS regions and accounts with EKS dashboard. This is a new capability that the team has launched recently. And joining me to talk about it is Shriram Ranganathan. Welcome, Sriram. Can you introduce yourself? Hey, Sruti. Nice to be here. I'm a product manager working on Amazon EKS, focused on multiple features like add-ons, upgrades, control plane, and multi-cluster.
0:48And EKS dashboard is one of the features that my team recently launched, and I'm happy to be here talking about this feature. Great. All right. Well, let's dive right in. So first, taking a step back, what are some of the challenges that teams and organizations typically face when they are managing multiple Kubernetes clusters and especially across different AWS accounts and across different regions? yeah that's a really great question uh so when we talk to our customers what we see is customers usually start off with a handful of clusters and as their business grows and their business needs grows they slowly expand to multiple different clusters either within the same account or in different accounts either within the same region or across different regions for various needs it could be for scalability needs it could be for data residency needs it could be for resiliency purposes, the ends are needless.
1:48And what we typically hear from our customers is as their fleet of Kubernetes cluster grows, they soon start losing visibility into the different clusters that are running within their organization. So governance becomes a problem and keeping up to date on security for all of those clusters also becomes quite a big of a problem. Because with kubernetes because of the open source nature you need to keep those clusters and any other affiliated resources up to date and without having the centralized visibility customers do not know what their exposure to is or what their current security landscape looks like so let's take the example if i have like 200 or 300 different clusters running across different accounts and regions how do i know which version of kubernetes those clusters are running on which accounts are hosting those clusters, which regions are running those clusters, what versions are they running on, and what is my timeline before which I should update those clusters.
2:47So that is a classic example of a visibility challenge that most customers run into. That is just one example, and that's not the only example. An other example of a challenge that we typically ask from customers is like Kubernetes has this plug-and-play nature. So by default, in order to make a Kubernetes is cluster production ready, you need to add a lot of add-ons like the networking add-ons, the storage add-ons, observability add-ons. And each of them come with multiple different versions. So let's say, for example, on an average, I'm using 10 add-ons per cluster and each of them are running on different versions.
3:23How do I know if tomorrow there is a CV for an add-on? What is my exposure across different clusters? So that is another challenge that we typically hear from customers. So this dashboard is aimed at answering some of those questions, basically giving you the 10 ,000 feet view of your entire Kubernetes clusters, the different managed node groups, the different EKS add-ons, and what are the metadata and properties associated with those so that you can quickly dive deep, figure out what is your impact, and quickly address those challenges. Right, right. So big challenges you're trying to solve are fragmented visibility across accounts, across regions.
4:02It, you know, kind of brings up this visual of trying to manage multiple like branch offices without a single headquarters. And so this is trying to give that centralized visibility. And then also it sounds like it also is a way to ensure consistent best practices across different accounts, across different regions and so on and so forth. That is right. It gives you essentially the single pane of class view, allowing you to govern and do operational planning so if you have to upgrade clusters you really need to plan because you need to upgrade the applications you need to look at what could break if i upgrade my cluster and plan so this gives you at least the visibility into okay what are the clusters that i need to immediately focus on what are the different properties are they meeting my internal guidelines so let's say for example i require all of the production clusters to be enrolled, let's say, for example, in zonal shift auto configuration.
4:56So how do I know which of my clusters are not following zonal shift practices for resiliency purposes? So I can quickly dive deep and figure those things out. So there are a lot of these dimensions across which you can look at your entire Kubernetes fleet and see if they are following your organizational standards, best practices, and other things. Awesome. Okay. So let's get into the specifics now, right? Can you maybe walk us through some key capabilities and features of the EKS dashboard and how they help solve these challenges? I know that a dashboard is such a visual sort of a service and a product, so it might be a little tricky.
5:38But if you maybe can call out specific features and what and why they were built, that would be really helpful. Okay. We'll start off with the most common request that we get, which is visibility across your entire Kubernetes clusters broken down by different versions. Customers definitely want to know if I have 200 or 300 different clusters running, what is the distribution of those clusters by different versions? So that anything that is at the tail end of the cluster version lifecycle, I want to keep them upgraded. That is one. the other use case that we are trying to address is EKS comes in two levels of upgrade policy support one we call as the standard support and then there is the extended support there is a price difference between what is in standard support versus what is in extended support so if you are not keeping your clusters up to date and you are enrolled in extended support I definitely want to know which are those clusters because I am paying additional for those clusters and I want to upgrade them first So how do I know which clusters are running on extended support?
6:44The other challenge that we are also trying to address is like, depending on the upgrade policy that your clusters are enrolled in, whether it is standard versus extended, if you do not upgrade those clusters within the end of support date or end of extended support days, EKS automatically auto-upgrades the control plane of those clusters. How do I know within the next 30 days, 60 days, or 90 days, which clusters are going to get auto-upgraded by EKS if I do not take an action today. And that is like an expandable time view. You can adjust the time frame to look at, okay, in the next six months, I want to see which clusters are going to get auto-upgraded.
7:22You can always adjust that. If you want the one-year view, you can always do that. So that is one more use case that I can think of. Like that, there are various other use cases at a cluster level. Some of the other use cases that we get is as an organization, we do not like any of our production Kubernetes clusters. to have Kubernetes API server have a public access. I want to see how many clusters have public access endpoints because I want to disable it. How do I easily identify it without diving into each cluster? Can you give me a broad level view? So that is one aspect of the things that we are trying to solve.
7:58The other thing that we are also trying to solve is a Kubernetes cluster comes in three different resources. One is the control plane itself. And then there is the data plane. which could have like self-managed node groups, or you can have nodes that are managed by Carpenter, or you can have managed node groups. Essentially, we are trying to provide a view from the perspective of managed node groups, because that is one of the APIs that we provide, and we have visibility into the metadata surrounding those managed node groups. So let's say, for example, you want to see the distribution of managed node groups by different Kubernetes version.
8:33You can view that through this dashboard as well. another use case is with respect to finding the distribution of different army families so let's say for example you want to find different node groups that are running different types of armies you can always do that be the bottle rocket army the al2 army the al2023 army you want to see the distribution of those you can always visualize that and dive into those specific node groups or let's say for example a particular version of an army has some kind of a cv you want to identify which are those node groups and upgrade them. You can see the distribution by different army versions.
9:10Click on them and it will give you the exact exposure of that army across different node groups and different clusters. The same thing can be extended to add-ons. Each cluster might be running 10 to 15 different add-ons. How do you know like which add-ons are running on which cluster? Which versions of add-ons are running on which clusters? I can always see that directly from the dashboard. And the good part is while it starts off with an aggregated view, You can click on that aggregated view and it exactly filters down to the set of affected resources. Awesome. Awesome. One follow-up question I have here is, does the dashboard also provide some cluster health metrics or not?
9:52Yes, it does provide the low severity cluster health issues directly through the dashboard. So let's say, for example, if your clusters are running out of IP addresses or if your node group has certain health issues or if your add-on has health issues in terms of insufficient number of replicas or any other health issues, you can visualize it directly from the dashboard and then you can take action on them to fix those. right awesome um now you also did kind of talk about two different types of eks modes the the one that where customers have to uh pay a little bit more can you just repeat that for a second because i had a follow-up on that sure so eks cluster has a defined life cycle by default we enter every cluster when it is released a new version is released it goes in starts off in standard support so that goes on for about 14 months and then after 14 months it enters into extended support for the next 12 months so the total life cycle is 26 months during the first 14 months you pay the standard support charges which is 10 cents per hour for cluster management fee beyond the 14 months if you have not upgraded the cluster and if it is enrolled in extended support, then you pay an additional 50 cents for the same cluster.
11:13So it's a total of 60 cents per hour starting from that period onwards. Right, right. So, I mean, it sounds like the dashboard can help teams forecast and manage these extended support costs. Does it actually give them this cost, like forecasted cost information, or does it just show, hey, these are the clusters that are on extended support, that are signed up for extended support and kind of leaves the leaves the bath to the customers that is a great question the answer is yes and yes to both it gives you the breakdown of clusters which are enrolled in extended support versus clusters which are just on standard support additionally we also forecast and show like if you leave these clusters in extended support for the next 30 days what is the additional control train management fee that you are looking at for 60 days what is that additional charges and we give the prediction for up till one year in that one year let's say for example within a specified time frame that you are interested in there might be clusters that are going to enter into extended support in which case we prorate and automatically adjust the calculation to give you a ballpark in terms of what is the cluster management fee that you are looking at one thing to notice this is just focusing on the eks charges this is not focusing on the additional nodes, EC2 instances that might be attached to your cluster.
12:36So this is purely focused on EKS charges. Right, right. That clarification is extremely important. Yes. Okay, so that was the one follow-up. And then the second follow-up I had was, you know, you did talk about sort of, you know, being able to upgrade versions and things like that. What role does Dashboard play in planning and executing these cluster upgrades? Is it primarily sort of this, you know, information that it surfaces about what version each cluster is on? Or is there any sort of capabilities around being able to plan for cluster upgrades at a certain cadence? Yeah, so there are two different things.
13:17One is obviously giving you a quick inventory of things that you should focus on. So let's say you are upgrading from 129 to 130 clusters and you want to plan and see how many clusters are there on version 129. you can simply go to the dashboard click on the clusters which are showing 129 it gives you the full inventory of clusters that you need to focus on but it does not stop there it also gives you for each clusters are there any upgrade insights which will affect the cluster when you upgrade it what i mean by that is upgrade insights is a feature which was launched about a couple of years back it tells you like if you upgrade the cluster without focusing on certain things will the cluster upgrade break are there going to be errors associated with the cluster are they going to be warnings associated to the cluster.
14:00That's called as upgrade insights. So the dashboard also tells you for each of the clusters, are there any error level upgrade insights? Are there any warning level upgrade insights? Are there any unknown level upgrade insights? So that you can look at those and then plan your upgrades. Makes sense. And then one last thing, can you maybe double click on the add-on management across multiple clusters? You mentioned that earlier, but it would be good to get one level deep in that. Sounds good. So you can think about add-ons as operational capabilities that you add to a Kubernetes cluster. The Kubernetes cluster itself, when you create a cluster, it's not production ready.
14:41You need to add a lot of these capabilities to your cluster to make it production ready. As an example, you need, let's say, an EBS CSI driver for you to use block storage. You need an EFS CSI driver to use some kind of file storage. You might use CloudWatch agent to basically have observability capabilities. Or you might install something like an OPA gatekeeper or Kiberno for policy management, other things. So there are a lot of these additional capabilities that you add. So you can think about add-ons as these capabilities that are adding additional functionality to your Kubernetes cluster.
15:16And each cluster might have 10 to 15 different functions like these. So let's imagine if you have 200, 300 different clusters and you have 10, 15 different add-ons running, all of them might be using different set of add-ons. How do you know like what add-ons are running on which cluster? So let's say tomorrow, if one version of, let's say, EBS CSI driver has some kind of a vulnerability and you want to upgrade it to the next version, how do you quickly identify without going into each cluster to figure out what is my exposure. I don't want to query 300 different clusters to figure out what version of EBS CSI is running.
15:56So you can come here, you can just go to the exact affected version, you can filter by that and see which clusters are hosting those that will allow you to immediately take an action. So you save a lot of time in terms of figuring out your exposure of different components across different clusters. That's awesome. That's awesome. Actually, that's a really good segue into my next question, which is it sounds like there's lots of great filtering, sorting kind of functionality available to create whatever view that is useful for the operator. What are some of the reporting and export capabilities that are available for teams that are wanting to do this sort of a deeper analysis or report it out to their, you know, to their leaders or whatever, what have you?
16:44Yeah. So while this feature is a console-only feature with no SDK or API access, but we do have an ability wherein you can download the entire data set directly from the console. So it's an export to CSV option, which will allow you to get hold of the entire data and then you can ingest it into any visualization tool of your choice. or if you are only interested in a specific set of affected resources, you can apply the filters within the dashboard and just export that filtered view. So you have both the options. So depending on your use case, you can choose to either download the entire data set or narrow it down to your specific filtering criteria and then export that particular data.
17:28And then you can ingest it with any other visualization tool if that's your preference. Right, right. And so you can do it, download it into a CSV file, and it could be filtering on specific regions or accounts. This filtering, at least in many cases, is also available in console. Like you can still do it there, but if you want it offline into a different tool, you can do that through the CSV. Exactly. Okay, so these are some really useful capabilities. And as you might imagine, organizations may not want everyone in the organizations to have access to EKS dashboard. So how does dashboard integrate with AWS organizations and what is involved in setting it up?
18:18How does it all work? Sounds good. Yeah, that's a really good question. So this feature natively works in integration with AWS organizations. so only if your accounts are nested under AWS organizations you will be able to use this feature so if there is an account that is a standalone account outside AWS organization this feature will not be available to such accounts so imagine an organization which has like hundreds of accounts nested under AWS organizations and you want access to this particular dashboard all you need to do is log in to the AWS organization's management account and there is a one-time enablement of this feature.
18:58You click on a button called as enable trusted access from the dashboard settings page and that basically allows EKS to create SLR for us to render the dashboard. So you can just stop there and start accessing the dashboard but as a best practice we recommend that you do not use the management account for accessing the dashboard because management account has a lot more privileged capabilities. So that's where the concept of a delegated administrator account comes in. A delegated administrator account can be any account within the AWS organization that you can choose and nominate as the delegated admin.
19:33What it essentially means is you are giving it some admin functions for the EKS service. In this case, the admin function is viewing the dashboard. It does not have any other privileges other than viewing the dashboard. And you can use the delegated administrator to start accessing this particular dashboard. So essentially two accounts get access per organization. The management account by default gets access once it enables trusted access. And once you enroll a delegated admin account, that account will also start getting the same view and you can start using that particular account for viewing the dashboard going forward.
20:07Awesome. Great. That is really helpful to know of how to operationalize this. So, you know, in closing, because we've been chatting for a while now, I'm curious sort of what else is on the roadmap for EKS dashboard. What else can you share about it? Sure. So after the initial rollout, we have been focusing on improving the widgets in terms of the verbiage and adding some more insights to the dashboard itself so any improvements to the widgets that will keep happening on a rolling basis other than that we also have a feature wherein we more many customers have asked us how can we reduce the scope of this particular dashboard by default it gives a view of the entire organization but i might have different organizational units within my organization and how do i reduce it so we are working on a feature called as scope delegated administrator, wherein you can define multiple different delegated administrators, each of which can have its own scope in terms of what visibility they get from the organization.
21:18So that will give you a restricted view of the organization. So if you have like different organizational units that just want to look at their clusters, it allows you to slice and dice by that dimension. So we are working with other AWS service teams to make that particular feature generally available, but that's something that's definitely on the roadmap. Awesome. One question I had, which I don't know if it pertains to the roadmap, maybe it is true today, but I know that it was likely at reInvent that we launched support for EKS hybrid nodes. Does the dashboard also give you visibility into everything that you said for those hybrid nodes sort of a setup?
22:05The answer is yes and no. In terms of whether you will be able to see which clusters are enrolled in hybrid nodes feature or auto mode features, that is work in progress. So we would definitely be able to provide a distribution in terms of how many clusters are using auto mode or how many clusters are using hybrid nodes. That's a feature that is going to come out probably within the next three months. It does not go into the specifics of what those hybrid nodes are made up of, similar to what you see with managed node groups. So the visibility will be restricted to just the cluster level view, which shows how many clusters are enrolled in hybrid nodes.
22:44So that's precisely what we are working on. Hybrid nodes, similar to self-managed node groups, is very difficult for us to figure out the different properties of an hybrid node because they are not exactly provisioned using our APIs. So we have limited visibility into those. But anything that is provisioned as a managed node group, we definitely have more visibility into it. And we can give you the high-level properties that will allow you to start planning your operational activities. Awesome. Okay, well, this was a really great sort of technical deep dive into what all the capabilities of EKS dashboard are.
23:21If our listeners want to learn more about how to get started with EKS dashboard, what should they use? Where should they go? Yeah, so we have like multiple launch materials that are available to you. You can just search for the EKS dashboard launch blog. And we also have a deep dive blog. Or the other place to get started is you can go to the EKS user guide. There is a dedicated page for EKS dashboard. You can refer to it in terms of how to set it up. and as you start using it if you have any feedback feel free to reach out to us our roadmap is public so you can go to github go to the containers roadmap and directly provide your feedback to the service team on this particular feature and we would be happy to engage you with respect to your feedback and see how what we can prioritize and when we can prioritize it or if you prefer to engage through your account team that's also fine please keep your feedback coming that is what will help us improve the product yes we are very customer obsessed as you just heard Sriam say so we will drop links for all of those resources on the on the podcast in the show notes for the podcast and yes please reach out with feedback through whichever mechanism best works for you all right well thank you so much Sriam for joining and thank you to our audience for tuning in.
24:48Until next time, keep on building.
From the publisher
In this episode, we'll explore how the new Amazon EKS Dashboard solves key challenges in managing Kubernetes at scale across multiple AWS accounts and regions. We'll discuss how it provides centralized visibility into cluster health, versions, and costs - enabling teams to improve governance, streamline operations, and optimize their Kubernetes infrastructure. Listeners will learn about key use cases like version lifecycle management, upgrade planning, add-on governance, and cost forecasting. We'll provide both a high-level overview for architects and managers, as well as dive into some of the technical details that will interest Kubernetes practitioners.
1) Whats New Post: https://aws.amazon.com/about-aws/whats-new/2025/05/eks-dashboard-multi-cluster-view-kubernetes-infrastructure-aws-regions-organizations/
2) EKS Dashboard User Guide: https://docs.aws.amazon.com/eks/latest/userguide/cluster-dashboard.html
3) Deep Dive blog: https://aws.amazon.com/blogs/containers/deep-dive-amazon-eks-dashboard-for-visibility-into-multi-cluster-operations-and-governance/
4) YouTube Video: https://www.youtube.com/watch?v=N1I1H2wA1Sk
