#710: Amazon S3: From Simple Storage to Smart Scaling

3 Mar 2025 · 48 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AWS Podcast Episode #710: Amazon S3: From Simple Storage to Smart Scaling

Episode Overview

  • Release Date: March 3rd, 2025
  • Hosts: Simon Elisha and Wali Akbari (Principal Solution Architect and Storage Specialist at AWS)
  • Focus: An in-depth exploration of Amazon S3, its new features, and best practices for data management and cost reduction.

Key Topics Discussed Introduction to Amazon S3

  • Definition: Amazon S3 (Simple Storage Service) is a highly available and durable object storage designed for cost, performance, and scale.
  • Usage: Customers utilize S3 for a variety of applications, including:
  • Machine learning training datasets
  • Data lakes and analytics
  • Backup and archival storage

Integration and Access Methods

  • APIs and Application Integration: Customers can access S3 using native APIs, command line interfaces, or through application integrations (e.g., S3 connectors).
  • File Protocols:
  • S3 File Gateway: Allows applications using SMB or NFS protocols to access S3.
  • AWS Transfer for SFTP: Enables SFTP access to S3 for data transfer.

Recent Updates and Features New Features from reInvent 2024

  1. Increased Bucket Quota: Default increase from 100 to 10,000 buckets per AWS account.
  2. Amazon S3 Tables: Managed Apache Iceberg tables for optimized storage and analytics.
  3. S3 Metadata: Automatic metadata capture for efficient querying of large datasets in near real-time.

Cost Management and Performance

  • S3 Intelligent Tiering: Automatically tiers data based on access patterns, reducing storage costs without manual intervention.
  • S3 Express OneZone: A high-performance, single-zone storage class for workloads requiring low latency.

Data Management Techniques

  • S3 Inventory Reports and Storage Lens: Tools for observing data usage and cost management.
  • S3 Replication: Automates the copying of objects across buckets, supporting disaster recovery and data sharing needs.
  • Batch Operations: Allows bulk operations on S3 objects without the need for scripting.

Security Features

  • Block Public Access Policy: Default on S3 buckets to enhance security.
  • Data Encryption: All new objects are encrypted by default.
  • S3 Object Lock: Provides data immutability to protect against accidental deletions.

Access Management Enhancements

  • S3 Access Points: Enable granular access control for different teams and applications.
  • Mountpoint for Amazon S3: Allows file-based access to S3 data without converting it to a traditional file system.

Best Practices and Recommendations

  • Tip for Cost Optimization: Consider using S3 Intelligent Tiering as a default for data management due to its flexibility and cost-effectiveness.
  • Use of Monitoring Tools: Leverage tools like S3 Storage Lens and CloudWatch for observing usage patterns and optimizing costs.

Conclusion

  • The episode emphasizes the versatility and evolving capabilities of Amazon S3, making it essential for cloud architects and developers to stay informed about these features for effective data management and cost control.

Additional Resources

  • [Amazon S3 Documentation](https://aws.amazon.com/s3/)
  • [AWS Storage Blog](https://aws.amazon.com/blogs/storage/)
  • Feedback can be sent to: AWS podcast at amazon.com

Key Takeaways

  • Amazon S3 is a foundational service for data storage, with a wide array of features designed for scaling and cost efficiency.
  • Understanding new features and best practices can significantly enhance how organizations manage their data in the cloud.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00This is episode 710 of the AWS podcast released on March 3rd, 2025. Hello, everyone, and welcome back to the AWS Podcast. So I'm going to share with you. Great to have you back. And we're doing a deep dive. We're doing a deep dive deep into S3. And the only person I could reach out who could help me for this was, of course, Wally Akbari, who's a principal solution architect and a storage specialist at AWS. G'day, Wally. How are you doing? Hey, Simon. I'm excited to be here on your podcast as a first-timer, so please take it easy on me. And look, obviously, love talking about data and storage. So let's dive deep into S3 and anything else that you want to talk about.

0:41Absolutely. Now, we did an S3 deep dive a few years ago. It's, in fact, one of our most popular subjects because S3 is kind of in there in many, many things. It is in many ways a magical thing when it comes to storing data at scale. It's one of our oldest services, not the oldest, but it is one of the oldest. And certainly probably the one that I'd say most people had their first experience of AWS with. So let's do the very high level. What is S3? What was it built for? And what purpose does it serve? All right. Awesome. Well, if we want to start at the very basic level, what Amazon S3 stands for.

1:24So Amazon S3 stands for Amazon Simple Storage Service. Look, it's a highly available and durable object store that's really designed for cost, performance and scale. Now, when data is stored in S3, it's stored in S3 bucket as objects, and they're based on unique key value pairs. It's a little bit different to the file world where you have a tree and a file, so to speak. And to put simply, S3 buckets are where you upload and download your data to and from. Now, to be honest, I know it says simple storage service, but it really is a simple AWS service to use, especially given the rich features that it has that our customers use to store their data for a wide variety of use cases.

2:06So we've got customers storing data on S3 ranging from, let's talk about machine learning, right? They're storing their training data sets to their inference-based models, all the way to data for their data lakes and analytics, and all the way to storing durable copies of their backup and archive assets, so to speak. Now, think about all those important backup, Simon, that customers want to store in a durable, available place that's secure. So that's really what S3 is at a glance and, you know, what some of our customers use it for. But, you know, how do our customers actually use Amazon S3? I've talked about how it's a great place to store different data types because it's so versatile.

2:49But it's, you know, customers can natively integrate with S3 using the S3 APIs or command line, or they can do it through application integrations into Amazon S3. Some apps have S3 connectors. But I really want to touch on some of the traditional applications and architectures that our customers have. And they want to integrate and use cloud technology, but look, their apps don't talk native S3 API. So I'll give you an example. Imagine your application uses the SMB or NFS file protocol for access, but you wanted these applications to access data in S3. Well, one method is you could spin up an S3 file gateway instance.

3:32This provides a file interface that applications can access, right? And that file interface is a connection into S3. So basically, your application has an SMB or NFS mounted on it, and the backend is S3, and you access those objects as files, right? Pretty simple. and I'll finish up with one other scenario, right? Simon, when was the last time you used the SFTP protocol? It's been a while. Yeah. Right? But SFTP is a robust protocol. It's been around for decades. And it's still widely used, right, in our customer architecture. So how does an app that uses SFTP integrate with Amazon S3 and maybe it's a way to integrate with the data, access the data, upload data.

4:24Well, this is where you can stand up one of our fully managed services. It's called AWS Transfer for SFTP. Basically, it's an SFTP endpoint that you spin up in a few minutes, present it to your SFTP clients, but that SFTP endpoint that you've spun up is actually backended by S3. So all your clients that use SFTP for data transfer now are actually uploading and accessing data from S3. Pretty cool, right? It's really as simple as that to modernize your traditional architectures with Amazon S3. And the thing that is amazing to me about S3, because I worked in storage for a long time, is it takes away the undifferentiated heavy lifting of storage.

5:09You know, storage is difficult because it's typically on some sort of recordable media, be it hard disks or SSDs or whatever comes next. The bad news is you have to refresh those pieces of hardware year on year or every three years or every five years, which is a royal pain in the whatever. And also you've got to manage capacity and you've got to always pay for more capacity than you need, et cetera. And so suddenly this model where A, you can store as much as you want, exabytes, have fun, no problem. You can store as many objects as you want. You can store objects up to five terabytes in size each, and you only pay for what you use as you're using it and you don't have to worry about capacity management or migration.

5:47It's a huge thing that in many ways we forget. But what we're going to talk about today is A, some of the new things that have come out recently through reInvent, but also some of the more advanced and detailed capabilities that maybe folks haven't really had a chance to poke at or have a go at. So what are some of the things that really popped up at reInvent 2024, Wally, that we can share? Sure, there was a fair few things, but But I'll try to keep to the top view. So first, look, Amazon S3 released increased the default quota for how many buckets you can have, from 100 buckets to 10 ,000 buckets by default per AWS account.

6:28Now, I just want to point out, just because you can create 10 ,000 S3 buckets per account doesn't mean you need to create 10 ,000 buckets. you know there's a saying keep it simple have the least amount of buckets and at the same time don't have one bucket to rule them all right um and by the way customers can actually increase a request for a quota increase beyond 10 000 buckets per account that's a lot of buckets right um another really cool uh capability or feature that we released uh it's called amazon s3 tables basically it provides fully managed apache iceberg tables and s3 now we keep talking about S3 buckets, now there's a new type of bucket.

7:07S3 tables is a new type of bucket called a table bucket. And it's purpose-built for managing and storing tabular data at scale using the Apache iceberg standard. And it's really optimized for performance. So, you know, you're probably thinking, well, how can S3 tables help customers with analytics applications, such as they may be using Amazon Athena, Amazon Redshift, Amazon EMR. Well, it makes it easier for customers who are actually using a self-managed approach with their tabular data, because if they use S3 tables, because it's fully managed, the underlying storage has been tuned for maximized performance.

7:48It does things like automatic compaction, snapshot management, and unreferenced file removal. and you talked about the undifferentiated heavy lifting. So, you know, don't worry about the storage and the scale anymore. And, you know, don't really think too hard about the, you know, how do I run the compaction and the snapshot management and the unreferenced file removal. So my analytics data is really, you can query it at maximum performance, right? So that's really what it's about, right? It is introduced to help those customers. And, you know, we talk publicly about that customers using S3 tables could see their queries accelerated by up to three times, right?

8:27Versus self-managed Apache Aspect tables in general purpose S3 buckets. And another one that was really... So you may not think this is a big change, but I feel this is a big change, right? So we released S3 metadata, which automatically captures metadata, which when you upload an object, there's lots of metadata, right? Now, we make that metadata creable via read-only table, and you can access that metadata in near real time in terms of minutes. So what does that mean? This is very cool. Very cool. Think about it. If you have hundreds of millions of objects or even billions, I've got customers with billions and billions of objects, right?

9:12Imagine you wanted to query that data, right? Even a list operation. Don't do a list operation if you've got hundreds of millions of files. Don't do it. um well you could use an s3 inventory report right which gives you a view 24 hours ago because it's generated every 24 hours and you can look at okay what type of data i had the names of the files you know use amazon theme to query that but what if you wanted something closer to real time uh reflection of your data assets that's where you would use s3 metadata and you guessed it or maybe not but uh But this S3 metadata is actually stored on Amazon S3 tables.

9:48So if you're looking - It all ties together. Yeah, it all ties together. If you're looking to query your metadata in near real time and S3 Invincial Report didn't do it for you and you didn't want to run a list operation, this is a win. Yeah, absolutely. Absolutely. Now, let's talk about cost because, again, in the storage world, storage always grows. It never shrinks. and so people get very wrapped up about cost per gig or cost per terabyte or what have you and as they should um and one of the things that's always been battled for in the storage world is tiering of storage you know i want to i'll spend more money on storing a piece of data for frequent quick access if it's worthwhile from a business standpoint but if i don't need it i don't pay top dollar but managing the tiering was always a challenge talk to us about what the state of the yard is today?

10:40Oh yeah, cost and performance are top of mind and cost. Look, I don't remember the last time I deleted all my photos, like my photos from on the phone just keep growing and growing, so to speak, right? And we tend to archive that data in a cost optimized manner. You never know when you're going to access that data. And look, in reality, when we talk to our customers in enterprise, there's a lot of requirement to keep data for a long period of time, right? Different retention policies, right? And how do you actually store that data in a durable and cost-effective manner and not have to worry about, like you said, tiering and what do I need to tier?

11:14So we released S3 Intelligent Tiering a few years ago. Now, S3 Intelligent Tiering is a storage class that actually automatically tiers your data between its frequent, infrequent, and archive tiers based on access patterns, right? Now, when you think about the object world, when you upload an object, it has a last, it has a created date. And we use algorithms to understand the access patterns to that data. So to put simply, you could upload 10 terabytes of data, for example, Simon, and out of that 10 terabytes, one file is very hot, and you upload these 10 terabytes January 1 in a particular year.

11:55And that one file is hot all the way to December. Now, S-Thinger Intelligent Tiering knows that that file is hot and the rest of the data is, let's say it's not accessed, it'll tier all the other files and data accordingly and leave that one hot file in the frequent access tier. So the way Intelligent Tiering works is it assesses your data patterns in the first 30 days. Everything's in the frequent access tier. Then if the data hasn't been accessed, then it moves it into the infrequent access tier, where it sits there for another 60 days. And if it hasn't been accessed again, it automatically tiers that into the Glacier instant access tier.

12:37These are all online storage classes, by the way, so millisecond access. Now, how's that, right? You don't have to worry about tiering, it automatically moves it down, saves you on cost. Because the access semantics for the object doesn't change wherever it is. The application is not aware and doesn't need to be aware of the tier. That's correct. And for those who are familiar with our S3 lifecycle policies, you could create that when you understand your object access patterns. So S3 lifecycle policies work on object-created date. So you create a rule saying, I want you to tier my data that's older than 180 days to this storage class.

13:17Fantastic. But storage S3 Intelligent Tiering is designed for situations where customers don't know their access patterns. And, you know, it's really hard sometimes understanding how your users interact with a vast array of data. So it takes away that guessing game for you. It automatically tiers your data down for you to save you on cost. And I think if you think about it, like my recommendation for customers is to just use intelligent tiering as your default option. That's where you start. And really, lifecycle policies comes into play if you know the nature of the data. So let's say you've got data that's for backups.

13:52And you know, well, this data has to be kept for 30 days in this tier. Then it's got to be 120 days and can't be deleted in this tier. And then it has to be deleted. That's when lifecycle policies come into play. That type of, as you said, that really well understood data classification. but I'd argue that probably 80 % of data, no one knows what the classification is in terms of use. So just use intelligent tiering and you don't have to worry about it, but you get the saving. Yeah, and that's it, right? So, you know, we have S3 Standard, which our customers used to start their data life, data journey on.

14:25Have a look at S3 Intelligent Tiering, right? If you have automatic tiering and access patterns that are worked out for you, sounds like a win. Yeah, exactly. Just happens. Now, we touched on, I guess, data management and observability of data. Well, let's go a bit deeper because there's a lot in there. Once you start picking away, well, what do I know about an object and what can I do with it? There's a lot. And there's a few different ways to approach it. So maybe give us a, if you want to do this, this is how you do a type view of how to use these different capabilities. Yeah, look, with Amazon S3, you know, we're talking about different scales of data.

15:03You know, some customers have terabytes, some have petabytes of data and billions of objects. How do you actually like, it's one thing to have data, but it's observability and then the data management's key. So the way I talk to our customers is have a way to look at a macro level, like organization-wide level or at the top level, and also have a way to look at the micro level. So let's look at the micro level really quickly. So if you wanted to work and understand your data at the prefix or name or tag level, that's where you'd use S3 inventory reports. I can create an S3 inventory report for a bucket or a specific prefix, and I can check the data there when I need to.

15:47Inspect it really at the micro level. But when you look at it at the macro level, what if you had hundreds of buckets? how frequently accessed some buckets or some prefixes are because that also then falls into security, right? Your security team may want to understand, okay, this bucket, for example, why is this bucket being accessed all of a sudden with hundreds of thousands of requests? We don't expect this. This is normally a dormant bucket, dormant data, so to speak. So this is where we released S3 Storage Lens, right? It gives you observability at the macro level, starting at the org level and you can drill all the way down from having a glance at all your buckets, how much data, how many objects you're storing at a glance, which is amazing, by the way, all that data available.

16:35You can go down to the bucket level, the prefix level, the account level. And on top of giving you that visibility, it actually gives you recommendations, some really intelligent recommendations, right? Firstly, it gives you, when you use S3 Storage lens, it gives you outliers, which are calculated using statistical analysis of the data, you know, of the last 30-day trend that, for example, we saw 5 million requests to this bucket or prefix in the last three days. You may want to check this out. Secondly, it gives you cost efficiency recommendations. Thirdly, it gives you data protection recommendations around best practices like, hey, you know, around encryption and replicating your storage.

17:15So that's really at at the macro level. And one of my favorite features of S3 Storage Lens, apart from the fact that there's a free dashboard configured for customers and they can optionally pay for the advanced metrics, which give you prefix level capability and longer historical reports. One of my favorite features is the top end dashboard for those who get to play with Storage Lens. It gives you a glance at the top end and you can pick that whatever end metric is. It could be top requests tool bucket or biggest bucket in terms of size and things like that, it is really, really cool, that top-end dashboard.

17:55Now, we talked about macro and micro-level observability, but then there's also data management, right? Especially customers want to create sometimes copies of data. How do I do that without having to write scripts and code? That's where we have S3 replication. Now, you can use S3 replication to replicate data that's newly created in S3 bucket to a different bucket in a different account in a different region to a different storage class. So many options there. But what we also provide now is if you have an S3 bucket and you turn on S3 replication, it'll prompt you saying, hey, Simon, you've got existing data in your S3 bucket.

18:36do you want to replicate the existing data as well on top of the newly created data? You can choose yes or no. So that's really an awesome capability, right, to manage your data and replicate it for different reasons. If you're activating it much later in the lifecycle of the bucket because your requirement has changed. In the past, you had to do some fancy coding to manage it yourself, and now you don't have to. Oh, definitely, definitely. it. And that replication also, you can sort of control or visualize the lag between one location and the other as well. So you know that data is being replicated within a good timeframe too.

19:14Yeah. And if you had really important application data and you want to have it replicated ASAP, we have a capability that you can enable called S3 replication time control, RTC. And that means that your data will be replicated. The majority of your data will be replicated in under 15 minutes. Now, that's pretty awesome when you think about cross-region, right? So it's one thing to be in the same region, but also cross-region. So that gives customers that added flexibility, right? I need, this is really high priority data in this bucket for this rule. I want this data to be replicated ASAP. Let's dive right down into the object level.

19:51So that's the sort of the most small component that we have within the S3 construct. And talk to us about versioning and tagging, because they're two very important concepts, So I think it's easy to forget they exist, but they're actually really handy. When you turn S3 replication on, it'll ask you to turn on versioning. That's something to remember. But look, versioning gives our customers peace of mind because they've used it for different reasons. Now, it could be for a, apart from keeping multiple versions, which, you know, makes sense. It's called optic versioning. But they use it almost as a, I'll say, backup mechanism.

20:28and it's not really a backup, but when you delete the active version of the object, it actually doesn't delete it. It's hidden from the view of folk until you go and actually delete it. So versioning lets, if you've got an application that generates lots of data and it's always creating the same file name, it gives you that peace of mind that you can go back to any particular version of the file, depending on how many versions you want to keep. So it gives you that peace of mind, if that makes sense. But really important - I always turn it on because I know I'm going to make a mistake. I always turned it on because I know I'll make a mistake at some point.

21:01Yeah. It's not a backup, but it gives you that peace of mind. But really important. It's like a rollback capability. Yeah, exactly. Like a rollback capability. But, you know, I tell our customers with versioning, just make sure you have the guard rails. Don't know to keep as many versions as possible. That's one thing, right? And S3 Storage Lens, the beauty of it will tell you about how many non-active versions you have, like in the data. So if you've got, someone's accidentally put the wrong object versioning policy on, you can use Storage Lens to kind of point in that direction along with other things for cost saving.

21:35Now, tags, tags are great. You don't want to search for data in S3 objects by the name or the prefix, right? So tags are a great way for cost attribution, for observability of data. But the one thing to remember is that, you know, create a framework for what type of tags you want to use and then apply it, right? Having two little tags may not give you the data that you need later on, but you can always add tags. You can always remove tags as well, but you also don't want to have too many tags, right? You don't want to have 20 tags for every object. It doesn't make sense. No, no, it's just the right amount.

22:14And again, I think it'll be interesting see how the use of tags evolves now that we have the metadata query capability too so again it's it's it starts to be you know it's our classic answer it depends it all depends what you're trying to do and that's the beauty of amazon s3 there's you know our customers have unique requirements even when it could be to observability some may use inventory reports some may use s3 storage lens some may use metadata and capture that data and build their own capability Some may want to leverage tags in a particular way. And it really helps them customize to meet their requirements.

22:52Absolutely. Now, we talked a bit about replication, which is more sort of a continuous replication. What about if I just need to move some data around on a one-off thing or maybe just periodically? Well, how good are your scripting skills summed? It depends if I'm using Amazon QDeveloper or not. well look the good news for you simon and some of our other customers who aren't um uh scripting gurus neither am i um look we've got a a s3 feature called s3 batch operations right uh it basically does what it is it performs batch operations and our customers use this at scale and to keep it simple you give batch operations an input file list like an inventory port that you've generated or your own custom one and say, these are the objects I want you to perform this next action on.

23:43It could be a copy. It could be, you know, change the tags on it. There's a few features that you can apply to it. And it makes it simple. You give it an input file. You tell it what action to perform, where to write the output to, maybe to a different bucket and away you go. So maybe, you know, if we go back to that replication question, I've got an existing S3 bucket. I turn on S3 replication for new objects. I've got a lot of data there, but I just want to copy some of the existing data, not everything. I don't want S3 replication to replicate everything. So you could get batch operations and give it that list or the prefix list and say, hey, just copy this to this other location.

24:25Now, another capability, and we talk about Amazon S3 and and there's a lot of integrations by other services into Amazon S3, is our AWS DataSync service, which really is an at-scale data movement service, right? So you can use DataSync, create a DataSync task to move data between S3 to, for example, EFS or EFS to S3, Elastic File System, to Amazon S3, or even between S3 buckets, right? So, you know, it depends on, again, your use case, but there's different ways to obviously move and create copies of data around. Yeah, absolutely. Now, one thing that's interesting, I think, about S3 is often we don't have to think about performance.

25:12You know, it performs pretty well and just does its thing really happily and you sort of sit and forget, which is great. But there are times, as you mentioned, where customers are like, no, I need the absolute maximum performance possible. What should they be reaching for? well they could reach out and leverage our new s3 express one zone storage class so it's it's high performance single zonal storage now for those who are familiar with abazon s3 they said wally did you just say single zonal yes i did uh now traditionally like um s3 standard and our other storage classes like s3 glacier are regional based services they're multi az now customers have said, hey, we want consistent single millisecond, basically, performance capability from an S3 storage class for our machine learning workloads and for our analytics workloads and so forth.

26:08So S3 Express OneZone gives our customers that capability. It's single zonal storage. Our customers can actually, when they're deploying their high performance compute stack, For example, they will deploy in a particular availability zone. They can then select the S3 Express One Zone Storage Class to also be in that same availability zone. So we're talking about minimizing latency at the storage layer and the access point to that storage, if that makes sense, Simon. And we're trading durability for speed in this case, because typically an S3 bucket is across all the availability zones in a region.

26:45There's multiple copies of the data. there's lots of overhead that takes place to do that. And what we're doing here is kind of shrinking it down saying, well, this is not for data that you're trying to keep for the next hundred years. This is for data you're processing very quickly. Durability, we're still designed for 11 nines, right? One is 11 nines across, let's say a region and its availability zones, let's say it's three AZs, like for S3 standard, three availability zones, but you can still have it designed for 11 nines of durability within a single availability zone. So how do you get the lowest latency, it's by localized access, right?

Read the full transcript

27:22So therefore, firstly, it's a zonal bucket. And secondly, it's a special type of bucket called a directory bucket. Okay. So we have general purpose buckets now, which is what a lot of our customers are used to. We have table buckets for Amazon S3 tables, and we have directory buckets for S3 One Zone Express. Now, let's talk a bit more about performance. So I mentioned, you know, as an S3 user, you don't kind of worry about performance per se. It just happens. But let's look under the covers a little bit. How is S3 actually managing performance? You know, if I'm accessing particular data heavily or how does it manage the access requirements?

28:04Well, that's the beauty. We do all that work for you, Simon, so you don't have to worry about it. But so I'll go back to just finish off something I just remember about S3 Express One Zone. So along with obviously the consistent single-digit millisecond latency, it's also designed for hundreds of thousands of access requests right to the directory bucket. Now, if we go to our general purpose buckets, for example, and performance that you're talking about. So they're designed for 3 ,500 puts and 5 ,500 gets per prefix, per bucket, per second. That's a long one there. And we always talk to our customers about in the past, you know, many, many years ago, probably in the previous podcast, you know, about partitioning your bucket.

28:51Yeah, and naming the objects and all that sort of stuff. It has unique prefixes, right? But guess what? We now have this, we've had this capability for a while. It's auto partition. so you know amazon s3 assesses what partitions are or prefixes are hot and you automatically you know do its magic on the back end to ensure that that data that's hot all right it uses access patterns oh this data is hot it's been hot for a while you know what it needs more performance and it makes the adjustments accordingly on the back end right so you don't have to you don't have to manage it you don't have to do the fancy the fancy object naming dance anymore oh you can if you want it's always fun right having those unique names but no the vast majority of our customers don't have to go through that and they just basically like you said undifferentiated heavy lifting is taken away from them they just go okay i need to create x buckets i need it for this purpose i need um this framework around the bucket and they're not really worried about performance, cost, or scale.

29:55Let's talk about what is job zero, our first priority, which is security. And there is no excuse for S3 bucket to be public. This is how I'd put it this day. Because even if you want the data in your S3 bucket publicly accessible for some reason, web serving, et cetera, it should be behind a CloudFront distribution anyway. But let's go back at the start. Let's talk about block public access and how that works and also some of the other things related to security for S3 that people may not be aware of. Yeah, and look, Simon, security is job zero for us, spot on. And what we did a few years ago is we released a, by default, block public access policy, which is enabled on all S3 buckets by default.

30:42So if you go to create an S3 bucket right now, that policy view we check, you actually have to manually go in and check it. So that's a failsafe. right there in a guardrail, right? Secondly, from a security point of view, S3 also encrypts all new objects by default, right? Using serve-side encryption, or you can pick your own encryption type that you want. So we stop public access directly to the bucket by default and we encrypt all new objects. But there's a lot more to securing your data, right? So you would need to have an AWS IAM or Identity Access Management policy created and also leverage Amazon bucket policies.

31:26So, you know, we're talking about securing your bucket from access at the access layer at the data layer as well. So there's a few elements there. And, you know, it's not just, okay, then you've stopped public access and you've got your IAM policies and your bucket policies. Fantastic. But what about malicious intent? What if, you know, it's something else from within the organization? People that have access. Well, we have a capability called S3 object lock, right, which you can enable in compliance or governance mode. What that gives you is data immutability capability for your S3 objects. So you could have a bucket policy that prevents or deletes, but maybe an admin's got access, right?

32:11You never know. So you could enable object lock for that really critical data. And that further enhances your security posture, right? So you've looked at the access angle, you've looked at the permissions angle, you've looked at the data angle to ensure your security posture is where it needs to be. Yeah, absolutely. And I think the other thing is that the ability to also track what's gone on from an API level through CloudTrail as well means that you can see what's going on internally, what those usage patterns are as well. Yeah, definitely. And look, you can go through CloudTrail, you can go through server access logs as well.

32:51And at the same time, they give you that logging capability. So, you know, you can go and triage that data. but again i want to go back to uh s3 storage lanes well how do you know right how many of your buckets are encrypted or your objects are encrypted and what's your security posture there and how much of your data is replicated and should it be replicated and you know which buckets all of a sudden have hundreds of thousands of requests coming in you know you have the different mechanism from the top level the macro down to understand oh our security postures should be better for this bucket.

33:25And then at the cloud trail level, you're inspecting on, hey, what's going on at the very granular level. So you've got both ends of the spectrum, Simon, which is really awesome to have. Exactly. One other thing I want to talk about in terms of accessing buckets is, now this service, as you mentioned, has been around for a very long time and the team have iterated on it continuously for customers. And so if at the start you were used to sort of, ACLs and bucket policies and stuff like that, that's not the way to do it anymore. So S3 access points really gives you a far better, cleaner, and more scalable approach to accessing your data.

34:00So let's touch on that briefly, because I think, again, I'm guilty of this. For some of the longer-term, I was going to say older users, longer-term users, we're sort of in the habit of doing it the original way, but this is a much better way to do things. Yeah. So we've got bucket policies, which are obviously at the bucket level. They're very top level, right? And S3 Access Grants is a way to give more fine level access to different applications and teams to that data. For example, you've got a bucket, it's got different data types, but you want to have maybe your machine learning team and your analytics team access the bucket, but at the different folders, if that makes sense.

34:46So it gives you that more granular access to the data, whereas the bucket policy gives you a more top-level approach to managing access to the data. It's really useful, I think, also if you have multiple defined customers or customer groups, internal or external, and need to access the data, it gives you a much more scalable way to manage that. And look, the name kind of says itself, it's an access point, right? you know if you for those who are familiar with amazon elastic file system efs and access points it's kind of similar to that right it's it's this it's um the same location but you give a different access point almost like a uh and obviously the policy attached to the access point means you have a particular access so we could be accessing the same bucket simon right but i'll use one access point you you'll use another and we get different views of the data exactly and it also That also means if you need to rotate credentials or turn off, let's say I can't have access to that data anymore, but you still do, then my access point gets disabled, but yours stays maintained.

35:51So there's no sort of, you know, big disruption versus if you have a single policy to control everyone, then it's like, oh my goodness, I'm going to redo it for everyone. Policies can get very long if you try doing it, very complex if you try doing that, right? And then at the same time, you don't always want to be mucking around with bucket policies for a single user, right? because then if someone accidentally makes a mistake, it could have a much broader radius impact. So customers have been asking for the ability to use S3 as a file system and there's been many ups and downs around that and obviously the semantics aren't the same and there's costs involved using gets and puts in a file system structure, but we do have something now which is specifically designed for customers to do that called Mountpoint for Amazon S3.

36:34Yeah, and Mountpoint for Amazon S3 has been a long ask from our customers. It's an awesome capability. So to put it simply, if you wanted a file view of what's stored in the Amazon S3 bucket from your local Linux client, then you would install the mount point package on your Linux client, and then you would just mount it like a normal, let's say, NFS mount point, similar to that, right? So you could go into the directory, do LS, you see all your S3 data as files. Now, really important, right? this doesn't convert Amazon S3 as an object store into a file system. Okay. They're two different things.

37:13Don't treat, you know, you shouldn't treat object stores as file systems because they have different characteristics. So this is a great way that our customers who use S3 to read data and write new data into S3 using a file protocol, this is a win for them. And this could be giving an example, you've got a few servers, Amazon EC2 instances, and you want to give them read access to the S3 data using a file interface, because maybe that's what the application needs. It doesn't talk S3 API. Then this is a great way all these applications can use amount point, access the data through a folder, and they have their own little cache, amount point cache as well.

38:02So there's some tuning parameters there. So that's really awesome. But if we look at the other end of the spectrum, again, you want to provide your applications that need data in Amazon S3, and it could be for generative AI inference. It could be for machine learning and training data sets. It could be for high performance compute examples, right? And you've got lots of Amazon EC2 instances or Amazon EKS containers and pods. Well, we have a service called Amazon FSX for Lustre and it has native integrations with Amazon S3. So FSX for Lustre is high performance parallel file system that is pretty self-explanatory.

38:45It's a high performance parallel file system designed for performance, right? We're talking about tens of gigabytes to hundreds of performance. Big performance. But that's a file-based protocol, Simon. But if you've got machine learning assets in the S3 bucket, well, how do I get my app that uses Lustre to access this S3 data? I don't want to write scripts. So when you spin up an FSX for Lustre file system, you can tick a box and say, I want you to import data from this S3 bucket and also export data back to this S3 bucket automatically. That's for new data, changed data, deleted data. When you spin up the FSX for Lustre file system instance, you see your petabytes, for example, of S3 data through the view of the Lustre file system, which is mounted on your compute host.

39:39Now, if you have a petabyte of data in S3, you don't need a petabyte of FSX for Lustre. It's a high-performance cache. It could be a few terabytes in size. And you go in there and basically any file that you touch, it'll pull it in from S3 the first time into the file system. Subsequent access is sub-millisecond, super fast. But the beauty is, what if you then create new files on that file system? You've done some work on the raw data, you've got export, you write it back to the file system, it automatically writes the Amazon S3 for you, exports it very quickly. So it just happens in the background.

40:16Yeah. And now you tie in S3 replication to this, Simon. What if you need to share data assets between different teams in different regions, right? So then you could have S3 replication created. And when that data hits the S3 bucket and you've got a rule to replicate new objects, it's replicated to a different bucket, same region, different region. For example, could be for backup, for DR, sharing, data sharing. So you can see how S3 becomes a centerpiece and apps integrate with it in different ways. So we've covered a lot of ground and there's even more we could cover, but we're not going to go too, too long.

40:53But let's touch on again, monitoring. We talked a bit about CloudWatch and CloudTrail and you talked a lot about storage lens, but how do I capture long-term the data or the metadata about my storage, my S3 storage? fantastic question so you know customers have used amazon cloud watch uh for a very long time you know there's s3 metrics in there you can even create custom dashboards using cloud watch which i love doing right you can create um a custom dashboard and then share it with other folk who don't even have adibus console access right which is amazing right so you know you give them it creates a username and passwords and uses our adibus authentication mechanisms and this could be app owners, right?

41:38So they can then see their S3 metrics, for example, get puts and other bits. But, you know, that's for self-service monitoring. But, you know, if they wanted longer term, again, we also publish storage lens CloudWatch metrics, if that makes sense. There are additional metrics that are published into CloudWatch, which customers can leverage as part of their CloudWatch dashboards. And for historical trending, again, I lean customers towards looking at storage lens. And I've said that so many times, we should say bingo, Simon. And with the advanced metrics, which are paid metrics on top of the base ones, you get up to 15 months worth of reporting, right?

42:22So you could go in there and look at the capacity, the total storage right or total requests for a bucket for example and it'll show you from x months ago till now you know and you got that graph which you know we love to see how what does this look like over time or how has this looked over time right so you know that's like capacity trends and performance like access trends um at the quantitative level how is it being used like what is going on yeah yeah that's right and um yeah we've got there's a lot of things we've got storage access analyzer and and there's too much to uh obviously discuss in the time you can you can see as much as you want about your storage is basically that's the end whereas at the start if we think about yeah again let's go back 15 years you know at the start you had you know list list objects and that was it that's what you got you know and then figure it out yourself well In the subsequent years, the team has worked very hard to give you all kinds of visibility.

43:23So you don't need to worry about that from that perspective. One other thing I want to mention, I'm going to come back with one last question for you, is that, again, folks who have used things for a long time will think of S3 as being eventually consistent. And in fact, some of the certification exams used to talk about, you know, what is the effect of eventual consistency? But now S3 is strongly consistent for gets, puts, and lists, as well as operations that change tags and ACLs or metadata. So this means you don't have to worry about that is my short message there. So good to bear in mind.

44:05That is an awesome capability. You think about strong read-after-write consistency for all applications without impacting performance availability. So that's the extra win there, right? So, you know, for new objects, deleted objects, you know, subsequent reads and lists are consistent, which is amazing. But I just want to talk about one other re-invent release that popped up and real quick is that we now also support conditional writes. Now, you may think, what are conditional writes? Well, you know, customers have asked for how do we check if an object exists, you know, without writing code so we don't overwrite the object and we don't want to turn versioning on, for example.

44:46Well, conditional checks allow you now to actually check if an object exists before they're uploaded. So this is awesome, right? On top of, you know, the strong read-after-write consistency. So it helps our customers actually continually optimize their applications. It's a nice little one that, again, if you're used to using DynamoDB, which has conditional updates, etc., it's a similar sort of concept. Very handy and fundamentally reduces the amount of code you have to create. Now, Wally, I've got a question without notice for you. And the question is going to be, besides using storage lens, which I get the point, is important to use.

45:26What's the one tip when you're talking to customers, you keep finding yourself giving them? What's the thing, the one sort of tip that keeps popping up as a really common thing for folks to do to get the most value? Wow. You saved the hardest question till last time. That's the way we roll. that's the way um to be honest talking to so many customers they all have um unique requirements and you know there's tips from you know cost optimization like i'll go with that right i say look at s3 intelligent tiering right if you've got all your data on amazon s3 standard right now have a look s3 intelligent tiering because s3 standard doesn't give you the tiering this this does effectively, right?

46:11And it's a two-way door, right? So you can always go back to S3. And for example, S3 Intelligent Turing will not tee objects smaller than 128 kilobytes in size, okay? They'll stay in the frequent access tee. So if you're using S3 Standard and you've got lots of small files and lots of large files as well, well, the rest of your data will tee down. So clearly, yeah, intelligent archiving is the thing to look at because that gives you a few things. It gives you cost control. It gives you potentially performance benefits as well. And I like your reference to a two-way door. So for folks that aren't familiar, a two-way door is a door that a decision that you can undo easily.

46:49And the beauty part here is we're not talking about, to emphasize this, an application level change. Like the interface 2S3 doesn't have to change. This is a backend cloud architect change, isn't it? Yeah. It's basically, it's, you know, another lifecycle policy or a copy or a batch, S3 batch operations job, there's a lot of ways, right, to move the data back, right? And that's why we make it sound like it's a two-way door, right? So it's designed to save you on time, cost, and money. Exactly. Exactly. Wally, thanks so much for sharing your insight into all the wonderful ways you can use S3. Thank you for having me, Simon.

47:30I hope the folk listening are as excited about data and storage and Amazon S3 as I am. And yeah, awesome. We've certainly learned a lot. And we do love to get your feedback, AWS podcast at amazon.com. And if you're wondering, yes, the podcast files are stored on S3 and they are served via CloudFront and the RSS feed is also on S3 and it is also served on CloudFront. So just saying, we use what we talk about. And until next time, keep on building.

From the publisher

Think you know Amazon S3? Think again. Discover game-changing new features like S3 Metadata and S3 Express OneZone with your host Simon, and guest Wali Akbari, storage expert at AWS. Learn pro tips for managing data at scale, reducing costs, and leveraging new capabilities - essential listening for cloud architects and developers.
For More Information about Amazon S3 visit: https://aws.amazon.com/s3/
And be sure to visit the AWS Storage Blog: https://aws.amazon.com/blogs/storage/

More from AWS Podcast

All 45 episodes
#710: Amazon S3: From Simple Storage to Smart ScalingAWS Podcast · 48 min
Listen in VO