#230 Jamie Lerner: How Quantum Solves AI’s Need for Unstructured Data Solutions

13 Jan 2025 · 53 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. Podcast Episode Notes

Episode Information

  • Title: #230 Jamie Lerner: How Quantum Solves AI’s Need for Unstructured Data Solutions
  • Host: Craig S. Smith
  • Guest: Jamie Lerner, CEO of Quantum
  • Sponsor: Oracle Cloud Infrastructure (OCI)
  • Release Date: [Insert date]

Episode Summary In this episode, Craig S. Smith speaks with Jamie Lerner about the evolution and future of data storage, specifically focusing on unstructured data management and how Quantum is addressing these challenges. Jamie discusses Quantum's innovative technologies, the important distinction between structured and unstructured data, and how their solutions support AI-powered workflows in various industries, including healthcare and media.

---

Key Topics Discussed

Introduction to Quantum

  • Overview of Quantum: A company specializing in the storage and management of unstructured data.
  • Historical Context: Originally founded as a disk drive company in the late 1970s, Quantum has evolved significantly in the past decade.

Unstructured Data vs. Structured Data

  • Definitions:
  • Structured Data: Includes easily organized information within databases (e.g., phone numbers, billing records).
  • Unstructured Data: Encompasses complex files like videos, CAT scans, or genomes that require specific handling due to their large size and varied formats.
  • Data Workflow Lifecycle: Unstructured data experiences a lifecycle that includes creation, modification, analysis, and storage.

AI and Automation in Data Management

  • Role of AI: AI technologies are increasingly integrated into data workflows, enhancing processes like analysis, tagging, and metadata organization.
  • Automation Policies: Quantum's tools enable organizations to automate data movement based on predefined conditions (e.g., duration of inactivity).

Data Sovereignty and Security

  • Importance of Data Sovereignty: Organizations need to know where their data is physically stored and the legal implications of that location.
  • Global Trends: Many countries are developing sovereign clouds to ensure data stays within national borders.

Innovations in Data Storage

  • Technologies Discussed:
  • Flash Storage Systems: High-speed storage solutions designed for rapid access and processing.
  • Tape Storage: Long-term archival solutions that provide cost-effective storage for massive data volumes.
  • Emerging Technologies: Innovations such as synthetic DNA for data storage and advancements in compression techniques.

Challenges in the Data Storage Industry

  • Exponential Data Growth vs. Budget Constraints: Organizations are accumulating more data but with flat operational budgets, necessitating more efficient storage solutions.
  • Competitive Landscape: The storage industry is dynamic, with numerous players ranging from startups to established companies like IBM and HP.

Future of Data Storage

  • Vision for the Future: Quantum is focused on developing high-speed, scalable solutions that cater to the growing demands of AI and data analytics.
  • Environmental Considerations: Ongoing exploration of how to store increasing amounts of data in smaller, more efficient formats.

---

Conclusion Jamie Lerner articulates a comprehensive vision of unstructured data management, emphasizing the necessity for versatility in data storage solutions as AI continues to evolve. The conversation reflects on the foundational role of infrastructure in enabling AI-powered applications across diverse industries.

Call to Action Listeners are encouraged to subscribe to the Eye On A.I. podcast for more enlightening discussions about technology and innovation.

---

Additional Resources

  • Quantum Website: [Quantum](https://www.quantum.com)
  • Oracle Cloud Infrastructure: [Oracle AI Offer](https://oracle.com/eyeonai)
  • Follow Craig Smith on Twitter: [@craigss](https://twitter.com/craigss)
  • Follow Eye On A.I. on Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)

---

Episode Timestamps

  • 00:00 - Introduction to Jamie Lerner and Quantum
  • 02:21 - Quantum’s Focus on Unstructured Data Storage
  • 05:19 - Structured vs. Unstructured Data: Key Differences
  • 07:52 - Managing Data Workflows with AI and Automation
  • 10:55 - Quantum’s Role in Long-Term Data Archives
  • 13:32 - Data Sovereignty and Security
  • 16:18 - How Data is Stored and Protected Across Mediums
  • 19:54 - Metadata in AI and Data Management
  • 21:29 - Quantum’s Role in Building Forever Archives
  • 24:16 - Tape Storage: Efficiency and Longevity
  • 29:11 - Innovations in Data Storage
  • 34:39 - Competing in the Evolving Data Storage Industry
  • 37:56 - Innovations in Flash Storage
  • 40:55 - Balancing Cost and Efficiency in Data Storage
  • 44:28 - The Future of Data Storage and AI Integration
  • 50:07 - Quantum’s Vision for the Future

---

This markdown file serves as a comprehensive overview of the discussed topics, key takeaways, and important insights from the podcast episode featuring Jamie Lerner.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Very large percentages of the world's digital archives are on quantum equipment. So things like the Library of Congress, the biggest archives of movies and television dating back to the first movies and television ever recorded on film and then ultimately digitized. And we call these sometimes 100-year archives, 500-year archives, but we design the technology to last centuries. We have a single customer who has 14 acres of that equipment. They've chosen tape because they want the lowest cost because it's so damn big. They chose tape to get the lowest cost per terabyte. And they've used enough tape to go to the moon and back eight times.

0:47Even if you think it's a bit overhyped, AI is suddenly everywhere from self-driving cars to molecular medicine to business efficiency. If it's not in your industry yet, it's coming fast. But AI needs a lot of speed and computing power. So how do you compete without costs spiraling out of control? Time to upgrade to the next generation of the cloud, Oracle Cloud Infrastructure, or OCI. OCI is a blazing fast and secure platform for your infrastructure, database, application development, plus all your AI machine learning workloads. OCI costs 50 % less for compute and 80 % less for networking. So you're saving a pile of money.

1:36Thousands of businesses have already upgraded to OCI, including MGM Resorts, Specialized Bikes, and Fireworks AI. Right now, Oracle is offering to cut your current cloud bill in half if you move to OCI. This is for new U.S. customers with minimum financial commitment. See if your company qualifies for this special offer at oracle.com slash IonAI. That's IonAI all run together, E-Y-E-O-N-A-I. So go to oracle.com slash IonAI to see if your company qualifies for this special offer. Quantum builds infrastructure for people dealing with large quantities of unstructured data. We deal with data storage infrastructure for that in many modalities.

2:33Think of it as an end-to-end architecture. So whether you're dealing with AI work or editing or processing work of unstructured data, or whether you're archiving that or backing up that data for long periods of time, we deal with all modalities of unstructured data in its whole life cycle. And so Quantum, the company name, has nothing to do, certainly, with quantum computing. But is Quantum a term used in data storage? No. Quantum was originally founded as a disk drive company in the late 70s, early 80s. And the founders chose Quantum to be their name. It might be time for a name change for our company, actually, at some point.

3:27But because what we do today is so different than even what we did 10 years ago. We've really, over the last six years, just completely reimagined our portfolio and really changed the end customers that we serve and the use cases that we go after are quite a bit different today than they were even five years ago. Well, the storage industry or storage question has always fascinated me since I've been involved. And I used to talk a lot to a company called LabelBox, and they have a platform for labeling data, a lot of it video. And I used to ask them, like, what happens? I mean, you know, we've already got mountains of data and we're creating mountains more.

4:27Where does all this data go? And I wrote a piece and what I discovered in that research, this is a number of years ago, that a lot of it gets ultimately transferred onto tape drives and is sort of in cold storage. but maybe you can talk first about where data goes once it's created. I mean, certainly there are databases, but they don't necessarily hold the data. They're like indexes, right? Yeah. So, yeah, can you, I mean, it's really a fascinating part of technology today, particularly with generative AI. So can you talk about how that's structured? Yeah. I mean, you know, let's start with structured data.

5:24Not so interesting. So you're talking about things like phone numbers, people's names, billing records. They usually go into a database and they live in that database for their whole life. And they may be edited or updated. Like your bill may be marked as paid or not paid. But unstructured data is totally different. So an unstructured data would be maybe a two-hour feature film, or it might be a human genome, or it might be a CAT scan of a human mind, and that CAT scan may have 100 ,000 images in it. Now, those don't really go in databases. They tend to sit as a file. And those files, unlike structured data, they live a life as part of a workflow, and that humans work on them.

6:13People work on them. They do stuff with it. So like when your x-ray image is taken in a visit to the doctor, that image may begin its life on an x-ray machine with some attached storage to that x-ray machine. Then that file may be moved offshore where during the night a doctor in another country studies that x-ray. And then that data might be enhanced in that that x-ray image now has notes attached to it. So now that x-ray says this x-ray belongs to this patient. This patient is a male and this patient is of this age and weight. And this patient had this complaint and this is the analysis. So that file starts to grow.

6:59It starts to have metadata. Then that file may go back to a doctor in another country where that data is kept on a very warm storage, maybe all flash storage as they go into surgery. But two weeks later, four weeks later, that may not need to be on very fast storage anymore. It might then be moved to a hard drive based system. And then, like you said, it might as it gets to be months old, it might go to tape. It might go to tape at a hospital. It might go to tape in the cloud, like things like Glacier are built out of tape. Amazon's Glacier uses tape, as most clouds do. Or it may go to another cloud like Google where they don't use tape and it sits on a hard drive forever.

7:49But it moves. And really what quantum's become really obsessed with is when do you move it? Where do you move it? and ultimately finding that storing the data isn't enough, that you have to store not just the data, but the metadata, all that enhancement data that goes around it. You have to provide tooling that says, these are the rules whereby I move data. If that x-ray image is three months and no one looked at it, that might be the trigger to automatically move it to the cloud or move it to a tape system or an archival system of some kind, or maybe move it to three or four places so it's very safe.

8:37You got four copies of it or three copies of it. So that policy engine needs to be there. That metadata engine needs to be there. The data mover technology needs to be there. And you need to have those modalities that follow you from the emergency room to analysis of that medical image, all the way through storing that medical image for the remainder of your life and maybe even possibly after your life. And that's what makes unstructured data so interesting because it has this life. And what's making that life even more interesting is in almost all of those life cycle steps, there's AI somewhere.

9:20Maybe there's a doctor who reviews that x-ray but before they review it there may be ai analysis of the bone break or whatever may be going on with that image more and more there's ai steps which make that infrastructure and the tooling all that more interesting and more to my point it's going to take a lot more than just storage to support these very modern work. And so physically, I take a video and I upload it to the cloud. So physically, those zeros and ones are being written onto a hard drive, let's say, in Google's cloud. But over time, it gets offloaded to some longer term storage if I'm paying for it.

10:26And how old, you know, ultimately, then it ends up on a tape drive somewhere, right? And so how old is the oldest digitized data that we have on tape drives and how how long will that last i mean is it is this going to be uh and how much space does that require so i mean we have you know very large percentages of the world's digital archives are on quantum equipment so things like the Library of Congress, the biggest archives of movies and television dating back to the first movies and television ever recorded on film and then ultimately digitized. And we call these sometimes 100-year archives, 500-year archives, but we design the technology to last centuries.

11:30And it doesn't mean it stays on one piece of equipment for centuries. What it means is it's designed to generationally be upgraded. So you might store something on a hard drive, but we keep multiple copies. And then what we begin to do is start upgrading those copies to newer and newer things. And the older copies, they tend to be retired on older equipment, older protocols. And we have customers that are in their 40th year, 50th year with us, where we are constantly upgrading, sometimes every 10 years, every 15 years. Some of our archival systems are designed to run 30 years. And it's pretty rare in computer equipment that tends to be fully depreciated in three years, almost always thrown out after five years.

12:30To have storage equipment that can have lifetimes in the 30s of years. And then around 30 years or 15 to 30 years, we begin taking it through generational upgrades. but it is designed to last hundreds and hundreds of years. Yeah. And where physically, where is most of that stuff? And what medium on Quantum's equipment is it? Almost all of our technology runs in three modes, runs flash, disk, tape, actually four modes, flash, disk, tape, and cloud. Almost every product we have, whether it's the world's fastest file systems, whether it's our cheap and deep archival, they have those modalities. And I mean, they exist in data centers, obviously.

13:20We have a lot of customers that want to have data sovereignty, have full control over the data. They want to know what nation it is in. They want to know what laws their data is under. I think that's becoming a really big deal for a lot of countries. It's not just where is my data, but what laws is it under? I, you know, having data in a foreign country that's under U.S. law or someone else's law can be really problematic for some people. So we have a lot of customers that want to own that equipment, want to know where it is, want to know what laws it's under and want to have full control. We have other customers that put their data into the cloud, and very often that's using quantum technology in the cloud.

14:10Usually our cheaper and deeper technologies is usually what the cloud customers are after or our extremely high speed file systems. But it's spread between the cloud and on premise. Our products are bought by everyone from cloud operators to national laboratories, militaries, governments, and large enterprise. And that question of wanting to know what country their data is, they're talking physically. Like, where is the tape drive on which this is stored? Do you have disk drive, flash drive, whatever medium it's on? I mean... And how do you track that? Well, if you own the system yourself, you know.

14:56The bigger issue is when I put my data into a cloud provider, many of them cannot guarantee you that that data will never leave their country. And then if the data does leave their country, they cannot guarantee you that it'll be your country's laws that govern it. And things like the Patriot Act do concern a lot of people and especially certain banking data, government data. They do want to know the sovereignty of that data. And that's why you're seeing a lot of countries create their own sovereign clouds. Germany has several sovereign clouds, France, Italy. you're just seeing uh india china you're just seeing countries start to say we're creating clouds for our country so that we know what laws we're under and we also know physically that it doesn't leave our borders or if it does we're very explicit about it yeah and and uh When you upload something to the cloud, it doesn't necessarily reside all of the data on one drive or on one.

16:17Isn't it? Almost never, right? I mean, the way you protect data is you shrink it and then you copy it lots of places, right? That's been the, you know, whether it's RAID or erasure coding, but the core technology of how almost all data storage works is make the data as small as you can. So you can put it lots of places, lots of locations, different pieces of equipment, like different hard drives, different computers, different locations, spread it around so that you have very little risk that if something bad were to happen, you always have one copy somewhere. And so when a child does that, they're worried there's a copy of it in a country we don't want or in a place that we don't know about.

17:03But it's always, the data is always kept together. I thought there are cases where the data itself is broken up into bits and those bits are stored across various places. Yes. A lot of algorithms shard the data. So you might take a file, break that file into pieces, and then put tons of pieces all over the place and reassemble it. Is it a technique that's used in object storage a lot? And how physically is there, I mean, I'm sure somebody, young people measure stuff all the time. Is there some metric that you could use to describe physically how large all of the world's existing data is? Because it's growing exponentially, right?

18:00Yeah. I mean, I think the reason data is growing is not entirely because it's higher fidelity, right? A movie today uses more storage than a movie 20 years ago because it's shot in higher definition. There's more frames per second. There's more fidelity in the, you know, a 4K camera, 8K camera have more fidelity than a high depth. But coupled with that, I think companies are realizing we'll never be able to analyze the data we don't have. And analyzing public data on the cloud, it's very hard to do that in a differentiated way. So companies are saying, my strategy is that I'm going to keep my data that no one else in the world has, my data.

18:59and I'm going to analyze our company's data and our customers' data. We're going to analyze it in a way that we can make our business more successful. And you used to have a strategy where after you kept data seven years, 10 years, you would dispose of it. And now people are realizing, oh, my gosh, we'll never be able to look back at that data. So people are kind of becoming data hoarders. Like, I can't analyze it yet, but I know I will be able to. So I want to keep all my financial data, all of my video surveillance data, all product data, every piece of data I have, because there will be insights I can get from it.

19:48So I think that is the first step to AI is keeping your data. your first step to AI is not usually doing AI it's I better keep my stuff and then I better organize my stuff you know just having a ton of stuff doesn't really help it's like what is it all like is this video surveillance video surveillance of what of whom where and that's where data enrichment and what you were calling labeling or metadata tagging like I better to start organizing my giant warehouse of old stuff. And then, and that's why you end up with these modalities where almost every modern organization that's doing analytics ends up with three modalities, a crazy high speed area, a working and organizing area, and then a giant, you know, cheap and deep warehouse to keep stuff forever.

20:46And one area is where you do your, the sexy part, which is, you know, developing models, training models. But the majority of work happens in this middle area where you're organizing, cleaning, making things of the same formats, categorizing things. And then again, all the work, old work, work that's completed goes into this, this kind of a, what I call kind of a forever archive. And I don't think anybody's ever contemplated forever archives because you threw your stuff out after 10 or 20 years. Now people are really, I think that's becoming a core part of an IT infrastructure is a 100-year, 500-year data archive.

21:28And that kind of a data archive. So, Quantum, you develop the tools or the platform to move data around between these different modalities. But do you also have data warehouses where you physically store data, or is that a different business altogether? Yeah, it's somewhat of a different business. We make some of the world's fastest file systems for doing the large language models, for doing analytics work. Those are all flash systems. Those are very high speed. But they have a lot of the capability, metadata tagging, these very large namespaces to be able to have billions, hundreds of billions of files.

22:22They scale massively. They're designed to saturate a super pod or saturate a analytic engine. We make those mid-range systems and we make these kind of cheap and deep archival, you know, forever archive systems. And we work with customers to understand what they're doing to figure out how much high speed do you need? How much mid-range do you need? How much archival do you need? And then what tools do you need to move the data, to tag the data, to protect the data? Sometimes they want this data protected on the cloud or protected at other locations or protected off site. And we help them build that end to end infrastructure.

23:08But we stop short of making the analytics software, making the, you know, the cube analytics software like you were talking about. we stop short of that and focus on the data storage, data movement, data curation software to manage all that. Although what kind of fascinates me is that the cheap and deep, as you describe, the long-term archival storage, and I just, I'm a visual guy, so I'm trying to imagine. It'll give you a sense. We have a single customer who has 14 acres of that equipment. And their infrastructure, they've chosen tape because they want the lowest cost because it's so damn big.

24:08They chose tape to get the lowest cost per terabyte. And they've used enough tape to go to the moon and back eight times. Wow. And then there's some sort of a robotic system that searches for the tape, pulls it out, brings it back and puts it on a machine if you want to retrieve the data. the data? Yeah, I mean, there's that piece of the architecture, probably the most, when you think about those 100-year archives, the most fascinating piece of that architecture is the software, right? The software that places it there, that keeps track of where it is, keeps track of how it's aging. Is the chemistry on the tape aging?

24:52Should it be brought to a newer system? Helps you search for things, metadata tags. So you could go to that archival and say, I want all cat scans of football players who died before the age of 50. And you could go find that. So it's not just a dumping ground of files. It's actually a card catalog, a very rich metadata of what all is inside there. Because what we learned working with movie makers over the last 30 years is they kept everything right. No one's going to throw away Star Wars. A single thing from Star Wars or Indiana Jones or whatever, The Godfather, whatever your favorite movies are, nothing gets thrown out.

25:39But what they found in the early days is they didn't build that rich metadata tagging. And you just ended up with sometimes hundreds of millions, sometimes billions of files from these franchises, these movie franchises. And it's almost impossible to do anything other than scroll through them to find something. What we've learned is it does take this curation to say, this is, yes, this is a clip of Star Wars, but it was shot in this year. It has these actors in it. This is what is happening in the scene. This is what they're saying inside it. So you end up with a much richer thing than here's a file that's in a Star Wars folder.

26:29Yeah. And that that's the software to do that. That's automated. Is that right? You feed the data. Yeah. I mean, we make software that takes everything people say and turns it into searchable text. So you can say, what are they saying in the video? We integrate facial recognition to say, what are the actors in it? If you've got a slate like those little chalk slates, we can read the slate and say, this is what the slate says. We can determine what camera was used to shoot it. What sound recording quality is there that we can automate? And then we allow humans to come over and say, you know, the director didn't like this angle.

27:15or someone can say, well, the director loved this angle, but the lighting was terrible. So I had to do a lot of work to it. And similarly, you could imagine with a genome, this is what makes this gene very interesting. This is a very interesting case. And a doctor or a physician of some kind or researcher may begin to annotate that. So it's a combination of AI generated metadata and human added metadata that starts to build these archives and these searchable archives. And then you, so one organization with 14 acres of, and I imagine inside it's racks and racks and racks of tape reels like that big or I have no idea in the world of tape I mean again the technology I'm describing you know I will tell you runs flash disc and tape for the tape side of it it's cartridges you know they're$40 cartridges that that store uncompressed 18 terabytes and much more compressed and deduplicated they last well curated they can last 30 years you know the lifetime of five or six hard drives.

28:34So they are very long life. They're very fast when they're sequential and super inexpensive architectures. And we've just launched a new deep archive box that has over 2000 tapes in a single standard rack, if you can believe that, and can store, you know over 30 petabytes in a single rack so they're just quite enormous items um they're just the the cartridge itself is how big yeah wow before yeah about four or five inches right and so inside those 14 acres for example there would be tens of millions on racks tens Tens of millions of tapes stuffed inside racks with little robots, moving them around to drives.

29:35So, you know, here we are in 2024, and this stuff has been piling up since the 50s, maybe.

29:48In another 50 years, I mean, how much space? This is what boggles my mind. I mean, are we going to live in a world where there are everywhere these giant data centers of, you know, tape storage? Maybe. You know, the other thing we're getting a lot better at is storing more data in a smaller place. Just aerial density, right? I mean, you think about how much data that may have taken a room to store could be stored now on a single microscopic piece of DNA. so you know and that technology is you know it's pretty hard to commercialize but it does make the point that there are ways that we know of now where we can store an enormous amount of data in much smaller you know on synthetic dna um on etched glass there's just cheaper ways there's more compressed ways to store data.

31:09And, you know, so I assume that as the it's kind of what's always happened is the data we store gets bigger. The mediums we store it on gets smaller and smaller. Right. I mean, just look at a hard drive. We used to be like, hey, I got a I got a whatever 10 megabit hard drive or I got a 10 gig hard drive. And now it's just like. So what? You know, It's just they get bigger and bigger and bigger and bigger, but they're not physically any larger. You still get a hard drive that's this big. It's hundreds of times larger now than it was 20 years ago, but it's not physically any bigger. So I assume we'll just continue to mash them down.

31:53But I think it's also safe to assume data center is going to eat a lot more power and use a lot more physical space, too. Yeah, yeah, it's fascinating. And then for Quantum, how do you work with a customer? Well, usually it starts by understanding the work they're doing. Right. Are they making a movie? Are they developing a drug? Are they doing analytic work? What kind of analytic work? And that's usually a workflow. Right. The data starts its life here. It moves to here. It moves to here. It moves to here. And it goes through a series of steps, a lot like a postmodern factory. It's really a data factory.

32:42Right. The genome gets shot out of the genome sequencer. Then the next step in the factory is a researcher does certain work to it. An analytic run is made. Learnings are found, you know, and it goes through this factory. A movie is made in a factory, right? It's colorized. It's sound edited. It's closed captioned. It's got, you know, all these steps are made to it. So we understand what is your data factory look like? And then we provide the infrastructure to run that factory from extremely high speed record setting, all flash file systems, mid-range object systems, and again, archival systems.

33:26And we make all of those modalities, you know, putting a lot of energy into the high speed part, right? It's usually, you don't get the archival piece and the mid-range piece unless you win the speed race for the high speed. So we put a lot of energy. We have a new system out called Myriad, which is setting some of the world's fastest SMB and NFS benchmarks. We're adding a parallel client, extremely high speed,

33:59parallel computing client to allow us to get even higher speeds with GPU direct built into it, with our global namespace built into it, to go even faster. So we're putting a lot of energy into our Myriad file system, which is just very scalable, very fast, and almost shockingly simple to use. And by winning the beachhead of speed, then we think we pick up the other pieces with the workflow. But I think you've got to win that speed race to get the full workflow. And is this a competitive industry, the storage, or is it like two or three big players? No, I would say it's very competitive from startups to long-term players like IBM, Hewlett-Packard, Hitachi.

Read the full transcript

35:00You've got long-term players. You've got new startups. So it's a pretty robust ecosystem and customers have a lot of choice. So you really have to you got to be coming out with new versions of your product all the time. And I think that's one of the things I'm most proud of at Quantum is the the amount of investment, the amount of innovation, The amount of unique technology that we have is probably why we've been around since 1979 and continue to operate as an independent company. It's just the innovation here. You just rarely see this much innovation in a company of our size and of our age. Yeah.

35:49And are you global or are you focused on the U.S. market? No, we support customers in, I think, slightly over 150 countries today. So a very large global footprint. A lot of that's got to be government because a lot of those 150 don't have much industry. Yeah, a lot of government customers, everything from government archives to, you know, disease modeling, water modeling, city planning modeling. So a lot of governments do a lot of AI modeling as part of the research. A lot of research institutes that are doing advanced AI models. drug development around the world, uses a ton of AI genomics modeling.

36:35So, you know, we look for that kind of advanced scientific computing. And then, you know, we've been supporting movie and television makers in every nation. Just about every nation has a news station, has entertainment, television, and we're in just about every country for their movies, television, and sports production. And a lot of this equipment, other than on the software side, the physical equipment, the storage medium, do you have partners that manufacture that? Does Quantum manufacture that? Yeah, we all of our software for file systems, object stores, backup, it runs on any hardware, any server hardware.

37:18So a customer can have us deliver to them on commodity hardware or they can buy their own. Now, you can imagine certain countries like India, China, they much prefer to use their local brands. Other countries aren't as specific. But the tape equipment, we do manufacture that because that is a electromechanically engineered thing that quantum designs and manufactures. But our file systems, object store, backup, those things are just software that we make with commodity hardware. Yeah. And when the tape cartridges that you're talking about, what is the, you know, we all have videotapes from the first generation recorders and I look at them and they're all pixelated and, you know, they break down.

38:16How do you manage that? So what makes tape cartridges so interesting, and again, we have millions of customers, sorry, thousands of customers who don't use any tape from Quantum, right? It's not the only product we make. But a tape cartridge is effectively a spooled up piece of cellophane with some chemicals sprayed on it. and the data is stored three-dimensionally in those chemicals. It's actually kind of amazing. And so we can study that chemistry. We have a device inside the tape unit that periodically looks at the tape and says, how's the chemical substrate holding up? Under proper humidity and temperature, it could hold up for 30 years.

39:11So we're constantly looking at it. And when we determine it's starting to decompose, either because of poor conditions or age, we'll move the data to a safer place for you and say, why don't you get rid of that old tape? Don't worry, we've already moved it for you to a newer one. And again, we talked about sharding, where we break up data and spread it around. What we do is we start spreading the data around, moving it from the older aging tapes to newer ones. And we do the same with flash and with hard drives where we can similarly, just in actually the same way, you're probably aware flash storage ages as well.

39:57As you read and write and read and write to flash, you wear it down. And a hard drive, similarly, as you read and write and read and write, it actually starts to, sectors start to die, parts of the drive, you know, get blemishes. and the drive starts to panic, we similarly, so whether it's flash, disk, or tape, we're studying it's wearing out. And as it wears out, we're moving it for you to newer, safer pieces. So you can literally just once every couple months, just take all the dead hard drives, all the dead tape, all the dead flash memory, SSDs, and you just throw them away, replace them with new ones, and the data is just balancing and loading itself.

40:47And you never even really have to think about it. You just have to cart away the aged out components over time. Yeah. And on the flash, you were saying that the business is like being as fast as possible on the front end that's really your differentiator i would imagine can you talk about flash storage and and what it is and how it works yeah i mean there's a lot to flash storage but what you want to do is get the very very fastest speeds and you're unable to do that taking an old hard drive system and saying I pull out an hard drive and I'm going to put in a piece of flash storage because the software is still going to treat it like a hard drive.

41:37So in our case, we found we had to rewrite a system ground up for flash. And what we were doing is starting to realize that the flash is so fast that we can spread the flash between many different servers and the server can't tell what flash is inside it and what flash is in its neighbor because the connectivity between the machines is so fast. And that allows us to build a enormous flash storage area where you're not bound by, you can just click servers together filled with Flash. And then you could begin to do some very interesting things. One of the first things we did is we made all that Flash system reside in Kubernetes.

42:31So it lives inside a container. So you can begin to say, okay, my storage system is a container. So what I can do now is have that container fail over like other containers. I can install it easily like other containers. I can move it around. So all of a sudden, our storage begins to behave like any other cloud workload. And then if it's an all-flash Kubernetes workflow, then that workload, that workload could run on different types of servers. All my servers don't need to be from the same vendor. I'm not tied to hardware. I'm also not tied to a cloud. So now my storage system can move across different vendors' hardware.

43:14It can move across different clouds, but behave as one. Single namespace, single operating unit. Then it becomes really, really interesting. And one of the things we used to spend a lot of time with in hard drive systems is, what if two people edit the same file at the same time? What do we do? What if they save at the same time. What we did is we started to achieve such high speeds that the odds of that happening became like the odds of two bullets colliding. It's just so unbelievably rare that we found we didn't need to do things like queuing, blocking, blocking and non-blocking transactions.

44:02We could just let the system rip and never put any of those speed barriers because the odds of a collision are so incredibly rare. And if one were to happen out of trillions of times, we could just redo it. And we just found when we took all those barriers away out of our system, the speeds that we started to achieve were just outrageous. And, you know, we started doing native RDMA directly from within Kubernetes, which people hadn't done before. We started saying, well, if it is Kubernetes, why couldn't we run on NVIDIA's processors? Why couldn't we run inside NVIDIA's Kubernetes instances instead of other vendors.

44:52And we just ended up being able to merge our Flash into clouds, into AI workloads in just totally unique ways. So we're still figuring out what all we can do with it. Our customers are just getting this myriad file system and starting to realize what they can achieve with it. And forgive me, this is going to sound ignorant, but that's part of why I like these conversations. I understand tape from what you described. Hard drives are generally a disk with some substrate on it that's being etched and read by a laser. Is that right? Yeah, a laser magnet. But yeah, it's magnetically charging the platter.

45:45Right. What is flash physically? It's transistors, right? It's this billions and billions of transistors that are electrically charged. So instead of it being a magnetic substrate, it's a transistor that is switched on or off. I see. I see. Yeah. And so it requires, you know, that's why it requires electricity. so to make it non-volatile a lot of work had to be done to say okay how do you power it off and keep those those settings um and or you know how do you reboot the machine and not lose everything so it's really when we achieve non-volatile memory where we could store data in memory yet powered off and power back on and keep that data that we've begun storing data onto it now it's incredibly fast.

46:41But it is, you can imagine, the cost of these transistors is much higher to make than a piece of cellophane with barium ferrite sprayed on it. So it's kind of a lot of what we're doing with this is we're managing the economics of how do we allow a company to store an obscene amount of data with a budget that is not growing. We always talk about data is growing exponentially. Well, no one's budgets are. So a lot of what we do in storage is come up with techniques. How do you deal with the fact your data is growing exponentially, your pressure to analyze it is growing exponentially, but your budget's essentially flat.

47:31So we have to get more efficient as storage providers to say, okay, how do we figure out almost a just-in-time model that you get just enough really expensive storage to do your analysis and get it off of that thing as fast as you can, right? Don't store data that you look at occasionally on this incredibly expensive medium. Shove it over on the cellophane or at least put it on a hard drive. A hard drive still similar to tape has a substrate on a platter that's a lower cost piece of magnetic chemistry. Very dense, very great aerial density and still less expensive than or in some cases less expensive.

48:17I do think, though, the flash guys are starting to press down into where they're hitting price parity with certain hard drive price points. And some of our products are all flash offering and our hard drive offering are the same price. So in our DXi product, which is a backup target, so you're like, hey, I'm backing up data, hard drives are good enough. Well, not really anymore. If you're backing up data and you're really worried about your cyber resilience, you don't care so much about the backing up of data. What you care about is I better be able to restore it fast. If something bad happens to my bank, my bank needs to get up and running fast.

48:58And so what we found is people want their backup appliances to be raging fast for cyber resilience, to be like, I know that I'm going to get attacked. I know that's going to happen. What I also know is I can restore from that attack in seconds. And so what we're finding is the fastest all-flash appliances we're selling, we're selling into the backup space with our DXi product, the T10, the T20. People just want extremely fast recovery times. And we're finding that those capacity points are actually about the same price as if you bought a hard drive. Wow. And they're physically much more compact, right?

49:43Yes. They're smaller footprint, about the same price as hard drives, and you get just a huge performance boost. both in the time it takes to backup your data, but the time it takes to restore it. And this business, I would guess, is growing along with the data. I mean, it must be a massive business. The business is growing. Now, the price per terabyte that people pay goes down. So they're storing more data, but we're always pushing the prices down. So we're always in the storage industry trying to figure out how to make money when what made a lot of money yesterday doesn't make that much money today.

50:26So you've got to be getting better and better and better all the time. And the other thing we manage is the world doesn't use as much tape today as it did in other days. So we have product lines that are shrinking and then product lines that are growing. So you're always managing your portfolio where you have, you know, you know, later stage adoption technologies, brand new technologies, and certain ones are growing like this and certain ones are tailing off and you're managing that portfolio and, you know, deciding, you know, where do you put your money in your portfolio? For us, we're really putting a lot of it into these high-speed all-flash file systems.

51:08The Myriad file system is really where we're putting a ton of our R &D effort and also in our object store, you know, which we really think. Object storage is really that cheap and deep technology and file systems are for the raging high speed. And it's really those two pieces of software are really the controlling elements of a modern architecture is file-based and object-based. Wow, that was quite a ride. Is there anything I didn't touch on that you'd like listeners to know? No, we covered a lot of ground. It's always good talking to you, and thanks for having me on. Even if you think it's a bit overhyped, AI is suddenly everywhere from self-driving cars to molecular medicine to business efficiency.

51:58If it's not in your industry yet, it's coming fast. But AI needs a lot of speed and computing power. So how do you compete without costs spiraling out of control? Time to upgrade to the next generation of the cloud, Oracle Cloud Infrastructure, or OCI. OCI is a blazing fast and secure platform for your infrastructure, database, application development, plus all your AI machine learning workloads. OCI costs 50 % less for compute and 80 % less for networking. So you're saving a pile of money. Thousands of businesses have already upgraded to OCI, including MGM Resorts, Specialized Bikes, and Fireworks AI.

52:48Right now, Oracle is offering to cut your current cloud bill in half if you move to OCI. This is for new U.S. customers with minimum financial commitment. See if your company qualifies for this special offer at oracle.com slash IonAI. That's IonAI all run together, E-Y-E-O-N-A-I. So go to oracle.com slash IonAI to see if your company qualifies for this special offer.

From the publisher

This episode is sponsored by Oracle.

 

Oracle Cloud Infrastructure, or OCI is a blazing fast and secure platform for your infrastructure, database, application development, plus all your AI and machine learning workloads. OCI costs 50% less for compute and 80% less for networking. So you’re saving a pile of money. Thousands of businesses have already upgraded to OCI, including MGM Resorts, Specialized Bikes, and Fireworks AI.

 

Cut your current cloud bill in HALF if you move to OCI now:  https://oracle.com/eyeonai



In this episode of the Eye on AI podcast, Jamie Lerner, CEO of Quantum, joins Craig Smith to discuss the future of data storage, unstructured data management, and AI’s transformative role in modern workflows.

 

Jamie shares his journey leading Quantum, a company revolutionizing the storage and management of unstructured data for industries like healthcare, media, and AI research. With decades of expertise in creating innovative data solutions, Quantum is at the forefront of enabling efficient, secure, and scalable data workflows.

 

We dive into Quantum’s cutting-edge technologies, from high-speed flash storage systems like the Myriad file system to cost-effective, long-term archival solutions such as tape systems. Jamie unpacks how Quantum supports AI-powered workflows, enabling seamless data movement, metadata tagging, and policy-driven automation for unstructured data like medical imaging, genomics, and video archives.

 

Jamie also explores the critical role of data sovereignty in today’s global landscape, the growing importance of "forever archives," and how Quantum’s tools help organizations balance exponential data growth with flat budgets. He sheds light on innovations like synthetic DNA and compressed storage mediums, providing a glimpse into the future of data storage.

 

Don’t forget to like, subscribe, and hit the notification bell for more engaging discussions on AI, technology, and innovation!



Stay Updated:

Craig Smith Twitter: https://twitter.com/craigss

Eye on A.I. Twitter: https://twitter.com/EyeOn_AI



(00:00) Introduction to Jamie Lerner and Quantum

(02:21) Quantum’s Focus on Unstructured Data Storage

(05:19) Structured vs. Unstructured Data: Key Differences

(07:52) Managing Data Workflows with AI and Automation

(10:55) Quantum’s Role in Long-Term Data Archives

(13:32) Data Sovereignty and Security

(16:18) How Data is Stored and Protected Across Mediums

(19:54) Metadata in AI and Data Management

(21:29) Quantum’s Role in Building Forever Archives

(24:16) Tape Storage: Efficiency and Longevity

(29:11) Innovations in Data Storage

(34:39) Competing in the Evolving Data Storage Industry

(37:56) Innovations in Flash Storage

(40:55) Balancing Cost and Efficiency in Data Storage

(44:28) The Future of Data Storage and AI Integration

(50:07) Quantum’s Vision for the Future



More from Eye On A.I.

All 266 episodes
#230 Jamie Lerner: How Quantum Solves AI’s Need for Unstructured Data SolutionsEye On A.I. · 53 min
Listen in VO