DNA's Potential to Store the World's Data

5 Jan 2024 · 21 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

a16z Podcast Episode Summary: DNA's Potential to Store the World's Data

Podcast Details

  • Title: a16z Podcast
  • Description: Discusses tech and culture trends, news, and the future, with insights from industry experts and leaders.
  • Episode Title: DNA's Potential to Store the World's Data
  • Description: Explores the efficiency of DNA as a data storage medium, its historical context, and its transformative potential in biotechnology.

---

Key Concepts

  1. DNA as a Data Storage Medium
  2. Storage Capacity:
  3. A single gram of DNA can store over 200 million gigabytes of data.
  4. The human body contains enough DNA to store 150 billion terabytes of information.
  5. Longevity: DNA can last for hundreds, possibly millions of years, making it a stable storage option.
  1. Biological vs. Traditional Computing
  2. Natural Intelligence: Nature's optimization in data storage is being leveraged for synthetic biology applications.
  3. Efficiency: DNA's storage density surpasses traditional storage technologies, making it a compelling option for data archival needs.
  1. Trends in DNA Sequencing and Synthesis
  2. Sequencing (Reading) vs. Synthesis (Writing):
  3. Sequencing has advanced more rapidly than synthesis, which remains a bottleneck in synthetic biology.
  4. Current challenges include the complexity of writing DNA accurately and efficiently.
  1. Future of DNA Data Storage
  2. Exponential Growth Potential: The idea of developing a “Moore's Law for DNA” to enhance the efficiency and cost-effectiveness of DNA synthesis.
  3. Applications: Potential applications in health, food, and materials engineering, driven by reduced costs in DNA synthesis.

---

Discussion Highlights

Historical Context

  • Richard Feynman's Lecture (1959):
  • Feynman emphasized the vast amounts of information that can be stored in tiny volumes, foreshadowing modern understandings of genetic coding.

Current Trends in Data Generation

  • Exponential Data Growth:
  • Predictions indicate that by 2025, humanity will generate 175 zettabytes of data.
  • AI's Role: The increasing convergence of computational power and data storage is critical for advancements in AI.

Technical Challenges

  • Error Rates in DNA Synthesis: High fidelity in DNA synthesis remains a challenge; efforts to improve this area continue.
  • Storage Retrieval Mechanisms: While encoding data into DNA is feasible, effective retrieval mechanisms still need development, akin to current data management systems.

Implications for the Future

  • Biological Computing: The potential for DNA to serve not just as storage but also as a biological computing medium.
  • Privacy and Security: The confidentiality of DNA data and its implications for personal data protection.

---

Key Takeaways

  • The use of DNA for data storage could revolutionize how we archive and manage information, offering an efficient, long-lasting solution.
  • Current advancements in sequencing technologies highlight a compelling future for synthetic biology, contingent on overcoming synthesis challenges.
  • The ongoing race to improve DNA synthesis may unlock significant advancements in various fields, echoing the transformative nature of software development.

---

Additional Resources

  • Articles:
  • [Save As: DNA Part 1](https://exo.substack.com/p/saving-our-story-in-dna-part-1)
  • [Save As: DNA Part 2](https://exo.substack.com/p/save-as-dna-part-2)
  • Follow a16z:
  • [Twitter](https://twitter.com/a16z)
  • [LinkedIn](https://www.linkedin.com/company/a16z)
  • [Listen on Spotify](https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX?si=3E8B3qT9TyiwAHJ7JnaKbg)
  • [Listen on Apple Podcasts](https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711)

---

Conclusion The episode provides a fascinating glimpse into the potential of DNA as a data storage medium, emphasizing the need for ongoing innovations in the field of biotechnology. As we stand on the verge of significant breakthroughs, the intersection of biology and technology promises an exciting future that could redefine data management and storage.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The Moore's Law is not a law of physics. It's a law of human determination. They will just will it into existence. The race is to be able to build the Moore's Law for DNA. That feels like one of the last big unlocks in synthetic biology. Actually, that's where you really quickly get humbled by how well nature has done it itself. A thousand years from now, humans will know how to read DNA. Biology has had a huge story of problem on itself. So it's dealt with it within the biological context, the fact that you look out in a world with trees and people and birds and so on, there's tons and tons of bits there and bites all over the place that's been there for millions of years.

0:40Life itself is literally digital. Hello A16Z podcast listeners, welcome to 2024. At the speed that many fields have been moving recently, whether it be AI or robotics or biotech, almost nothing feels impossible. And that's why we're kicking off this year with a topic that does sound outlandish, but might actually be within our field of view. That topic, DNA as data storage. Scientists have estimated that the average human body has trillions of cells with billions replaced daily. Now, that's an incredibly efficient machine, driven by the genetic code in every single one of us. And of course, that's DNA.

1:21This blueprint has also evolved and become optimized for space and longevity. You cannot see it with the naked eye, yet it can last for hundreds, maybe even millions of years. And get this, the storage capacity of a single gram of DNA is over 200 million gigabytes. The amount of DNA in your body, 150 billion terabytes, capable of storing every single movie released in the 21st century, billions of times over. or equivalent to thousands of data centers, except requiring weightless energy and lasting much longer. So it should be no surprise that humans wanted to leverage this, quote, natural intelligence built into all of us.

2:03And the ever -increasing demand for data only pushes this trend further, with some researchers estimating we may even run out of storage this decade. And I have a funny feeling that software is not done eating the world. Today, you'll learn from A16C General Partner Vijay Ponday about the fascinating world of DNA as a storage mechanism and how advances in DNA sequencing, aka reading, and synthesis, aka writing, have let us hear. As Vijay says, this next wave of biological competing will be the mother of many new exponentials to come. As a reminder, the content here is for informational purposes only.

2:42should not be taken as legal, business, tax, or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any A16Z fund. Please note that A16Z and its affiliates may also maintain investments in the company's discussed in this podcast. For more details including a link to our investments, please see A16Z .com slash Disclosures.

3:09Alright, first I want to kick off this episode with a quick story that I originally stumbled upon via an article called Save As DNA. We'll of course drop that link in the show notes and it's written by Toby, a writer for EXO, and he shares the Nobel Prize -winning physicist Richard Feynman's 1959 lecture, called There's Plenty of Room at the Bottom, an invitation to enter a new field of physics. In it, Feynman calculated that all the volumes of the encyclopedia could be written on a quote, cube of material 1 -100th of an inch wide, to which he followed, so there's plenty of room at the bottom.

3:46Remember this is in 1959. He even postulated the idea of swallowing a surgeon in the form of a swallable robot. Knowing he was well before his time, he quipped, this fact that enormous amount of information can be carried in an exceedingly small space is, of course, well known to the biologists and resolves the mystery which existed before we understood it all this clearly. Of how it could be that, in the tiniest cell, all of the information for the organization of a complex creature such as their selves can be stored. All this information, whether we have brown eyes or whether we think it all, or that in the embryo the job bone should first develop with a little hole in the side so that later a nerve can grow through it, all of this information is contained in a very tiny fraction of the cell in the form of long chain DNA molecules in which approximately 50 atoms are used for one bit of information about the cell.

4:41Now, Feynman ended his lecture with a $1 ,000 prize to the first person to take the information on the page of a book and miniaturize it to $125 ,000 of its size. So that's 25 ,000 times smaller to be read on an electron microscope. Now, it ended up taking 30 years for Tom Newman to claim this prize. So you can get a sense of just how long it took for more people to widely understand the intelligence embedded at the atomic scale. All right, now let's bring in VJ to get us up to date on where we really are in the trajectory today. I'd love to just start out by getting your take on a few trends that are emerging and you could even say colliding.

5:23So the first one is around data storage. The world is undeniably becoming more digital. As Mark likes to say, software continues to eat the world. Do we realistically have the data storage we need for this increasingly digital world? Or what are you seeing there? Well, yes, so from a compute point of view, when we talk about Moore's Law, we often talk about just from a calculation point of view. Can we compute more and more? But keep part of AI and compute today is data. And I think what a lot of people forget about is actually storage has been exponentially increasing over time. And if you think about it, your laptop now might have a terabyte, wasn't that long ago, you're happy to have 100 gigabytes and then 10 gigabytes and so on.

6:04So that exponential increase in storage, just due to technologies like originally, just technologies now SSD technologies has enabled this exponential increase in storage, which in turn actually is a key part of AI. AI is this confluence of exponential increase in compute, meeting the exponential increase in data. Exactly. And I think to your point, many people don't realize how many zeros follow the number of bites that humans as a species are not producing. One report predicted that by 2025 humanity is set to unleash 175 zeta bites. So that's 175 followed by 21 zeros. Another exponential trend, or at least vastly advancing trend, is around genomics and DNA sequencing and synthesis, so sequencing being reading, synthesis being writing DNA.

6:53Maybe one interesting thing I'd love to get your take on is people may be are familiar with the sequencing graph, where that also has exponentially declined similar to Moore's Law, but synthesis hasn't quite followed as much of an exponential trend, despite us working on both for at least a few decades. Yeah, in the genomics field, reading actually is in the end a lot easier than writing, in part due to a variety of very clever technologies on the reading side. The writing side is actually much more complicated for somewhat technical reasons. The reason why reading actually is somewhat simplified is that the way most the reading is done is you have a long key CDNA and it's chopped up into little bits and then it's read as little bits and then put back together like this massive jigsaw puzzle.

7:38That's so called shotgun sequencing, that was invented in the 90s, and that really was a huge advance in the reading. And sequencing now is the successor to those types of technologies. Writing there's no equivalent trick yet. And you can imagine maybe you write little bits and then you put them together, but putting that together in real life is a lot harder than to put the pieces together on software. And that's maybe the fundamental asymmetry. And is that changing? Are there new innovations that make you confident that maybe we'll see an inflection there? Yeah, there are many companies in the space, many researchers pushing it for a variety of applications.

8:13Obviously in biology, having DNA is the starting place for any sort of synthetic biology operation. That synthetic biology is this whole space where we can actually engineer biology, and that it starts with the writing into it. And so most synthetic biology companies are really bottlenecked by the speed and cost of writing DNA. And then when you think about like RNA vaccines or so on, like we've all took COVID vaccines, that's writing a sibling of DNA RNA. And so there's been a huge effort for doing that. The needs are great. And it bottlenecks so many key things that I think that's really encouraged a lot of people to move into the space.

8:49One application that I just thought was so fascinating was storage. So coming back to that first trend, at least to me was not intuitive to say, let's use this building block that's in all of our bodies for data storage. How did we get there? And is this really a potential solution? Yeah, it's fun for our arteries. So one reason is that super dense. Disc or any other technology is not going to be near these densities of DNA. Each bit in DNA is just not that many atoms when it comes down to it. The density actually will be really hard to beat that with other technologies at least for a while And secondly we have technologies to read it super fast and finally I don't know if you have any old storage technologies like zip -diffs or old sadadists Can you read that I mean even like usb sometimes people even have usb a To read now the nice thing about DNA is like a thousand years from now humans will know how to read DNA That I will have no doubts so that ubiquity and significance of DNA from biological point of view it will mean that actually we'll always know how to read it.

9:50And that's actually truly interesting. Oh, and I forgot that other kicker is that it can last for a thousand years. So if you wanna have something in a vault, like a copy of a great movie, like the Godfather to be there for safekeeping, you could have that in DNA. Yeah, I think those are great points. And just to add a number to this, the storage capacity of a single gram of DNA is around 250 million gigabytes, which is wild. I mean, to your point about efficiency, that is just crazy when you compare it to some of the man -made alternatives. Maybe we can compare and contrast the DNA and data storage concept with what we actually use today.

10:30Are there other things you'd call out there in terms of whether this is truly viable or any other aspect of whether the capabilities are really there yet? Yeah, so I think there may be other intermediates. So archival storage is probably the first application. But then also what's interesting is that people are coming up with more and more compute elements that you can encode in biology. So biology can do a little circuitry and so on. And so the nice thing about DNA as a storage medium is that it's very compatible with our products. And so you can imagine new types of therapeutics that actually use some degree of DNA as storage such that these elements are doing some very simple version of compute.

11:06That's very early, but I've been around playing with computers since the 70s and those things are pretty early then too and look here we are. I think things start simple but the part that gets us excited about something like biological computing is how compatible it would be with us and how far it could grow from here. Yeah and if we compare it to software and this idea of when we save something on our computer a lot of people are familiar with zeros and ones and that being encoded into bits maybe you could break down what the equivalent is when we're talking about DNA like how can we actually get biological computing with this structure?

11:43Yeah, so in DNA it's like a long molecule like a single rope and it's comprised of DNA bases and each base could be one of four possibilities. And while bits are two possibilities, the bases get four possibilities so you can code two bits with each base that way. And then the other key thing about it, and this was the huge revelation that Watson quick published is that DNA has a structure where one base will connect in with a complementary one. And that forms a double helix. And this double helix is very stable. That's the thing that can last for a very long time. But also it's error correcting because you have a redundant copy, essentially a complimentary copy in there.

12:22And so that also is very appealing. In the end, biology has had a huge storage problem. And so the fact that you look out in a world with trees and people and birds and so on, And there's tons and tons of bits there and bytes, all over the place that's been there for millions of years. And so it has all the same compute problems of how to store, how to error correct, how to read and write quickly. And so it's dealt with it within the biological context. And that's partially what also makes DNA interesting. Yeah, maybe something else that comes to mind is that storage is not just about writing, but it's about that retrieval side as well.

12:55At least if we're thinking about, I write something up in Google Docs and I want to save it and I want to be able to bring it back and share it. Can we do that with DNA? We're talking about encoding all this data into DNA. Do we really have the retrieval mechanisms to do that effectively? Or are we talking about a different kind of storage? In principle, you could do that. And what you would do is you could have a drive equivalent where you have file names, which are little bits of DNA at the end. And then if you want to retrieve that file, you'd get the complimentary part to that. And so you pull out that strand, and then you just read that strand.

13:30And so that's just even a simple example. People have come up with very clever approaches. The one thing about this is that this is much more in the hierarchy of computers where you have rammed to SSDs to tape drives. There's a much more in the tape drive side of the thing. So you could get a lot of bandwidth, but probably not very good latency. It would take some time to read all that stuff. But also, and this starts to get into James Bond like territory, but people or also realizing it's a way you can move a lot of data very discreetly. And so you can imagine like injecting somebody with something and they work across the border and there's nothing to scan for and that could easily be like terabytes of data.

14:08I hadn't considered that but one aspect of storage is security. Are there any other things that are tough of mine there for you in terms of if this were a new storage mechanism, security is such an important aspect of that when we think of software, where we're thinking of now digital biology. Are there any other implications of that? I think it's all the same things as we deal with with any sort of cyber security. And so I think people keeping things private is generally not a bad thing. And so I think actually I would flip it and say that it actually is an interesting scheme for privacy. But we're really dipping into sci -fi here.

14:42That's not something people are doing now. But that something in principle that could be very plausible. Yeah. Well, coming back from sci -fi, grounding ourselves, When we're talking about this one application of potential DNA synthesis, is this something that's already in motion or even commercially viable? Our company is on the ground actually producing this technology and also do they have customers yet? Yeah, so there's numerous companies of various stages. Some have been around for many years, some that are startups that are producing DNA and like any commodity you can actually go to the web, upload a sequence and get your DNA.

15:15So that's there. I think what the race is to do is to be able to build the Moore's Law for DNA, to build that exponential decrease in cost. And that, as you said, point out the beginning actually hasn't been there quite yet. There's been a decrease, but maybe not a true exponential decrease. And so if we can do that, that feels like one of the last big unlocks in the synthetic biology that we've got the read and actually even got the edit with CRISPR. and so people are editing all the time. I really just don't have that right part. I think if we can enable that, so much instead of thanking biologists ready to go.

15:50Quick interjection, just in case you're looking for a real -world example of how all this can be applied. So spider silk has long been known for its strength. In fact, it's five times as strong at the same weight of steel, but it turns out that you can't just grow spider silk. If you try to farm spiders, unlike silk worms, they will actually all just kill each other. But a paper published in September showed how scientists using CRISPR were able to genetically engineer cell forms with spider genes that not only didn't kill one another, but were able to produce fibers six times as tough as Kepler.

16:23And of course, we're just getting started here. Now back to Vijay Deschadlight on what's between us and these potential applications across health, food, materials, and more. What do you think is the biggest bottleneck if you could point to anything? For now, this has largely been a sort of a chemical problem. And so people have been handling it with various chemistry methods. There's different types of ways people synthesize proteins, which are analogous long chain molecules and people have tried to extend that to DNA. And those things work well, but you can imagine you have to make this thing perfect.

16:54And so as it gets longer, it gets exponentially harder to have higher fidelity. And so that's why people typically have been selling really short ones. And then maybe you can try to combine the short ones. but it is sort of a different type of exponential problem that it's hard to do it really well without errors. And actually, that's where you really quickly get humbled by how well nature has done it itself, that it solves this problem in its own ways. And so that's still something that I think to do it at scale with very low error is the holy grail. Yeah, I mean, the more I research this topic, the more I just appreciate it.

17:26Oh my gosh, this is so efficient. Our body's storage mechanism is just incredible or nature, really, for that matter. trying to stay away from sci -fi again, but I'd love to get your take on just where we go from here. I know it's impossible to make a true prediction, but just given all the things you've seen, what kind of timeline do you think is hopeful? It's impossible to really know, but there are actually now many companies around the hoop that are doing exciting things, and there's a great need for it. So the combination of the market maturing and sythag biology and companies pursuing lots of different approaches.

17:59It has the right elements of what we want to see in terms of new type of tech company. But, you know, this is everything where that advance has to really be there. But what's been really unique about sequencing is that there's been advance after advance. Not unlike what makes computer chips work as Moore's Law is not a, not like a law of nature, law of physics. It's a law of human determination where people have been just trying and they'll do this lithography of that lithography. And there's an entire that protect the transistor and they would just will it into existence. We've seen the analogous willing into existence in the sequencing part.

18:33On top of the platform that was compatible for that, we've yet to sort of get started. And I think if we can build a platform such that we can enable that human willing to existence sort of phenomenon, that's really what's been missing. Right, and maybe to get listeners excited about that potential unlock. You touched on some of these earlier, but if we are able to generate low -cost at -scale synthesis of DNA. What is that unlock, whether it's materials or food or health? What are the applications that maybe excite you most? I think the most foundational statement is that it really unlocks large -scale engineering of biology.

19:13And so that shift from biology as let's tweak and experiment and just discover two, let's build things. And it's that building part that really the DNA part is central for because once we get the DNA we can actually Now CRISPR edit it into any type of system and then the key part is actually not just building it But then building it quickly so we can have fast iterations I think what's really great about programming is that you can compile and run your code like in minutes or seconds and Get those fast iterations once DNA synthesis can get to that point then we'll see faster relations since the technology will see that engineering cycle kick in and it will be the mother of many new exponentials to come.

19:54We really need that platform the same way we saw software. Yeah, I think the last thing I'll add is that you started by mentioning how a lot of our life is becoming digital. Irony is that many ways it always has been. That life itself is literally digital. That we may be coming full circle to adopting these technologies for new advances in engineering biology. but the super fun thing is what we're talking about is when CSM Bioconferencing. Yeah, I think that's a wonderful place to end off. Thanks, Vijay. Fantastic. All right, there you have it, DNA as data storage. Yet another example of science fiction actually may be just being science reality.

20:31Now clearly there are still hurdles along the way, but hopefully this episode got you amped about the possibility to come and also maybe an appreciation for just how efficient and our own bodies are. And by the way, if people want to hear more from VJ, you can listen as he hosts our sister podcast, Raising Health. Now, Raising Health was previously called by our reads world, and they actually just relaunched. So again, if you want to hear more from VJ and the wonderful guests on Raising Health, make sure to go check out that feed. All right, we'll see you next time.

From the publisher

Nature’s blueprint – DNA – is an incredibly efficient machine. You cannot see it with the naked eye, yet it can last for hundreds, maybe even millions of years. Plus, the storage capacity of a single gram of DNA is over 200 million gigabytes! As the cost of DNA sequencing (reading) and synthesis (writing) comes down, scientists are looking to our very own biology for applications reaching as far as data storage. Learn more about this fascinating world with a16z General Partner Vijay Pande, as he says, this next wave of biological computing will “be the mother of many new exponentials to come.”

 

Resources: 

Save As: DNA Part 1: https://exo.substack.com/p/saving-our-story-in-dna-part-1

Save As: DNA Part 1: https://exo.substack.com/p/save-as-dna-part-2

 

Stay Updated: 

Find a16z on Twitter: https://twitter.com/a16z

Find a16z on LinkedIn: https://www.linkedin.com/company/a16z

Subscribe on your favorite podcast app: https://a16z.simplecast.com/

Follow our host: https://twitter.com/stephsmithio

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.

Stay Updated:

Find a16z on X

Find a16z on LinkedIn

Listen to the a16z Podcast on Spotify

Listen to the a16z Podcast on Apple Podcasts

Follow our host: https://twitter.com/eriktorenberg

 

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

More from The a16z Show

All 489 episodes
DNA's Potential to Store the World's DataThe a16z Show · 21 min
Listen in VO