In short
Podcast Summary: Eye On A.I. - Episode #171: Anna Marie Wagner: Harnessing AI for Synthetic Biology
Episode Overview In this episode of *Eye On A.I.*, host Craig S. Smith interviews Anna Marie Wagner, the Senior Vice President and Head of AI at Ginkgo Bioworks. The discussion focuses on the intersection of artificial intelligence and synthetic biology and explores groundbreaking projects in the domain, such as engineering self-fertilizing corn and developing RNA therapeutics.
Key Topics and Discussions
- Introduction to Anna Marie Wagner
- Anna Marie has been with Ginkgo for about five years, overseeing corporate development and recently transitioning to lead AI initiatives.
- Her role includes negotiating collaborations, notably with Google, and working closely with scientists applying AI across various disciplines.
- Overview of Ginkgo Bioworks
- Founded 15 years ago by five co-founders from MIT, Ginkgo aims to revolutionize biotechnology by treating biology as an engineering discipline.
- The company focuses on creating a genomic library that can be used to engineer custom proteins for various industries.
- Applications of Ginkgo's Technology
- Ginkgo caters to various sectors, including:
- Biopharma: Collaborating with major pharmaceutical companies like Pfizer and Merck.
- Agriculture: Projects include nitrogen-fixing microbes for crops in partnership with Bayer.
- Specialty Chemicals: Engineering products for companies focused on sustainability and efficiency.
- The Role of AI in Biotechnology
- AI is employed in Ginkgo’s operations for tasks like protein engineering, DNA design, and optimizing biological processes.
- The company collects data from both successful and failed experiments, which is critical for training AI models.
- Specific Projects
- Corn Nitrogen Fixation Project with Bayer: Aiming to engineer microbes that can live inside corn and fix nitrogen autonomously to reduce fertilizer use.
- RNA Therapeutics: Exploring new methods for drug discovery and development using AI-guided approaches.
- Biosecurity and Ethical Considerations
- Discusses the importance of biosafety in genetic engineering, especially in agriculture.
- Highlights the need for infrastructure to evaluate and respond to potential biological threats.
- Innovations in Material Science
- Ginkgo is exploring biodegradable materials and alternatives to traditional plastics made from fossil fuels, emphasizing the potential of biology to create sustainable solutions.
- Collaboration with Google
- Ginkgo has a nearly $300 million cloud partnership with Google, focusing on building AI foundation models specific to biopharma.
- The collaboration seeks to leverage Ginkgo’s vast genomic data to enhance AI applications in biotechnology.
Key Takeaways
- Synthetic Biology Potential: The podcast illustrates the transformative potential of combining AI with synthetic biology to solve complex global challenges, such as food sustainability and healthcare.
- Importance of Data: Ginkgo's unique position in the biotech ecosystem is attributed to its extensive genomic data collection, which is pivotal for training robust AI models.
- Collaboration Over Competition: Anna Marie emphasizes the importance of collaboration within the biotech community to harness the full potential of AI and drive innovation forward.
- Future of Biotech and AI: The episode portrays an optimistic view of the future intersection of biotechnology and AI, suggesting that ongoing advancements will fundamentally change various industries.
Conclusion This episode of *Eye On A.I.* provides valuable insights into how AI is revolutionizing synthetic biology and addresses the potential benefits and ethical considerations that come with these advancements. Anna Marie Wagner's perspective on Ginkgo Bioworks showcases how technology can be harnessed to address pressing global issues while fostering innovation and collaboration in the biotech sector.
For more insights and future episodes, listeners are encouraged to engage with the Eye on A.I. community and stay updated on the latest developments.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00So customer comes to us, they've got a spec, and again, this is across all different industries. So it's as varied as I want my corn to fertilize itself, you know, and I'm interested in engineering a microbe that will live inside the corn and fix nitrogen out of the air. And so I can reduce my nitrogen addition by 50 percent while not harming plant yield. Like that could be the spec or the spec could be, again, something as specific as I said, catalyzing a particular reaction. Or it could be I want to discover a new RNA therapeutic. Hi, I'm Craig Smith, and this is Eye on AI. In this episode, we explore the dynamic world of synthetic biology with Anna Marie Wagner, head of AI at Ginkgo Bioworks.
0:48Anna Marie shares how their platform leverages AI to push the boundaries of biotechnology, touching on projects like self-fertilizing corn and the quest for biodegradable materials. We delve into Ginkgo's collaboration with Google, the nuances of building foundation models for biological data, and the broader implications for biosecurity and sustainability. The conversation provides a window into a future where biology and machine learning converge to innovate and heal. I hope you find the conversation as fascinating as I did. Introduce yourself, Anna Marie. tell us a little bit about who you are and where you came from and what you do.
1:37Sure. Thanks for having me on. So I have been at Ginkgo for about five years. I wear a number of different hats. I am our head of corporate development. And so in that role, I oversee M &A, licensing strategic partnerships. And that was the role I joined to lead. But as part of that, I negotiated our collaboration with Google, which I'm sure we'll talk about later, and started working very closely with all of our scientists across the business that are deploying AI across many different disciplines, and got candidly just really excited about everything that we were doing in AI and the potential for AI in our field, and so took over a new role about five months ago or so as our head of AI, overseeing our AI strategy and commercialization efforts, so really working with our customers.
2:24yeah and so tell us uh about ginkgo bioworks and and what what you guys are focused on um i mean that's it's there's so much happening in in biotech and ai yeah absolutely yes it might it might be helpful if i give you just a little bit of the history not just of the company but of the backgrounds of our founders because it does help elucidate some of the philosophy by which we've approached the business. So Ginkgo itself was started about 15 years ago, five co-founders, all of whom are still with the business now 15 years later, but they've actually been working with each other for over 20 years.
3:02So they all met at MIT. Four of them were doing PhD programs in the newly created sort of bioengineering field or synthetic biology field. And the fifth was a professor. His name is Tom Knight. He had been at MIT for 40 odd years, originally in the electrical engineering department. So he built one of the early semiconductors, worked on LISP, built the first real-time debugger. Eventually, as semiconductors came onto the scene, he taught the semiconductor design class at MIT. And it was around the early 90s he got interested in biology for a couple of reasons. One is as a substrate for nanoscale precision, which as a semiconductor designer was something he thought a lot about as he thought about the future of Moore's law but also because biology is a programmable physical discipline and you know so it runs on code ACTG not zeros and ones but we could read that code right with DNA sequencing we can write that code with DNA synthesis and he's like we should be able to program this stuff And so he's approaching the field like, well, what would it mean to program the physical world?
4:18That's pretty neat. And so to his great credit, he's in his mid-40s. He starts taking undergrad biology classes. He builds a wet lab in the computer science department, which, as you can imagine, the rest of the computer science department was really thrilled about. And he started assembling all of these engineers from different disciplines, you know, mechanical engineering, chemical engineering, computer science, electrical engineering, obviously, to start thinking about what would it mean to make biology an engineering discipline rather than just this artistic field of discovering nature. It's a very different way of thinking about the field.
4:56And there were a few observations relatively early on. One was, you know, something we needed to build, first of all, was like the layers of abstraction, right? So to be a programmer of biology, the first thing you were told to do is pick up a pipette. It was sort of like being a programmer of computers in the 60s. When to program a computer, you needed to understand the electronics, not just programming languages didn't exist, right? And so we are still living in that era of programming biology, where to be a programmer, you had to understand the fundamental mechanics of the machine. And so he's like, okay, how do we build those layers of abstraction?
5:36That was kind of one key question. Another key observation was because this is a physical discipline, physical disciplines tend to benefit from scale economics. And so how do you have scale economics and biology? You're going to invest in the sorts of things that drove scale economics and things like industrial manufacturing, right? So automation, miniaturization, things like that. And that led to, I think, business model decisions ultimately, you know, 15 years later when Ginkgo got started around being a platform that was going to serve many different areas of biological design rather than focusing on a single product, which wouldn't allow us to make a the kinds of investments at scale that we knew were going to be needed in order to actually drop down the cost of doing this engineering work sufficiently to make it scalable.
6:27And then the third piece and final piece that he really noticed or question he asked at the time was, well, where are the code libraries? I'm going to program some biological code. Where's the library? Like we have this stuff in computer science. Where is it in biology? Didn't exist. All of the code that had been created was happening in silos in small biotech companies that were making their individual drug. And they're religiously protective of it because the only thing that creates value in biotech is IP. So everyone is protecting their IP. And so there was a ton of recreating the wheel that was happening in the biotech field.
7:01And so that, again, led Ginkgo to make a very different kind of business model decision, which is that when we create IP, we want that IP to be shared broadly. We know that if we engineer a protein that's useful for one domain, it's very likely useful in many, many other domains. And we've seen that time and time again in the history of Ginkgo. But that was quite countercultural in the field. And so I just want to give that history because it helps explain what Ginkgo is today, which is a horizontal platform with a very large, 300 ,000 square feet or so of infrastructure, you know, largely automated labs that is acting as a platform to enable R &D to happen more efficiently, faster, ideally higher probability of success for customers across many, many different industries.
7:50And then the business model alongside that is we get paid sort of like AWS, you know, service fees for doing the work on the platform. And then more like an Apple App Store or something, we're getting paid on success. So if the customer's program actually works and they sell product, we get a royalty on that. And what kinds of products? Because is this drug discovery? Is this, you know, engine, gene editing of plants? Is it, yeah, sort of give me a scope of what you guys covered. Yes. So yes, all of the above. So sort of counterintuitively, we got our start outside of pharma, even though it's obvious to everybody that pharma is the biggest market for biotech.
8:33Had we walked into Pfizer 15 years ago and, you know, we were like five people and some robots, you know, they would have laughed us out of the room. They were way better at doing biotech R &D than we were 15 years ago. Right. And so we started in areas like flavors and fragrances, you know, clearly a biological product. Right. These are specialty compounds that are extracted from flowers and fruits and things like that. But the companies that were making these things had no real biotech expertise. They had supply chain problems, they had cost of goods problems, and they didn't really know how to solve those.
9:07And so it was a nice interface between our capability set and our partner's capability set. And so that was where we had our first product. So things like companies like Roburte and Givadon, you know, very large European flavors and fragrances companies where we got our start. And then, you know, sort of slowly we marched up the credibility curve where we then started working in more and more sophisticated end markets because we had the data to show. all right, we successfully delivered on products in this area, programs in this area. Let's go tackle the next area. So specialty chemicals, food, agriculture, ultimately pharma.
9:42Today, our business is split about a third, a third, a third between biopharma. So customers like Pfizer, Merck, Nova Nordisk, Biogen, all customers of ours across many different modalities in biopharma and everything from discovery to manufacturing. And then, you know, in the chemicals industry, We work with Sumitomo. We work with Solve, some of the big chemicals companies. And then in agriculture, we work with Bayer, Syngenta, and Corteva, three of the biggest ag companies in the world, both on things like traits, but also on things like microbes that live in the soil and help fertilize your plants for you in the soil.
10:21So really interesting variety of work that is running on a common platform, which, again, is kind of unheard of in biotech land. Yeah, that's fascinating. And just to clarify my understanding, because this is way out of my area of knowledge, in pharma, there's a lot of talk right now about small molecules as opposed to biologics. Are you guys focused on biologics or do small molecules, is that part of the mix? We do everything. So yeah, we have small molecule programs and we have sort of complex, so everything from small molecules to large molecules to complex therapies like cell and gene therapies, we have programs in.
11:12On the small molecule side, typically where we have a particular advantage is when you're looking at what are called natural products. So these are small molecules that have otherwise been created in nature. So, for example, a lot of our antibiotics come from microbes that live in the soil because all these little bacterias are fighting each other. And so if you want to fight that bacteria, it turns out some other bacteria is probably already successfully fought it away. And so we found a lot of that chemistry by looking at biology. And what's interesting is that quite often these sort of magical molecules are quite hard to synthesize chemically.
11:48They're quite complex chemistries. And so even though they're small molecules, making them with traditional chemical synthesis pathways is hard. And so you also need biology to produce them. So that's one version. And then more on the manufacturing side, similarly, so for example, Merck is one of our customers in this space. They're producing small molecules. They already sell these drugs. Right now, they might have some really expensive or messy or inefficient chemical synthesis stuff. And so one of the project areas we work on is something called biocatalysis, which is can you do a chemical synthesis step using a protein, using an enzyme?
12:29And so we would discover and engineer an enzyme to catalyze that chemical reaction more efficiently than the current chemical synthesis pathway that they're using. So those are a couple examples of small molecule products. But the industry writ large is quite focused on biologics and cell and gene therapy as well. Yeah. Where does AI come in to the process, to your workflow? So, again, I'll give you a little bit of history. Ginkgo, historically, I'd say for the first 14 years of our 15 year history, Ginkgo was sort of accused of being this like brute force experimentalist. Like, we're just going to solve all the biological problems by throwing scale at it.
13:11And remember, biology is this, like, artistic discipline. So, you know, really good biologists sort of looked down their nose at this brute force scale that we were applying. And to be honest, the way that Ginkgo originally thought about data, and remember that those code libraries, was around the successful experiments. Like, we're going to reuse debugged code that we know works. We're going to reuse those across applications. But as it turns out, the real magic of operating at the scale that we have operated at and that we've built is that we are not just studying the success stories. We're also collecting a whole lot of data on the 99 % of experiments that did not generate success.
13:53And as you well know, when you're thinking about training a model, you want to have, yes, examples of what good looks like, but you also want to have a hell of a lot of examples about what bad looks like. And you can learn just as much from that, if not much more, just because there's candidly a lot more failures than successes in biology. And so now as we think about what is our role in AI as applied to biology, the rest of the industry really hasn't been designed or set up in a way to collect that kind of data. You know, real labeled data on how biology functions at scale. Like that is a pretty unique infrastructure that we've built.
14:34So to give you a couple of tactical examples, our protein engineering team, big users of AI, right? So we've deployed several AI models inside Ginkgo, sort of homegrown, and we, you know, we'll use anything that's out there. We're sort of ambivalent as to what the architecture is, because the thing that we really bring to bear is data, right? We have a much, much larger genome collection than is publicly available. Our proprietary metagenomic collection is about 10 times the size as the public databases. And then we have all this labeled data on how those proteins actually function in the lab against a battery of tests.
15:10And so now when we're designing a protein for a customer, we can deploy all that data into these models and fine tune them to yield better results for the customers. So protein engineering is a big area for us. DNA design is another big area for us. So, for example, trying to improve the expression of a particular gene sequence by engineering promoters and other kind of non-coding regions of DNA. Those are kind of easy use cases for us today. But the reality is anywhere we're generating data, that data is useful for training increasingly complex models over time. And so you can certainly imagine this getting to a point where we're creating models about how an entire cell functions.
15:54Right now we're starting with the building blocks, proteins, DNA, RNA. Then you get into things like systems biology and pathways and things like that. Then you get into broader cellular function. Then you could imagine trying to predict how entire ecosystems behave together. There's an immense amount of complexity that is beyond the scope of what anyone can really do right now. But even the advances we're seeing just on those building blocks are quite remarkable. And so the AI that you're working with is primarily a search function. You're looking for attributes in a database of molecules, or are you using it to generate new molecules that have certain properties?
16:44Yeah, so it'll be both. So our projects, so again, let's take one of those enzyme engineering projects that would typically start with an ML guided search campaign of the metagenomic landscape that we have to identify interesting starting points. And then we would use those starting points, generate data about them, and then in a kind of a closed loop system, generate new designs that we think are interesting and what's important about the way that we operate, again, because we have the labs and we can actually generate this data at scale in a pretty quick way, is rather than on the first round saying, all right, we got one shot, let's make the best thing we can make, we can really design the experiments to maximize learning.
17:30And typically, we're optimizing across many different parameters at the same time. So we will generatively design new proteins, right, which are going to be based on a scaffold that we've found in our databases, you know, from an existing piece of biology. And that's useful because you want to, you don't want to, like, there's a reason all that biology evolved. And usually it evolved because it functions well, it grows well it's soluble like they're going to be a bunch of features of the biology that's useful and you want to preserve um so we'll start with that emma guided search we'll then generate a library that maximizes our ability to learn across the many different parameters we care about so it might be uh activity it might be selectivity um so if you are off target you know kind of interactions um and then we would design subsequent libraries to maximize across those dimensions in parallel But those are also using generative models to create new protein designs in that case.
18:27Yeah. Do you know the company Insilico, which is a drug discovery company? They're using AI to generate or discover new molecules, small molecules. But they also have an automated wet lab in which they can synthesize these molecules or at least samples. And then they send it out to a larger lab for production. Is that similar to what you guys are doing other than the fact that they're focused on developing drugs? And you guys are really a platform for anybody to use? So I think there's a lot of really great research that's happening across that intersection of AI and biology. I think what's interesting about Ginkgo is that, again, we're quite ambivalent as to what architectures are going to end up working best.
19:27And we tend to be real beneficiaries of that kind of innovation that's happening elsewhere in the industry. And interestingly, I honestly see more opportunities to partner than to compete, right? And so we can, in many ways, operate in a couple different venues with some of these AI-focused companies. One is to help their platform reach more customers. We just integrate the model. Today, our protein engineering team uses, I don't know, seven or eight different models in all of our campaigns. So we'll integrate different models into our work, and we'll lean on the ones that then start yielding the best results.
20:06And we're obviously building our own internal models, but there's benefit to having that diversity. And then the second is, you know, as you've seen in the large language model space with human language specifically, there are companies that have specialized in building architectures and companies that have specialized in building data. And so take, for example, OpenAI, building GPT-4, but then ScaleAI, which created a lot of training data and allowed them to do reinforcement learning so that you got ChatGPT out of it. One way to think about Ginkgo is that we're like a ScaleAI for the broader kind of biotech ecosystem as it relates to AI.
20:45We're creating that data that's going to make all of these models better. And so our goal is to get a customer result. We don't have our own pipeline or anything else. And so we're very happy to use the best of what other folks are developing. And where folks have labs, they tend to be quite narrowly focused on the questions that they're interested in for their own model or their own pipeline. And so there do tend to be still really great collaboration areas to round out the type of data that you would want, or to the point you made earlier, once you start thinking about scaling this up, making it more industrial, wanting to really manufacture this stuff, how do you do that?
21:20And can you walk me through an interesting, because I'm sure you have many case studies at hand, an interesting case study of sort of what happens with the client when they come to you guys, what's the process, and then the workflow and the work product as it is. Yeah, it is quite interesting because we're often compared to a contract research organization, a CRO, and in many ways, we are like a CRO. But traditionally, the way a customer would interact with a CRO is they would say, I need this study done. Can you please do this study for me? It's an outsourced service, but the customer knows the work that they want to get done.
22:11When customers come to Ginkgo, certainly we're providing them a lot of services, but they're not dictating, all right, I want you to run this assay on your mass specs, and I want you to run whatever, this fermentation process in your labs, etc. They're coming to us with a problem. I need that, I need to replace this chemical step with an enzyme, and it needs to cost less than this, it needs to catalyze at this efficiency, it needs to have this level of specificity, like they're giving us specs of a product. They're not giving us a research plan. We develop the research plan. And it quite often looks very different from what the customer either might have tried internally or what they would be able to do anywhere else because of the scale of the platform we've built.
22:52We're able to try a lot more breadth. And quite often what we see is you're trading a local optima for a global optima, right? When you can only afford to search a relatively narrow space, you have to make pretty safe bets. You can't afford to try off-the-wall things. When you can, you know, increase the scale you're operating at by three, four, five, six times, sorry, the order of magnitude by three, four, five, six times, then you can suddenly start to ask much more interesting questions. And we very often see that the winning design for a particular program might look absolutely nothing like their best, their best design coming into the program.
23:37And so, yeah, so customer comes to us, they've got a spec. And again, this, this is across all different industries. So it's as varied as I want my corn to fertilize itself, you know, and, and I'm, I'm interested in engineering a microbe that will live inside the corn and fix nitrogen out of the air. And so I can reduce my nitrogen addition by 50 % while not, you know, not harming plant yield. Like that could be the spec or the spec could be, again, something as specific as I said, catalyzing a particular reaction. Or it could be I want to design, I want to discover a new RNA therapeutic. I don't even know what the right design should be.
24:16I don't know if it should be circular. I don't know if it should be linear. I don't know how to deliver it. You know, you've got to work on a whole package for some of these projects. And the customer has got a product idea in mind. And then it's our scientist's job to bring together the pieces of technology that are necessary to deliver on that. Yeah. Actually, of those examples, the corn is the easiest one for a lay person to understand. Was that a real example? That's a real program we're working on with Bayer. Yeah. Yeah. Can you – so how do you start and what's the process that you go through?
24:55Yeah. So that is it. that is a real moonshot problem. And this is a like, you know, it's like a$70 billion a year industry in nitrogen fertilizer. So it's a really cool problem to work on. It's also, by the way, fun fact, nitrogen fertilizer alone is something like 5 % each of global energy consumption and global greenhouse gas production. So I mean, this is like, if you can reduce the amount of nitrogen fertilizer that is needed to feed our planet, like this is hugely, hugely, hugely valuable. Okay, so where would you start? So there's sort of two components to that. This is a drastic oversimplification, but simplistically, there are two components to this challenge.
25:36One component is how do I get the bug to live happily inside corn or wheat or rice or, you know, you name it, your cereal crop of choice. And so there you're engineering, you're trying to find bugs that live, And there already are, by the way, lots of little microbes that live inside plants, very happily. And then the second key question is how do you get that bug to fix nitrogen? Well, it turns out there are bugs in the world that do that naturally. Like you don't need to fertilize soy. You might be like crop rotation. You might remember from like middle school, we used to rotate crops. And the reason you rotate crops is some of those crops, the legumes, they fix nitrogen.
26:15They replenish the soil with all of those nutrients because they've got these little microbes that live in their root structures that fix nitrogen. They do Haberbosch, basically. And so now the question is, can I figure out the engineering circuitry of the bugs that do Haberbosch? And can I put that circuitry inside the types of bugs that are very happy living inside corn or inside rice? and so you're you're and you I suppose you could do it either way but like you could try to make those bugs be happier with corn you know they're they're obviously many different ways to to solve the type of problem but you're you're figuring out what is the what is the DNA that codes for Habermasch basically and then you're trying to figure out what is the DNA that codes for corn lovingness and you're trying to get both of those things to exist in the same bug and then you've got all sorts of things like how do you deliver it is it a seed treatment is it a soil additive you know like all sorts of other other things but um that would be the the basics of that type of a program right so so uh let's take the the which sounds simpler uh or maybe not but uh figuring out the uh the genetic mechanism for uh fixing nitrogen in the soil yeah walk me through how theoretically how you would do that yeah yeah so it's a it's a metabolic pathway right so the the microbe is eating something in the air right it's like what's in the air there's some nitrogen floating around, there's some oxygen floating around, there's some carbon floating around, right?
28:00And these bugs are ingesting those molecules. And then there are a series of proteins that are coded in the DNA, and those proteins catalyze certain chemical reactions. And so many of those chemical reactions, we sort of understand. Like, we understand generally, well, first, you know, all right, the plant eats the carbon dioxide, and then the carbon, sorry, I'm not a biologist anymore. I don't even remember these. The carbon dioxide is broken down into something else. We do understand these metabolic pathways in general. And so then we can start identifying the proteins that are responsible for catalyzing those reactions.
28:36And then it's a question of, so you generally have a decent starting point as to what the set of chemical reactions needs to be. Sometimes, by the way, you can find more efficient pathways. So like, hey, what if I took out these two steps and instead just did that one that kind of circumnavigated the chemical, chemical steps. But typically, you'll find, okay, there are, I don't know, let's say four steps involved, making this up, four steps involved in converting, you know, carbon dioxide plus, you know, whatever else into ammonia, effectively. Here are what those steps are. Here are the proteins that code for those in the microbes that live near soy.
29:18Then there's a question of like, you know, load balancing, basically, like how much of these things do I need to get expressed? And so that's where understanding things like promoters, which are like non-coding regions of DNA matters, because you need to make sure the right amounts of these proteins are getting expressed in your new bug. You need to make sure your new bug isn't getting killed because of these new pieces of DNA that you're adding or any of the intermediate molecules that are floating around in the cell now that are being catalyzed. So the hard part is not typically, let me figure out the proteins that are responsible for this.
Read the full transcript
29:51It's let me figure out how to engineer those proteins, those pathways that make those proteins and catalyze those reactions into an organism that has never had to do that kind of work before. Because you might then find, oh, well, it turns out that this intermediate compound is building up in the bug and the bug is dying. And so you're not ever making the ammonia. And so those are the, so then you might have to engineer a bug that is more resistant to having that compound built up in it, for example. We have this all the time, actually, where in the industrial chemical space, many of the products we need to make are acidic.
30:28And a lot of bugs aren't very happy living in very acidic conditions. And so then we need to actually engineer the bugs, not just to make more of the acidic thing, but to be able to survive in this broth of acid that they're now living in when they're successfully making that. And so you end up again with these really complicated multi-parameter optimizations. And many times going in, you don't even know how many parameters you're going to ultimately need to optimize over. It's why they're hard problems. And it's why like AI is so important. It's like the human mind sometimes just can't even comprehend all this kind of, the amount of data that we are generating.
31:05And to be able to use some of these tools to find bits of signal and all that noise is really quite critical. And then, so you're using AI on this tremendous amount of data to find an optimal solution. Once you found the optimal solution, then what's the next step? Or is that as, are you done with your work then? It depends a bit on the customer. So So some of our customers have their own internal manufacturing and they're quite sophisticated as it comes to scale up and further development. And so we would just send them a tech transfer package that would probably have the organism and some instructions on how to manufacture and things like that.
31:56We'd give it off to them. Some of our customers have no ability to manufacture biological products themselves. And so we would partner on their behalf with a contract manufacturer and do a similar thing. Tech transfer the organism with the manufacturing conditions. That tends to be basically the final package is an organism that produces whatever it is, you know, that the customer is interested in, along with effectively a recipe for how to manufacture it. And, you know, there's a lot of concern about genetic engineering because, you know, it's not clear what the implications might be once organisms are released into the wild and start interacting with other organisms.
32:48How do you, I mean, presumably safety is something that's part of your remit. How do you, is that again, are you using AI to explore, to sort of look forward, to explore potential outcomes as an organism interacts with other elements in the environment? How do you deal with that? Definitely a major focus, certainly in the agricultural field to that point. If you're releasing something into the wild, there is a lot of focus on can you control it. The good news to set everyone at ease is usually the problem is the opposite, right? Like your bug dies because you've just engineered it to be really metabolically inefficient, right?
33:41It's busy making a bunch of nitrogen, not busy reproducing and surviving. So usually the issue is the inverse. But yes, we obviously spend a lot of time thinking about biosafety. And specifically, we've thought a lot about biosecurity, you know, even well before the pandemic, where suddenly everyone started caring about, you know, the risks of biology. Our view was that it's whether it's nature made or man made, we are made of biology. And so there's no question that, and we eat biology and most of our materials come from biology. There's no question that biology is impactful, but there's also no question that we are very vulnerable to biology.
34:23Our food system is vulnerable to biology. Our bodies are vulnerable to biology. Our environment is vulnerable to biology. And so it is somewhat preposterous, candidly, that we do not today have a good understanding of what biology is floating around us at any given moment. Like, we have radar for the weather. We don't have radar for what pathogens are floating around my kid's school. I'm sure there are a lot, by the way. Had we had the equivalent of radar for biology, you know, four years ago, would COVID have spread as far as it did? Maybe not. Like, maybe you could have responded before people started dying in hospitals.
35:02Like we really, that should have been a wake up call for us to start wondering about this. What's happening is at the intersection of AI and biology, there's a lot of fear, right? Like is this intersection going to create a moment where suddenly it is easy to create bioweapons or, you know, create products that are so powerful that we can't control them, et cetera, et cetera, et cetera. And, you know, in general, again, I would like biology is still really hard. We can all relax a little bit. But I do think that it is not enough to rely on things like red teaming and building safer models and relying on good scientists to do good science.
35:46Like we have to build an infrastructure to help us figure out when things have gone wrong. Again, whether those things are manmade or nature created, like there's a lot of biology out in the world that humans haven't touched, most of it, in fact, that is still very, very, very dangerous to us. And so we really need to start building up an infrastructure to detect and respond to biological threats. And some of the work that we've done, like with IARPA, which is basically the DARPA for the intelligence agency, and with the Defense Department, is around this type of thing. So like, if you find a new sequence of DNA, can you tell even if it's manmade or if it was naturally occurring?
36:28Like, it actually matters to be able to answer that question, right? Can you figure out what it does? Is it something you should be worried about or is it totally harmless? Like, these are important questions that we need to be able to answer. And again, today, biology is kind of a big black box to everybody. This is cutting edge technology and very much an area where we've deployed AI. Are you working on new materials? Do any of your clients, are any of them looking at new materials? I mean, you know, plastics is an issue that people are concerned about, the proliferation of plastics in the environment.
37:11And, you know, I keep expecting these. I've talked to a lot of people in new materials labs who say that there are solutions being developed for biodegradable plastics or microbes that eat plastics and that sort of thing. but it doesn't seem to have arrived commercially. So do you do any work on that? We do, yeah. And I've got a lot of thoughts on this topic. I think, so maybe a philosophical lens or maybe just a framing lens on this. I do like asking the question, what can't biology do? And it's a very small list. It's like plastics are obviously something biology can do. They are made today with fossil fuels, like fossils.
38:11Those are biology, right? It's dead biology. Over billions of years, it turns into oil. We make petrochemicals, i.e. plastics, out of it. So it is very clearly a biological product. There are lots of companies today working on being able to replace traditional petrochemicals with biologically derived versions of those petrochemicals, and they tend to be biodegradable versions of those chemicals. of those chemistries. The issue has historically been cost. We are not pricing in the externalities of extracting oil into the price of oil. And so this has either a social solution or a economic solution, right?
38:49So, or I should say a technical solution. So an economic solution or a technical solution. The economic solution is the price. If the price of oil is higher, then a lot of technologies that are already on the market suddenly become economic. Okay. The second solution is technical, which is like today, most of these fermentation processes that are creating biomaterials are relatively basic. You've got a big 50, 100, 200 ,000 liter stainless steel tank. You got a bunch of bugs in there. You're feeding them sugar and out comes a petrochemical. Well, if sugar is more expensive than oil, you are never going to get a petrochemical, a fermented chemical that is cheaper than a synthetically derived chemical.
39:35So there's where like a technical innovation can really solve it, right? Well, what if we start feeding these bugs carbon dioxide instead of sugar? All you need to do is figure out how to get these things carbon, and you can make a lot of different products with it. It just turns out that's not an efficient process today. It's like it's a pretty hard technical challenge, but there are a lot of folks working on that one too. And so I suspect that both of these will advance kind of in parallel, where it's like, we will start more and more and more realizing the need to have alternate solutions beyond, you know, current oil-derived products.
40:09And the technology is advancing now at a clip where we will probably make, you know, increasingly make biologically derived commodities that are cheaper than their synthetic counterpart. And that will be a really interesting moment when that occurs. Yeah. And from your position in this world, you know, from the layman's point of view, things seem to be moving very fast, at least on the research end.
40:46Do you, I mean, you know, I've been paying attention to machine learning for a while and i could see that uh you know from from the advent of deep learning and then the transform algorithm that this was going to change everything uh and the people around me i you know i would tell them this you know this is it's it's it's going to be amazing and everyone kind of in my family sort of humored me, but weren't that interested. And now, you know, everybody is talking about AI. Do you feel that way in computational biology or synthetic biology or biotech or however you want to refer to it, that there's a lot of stuff happening that hasn't quite hit the public sphere yet.
41:54And the day will come soon when everyone is talking about it. I mean, counterpoint, we all know what PCR is now. We all know what RNA is. You know, we've all been gene edited for the most part at this point, right? Because we got the COVID vaccine. You know, so I do think that COVID created like a silver lining of COVID. Like it did create an understanding of both biological risk and the opportunities that come from cutting edge biological research. And so I do think it is changing. I think we're behind AI, obviously. And I would say candidly, like it is starting to change now at a fundamental level, not just a perception level.
42:41But for the past 40 years, biotech innovation has moved slower than I think any of us hoped it would have. You know, certainly has been impactful, but probably could have been a lot more impactful if the industry was organized a little bit differently, had slightly different incentives and had better tooling and stuff. and i think what's what's happening now is the intersection of a technological shift and a cultural shift in the industry which i do think is going to allow innovation to progress quite a bit faster than it has over the past you know 40 or so years yeah uh and and for uh ginkgo labs or do you call yourself a lab i guess uh we call it the phone we call our lab the foundry uh we the language from the semiconductor industry, but yes, they're wet labs.
43:34They're very large wet labs. Is that kind of platform scalable to the point that it'll accelerate progress, or is it still such a capital intensive and expertise intensive uh enterprise that that you you you you know there there will be these specialized companies like ginkgo bioworks uh how how will this scale yeah so look i i think it's i think it's both you know i think we we have certainly created a level of scalability that is unheard of, you know, historically has been unheard of in the space. The challenge is twofold. One is decoupling kind of scale of research from humans, right? And so that's sort of the first step, and we've made a lot of progress there, but most labs scale with the number of PhDs they hire.
44:37That is not scalable. They're excellent PhDs, but you're hiring people to do manual, you're hiring highly trained scientists to spend most of their time moving clear liquids around a lab. Like that is crazy. But that is what 95 plus percent of research looks like today. So Ginkgo has largely abstracted away the, again, that physical process of programming, which now our robots do from the scientific process of designing a program. The second step then is how do you scale the underlying infrastructure that's doing the programming? And so this, maybe the analogy I would give is like when you go from, you know, I don't know, vacuum, I'm not an electrical engineer, but like vacuum technologies to like microprocessors or something like, we are still using these relatively crude robotic devices to do that work.
45:29Now, that's way more scalable than a human, but it's still like a physical thing that is reasonably, you know, clunky. Many of our most advanced processes right now, we're using biology to solve the engineering process, right? So like biology, again, it operates at nanoscale. It's really quite an efficient little machine. And so we can sometimes create a biological assay that abstracts away the need to do a lot of the physical experimentation. So, for example, pooling millions and millions of designs into a single well of a plate where every single design has a barcode written in DNA. And so then instead of having to do a million different wells on plates, you have one well with a million designs.
46:17And then you can figure out which design was the one you liked by sequencing the barcode. You know, that's where you get like remarkable step changes in your throughput. Now, today, there is still a tradeoff between the kind of quality or depth of the data that you're going to get out of a pooled process like that versus the arrayed format that's more traditional. But those sorts of advances are the things that I think, again, get us the step changes in scale that are quite useful. Yeah. And what's the collaboration with Google? Yeah, so we have a nearly$300 million cloud collaboration with Google.
46:55So they are certainly our preferred cloud provider for building and training our AI models. But then what was interesting about collaborating with Google is that they really saw Ginkgo as an ecosystem enabler. So if they go make a great cloud deal with, pick your favorite biopharma company, that doesn't help them go get a deal with some other big biopharma company other than maybe it's a logo on a page that's a nice reference. You know, that's about it. And the way they looked at Ginkgo is if Ginkgo is building foundation models for AI, for AI foundation models for biopharma, and if we're successful in that, then that helps bring the rest of the industry on to Google and into using more of these computational tools.
47:42Because candidly, it's just not that big a cloud market today. And it should one day be the biggest cloud market, I would argue. And so they're funding over about$50 million of basically R &D and model building at Ginkgo for us to build foundation models with. So build the team, do the training, et cetera, which was neat. So you are building foundation models on Google infrastructure. And from scratch, are you fine-tuning models that Google provides or open source models or something like that? Yeah. So, again, our goal is to deliver the best technology to our customer. So, you know, we are students of the architectures that are developed elsewhere in academia, that are released into the open source community, that other companies have.
48:42And even today, we have brought many of those models into Ginkgo, and we have fine-tuned those models using our data, whether that's labeled data in a specific domain where we want to answer a very specific question or broader unlabeled data sets that are minimally labeled data sets that help understand a wider protein space. So I would say yes, and. So we are building foundation models for proteins, DNA, and RNA. But what we end up deploying to the customer is typically a fine-tuned model that is relevant for their particular question or particular area of focus that would be trained on the labeled data that we're generating in the labs.
49:22So both. Our theory, though, is that those fine-tuned models are a lot better if they are trained on top of a stronger foundation model. In the same way that ChatGPT got a lot better when GPT-4 came out, you know, we think those foundations should evolve and we want to be building our fine-tuned applications on top of the best foundations possible. And the theory there being data is the missing link right now in building strong AI models, and that's what we have a lot of. So we'll be students of architectures that other folks are developing and practitioners when it comes to ingesting a whole heck of a lot more data into those models.
50:01Yeah, and maybe I'm wrong, but it just, it seems that building a foundation model is an entirely different domain expertise. And that it would make more sense. I mean, you know, you've got, if you're partnering with Google, they've got DeepMind. and these different arms of their company that are building powerful foundation models. And it's unlikely that you're going to be able to build something to compete with them. What's the logic behind building your own foundation model? Yeah. So where I would where I think the distinction to draw is in when folks think about foundation models right now, they're thinking about human language foundation models.
51:04And there is, you know, so what what are the ingredients for a good foundation model? A lot of data, talent and a lot of compute. Right. And so the data is largely out there. Right. Like there's a common crawl. really. We can all get that data. And some folks have proprietary data, but that has largely been democratized in human language. Compute is a money problem. Let's just assume, you know, lots of folks can spend lots of money on compute, but some folks will decide to spend more, right? So that's one factor. And then there's talent. And I do think that that really matters. In the biological domain, you can't take the data piece for granted in the same way that a lot of folks are taking it for granted, I think, right now in human language, or where the money sort of dwarfs the data, because the real question is, do you have enough money to actually ingest as much data as is out there?
51:55And right now, there's candidly just not that much public data and biology relative to what you would really want. I do think, though, to your point, like, we're very open to collaborating on this space. Like, I would not argue that, you know, on those three axes, like, you know, compute, talent, and data. Like, we didn't invent transformers, right? Like, we've got really, really terrific application scientists who are great practitioners of AI, and will probably be a lot better at thinking through how this data needs to get ingested and things, because there's specialty knowledge on the kind of the ontology of biology, if you want to think about it that way.
52:35But we should absolutely be collaborating with leading thinkers, whether that's at DeepMind or academic labs or OpenAI or folks like that, around what types of architectures are going to be game-changing in this field. And I think we're very open to collaborating on building these models. Our view is just that you're going to need that data. And even somebody like OpenAI, right, they had to partner to get a lot of that data with scale. And so Ginkgo does, to some degree, play the role of scale AI in this ecosystem. system. And I think it does remain to be seen whether we're going to have to build our own foundation models or whether we can serve that kind of data, serve as that data partner role and collaborate with other folks on building models of that scale.
53:20Yeah. Well, that's fascinating. It is just mind bending to think about the potential of biology. Like if we actually just understood the stuff and, you know, the reality is we don't, but the reality is also we're learning really, really, really fast. And so I think the pace of change is going to be pretty mind boggling, I think. And this is a very powerful substrate to be learning about. And so it's an honor, if nothing else, honestly, to be able to work in this space. That's it for today's episode. I want to thank Anna Marie for her time. If you want to read a transcript of today's conversation, you can find one on our website, IonAI.
54:02That's E-Y-E hyphen O-N dot A-I. In the meantime, remember, the singularity may not be near, but A-I is already changing our world, so pay attention.
From the publisher
Join host Craig Smith on episode #171 of Eye on AI, for an enlightening conversation with Anna Marie Wagner, SVP, Head of AI at Ginkgo Bioworks, renowned for their work in synthetic biology and its integration with artificial intelligence.
In this episode, we explore groundbreaking projects like engineering self-fertilizing corn, pioneering RNA therapeutics, and developing sustainable materials.
Discover how Ginkgo leverages its genomic library and AI to design custom proteins, pushing scientific boundaries and paving the way for revolutionary industry applications. We'll also tackle the vital topics of biosecurity and ethical considerations in biotech, highlighting the balance between innovation and responsibility.
Anna Marie Wagner's insights provide a glimpse into the transformative potential at the crossroads of biology and technology, making this episode a must-listen for anyone intrigued by the future of scientific advancements.
Tune in to Eye on AI for this enlightening discussion, and don't forget to rate us on Apple Podcast and Spotify if you enjoy the episode.
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Preview and Introduction
(01:37) Anna 's Role at Ginkgo and Her Journey in AI and Biotech
(02:25) Ginkgo Bioworks: History, Founders, and Philosophy
(08:05) Who Can Use Ginkgo's Services?
(10:30) Small Molecules, Biologics, and the Scope of Ginkgo's Work
(12:48) The Role of AI in Biotechnology
(16:46) Enzyme Engineering and ML-Guided Searches
(19:11) Ginkgo's Unique Approach to AI and Biotech
(24:33) The Corn Nitrogen Fixation Project with Bayer
(27:20) Genetic Mechanisms for Nitrogen Fixation
(33:19) Addressing Biosafety and Biosecurity Concerns
(36:47) Exploring New Materials and Environmental Concerns
(40:27) Innovations in Biotech
(42:01) Increasing Biotech Awareness
(46:46) Collaboration with Google
(48:06) Building Foundation Models on Google's Infrastructure
(53:20) The Potential of AI in Biology




