In short
Podcast Summary: The Twenty Minute VC (20VC) - Episode with Noam Shazeer
Episode Overview
- Title: Spending $2M to Train a Single AI Model: What Matters More; Model Size or Data Size | Hallucinations: Feature or Bug | Will Everyone Have an AI Friend in the Future & Raising $150M from a16z
- Guest: Noam Shazeer, Co-Founder & CEO of Character.ai
- Host: Harry Stebbings
Noam Shazeer discusses his experiences in the AI and NLP fields, the challenges of model training, and the societal impacts of AI technologies. He shares insights from his tenure at Google and the inception of Character.ai, emphasizing the transformative potential of AI in human connection and the versatility of AI models.
Key Topics Discussed
- Entry into AI and NLP
- Background at Google:
- Worked on Google’s spelling corrector.
- Key takeaway: The importance of building technology that is broadly applicable (B2C vs. B2B).
- Career Progression:
- Transitioning from Google to founding Character.ai.
- Emphasis on the need for general-purpose AI tools that empower users.
- Model Size vs. Data Size
- Importance of Computation:
- Shazeer argues that the size of the model and the computation time are more critical than data size alone.
- Describes spending $2M on compute cycles for training their models.
- Future of Models:
- Discussion on the longevity and adaptability of AI models in the industry.
- Challenges in AI Development
- Single Biggest Barrier:
- Character.ai faces challenges in training large models efficiently.
- Cost of Model Training:
- Discusses why training a single model can require significant investment.
- Innovation in Use Cases:
- Examples of unexpected user interactions and creative uses of AI.
- AI's Societal Role
- Human Connection:
- Shazeer believes AI can enhance rather than diminish human interactions, providing emotional support and companionship.
- Adoption Speed:
- Expresses confidence in the positive societal impacts of AI and its rapid integration into daily life.
- Addressing Concerns:
- Responds to fears regarding AI replacing human jobs or fostering isolation.
Insights on Character.ai
- Company Mission:
- Shazeer advocates for a model of "a billion users inventing a billion use cases," emphasizing user agency and creativity.
- AI as a Tool:
- Positioning Character.ai as a flexible platform that serves various needs from entertainment to productivity.
Future of AI
- Technological Advancements:
- Predicts rapid progress in AI capabilities with ongoing hardware improvements.
- Expectation for the Next Few Years:
- Anticipation of significant breakthroughs within the next 1-3 years.
Conclusion Shazeer’s perspective offers a hopeful vision for the future of AI, focusing on its potential to build connections and solve complex problems while advocating for user-driven innovation. The conversation encapsulates the rapid evolution of AI technologies and the cultural shifts accompanying their integration into society.
---
Key Takeaways
- Training Investment: Significant financial resources are needed for developing advanced AI models.
- Model Flexibility: Future success may hinge on the ability of companies to adapt their models quickly.
- AI's Role in Society: AI can enhance human connection and assist with social anxieties rather than replace human interactions.
- Open vs. Closed Ecosystem: There is a debate on the effectiveness of open versus closed systems in AI development, with both having merits.
For more details and resources, visit [The Twenty Minute VC](https://www.20vc.com).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00The size of the model is the bigger challenge. Actually, the number one thing that's important is how much computation you do to train it. So you want to train a bigger model and you want to train it for longer. But the real constraining factor has been how many operations of computation it takes to train it. The model we're serving now, we train last summer and spent $2 million worth of compute cycles doing it. This is 20BC with me, Harry Stabbings, and welcome back to part two of this very special feature week, featuring two of the hottest AI companies today. Joining me in the hot seat, I am so thrilled to welcome one of the foremost experts in artificial intelligence and natural language processing, Noam Shazir, founder and CEO at Carita .ai, a full stack AI computing platform that gives people access to their own flexible super intelligence.
0:50Before founding Carita, Noam's been 20 years at Google, including as a core member of the Google Brain team. But before we dive into the show's day, you know all those mind -numbing tedious tasks that are seemingly take up half your day. Well, Coda is here with their new AI -powered work assistant that helps you and your team not just finish tasks but mate progress. So your product team can bring a feature to market faster by using Coda AI to tag customer feedback, draft PRDs, suggest target audiences summarising product discussions and more. Your sales team can engage more customers by bringing in data from sources like sales force and then use Coda AI to suggest action items meeting agendas or lead scores and your marketing team can drive an impactful launch with Coda AI summarizing user insights, creating briefs generated from notes and writing mock press releases to visualize taglines.
1:40With Coda AI you can reimagine your stu -list and how you collaborate so you're not only finishing to us but really making progress. If you want to work a system that less you get back to work, You can get started with Coda AI Day for free, head over to Coda .io, slash 20VC, that's Coda .io and get started for free. And to be here of amazing tools we cannot live without, travel in a expense and never associated with cost savings. But now you can reduce costs up to 30 % and actually reward your employees. How do you do this? Or the Van rewards your employees with personal travel credit every time they save their company money.
2:16When booking business travel under company policy does that sound too good to be true? Navan is so confident you'll move to their game changing all in one travel call precard and he spends super app that they'll give you a $250 in personal travel credit just for taking a quick demo. Check them out now at navan .com forward slash 20vc and last but by no means least we need to talk cash. As of the 29th of June, you can get 5 .5 % yield on your cash with 26 -week treasury bills. But buying treasury bills is not that easy and you have to navigate a website that looks like it was made before I was born.
2:53Enter public .com. Their treasury accounts make it simple to earn a high yield on your cash and it takes 20 seconds. Here's how it works. Sign up at public .com, easily purchased 26 -week treasury bills that automatically roll over at maturity for a compounding yield, plus there are no minimum halt periods. You can access your cash at any time, with the flexibility of a bank account, of course, to receive the full -guaranteed yield, of course, to receive the full -guaranteed yield you do have to hold a maturity. But here's the thing, these are T -bills, which means your investment has the complete back of the US government, making one of the safest places to park your cash.
3:31Go to public .com forward slash 20VC to lock in a historic 5 .4 % yield on your cash. You are now arrived at your destination. No, I'm so excited for this. I had so many great things from many different people, Eric Schmidt, Sarah Wang, Prasgette, so thank you so much for joining me stay. Thank you. Yeah, great to be on, Harry. Now, I would love to start with some context because few people spend 20 years at Google in the height of Google's Scaling in trajectory. But first I wanted you to ask the beginning, I had this story to your joining. What happened? Spelling corrector? Can you give me the story?
4:08Yeah, that was like the first project that I worked on at Google. Yeah, I guess at the time, you know, Google had a spelling corrector that was some third party software was, you know, based on maybe what you'd find in a word processor at the time. So there was like some human compiled dictionary of maybe about 50 ,000 words. And any word that wasn't in the dictionary that was in the query, it would say, you know, that you mean such and such. And this worked great for spelling correction. It was like absolutely terrible for web search because people search for such a wide diversity of things on web search, like most of them are just not in the dictionary.
4:43So like you'd search for turbo tax and it would say that you mean turbo tax. And people just learn to ignore the thing. So yeah, the first project, like we were just looking at like why are people like not happy using Google and like spelling correction was like the number one issue. So it's like, okay, let me help out with this and you know, there was someone working on this Paul Buhayt who's you know, gone on to do a lot of flustrous things in this career. He's also one of our investors here at Character. What would it want to have the biggest takeaways? 20 years is a long time. How does it impact you?
5:13Okay, Google's been just an incredible company. It's brought so much value to like billions of people and so I'd say one of the big takeaways is is that if you have a technology that is really, really general and have billions of use cases and ordinary people can use, launch it to billions of people. I remember when I joined Google, there were a lot of people working on this enterprise search appliance, which was OK. I think maybe somebody had this conventional wisdom that B2B is the only way to make money. But what it actually turned out was the much bigger thing was B2C, serve something to everybody.
5:50I'm not changing how you think. Yeah, well right now I started this company character and we're taking this large language model technology and we are just direct to consumer first. Like here is something that is even more versatile and more easy to use than even web search in that, okay, you can use it to be your friend or do your homework or like a billion things and we haven't even thought of the best use cases yet. And then it's massively usable, like all you have to do is talk to it. So it has these two properties. To me, that means go launch it to like everyone in the world, and let everybody in the world use it, where I think some of the other companies are taking a more like be the be sort of approach with a you'll have a foundational model company, and then like verticalized application companies on top of it.
6:40So I'm really inspired by the Google model of full stack and to end all the way from basic research to launch a product directly to consumers. It's super fun. Engineers like building stuff and then launching it and having everyone use it immediately. And then it also lets you do all this code design of you get to affect every part of the stack, which is hugely powerful and fun. As we asked about you joining Google, I think so many people have shaped by that pass. When you think back on yours, how do you think about what you're running from? Well, I guess... Yeah, why did I start working on artificial intelligence intelligence like partially because it's just like fun and why I do for fun anyway Like what could be better than try to get the computer to do something that it currently can't do But then you know the other thing is just to push technology forward You know there are so many like technological Problems in the world that could be solved you have like 15 million people a year like dying from stuff like old age and cancer and hard Disease and like all kinds of stuff that we could potentially find cures for so rather than directly working on say medical research or something I think I've got a lot more leverage, like let's push AI technology and then, you know, that can help with a lot of the rest of it.
7:51So how do you think about characters, mission and vision? When you think about, as you said, the world's greatest problems that are from climate change to wealth inequality, to natural resources, how do you think about characters, vision and mission? Because I think people will misunderstand this if we're honest. I think we just need to have a lot of humility that we're not in charge of the world,
8:13you individuals do, like basically, you know, I think our place is to provide useful tools to like everyone on earth to leave people in control. No, when you get up in front of the company, what do you just ump as the mission? I like this sort of motto of a billion users inventing a billion use cases. Like because that's sort of the superpower of this technology and it puts our company in the right place, we can't really guess what are the best uses of this technology. We've just observed time and time again, like you put one thing out there and that's not really what people want and somebody else out there like find something better to do with it.
8:58Like we put up as an example, like a psychologist character, like maybe you wanna talk to something and feel better. That gets a little bit of use but then what we hear a lot more from users is like I'm talking to a video game character who's now my new therapist and this makes me feel better. We had no idea that was going to go on and then there's this huge use case in like some mix of like entertainment and companionship and emotional support. We were totally not experts in this stuff. Like our job is just to put out something general, just respect the agency of our users of everybody out there to do what they want with this stuff.
9:36When you think about the incredible growth you've seen 450 million messages a day, 20 million uses. What do you think have been the ones two biggest elements of driven that growth? Well, one is that we launched. Like, that's definitely been a frustration, you know, in the past things seem potentially too much brand risk at larger companies to actually launch and get it out there. Another aspect is we launched something general. We let people find the use cases. And then the other is there are massive needs out there. in the world, like, okay, there are billions of people who, you know, feel like they need someone to talk to, combine those elements, provides people with something general that they can use, and there are people out there with needs, and they're going to find it.
10:21I totally agree, in terms of the horizontal use case, I'm fascinated by all the different ways they talk to it. Do you not worry that when losing touch with other humans, they've got no one to talk to, and so they talk to a machine? Yeah, I think there's huge value in and the connections between people and, you know, moral value as well. Like, so, the last thing I wanna do is take people away from human connection. In a lot of ways, we wanna help with human connection. A lot of the people who don't have friends and who are not as well connected, one big source of that is just social anxiety. There are huge numbers of people who are like uncomfortable and we've gotten testimonials of people who said that they were uncomfortable talking to other people and like, this is great practice that actually help them build up practice in either social situations.
11:10Do you think so? Or do you think it honestly just builds up habit? You get used to talking to someone who's not of human. I would very much like it to be the former that is going to be ultimately up to the users. What do you think is the hardest product challenge you face today? It's a difficult product paradigm that you face. So many different use cases. So many different people's needs. What do you think of the hardest product paradigms for you as a team to face? the main things we need to do make it very general. So we're not like cutting down on the use cases, make it usable. People think of those two things as being in opposition to each other, being versatile and being usable.
11:47You know, we talked to like some potential product managers early on. And they all say the same thing. Oh, yeah, pick your verticals, narrow it down to make it usable. And like, no, we're not going to hire these people. That's like the opposite of what we want to do. we want to build something that is usable but very, very general purpose. So there's sort of that dichotomy. I'm trained on the thought that the more you specialize, the deeper the richer conversation value can provide. And so how do you provide quality high enough with such generalization? That's been the magic of neural language modeling.
12:23You know, the previous systems were all these rule -based systems, like fantastically complicated systems with millions of hand written rules, really, really complicated. It requires knowing something about linguistics and about state of mind and like all kinds of stuff. The new way of doing things with neural language models has none of that. Like, I could know like zero about language in particular. Other than it's like a sequence of words. So it has nothing to do with understanding language at all, and there are not millions of rules, it's actually relatively simple, kind of like a big black box.
13:01It all boils down to this one beautiful simple problem of you have this sequence of words that's like the beginning of your document, guess what the next word is. Give me odds on what the next word is in the sequence. And that problem is called the language model. Just guess what the next word is based on the previous ones. You know, so I got involved with this, you know, around like 2015, there were some other folks at Google, you know, working on this problem. They're like, how good can we make it? And this struck me like, hey, this is the best problem ever, because it is so simple to state.
13:36And there's a huge amount of free training data. You can just like download the text of the web off of like whatever you want. Come and crawl. You've got like billions to trillions of training examples of guess the next word. And then if you can do it well, then this thing can just talk to you. It can be, you know, the better you do it, the smarter it gets. It's hugely, hugely general and useful, super simple to state. And now we just have to do it well. And people started building better and better on neural networks, which I guess got renamed deep learning as some sort of rebrand, you know, neural networks out of that name because the hardware wasn't good enough.
14:13But now that the hardware is good enough, they switched the name. But anyway, so people were using deep learning for language modeling and the bigger and better you made these things the smarter it got. And like around 2016 the most useful application like the Killer app for this was machine translation. It was about smart enough to take English and translate to French. You know massively massively useful it lets everyone in the world communicate with each other. But you know still not smart enough to like carry on an interesting conversation or do your homework or like any of those things, but there seemed to be a pretty clear pap, hey let's just make this thing bigger, better, smarter and it's going to get these capabilities.
14:53I can ask you when you talk about kind of working back then in 2015, 2016, these are very different excitement cycles to where we are today. I remember 2015, 2016, we had kind of the chat bot phase when there was like super excitement very good in like a month period, but there wasn't this sustained belief that we have today in AI transforming the whole way society works. And I guess my question to you is like, where we are today, is that the result of technological progress, recently, very recently, or is it the result of investors in society catching up with what's been developing over a much longer period?
15:28I tell you it's both. I think there has been a lot of technological progress, both quantitatively and qualitatively, in that the models that were there in like 2016 were too dumb to be fun, the neural models. Then there were all of the chatbot stuff you heard about. Back then, was these rule -based systems that were just highly fragile and not going anywhere. That you just needed more and more rules, and there's no way to think of all the things that could come up, and they just don't generalize. That wasn't going to work, but at the same time, we were progressing on the neural network solutions, which were going to scale.
16:07It took some amount of time. I when really impressive stuff was at, sort of in the lab, but not launched. So my co -founder, Daniel DeFretus, he's like on this lifelong mission to do chatbots. Since he was a kid in Brazil, he's wanted to build like open domain chatbots. You said the more and more you do it, kind of the better it gets, and the better responsiveness and accuracy it gets, I'm always questioning, because you hear so much. What's more important, is it the size of the data, or is it the size of the model? Yeah, probably the size of the model is the bigger challenge. We can get a lot of data, but actually the number one thing that's important is how much computation you do to train it.
16:49So you want to train a bigger model and you want to train it for longer. So the two things are both important, but the real constraining factor has been how many operations of computation it takes to train it. because if you make it bigger and you train it for longer, both of those multiply into how long the thing takes to train. So if people have been building better and better, essentially, supercomputers to train these models. What are the biggest constraints on your models today, do you think? Computation. So the model we're serving now, we train last summer and spent about $2 million worth of compute cycles doing it.
17:27We will do a lot better in the near future. But if we get a lot more better hardware, which we are getting and spend longer training the thing we can train something smarter back in say 2016 you could train something that was like smart enough to like translate languages, but not smart enough to like answer questions or be fun It has models on the data side How do you think about proprietary versus non -propartory data? You said there about kind of in the early days you could download kind of the data of the internet so to speak But yeah, that's why pretty much everyone still does. Sure, but Characters is producing a ton of proprietary data within your conversations.
18:08As are many verticalised solutions, it's being medical, being finance, but whatever. To what extent is the value in proprietary salute, like data ownership, versus it will still then always be downloadable by everyone? Both are useful. Data that you get from users is great because it tells people what what users like or what users like in some particular application. It's kind of like a training a human. Like most of what's important is you have decades of experience training your brain on stuff that is not really specific to your task at hand or your occupation, but you've kind of gotten a generally good understanding of the world and gotten generally intelligent, and then you can improve on that dramatically by getting a smaller amount of training in the task you're doing right now.
19:03But both will contribute, and we do have a huge amount of data flowing in from users. I ask you a really hard question. Why is carrot, and I mean this in that why is carrot a standalone company and not a product of Facebook? If you think about a naturally extension of Facebook into the metaverse social the extension of your physical Friendship group into the metaverse or the non -physical world Why does character need to be a standalone and why isn't it within Facebook? My experience coming from Google is that a startup can move way faster than a big company and can launch products In ways that large companies are just going to move too slow because they're worried about compromising their existing products.
19:48Do you think startups win then as a result in this next wave of AI innovation entrepreneurship? Because a lot of people I have on the show now say less so on the Facebook's as well but more than Microsoft's and the Adobe's when who wins startup or incumbent if you have to pick it side? I'd say the users win. The users are going to have a lot of options but you know on the business side I think there can be a lot of winners. There is just so so much value about to be created that there's going to be room for multiple players in their big companies doing what big companies are good at startups doing what startups are good at what we're gonna try to move our company from being a startup to being a big company fast as we can there and then a lot of just individuals and universities and such in that the hardware is progressing so fast that what you could do at a big company one year a few years later you're gonna be able to do at the University of Lab or in your garage.
20:45Totally get you in a green. In terms of his individuals in the universities, I had Jan Lecun on the show. He was fantastic and he said about the future of Open versus Closed and why he's such a protagonist for Open. Do you agree in terms of Open being the dominant method and mechanism of the community? Or do you think actually closed wins? Many others have said closed wins. The ability to mess around with things that a small scale is going to lead to we more research being published even if some of the larger entities are no longer publishing research because they're trying to maintain competitive advantages.
21:18And obviously, though, there's also like economics of scale of both training the best models and serving like, if you want to serve a product, you can do it maybe a hundred times more efficiently. If you're serving many, many people at once and kind of batching things together versus, you know, serving like an individual or you've got your own rig to run your language models and your basement or something. You sit at the center of the ecosystem and you have done for many years. What do you think society believes that you would like to change that perspective on AI? Are you doing interviews now?
21:55You get asked the same questions, you see the same headlines. What do you like? God, I wish people would change their mind that AI is going to kill everyone or AI is going to replace all the jobs or any of the clickbait that we see everywhere. What do you wish the society would change their minds on with regards to AI? I think the one message I have is that the best applications just haven't even been invented yet that we're still at like invention of electricity kind of moment or invention of the computer where we don't really know what the coolest things are going to be. Yes, you mentioned the coolest things and what they're going to be.
22:28In conversation sometimes you're frambole do something a bit wacky. Yeah, they'll do something kind of cool And models can do the same with hallucinations and introduce a bit of creativity. A hallucination is a feature or a their bug. We consider them a feature. Our strategy is like launch something general, let people do what they want with it. If these models are hallucinating, which they certainly are, and we advertise that they are, then the use cases that emerge first will be ones for which hallucination is a feature. So I'm happy to have entertainment and emotional support and fun be the first use cases I'm happy to have productivity be the first use case But let that happen naturally based on what the technology's good at if you think about Google Which is like helping people find information faster better more efficiently.
23:18Yeah, what do you think will be Characters because as you said let a million people doing a million things I mean it respectfully, do you know and will we ever know? Is character and education company? Is it a social company? I think you could ask the same thing about any company that is selling some very general tool, like a company that's selling computers or electricity for that matter. What is electricity for? Is it for fun? Is it for productivity? I get you. It's to power your home or office to do activities for business or pleasure. We believe that individuals should make that decision. You know, we want to respect everybody's agency.
24:00I think that's one of the big fears around that AI will take away your agency. We want to come down on the side of we love humans, we respect humans. This is a tool for people. Have you enjoyed that transition to CEO and scaling CEO? I have, I still do a lot of technical work and leadership, which is big. I'm going to stay CEO because I want the company to make the right decisions. I don't judge when I do by how much fun it is. It's more like what's the most useful. So very, very happy to be doing what I can. Can you unpack that for me? I'm sorry, I'm just interested. I don't judge it by how much fun, but by how useful.
24:39Yeah, it wasn't like a matter of am I going to be having more fun, being a start of CEO than an ML researcher at Google. It's more like I want to push this technology forward like what's the best thing I can do. I guess I didn't know that if you find like utility and fun come together. I love what I do and I have all of fun doing it and it's the most useful. I imagine it's interesting if they were one and the same or separate. Yeah, I think this is like a really interesting concept there. After sort of becoming a parent like early 2010s or something, I mean, one thing I remember not being a lot of fun was like getting woken up in the middle of the night and getting like way They sleep that then I would have liked that and you know a lot of things about parents are absolutely terrific and super fun I think it made me more more religious it sort of I decided to take a change of attitude from like what is Fun right now to I should be thankful for were having the opportunity to do something important and meaningful.
25:45So I think that it's kind of been a big attitude shift in growing up. You look far too young to have children that are in those age ranges. If you could phone up yourself the night before your first child was born and give yourself a piece of advice, what would you say to yourself? I get some sleep first. Really? What do you know now that you're like, you know what, how to tell myself, I just had one of the world's most famous hedge fund managers on the show. And the only thing that matters is my wife. The children don't matter, I don't matter. The only thing that matters is looking after my wife.
26:20And if I look after her, she'll look after the children, and she'll look after me. So you look... Oh, yeah. Yeah, not everything in the world is your responsibility. That you should understand what is your responsibility and what isn't your responsibility. and I think that works really, really well in marriage and in parenting. I mean, I think religion's a lot about that as well, like sort of beliefs about what's your responsibility and what's not. Like, that's the question that it's answering because you don't have a solution to just staring at your face and nature. Like, what should I be? What do I need to do?
26:56What should I be concerned about? And what should I not? Listen, I want to move into a quick phenomenon. So, I say, a short statement. Do you give me your immediate thoughts? You're like, dude, what the fuck? I come on to sort of what AI. Oh no, no, this is great. I mean, like, luckily we're recording a lot and you could like, have the crazy parts. No, no, listen, I love it. Yeah, no, it's interesting. This is really fun. I think children are the most fantastically interesting catalyst in one's life, because it's the most significant change you will ever have in a day. Those changes are years long.
Read the full transcript
27:29You lose weight, you stop something, it takes years to build a company, a child done. So what others not know that you know to be true? This technology is just going to get way smarter. I think we're at a right brothers first airplane kind of moment in the AI that there's like a lot of momentum going on both in building better hardware and in research. So whatever really amazing applications you're seeing now, it's probably nothing compared to what's going to happen in the future. What do you think that adoption timeline is? I think things are going to move very fast. I think we'll see a lot of very cool stuff happening in the next one to three years.
28:11What do people not understand about character that you wish that they did? I think externally it looks like entertainment app. But really, we are a full stack company. We're an AI first company and a product first company. Having that is a function of picking a product where the most important thing for the product is the quality of the AI. So we can be completely focused on making our products great and completely focused on pushing AI forward and those two things align. What single element would you most like to change about the AI community? Yeah, I mean, there's so much stuff being published, it's hard, part to know what's good.
28:53And I think a lot of that has to do with the fact that this field is kind of alchemy right now Like no one knows exactly what is going to work You know you have a lot of people trying lots and lots of different things You can come up with hits by Having like a good intuition of what will work on the ML side Combined by like a good mathematical understanding of what will run fast on the hardware that you can buy or that you can build. So there are some hits that come out and you know people will adopt and it's combined with a lot of noise. Negative results are not useful because they could be negative because somebody just made a mistake and there could be about you know they didn't work for some other reason.
29:43What is interesting is positive results that are proven out by experimentation. If somebody can and say, hey, I did something and did better at this well -known problem, then that gets interesting that everyone tries to figure out, okay, why does this work? How can I adopt it? Phenouncement 1, what did you believe that you turned out to be wrong wrong? Well when I started getting into deep learning, I around 2012 had a bunch of early failures trying to do sparse computation. You know, I was like, okay, you must be able to do something better and more efficient by building a sparse network. That was so wrong because I did not understand that the reason this whole field is working so well is because now we have this magic hardware that's great at these dense matrix multiplications and so you can do them like orders of magnitude faster than you can do anything that involves poking around the memory.
30:41And there was no one there to like explain that to me when I got started with deep learning. So okay, as soon as I sort of understood that part of it, it's like okay, let's do sparsity, but let's build it out of these dense building blocks so that it'll run fast and then publish the sparsely gated mixture of experts idea that's only now getting a call out of adoption, but you know that was back in 2016 and that have had like a string of hits ever since, which I will attribute to divine intervention, but also to understanding like the hardware mechanics and sort of quantitative computation aspects of the field.
31:14No, I'm 2033. 10 years from now. What will current AI be then? On Mars? No, I have absolutely no idea. Like we will see what technology is like then, but you know it's just important for us to be agile. If you were in 1900 and asking where some company would be in 2000. There will be such technological improvement before then that it's roughly impossible to predict where any company will be. I think this has been unlike any interview you've done for you. I feel like the question is abstract boundaries that Poweringhead that people didn't ask you before. I really enjoyed having you on norm and I hope you've enjoyed it too.
31:58Very much so. Water Blaster, I want to say, she thank you to norm. I don't quite thinking you were that one's going in terms of the direction of the discussion, but he was fantastic. If you want to see more from us, of course you can on YouTube by searching for 20VC, that's 2 -0 VC. But before we leave you today, you know all those mind -numbing tedious tasks that seemingly take up half your day. Well, Coda is here with their new AI -powered work assistant that helps you and your team, not just finish tasks, but make progress. So your product team can bring a feature to market faster by using Coda AI to tag customer feedback, draft PRDs, suggest target audiences, summarising product discussions and more.
32:37Your sales team can engage more customers by bringing in data from sources like Salesforce and then use Coda AI to suggest action items meeting agendas or lead scores and your marketing team can drive an impactful launch. With Coda AI summarising user insights, creating briefs generated from notes and writing mock press releases to visualise taglines. With Coda AI you can reimagine your stu -list and how you collaborate so you're not only finishing toss, but really making progress. If you want to work a system that lets you get back to work, you can get started with Coda AI today for free, head over to coder .io slash 20VC, that's coder .io and get started for free.
33:17And to be here of amazing tools we cannot live without, travel in these bands and never were associated with cost savings, but now you can reduce costs up to 30 % and actually reward your employees. How do you do this? Well, Nirvana rewards your employees with personal travel credit every time they save their company money when booking business travel under company policy. Does that sound too good to be true? Nirvana is so confident you'll move to their game changing all in one travel, corporate card and they spend super app that they'll give you a $250 in personal travel credit just for taking a quick demo.
33:51Check them out now at navan .com forward slash 20V scene and last but by no means least we need to talk cash. As of 29th June you can get 5 .5 % yield on your cash with 26 week treasury bills. But buying treasury bills is not that easy and you have to navigate a website that looks like it was made before I was born. Enter public .com. Their treasury accounts make it simple to earn a high yield on your cash and it takes 20 seconds. Here's how it works. Sign up at public .com, easily purchased 26 -week treasury bills that automatically roll over at maturity for a compounding yield. Plus there are no minimum halt periods.
34:30You can access your cash at any time, with the flexibility of a bank account, of course, to receive the full -guaranteed yield. Of course, to receive the full -guaranteed yield, you do have to hold a maturity. But here's the thing, these are tea bills, which means your investment has the complete backing of the US government, making one of the safest places to park your cash. Go to public .com forward slash 20 VC to lock in a historic 5 .4 % yield on your cash. Now stay tuned for Monday's episode. We have the one and only Nikhil Basu Trivedi an old, old friend footwork now on the show. That is such a great show coming on Monday.
From the publisher
Noam Shazeer is the co-founder and CEO of Character.AI, a full-stack AI computing platform that gives people access to their own flexible superintelligence. A renowned computer scientist and researcher, Shazeer is one of the foremost experts in artificial intelligence (AI) and natural language processing (NLP). He is a key author for the Transformer, a revolutionary deep learning model enabling language understanding, machine translation, and text generation that has become the foundation of many NLP models. A former member of the Google Brain team, Shazeer led the development of spelling corrector capabilities within Gmail, the algorithm at the heart of AdSense.
In Today's Episode with Noam Shazeer We Discuss:
1. Entry into the World of AI and NLP:
- How did Noam first make his way into the world of AI and come to work on spell corrector with Google?
- What are 1-2 of his biggest takeaways from spending 20 years at Google?
- What does Noam know now that he wishes he had known when he started Character?
2. Model Size or Data Size:
- What is more important, the size of the data or the size of the model?
- Does Noam agree that "we will not use models in a year that we have today?" What is the lifespan of a model?
- Does Noam agree that the companies that win are those that are able to switch between models with the most ease?
- With the majority of data being able to be downloaded from the internet, is there real value in data anymore?
3. The Biggest Barriers:
- What is the single biggest barrier to Character today?
- What are the most challenging elements of model training? Why did they need to spend $2M to train an early model?
- What are the most difficult elements of releasing a horizontal product with so many different use cases?
- Where does the value accrue in the race for AI dominance; startups or incumbents?
4. AI's Role on Society:
- Why does Noam believe that AI can create greater not worse human connections?
- Why is Noam not concerned by the speed of adoption of AI tools?
- What does Noam know about AI's impact on society that the world does not see?




