20VC: Why Data Size Matters More Than Model Size, Why The Google Employee Was Wrong; OpenAI and Google Have the Advantage & Why Open Source is Not Going to Win with Douwe Kiela, Co-Founder @ Contextual AI

30 Jun 2023 · 42 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: The Twenty Minute VC (20VC) - Episode with Douwe Kiela

Episode Overview Title: 20VC: Why Data Size Matters More Than Model Size, Why The Google Employee Was Wrong; OpenAI and Google Have the Advantage & Why Open Source is Not Going to Win Description: Douwe Kiela, Co-Founder and CEO of Contextual AI, discusses the challenges and innovations in AI and language models following a $20M funding round. He shares insights from his career and the future of AI in enterprise applications.

Key Guests

  • Douwe Kiela: CEO of Contextual AI, adjunct professor at Stanford University, former head of research at Hugging Face.

Key Topics Discussed

  1. Founding Contextual AI
  2. Douwe’s journey into AI began with a background in philosophy, which aids his perspective in AI.
  3. Key lessons learned from working with Yann LeCun at Meta include focusing research on real-world applications.
  4. Contextual AI was founded to address the limitations of existing language models in enterprise applications.
  1. Challenges with Foundational Models
  2. Major issues include “hallucination,” attribution challenges, and data privacy concerns.
  3. The current landscape of foundational models may not yield a single dominant player.
  4. Douwe praises OpenAI's data acquisition strategies as a key to their competitive advantage.
  1. Importance of Data Size vs. Model Size
  2. Douwe argues that data size is more crucial than model size for training effective AI models.
  3. The efficiency of training smaller models on larger datasets can yield better results than larger models with less data.
  4. Proprietary data is vital, but startups can still thrive with open data models.
  1. Regulatory Landscape
  2. Douwe expresses concerns about over-regulation in Europe, which could hinder innovation in AI.
  3. The gap in understanding between regulators and tech innovators complicates the establishment of effective regulation.
  4. Elon Musk's petition to pause AI development is seen as self-serving, with a minimal chance of existential risk from AI.
  1. Business Adoption of AI
  2. Douwe notes that while enterprise adoption of AI is growing, hurdles such as compliance, data security, and efficiency remain.
  3. He anticipates a gradual adoption of AI technologies, driven by ongoing experimentation and innovation.

Key Quotes

  • "Data size matters even more than model size."
  • "It's still very early innings in AI. We haven't settled on a lot of things that need to be solved before the technology is really ready."
  • "We need to invest a lot in educating regulators about AI."

Insights on the Future of AI

  • Douwe believes the next big wave of AI will focus on enterprise applications and tailored solutions rather than one-size-fits-all models.
  • He emphasizes the need for a transparent understanding of AI capabilities and risks, advocating for an informed and balanced perspective in the industry.

Conclusion This episode provides a deep dive into the evolving landscape of AI through the eyes of Douwe Kiela, highlighting the intricate balance between data, model efficiency, and the importance of understanding the regulatory environment. As AI continues to shape the future of technology, the insights shared suggest a path forward that embraces both innovation and caution.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Data size matters even more than model size. If you train a smaller model on more data for longer, then you get a better model. There was this Google memo from an internal Google employee who had written that open AI and Google have no mode. For me as an AI researcher, when I read that memo, I was like, this person has no idea what they're talking about. Welcome back to 20VC with me, Harry Stabbins. And stay with Delphur, the Into the World of AI with an expert who's been in the industry for the last 15 years. Dal Keeler, co -found and CEO of Contextual AI, building the Contextual Language Model to power the future of businesses.

0:37Last month, Contextual closed a $20 million funding round, including Bain Capital, Sarah Gwo, Eli Gil and 20VC. He's also on a junk professor in symbolic systems at Stanford University. Previously, he was the head of research at Hugging Face and before that a research scientist at Facebook AI Research. But before we dive into the show's day, did you know that over 50 % of your day is filled with tedious tasks? What would you do if you got half a day back? Well, you can, with Coda, the all -in -one platform that changes the way your team works. Together, and Coda just introduced an AI -powered work assistant to take the busy out of work.

1:13With Coda, your team's important workflows and content already live in one place, and Coda AI helps each team member focus on the highest priority work, even as priority shift across the team. By taking over a curry, and there's be honest, very tedious work often, Coda AI empowers team members to prioritize longer -term and strategic work to accomplish goals. Coda AI not only makes space for collaboration in each person's workday, but can make it easier to stay in the loop and share informed opinions. So if you want to work assistant that lets you get back to work, you can get started with Coda AI's day for free, head over to coder .io -2 -0 -VC, that's c -o -d -a -dot -io, and get started for free, coder .io -2 -VC.

1:59And speaking of tools we cannot live without like coder, we have to talk about Brexit. The all -in -one financial stack, trusted by founders. Founders have to think globally in order to open new markets, unlock cost savings, and gain access to new talent. That's what we're having the right financial stack is more important than ever, and that's where Brexit comes in. With Bracess you get a high -limit corporate card, a high -year business account, with up to $6 million in FDIC protection and Bill Pay all build with a global first mindset. Bracess enables you to operate in more countries and cancies than any other provider, so you can pay vendors, run payroll and make international payments faster.

2:35We all know that's a must have for start -ups at any stage of growth, so are you ready to learn more about Bracess' global first solution? Visit Bracess .com forward slash 20VC. that's b -r -e -x .com slash 2 -0 -e -c, and finally Angel List is fast becoming the center of the venture ecosystem. The startup world is just buzzing about how fast they've been shipping products that meaningfully improve the lives of both startups and fund managers and their investors. For startups, Angel List reduces the friction of camp table management, bank and fund raising, all in one place. Teams can focus on scaling and let Angel List handle the rest.

3:11Thousands of startups have moved their camp tables to Angel List in the past year. Anjelist also supports large ranch funds and their teams with an automated software first approach and the best customer service in the industry. Fun managers can focus on making great deals while Anjelist handles reporting, taxes, compliance and more. So if you're ready to scale your startup or fund with the platform at the centre of it all, visit AngelList .com for slash 20VC to get started. It's a bit... You are now arrived at your destination. Now I am excited for this and we've chanted before we've known each other for a while, but thank you so much for joining me today.

3:51Yeah, thanks very much for having me on the show, I have a big fan. Oh, it's very, very kind of you. I literally paid you $25 ,000 to say that. But my question to you is, it's such a hot space, and there's very few people who've actually been in it for a while. You are one of them. How did you first make your way into the world of ML and NLP first? My journey has been a little bit unusual actually. So when I was in high school in the Netherlands, I wanted to be a cool kid during the day, but at night I was secretly fascinated by computers. So I taught myself to code. So then by the time I had to go to college and go study something, I thought I already knew everything about computer science.

4:29So I decided to study philosophy instead, radical departure from what I had been interested in at the time. but it was fascinating. I use it still every day, I think. But then at some point in my career, it became clear that I had to start making money. So I needed a real job. Then philosophy is not really a real job. And I did some logic in between foundations of math, which is also not really a real job. So I decided to study computer science after all. So I went to Cambridge, in the UK. And so that's really where I started doing NLP, natural language processing. One of my internships was done at Microsoft Research in New York with a very famous researcher called Leonbo -2, who is Young Lecunse kind of, I wouldn't say sidekick because that doesn't really do justice to what he's done.

5:11So one of the Godfathers of Deep Learning and I had the opportunity to work with him that was really an amazing time. So afterwards when Jan and Leon started fair, Facebook and I researched, I joined that out of my PhD and that really kicked off my career. I actually supposedly was part of the prep for my interview with Jan. amazing amazing person but I do want to talk about that five years that you spend at Facebook say I research team it's such a transformational team as you said with incredible individuals like Jan and Neon what are your biggest takeaways from that experience and how did it impact how you think today I'd learned so much there mostly around how to focus your research direction so I think initially I was doing all kinds of weird stuff and it took me a while to figure out that a very clear real -world application for the research that you're doing makes it much more valuable than going off on the tangent and maybe being a bit too far ahead of the rest of the field.

6:02A special place and I think actually they don't get enough credit for the impact that they've had on the world. Facebook or meta in general actually. So almost every web app in the world runs a react which is an open source project coming out of meta. When does contactual come from? What was that aha? I've got to do this. This is now the idea and the time. Yeah, so this really started at the beginning of the year. We just saw This great need after Chatchy PT have gone viral. We saw this great excitement in the world that at the same time a lot of disappointment about it not being quite ready yet for real world adaption in enterprises where you actually want to use this technology.

6:38So we decided that now really is the right time to build a company to try to tackle that and we think it's still very early innings in the game. So I think a lot of people sometimes think that the game has been played but is just getting started. Okay, so you said that about not being ready for like traditional adoption. I think a lot of the general public would say, absolutely, it's cool. What makes it not ready for general adoption do you think? So there are a couple of just really big issues. Hallucination, these models make things up with very high confidence. Attribution, we don't know why they're saying what they're saying.

7:09We can't really trace it back to anything. There's compliance issues, so we can't really remove information from them. Tricky from a GDPR perspective, for example, we can't revise information, we can't keep it up to date. There's massive data privacy issues where you have to send your very valuable Company data if you're an enterprise you have to send that to somebody else's servers These models are also quite inefficient still so you can make them much faster What we are building at contextual is a different kind of language model We're really thinking about this as the next generation of language models where we think about it from first principles for Enterprise use cases and what that means is that we want to solve all of these problems by being a bit smarter about the architecture and the architecture where specifically basing it on this retrieval augmented generation which is something that me and my colleagues at Fair came up with in 2020 and what you do there is you decouple the memory from the generative capacity of the large language model and this allows you to ground the generations from the language model in the things you've retrieved in your memory essentially.

8:13So you get much less hallucination, you get attribution for free, you can always update the memory so you can remove information on the fly, you can add things, you can revise them, you can have a stream of memory. It's much more efficient because you're compressing a lot of the compute inside the memory, and you have a very clean separation between the data plane and the model plane as they call it, which means that you can have better data privacy guarantees. I have so many questions for you. First off, you said that about an attribution of knowing where it comes from and being a bit of a black box.

8:43I had someone on the show the other day and they said, open it all closed. You still don't really know what's going on in the core foundational model layer. The open or closed is not really the point, you still don't know. Is that true? And how do we think about actual true transparency of knowing what's going on and why it's producing what it is? So we're not going to be able to know why a neural net what does what it does at the scale that neural networks operate. So this is kind of like your own brain, right? I think your behavior is relatively predictable. So that goes for every human, right?

9:13We all like a predicted shutter's actions, but I have no idea what's going on in your brain. And I will never know. There's no way I can. The only way I can kind of find out is by asking you. But if you train the architecture the way we are training it right now, then at least you make sure that the model has learned to rely on the information that it finds. And that gives you much stronger attribution than if it's just predicting the next word based on what it has seen before. because it doesn't have this ability from birth, basically, to find relevant information and ground its generation on that thing it found.

9:47Another thing that I have to ask, you mentioned the word hallucinations that I had e -mad at stability on the show, and he said, hallucinations are a feature, not a bug, which I thought was a very tweetable statement. You agree hallucinations are a feature, not a bug. It's a great quote, but as always, it's a bit more nuanced than that. I think in some cases it is a feature if you want to use a language model for creative writing and if you want it to be really really creative Then you probably want it to hallucinate. So in a way it's a spectrum of ground it and hallucination If you really care about the language model doing the writing and you want to deploy it in an enterprise critical situation Then you really don't want it to be creative.

10:25You don't want it to hallucinate. You just wanted to do what it has to do But if you want to use it for a creative writing exercise then sure you can have it to hallucinate because you're going to revise whatever it gives you anyway. You also mentioned the multi -model aspect. Another guest on the show said that the winners in Startup Land will be determined by those who can bluntly switch models faster than anyone else, and that today's models will be unusable in the year. Do you agree with them? At this particular point in time, probably yes, but just because the field is moving so incredibly quickly.

10:56And so I think that in the next year we're going to see lots of other models coming out and if you can have a language model, agnostic AI company that relies on language models then that would give you a competitive advantage. But at the same time it also runs the risk of you relying on other people's language models. It's a bit of a tricky situation. Do you think there will be many more startup language model companies? This technology is really going to change the world than every aspect of it. So there are definitely of the couple of incumbents, but they are also focused on very specific parts of the market.

11:30If you look at Anthropical Open AI, I think they're really chasing for this idea of AGI and they're relatively consumer facing if you focus less on this idea of artificial general intelligence and you want to have something a bit more like artificial specialized intelligence where you just need the model to do what it needs to do. You don't need it to know about Shakespeare or quantum mechanics or things like that. You just need it to solve your business's problem. In that case, I think there's still a lot of room for innovation and that's where we are trying to innovate. We've seen size of model matter less and less it would seem.

12:03How do you think about the importance of size of model today? And does it matter as much that you used to and will it matter even less with every year, month, day? Yeah, great question. I think Sam Aldmann at this interesting quote where he was saying that he thought models would stop growing in size. GPD4 kind of hit this ceiling. I think that's probably right, but not really because size doesn't matter It's just that data size matters even more than model size And I think the llama paper out of meta really brilliantly showed this if you train a smaller model on more data for longer Then you get a better model So you get more bang for your buck if you train it on more data rather than having more parameters But in an ideal world if you had infinite compute budget and infinite data then you would train the biggest possible model because that's the most likely to give you emergent capabilities as we call it an in the field.

12:55Does that take more time then? If you have smaller models with more data, given you need to feed more data through the model, does it not take more time than if you needed less data to get through the model? It depends, so it's a trade -off here, right? But these big models also need a lot of data, so it really is a function of the number of GPUs that you have available. Let's say you have a 1000 GPUs you can choose to train a huge model on relatively little data and it will be okay but it will be under -trained. So you have some sort of optimal point where you can train the model to perfection.

13:26In the field we were underestimating where that optimal point is and it seems that data is much more important than model size when it comes to what's optimal. If we take this to a net it's like layer deeper. Data more important than model size. What does that mean then in terms of who's Vontaged. Does that mean that startups are more advantage that she than incumbents? I don't understand who's more advantage than that case. Is it incumbents? Because they already have existing massive data modes? It depends on where the data comes from. I think incumbents definitely have an advantage there, but only some of them.

13:58And a lot of the data is just freely available on the internet. And so the Lama model was not trained on any proprietary data. It was just trained on open data on the web. And there's a lot more data to be had there. And as the society were generating a ton of data every day to add to that big pile of data. So you can really train very high quality language models just on public data on the web. But I think if you look at the secret sauce to a lot of these other models, like Y is GPT -4 so awesome. A part of that is that they went through enormous lengths to get like special data that nobody else has.

14:34So allegedly they did this whisper project where they're very good at transcribing audio because that would allow them to transcribe like all of the podcasts in the world which gives you very high quality language. If you can train on that language but nobody else has it that puts you in the position of advantage. How important is proprietary data? The main reason I would say YVC is turning down startup AI companies is because they do not have a proprietary data set to operate against and they are defined as like a thin layer of generative AI on top of a foundational model. How important is proprietary How much we dated do you think four startups in a rating in space?

15:09If you want to build a deep tech AI startup, then you really want to get a big data flywheel going. You want to start with a lot of data and then have a way to generate lots more data and that data is going to be your mode. But I think one of the interesting things about these large language models is that they're incredibly simple efficient or data efficient. So you can do cool things with them with relatively little data that just previously just wasn't impossible that unlocks all kinds of possibilities that just didn't exist even a couple of years ago. So on the one hand, yes, you need lots of data if you want to build like big AI first things, but at the same time, if you want to do a startup that builds on top of this technology, you need very little data to get started, a bit of attention.

15:53But one of the use cases I've been seeing now for GPT -4 is actually that people are using it to generate data and then they're training on that data with cheaper models. So GPT -4 might end up disrupting not knowledge workers necessarily, but it might just disrupt mechanical Turk and is just an annotator on steroids. And you can use all of that data to get much more custom models that you can then deploy very cheaply on specialized use cases. That's a quite interesting development. Pre -trained data changes a lot. Can you just help anyone who doesn't know on Sam? What is pre -trained data? How does it change the game for a lot of companies that don't have existing data most.

16:31Maybe it's useful to kind of go through the steps if you want to build your own chat GPT like what do you need. So the first thing you need is a core pre -trained model and this tends to be just trained on the web, the task you're training it on is just next word prediction. Then once you have that core model, then you want to do supervised fine tuning. So essentially you want to fix the user interface to that model because the model doesn't really know how to follow instructions for example. So you want the model to listen to you, but it has only been trained on predicting the next word So it doesn't really know how to do that.

17:03So that supervised fine tuning. That's also proprietary data You can get a much better model out of that and then the final step is RLHF Reinforcement learning from human feedback where you get this feedback loop to make the model even better for your specific use case Even if you don't have signal at the word level you just have signal at the sequence level So you can tell it like okay, that was a good response or that wasn't a good response but you can't tell it like what did you do wrong necessarily. If you go through those three steps, then you get a chat GPT. It's as easy as that. Which company do you think has the best data acquisition flywheel?

17:36When you look at them today, who do you admire and respect most? Open AI. They haven't even really trained as far as I know on the data that comes out of chat GPT going viral. And so they had chat GPT. When viral, this led to just giant, giant data mode that they haven't even really used yet. So in terms of data modes and maybe you've seen this come by actually there was this Google memo from an internal Google employee who was written that open AI and Google have no mode for me as a AI researcher when I read that memo I was like this person has no idea what they're talking about. Why are they rolling down on the stand?

18:13These places have a giant mode because as I said this is really all about data and And OpenAI has this very deep understanding of how people want to use language models. Basically nobody else has. And they have this giant economy of skill where they can serve up language models very cheaply because they get so many requests coming in at the same time. So they have a giant mode. So I'm a big fan of open source, right? I would like it to be true that with open source we could just keep up with all of that, but I think that's just incredibly naive. What do you think of the biggest challenges that they face?

18:46because I think we will dismiss Google quite significantly if I'm honest. And then Bard came out and was pretty impressive. How do you evaluate Bard and Google's display actually? Language model evaluation is a whole separate topic. It's a super interesting question. Actually, I've been fascinated by AI evaluation for a really long time. And the answer is we don't really know how to evaluate the quality of these models anymore. So what we've seen people do in the field now is they're using GPT -4 to evaluate the quality of other language models. That just feels wrong. So I think there's a giant opportunity in the market actually for a startup or several startups becoming like the Moody's or the S &P The folks who evaluate the quality of AI for specific use cases because nobody really knows It's really the Wild West out there and one of the big problems for example is data contamination where a bunch of these language models are Trained on the things that they are being evaluated on so GPD4 looks like it's an amazing coder but it might also just be trained on the data that is evaluated on, which means that it's not actually that great of a coder.

19:50How do we know a yard state for progress or measurement? What is the right way to approach AI, measurement and effectiveness? So there's the Stanford Helm project, the holistic evaluation of language models and these benchmarks where you look on static test sets, how good language models are, how good they are on this static test set. But I've been arguing for a long time that that's just completely wrong anyway and we need to do something that's much more dynamic. So ideally what you want is to see how easy is it for an adversarial person to mess with your model. The harder that is, the better your model is.

20:27And so it used to be very easy to come up with adversarial attacks where these models would just completely mess up and it's getting harder and harder over time. So success rate of an adversarial attacker, that's something that can keep evaluating over time. So we need humans to evaluate these models by trying to break them. When you say about kind of the adversarial entrance and abilities, does this mean we'll have like an entirely net -esgeneration wave of cybersecurity companies around model protection? Oh, absolutely. Yeah. That's completely going to change everything. So these models can also get contaminated with data itself, right?

21:02So we need the security layer on the generations of the models. You can do all kinds of prompt injection attacks. These models right now, they're still mostly just producing language, right? But they're starting to produce code and instructions and actions. And when that happens, then you can mess with a model to get it to produce actions that you really don't want it to produce like removing your entire database and things like that. So that definitely is something that you want to think about. My question there is, OK, but is that built by contextual? Is that built by OpenAI in terms of a data contamination checker in terms of a health checker?

21:35Or is that a next generation semantek? You name it security company external that you provision internally. There are lots of startups looking at this opportunity. Standard security companies are also looking at this right now. So it's very obviously a very big new threat surface. But I don't think that the actual foundation model builders, like open AI and contextual, are going to build that technology in -house. It's probably going to be an external audit. How much of a concern is data contamination at this stage do you think? If you're sitting in open AI's day, how do you think they discuss data contamination?

22:11I think they're aware of it. They're starting to invest more in evaluation. They probably should have done that sooner. They have an army of annotators now, human annotators who are checking their models, so they probably have a pretty good sense of how good their model actually is. But obviously they're not going to share that with the world. I had Jan McEun on the show who we discussed earlier. Yeah, obviously a very big proponent of kind of open models. Where do you sit in terms of the model that rules for the next five, 10 years? And is it different for the model that rules for the next five years versus that that rules for the next 10?

22:43So the way I think about the language model space is kind of as a pyramid. So at the top of the pyramid, we have these frontier models. So these are GPD4 and entropic models and things like that that are much better than everything else, but also much more expensive and much bigger than everything else. And then at the bottom of the pyramid you have open source models, anybody can train on them, anybody can fine tune them on their data. That's a very fruitful area for research. But I think the most interesting part is kind of the middle piece of that pyramid, where you have the most bang for your buck.

23:18So that's from a business perspective, the most interesting part where you have mid -size models that have capabilities that you don't really see at this bottom of the pyramid, that you can monetize it in various ways. It's not going to be the case that there's just one model that wins everything. It's going to be lots of models at different layers of this pyramid being used for different kinds of applications. So if you have very strong AGI requirements, you probably want to have a frontier model. If you care about it a bit less, maybe you want to have artificial specialized intelligence. If you care about it even less, then you can just take an off -the -shelf open source model.

23:52So there is always going to be a place for open source. But why I don't think that open source models will move up that pyramid to the frontier is because they're just too expensive and this whole flourishing that you see right now of open source models that basically comes from meta's generosity in giving Lama away for free and if they had that then you wouldn't see that. We saw Elon's petition, he mad at stability very much was in favor of it on the show. Where do you sit in terms of Elon's petition and how did you read it? This is the petition where he asks everybody to stop working on AI so that he can catch up.

24:25A lot of the narrative in the media right now is really driven by self -interest from a bunch of folks in the field. So the whole kind of existential risk debate, I think it actually comes from a very good place and a lot of people are worried about this and I think there is a non -zero probability of AI extinction risk. So it's something we need to think about and a lot of smart people are thinking about this like Benjo and Hinton and all of these folks. For me, she said there's a non -series charm on stuff. So there is just a non -zero, but very, very, very small chance. There will be some sort of paperclip maximizer scenario.

Read the full transcript

24:59Have you heard of this paperclip maximizer? No, I hardly know. So if you give a very intelligent system and instruction like you need to make as many paperclips as possible, then it's going to turn everything into paperclips and it's going to basically destroy the planet and turn everything into paperclips because that's its objective function. In the process of maximizing paperclips, it will destroy everything else and turn the whole universe into paperclips. You can see why I'm saying that these are very, very small probabilities. So probably the chance of me getting hit my lightning like right now is much higher than that happening.

25:34I think one of the issues I have with the whole debate around existential risk is that it's really a tiny probability, but we're pretending like it's a massive issue. And I think there are much bigger risks, like nuclear your war and pandemics and climate change. And those are things we should be focusing on much more. So the people who are pushing this narrative are really the people who are benefiting from this being the narrative. So these are the incumbent AI companies who want to have either the market regulated because then they benefit because they can deal with the regulation but small companies like mine can't or they are the folks who benefit from kind of fear mongering in the broader public where people start having a lot of respect for AI's capabilities and want to use it everywhere because AI is so smart it might even kill us all.

26:18There's a lot of dubious motives behind to seeing there. You mentioned regulation that. I think a big question for me is like I don't think the chasm has ever been greater between private company knowledge specifically around AI and then also the regulators knowledge which is significantly behind. How can effective regulation be set with such a large chasm between private sector knowledge and regulation knowledge. Yeah, we have to invest a lot in educating regulators. The AI community has been terrible at this. And the broader populist just needs to understand much better what AI is and what it can do and what it can't do.

26:54It's been slightly self -interest driven, I think, in that a lot of folks in AI have just wanted to keep the technology for themselves. And that's why they haven't really invested in educating the rest of society. That's really a huge issue. A bit of a side point there, but I think the people who tend to write the regulation, they generally don't really understand technology all that much anyway. I think the main risk, so speaking of Europeans, what Europe is going to try to do is over -regulate everything and just completely destroy innovation. I'm very worried about that. I think the US is traditionally much better at not over -regulating markets to let innovation thrive.

27:31I hope it stays that way. Nobody really benefits from over -regulation here except the incumbents and those are to big tech companies already lobbying for a regulation anyway and spreading fears of AI existential risk and things like that. I'm quite concerned by the EU's regulatory stance around AI. What they've suggested so far pretty much makes it impossible for most startups to use any models that aren't owned and operated by themselves. How do you think about what happens with the EU regulation? The European Union just has a tendency of really killing innovation because they think they can lead through regulation.

28:06And that's really the only thing that they're really good at. It's a very dangerous thing when everyone wants to be the leader. And the best way to be the leader is to be the first and the strongest and the hardest. And when you apply that to regulation, it's like, It's even worse, I think. Oh, fuck yeah. Keep going there, there won't be anything to regulate. But I wanted to want B2B adoption before we do a quick fire. We had E -Math on the show, I get as I said, and he said, hey, business is on really adopting it yet. It's going to be a tidal wave of adoption next year. So I guess the first question is, what are the fundamental blockers for businesses adopting AI today?

28:40Iman is really great at speaking in quotes, by the way. That's an under great quote. Yeah, I think the tidal wave is coming. There are just big problems that we have to overcome, and these are the things I just talked about. So hallucination, attribution, compliance, up to dateness, data privacy, latency, and I think the whole field is moving in this direction of just making everything ready for enterprise usage. This is going to happen. Do you agree with him out on the timing next year will be the year when enterprises adopt AI at scale and it will be a freaking train in his words? I think it's already happening.

29:16I was at an exact event at Google earlier this week and there were all of these Sea -level folks from all companies across the world and they were all talking about how they're using AI and Everybody's experimenting with it and it's starting to make it into production already in various places I don't think we have to wait for a year It's already happening, but there are just some big hurdles that need to be overcome and they will be overcome very quickly Do you think it will be a fast or a gradual adoption? I'm always aware that excitement is very quick, but actually any new technology cycle, it always takes a little bit longer than one thinks actually.

29:49Yeah, it will be gradual, I think, and a lot of work is right now actually going into finding the reduce cases for this technology, because people are now starting to think that they can use GPT -4 for anything. That's just not true. People are trying to experiment with the reduce cases for different types of models. I spoke to a big, big European company the other day, And they said the biggest challenge for us is we have millions and millions of lines of kind of transaction data There is no freaking way we are letting any of our transaction data go anywhere off -prem Like it has to be so secure on -prem.

30:24It is our lifeblood How do you think about security of large enterprise data and willingness for it to go into models like open AI without Security loss. Yeah, absolutely. That's really one of the big questions, right? and that's why what we're building has this very clean separation between the data plane and the model plane because then you have much more control over where the data goes. Can you just talk us through that? What does that separation mean? Yeah, so traditionally in cloud deployments you talk about the data plane and the control plane and so control really is okay, you're a startup or a company and you want to be able to deploy models to your customers VPC, virtual private cloud, but you don't want any of their data leaving their VPC because it's their data.

31:09So that's the separation between the data plane and the control plane. So in this new setting where we have language models, there's this model plane and it's kind of unclear where to put it. So you could put the model inside the customer's VPC and then you get full data privacy basically, but you have no control over what that model is doing at all. You get no feedback, you get no learning. So you want to find interesting hybrids where you can and respect data privacy keep the data playing inside the customers VPC, but put the model somewhere else. So you can do that if you have a decoupling between the retrieval part and the generative part, which is what we are building.

31:44When you look at all the VC fund rises, do you look at it and go, this is getting crazy? Not so much. I think some of the rounds were pretty big, but I think it's also justified just because this stuff is really going to change the world tonight. So one that, if it's right, has massive payoff. There's this narrative, I think, in the VC community that there are these crazy rounds happening, but I think they're happening for a good reason. So I haven't really seen any companies come by where I was like, wow, why are they getting this much money? There's a few of them where I thought they really have to live up to massive expectations now and they need to actually start making real revenue now.

32:21At some point, there's going to be a disillusionment with the technology and then funding my dry -up and then these places are really in trouble. Can I ask you, when you look at the incumbent stats, who do you think has the strongest strategy execution to date? Is it Facebook? Is it Google? Is it Microsoft? Is it Apple? I've been very impressed actually by how Microsoft has managed to turn everything around by strategically collaborating with a better AI lab in the shape of open AI. And they've just really turned that into this narrative where Microsoft is an AI leader. And they really weren't an AI leader even a few years ago.

32:57So that's been impressive. On the flip side, who do you think's not done well and not adjusted to the knee landscape? I'm still very curious to see what Apple will do. They had this great vision of having like Siri on your phone and things like that. So that's like one of the first personal assistants. So if that could be a super powerful language model, then that could do very interesting things. But so far I haven't really seen interesting things coming out of Apple. I want to do a quick fire on. I've peppered you with questions anyway. But this is like a more structured pattern with 60 seconds per one and so it's much more you know informed And the questions will actually be on schedule unlike the rest of the you know last 45 minutes.

33:33Yeah, sounds great So what do others not know that you know to be true? So I think others are underestimating how early it still is AI feels like we've made so much progress that is very hard to enter the market right now And I think that that is just not true it's still very early innings. We haven't settled on a lot of things that need to be solved before we can really have this technology be ready. It's still very early. Me realizing that gives me a competitive advantage, hopefully. What do you advise founders who are building AI companies not in the valley? Do they need to be in the valley?

34:09No, absolutely not. The valley is kind of a dangerous bubble in a way where there's this giant echo chamber happening. And I think if you look at a lot of great AI companies, they're not in the valley and they don't have to be here Maybe they should have an office here because there are great universities to recruit from and things like that But I really don't see a reason why you would have to be here. What would you most like to change about the AI community? Hi, I think there's way too much hype. It would be good if the community at least acknowledges that and tries to really pay attention to the things that matter like having technology that actually works and not just jumping on the next hype train, and there's an auto -GPT thing that is going to change the world, but doesn't actually work.

34:50So there's a lot of debate right now that is just driven by Twitter and just sound bites and quotes, like, am I? Where I think it would be better to be a bit deeper and think a bit more carefully about what we're doing. Some fantastic quotes, though, aren't they? I mean, credit work. I'm really jokingly. Okay, do you agree that some of the biggest businesses to be built in AI over the next years will be services businesses, fall large enterprises helping with AI implementation. Yeah, totally. I don't know if those are going to be new businesses or existing in combat, so the hyperscalers are also trying to play that role.

35:25There's just so much demand right now for AI in any kind of enterprise and it's still very hard to get it right, so a lot of these companies are just looking for help. And so there are just opportunities there. How do AI and philosophy help each in your day -to -day role. I'm very happy that I studied philosophy. So philosophy is really about conceptualizing anything and any arbitrary level of abstraction and that ability you can use anywhere. So for AI in particular, I think that philosophy and AI are an interesting combination because philosophy is about the stuff that you can't really do science about yet.

36:01So at some point, the things that people are philosophizing about now, they become scientific questions that just have answers or hypotheses, and then there are no longer philosophy. So natural philosophy that used to be a thing, we now call that physics and mathematics. And I think with AI, there are lots of questions that we still don't even really know how to ask yet. And philosophy is great for thinking about those kinds of questions. What's the strongest belief that you had, which turned out to be wrong? The strongest belief is that I really underestimated how important scale is in artificial intelligence.

36:35And I think this is really one of the things that OpenAI has excelled at. If you throw an order of magnitude more compute and data at AI systems, then they just become much, much better. And if you keep that scaling up, you have these scaling laws that we know about now. I really underestimated this. And for a long time, I was just saying, like, oh, yeah, look at these silly OpenAI researchers. They're just scaling things. They're not inventing new algorithms. That's not cool. And I was very, very wrong. I'm gonna apologize in advance for this one, okay? I'm sorry. I'm asking it anyway. What do you think the timeline is for super intelligence?

37:10So so despite Nick Bosterum's book, I still think that super intelligence is actually very ill -defined and in many ways We have already achieved super intelligence and so in the 50s We achieved mathematical super intelligence so computers in the 50s were already better calculating stuff than humans. I don't think that that really is a well -formed question. If you're asking about the AGI and I think AGI itself, a lot of people, this is a mistake I made where I thought AGI kind of meant artificial consciousness or something like that, which also doesn't really have a meaning. But if you look at how open AI and anthropic and these places define AGI, systems achieving capabilities that allow them to effectively do the work of humans for the majority of economically valuable human tasks, then we're not that far away.

38:00And so I think in the next like five to ten years that sort of economic displacement is likely to happen. Was the most painful lesson that you've learned that you're also pleased to have gone through? I think when I was young I maybe put ambition before people sometimes. I learned is the hard way I think where I just didn't have enough empathy for the people I worked with. And as I grew older and more mature, I realize more and more just that is really all about people and working with fantastic people and doing cool things together and changing the world together. So I'm very happy to have learned that lesson because that makes me much better at my job right now.

38:37Final one, Tania's time. If all the stars align, where's Contactual then? So if all the stars align, open AI and Anthropic and all of these places, they had this great first generation technology. So they're kind of like the lycos and altavista of search engines and the technology we have is more like patreon And that would make us the the Google of language models with the right technology at the right time and with the right Execution so that's what I would hope for now. I've absolutely loved doing this I can't thank you enough for putting up with my wayward Sometimes very naive questions. I can't thank you enough for letting me invest in you And I really appreciate the time stay my friend.

39:16Thanks for having me What a show I have to say I just love diving to AI with some of the best in the world, but if you want to see more from us of course you can on YouTube by searching for 20VC. But before we leave you today, did you know that over 50 % of your day is filled with tedious toss? What would you do if you got half a day back? Well you can with Coda, the all in one platform that changes the way your team works, together, and Coda just introduced an AI powered work assistant to take the busy out of work. With Coda, your team's important workflows and content already live in one place, and Coda AI helps each team member focus on the highest priority work, even as priorities shift across the team.

39:57By taking over a curry, and there's be honest, very tedious work often, Coda AI empowers team members to prioritize longer -term and strategic work to accomplish goals. Coda AI not only makes space for collaboration in each person's workday, but can make it easier to stay in the loop and share informed opinions. So if you want to work assistant that lets you get back to work, you can get started with Coda AI today for free, head over to coder .io -20VC, that's c -o -d -a -dot -i -o, and get started for free, coder .io -20VC. And speaking of tools we cannot live without like Coda, we have to talk about Brex, the all -in -one financial stack, trusted by founders.

40:38Founders have to think globally in order to open new markets, unlock cost savings, and gain access to new talent. That's what we're having the right financial stack is more important than ever and that's why Brexit comes in. With Brexit you get a high limit corporate card, a high yield business account with up to $6 million in FDIC protection and bill pay all billed with a global first mindset. Brexit enables you to operate in more countries and currencies than any other provider. So you can pay vendors, run payroll and make international payments faster. We all know that's a must -have for startups at any stage of growth.

41:11So are you ready to learn more about Bratis' global first solution? Visit Bratis .com forward slash 20Vc, that's B -R -E -X .com slash 20Vc. And finally, Angelist is fast becoming the center of the venture ecosystem. The startup world is just buzzing about how fast they've been shipping products that meaningfully improve the lives of both startups and fund managers and their investors. For startups, Angelist This reduces the friction of camp table management, banking and fundraising all in one place. Teams can focus on scaling and let Angelist handle the rest. Thousands of start -ups have moved their camp tables to Angelist in the past year.

41:47Angelist also supports large ranch funds and their teams with an automated software first approach and the best customer service in the industry. Fun managers can focus on making great deals while Angelist handles reporting, taxes, compliance and more. So if you're ready to scale your startup or fund with the platform with the Sandra of the world, visit angelless .com for slash 20VC to get started. Now next week we are taking a week off. Yes, I am going away with my family to the British coast so we will not have any shows next week and this will be a first in a very long time but even I need a rest sometime so I so appreciate your support and stay tuned for more episodes in 10 days.

From the publisher

Douwe Kiela is the CEO of Contextual AI, building the contextual language model to power the future of businesses. Last month Contextual closed a $20M funding round including Bain Capital, Sarah Guo, Elad Gil and 20VC. He is also an Adjunct Professor in Symbolic Systems at Stanford University. Previously, he was the Head of Research at Hugging Face, and before that a Research Scientist at Facebook AI Research.

In Today's Episode with Douwe Kiela We Discuss:

1. Founding a Foundational Model Company in 2023:

  • How did Douwe make his way into the world of AI and ML over a decade ago?
  • What are some of his biggest lessons from his time working with Yann LeCun and Meta?
  • How does Douwe's background in philosophy help him in AI today?

2. Foundational Model Providers: Challenges and Alternatives:

  • What are the biggest problems with the existing foundational data models?
  • Will there be one to rule them all? How does the landscape play out?
  • Why does Douwe believe OpenAI's data acquisition strategy has been the best?

3. Data Models: Size and Structure:

  • Why does Douwe believe it is naive to think the open approach will beat the closed approach?
  • What are the biggest downsides to the open approach?
  • Does the size of data model matter today? What matters more?
  • How important is access to proprietary data? Are VCs naive to turn down founders due to a lack of access to proprietary data?

4. Regulation and the World Around Us:

  • How does Douwe expect the regulatory landscape to play out around AI?
  • Why is Europe the worst when it comes to regulation? Will this be different this time?
  • How does Douwe analyse Elon's petition to pause the development of AI for 6 months?
  • Do founders building AI companies have to be in the valley?

More from The Twenty Minute VC (20VC): Venture Capital | Startup Funding | The Pitch

All 520 episodes
20VC: Why Data Size Matters More Than Model Size, Why The Google Employee Was Wrong; OpenAI and Google Have the Advantage & Why Open Source is Not Going to Win with Douwe Kiela, Co-Founder @ Contextual AIThe Twenty Minute VC (20VC): Venture Capital | Startup Funding | The Pitch · 42 min
Listen in VO