In short
Podcast Episode Notes: 20VC with Richard Socher
Podcast Title: The Twenty Minute VC (20VC) Episode Title: 20VC: Does Value Accrue to Incumbents or Startups in the AI Race, Why Model Size Matters More Than Data Size, Why Artificial General Intelligence is Far Away, Why Carpenters Will Be Paid More Than Software Engineers & Future of Jobs with Richard Socher Host: Harry Stebbings Guest: Richard Socher (Founder and CEO of You.com, former Chief Scientist at Salesforce)
---
Episode Overview
In this episode of The Twenty Minute VC, host Harry Stebbings interviews Richard Socher, a prominent figure in AI and natural language processing. They discuss a variety of topics surrounding artificial intelligence, including the dynamics between startups and incumbents in the AI space, the importance of model size, the future of jobs, and the potential for Artificial General Intelligence (AGI).
---
Key Themes and Discussions
- Richard Socher's Journey in AI
- Background in AI: Socher's journey began in 2003 with a focus on linguistic computer science, transitioning to deep learning and natural language processing.
- Key Lessons from Salesforce: Insights gained from working under Marc Benioff and the impact of five years at Salesforce on his approach to AI.
- Model Size vs. Data Size
- Importance of Model Size: Socher argues that model size is critical for training effective AI systems, countering opinions that prioritize data size.
- Misconceptions in AI Models: Many startups are seen as merely wrappers around larger language models (LLMs).
- Hallucinations in AI: Discussion on whether AI hallucinations (incorrect outputs) are a feature or a bug, emphasizing context in evaluating AI performance.
- Value Accrual: Startups vs. Incumbents
- Where Value Lies: Socher believes value increasingly accrues to incumbents with established data, but acknowledges that startups can innovate rapidly.
- Role of Data Access: Large incumbents have substantial advantages because they possess vast amounts of proprietary data.
- Open vs. Closed Ecosystems
- Favoring Open Models: Socher is optimistic about open-source AI models advancing due to community collaboration.
- Challenges with Open Ecosystems: While open models are promising, coordination and funding remain significant hurdles.
- Future of Jobs and AI's Role
- Shifting Job Landscape: Discussion on how AI will change job markets and the importance of adapting skills in the workforce.
- Carpenters vs. Software Engineers: Socher posits that physical jobs (like carpentry) may command higher wages in the future as AI automates more digital tasks.
- The Path to AGI
- Overhyped Expectations of AGI: Socher cautions against unrealistic timelines for achieving AGI, highlighting the complexities and breakthroughs still needed.
- Barriers to Progress: Identifies difficulties in defining and achieving AGI, including the need for models with autonomous goals.
---
Key Takeaways
- AI's Impact on Jobs: Automation will lead to significant shifts in the job market, with a strong emphasis on upskilling workers for new roles.
- Model Development: Both model size and data quality are crucial for effective AI systems; they are not mutually exclusive.
- Startups vs. Incumbents: While startups may introduce innovative ideas, large organizations often have the data and stability to execute AI strategies effectively.
- Future of AI Models: The trend toward open-source models suggests a democratization of AI, potentially allowing smaller players to compete effectively.
- Realistic Perspectives on AGI: While progress in AI is rapid, the path to true general intelligence remains fraught with challenges and uncertainties.
---
Conclusion
This episode offers valuable insights into the evolving landscape of AI, emphasizing the need for strategic thinking in both startup innovation and corporate adaptability. Richard Socher's extensive experience provides a nuanced perspective on the interplay between technological advancement and societal impact.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00The future is already here. It's just not equally distributed. And I think it will meaningfully change so many jobs. The more data there is about a job, the more that job is likely to be automated. But at the same time, there are still a ton of jobs or no one collects data. What that means is that the tasks that are physical are getting more and more expensive and they're gonna become the new bottleneck. And then they're gonna slow down overall progress so that the GDP can like 100x because of AI, It can only increase less than we think it might when we're even the bubble. This is 20VC with me Harry Stebys and I'm so excited for the show's day as our guest is a true AI OG is being widely recognized as having brought neural networks into the field of natural language processing, inventing the most widely used word vectors, contextual vectors and prompt engineering.
0:49I'm so excited to welcome Richard Socher, found and see of you .com, previously Richard served as the chief scientist and EVP at Salesforce. Before that, Richard was the CEO and CTO of AI Startup Metamined, which was acquired by Salesforce in 2016. And check this out, Richard has over 150 ,000 citations, bloody hell as more than I've done podcasts. But before we dive into the show's day, you know all those mind -numbing tedious tasks that seemingly take up half your day. Well, Coda is here with their new AI powered work assistant that helps you and your team, not just finished tasks but make progress so your product team can bring a feature to market faster.
1:27By using Coda AI to tag customer feedback, draft PRDs, suggest target audiences summarising product discussions and more. Your sales team can engage more customers by bringing in data from sources like Salesforce and then use Coda AI to suggest action items meeting agendas or lead scores and your marketing team can drive an impactful launch. With Coda AI summarising user insights, creating briefs generated from notes and writing more press releases to visualize taglines. With Coda AI you can reimagine your stu -list and how you collaborate so you're not only finishing to us but really making progress.
2:02If you want to work a system that lets you get back to work, you can get started with Coda AI's day for free. Head over to coder .io slash 2 -0 -VC that's coder .io and get started for free. And to be here of amazing tools we cannot live without, travel in these bands and never were associated with cost savings, but now you can reduce costs up to 30 % and actually reward your employees. How do you do this? Well, Navan will reward your employees with personal travel credit every time they save their company money when booking business travel under company policy. Does that sound too good to be true?
2:36Navan is so confident you'll move to their game changing all in one travel, corporate card and he spends super app that they'll give you a $250 in personal travel credit just for taking a quick demo. Check them out now at navan .com forward slash 20VC and last but by no means least, we need to talk cash. As of 29th June, you can get 5 .5 % yield on your cash with 26 -week treasury bills. But buying treasury bills is not that easy and you have to navigate a website that looks like it was made before I was born. Enter public .com. Their treasury accounts make it simple to earn a high yield on your cash and it takes 20 seconds.
3:16Here's how it works. Sign up at public .com. Easily purchased 26 -week treasury bills that automatically roll over at maturity for a compounding yield. Plus there are no minimum hold periods. You can access your cash at any time with the flexibility of a bank account, of course, to receive the full guaranteed yield. Of course, to receive the full guaranteed yield you do have to hold to maturity. But here's the thing. These are T -Bills which means your investment has the complete back of the US government, making one of the safest places to park your cash. Go to public .com forward slash 20 VC to lock in a historic 5 .4 % yield on your cash.
3:55You are now arrived at your destination. Richard, I am so excited for this. I have been looking forward to this one for a while, so thank you so much for joining me first. Thanks for having me. It's a great podcast. But you are very kind. I appreciate all ego inflation, but I'd love to start stay with a little bit of context because you've been in the world of ML and NLP for a long time How did you make your first forays into the world of ML and NLP and what did that look like? Boy, that goes back to 2003 is when I started linguistic computer science at Leipzig University kind of the forays and the early days of natural language processing But then I actually felt like there wasn't enough math in it and I switched to computer vision did my master's in medical computer vision and a lot more sort of statistical and pattern recognition statistical learning.
4:42And then during my PhD, I saw some folks work in deep learning and neural networks for very small images, you know, 32 by 32 pixel images of digits and things like that. And I thought, couldn't we use these ideas that they're using and vision for natural language processing? and that sort of started down the road of contextual vectors and inventing prompt engineering and trying to train a single model for all of NLP. Given how steep you are in the community, in the technology, we've obviously seen the recent hype cycle around AI, and I really wanted to start with a landscape evaluation of what the historical context you have is now a fundamental shift in AI, or is it the result of for hype cycles doing that work.
5:29It's a great question, and it's actually so difficult to navigate for a lot of people who are not deeply in it because we are at the beginning of an exponential improvement in a lot of different capabilities. At the same time, some people think the exponential will just keep on going and we keep having these major breakthroughs, and I do think there's some inflated expectations also. People think they have to use a chatbot for every single thing out there, just like they used to think we have to use speech recognition for every single thing out there. And you know, back in the day, when speech recognition finally started to work, people thought, oh, I'm gonna have like this restaurant recommendation engine in speech, that's actually probably not a good user interface, right?
6:08So long story short, the entire level of the I capability is rising massively, but then on top of it, there's sort of small waves of inflated expectations. You mentioned before, something that was really interesting, I wanna start that, but you've been on a quest for a single pre -trained model. I just wanna make sure everyone follows along with us in this conversation. For those that don't know, what do we have today? And why is that maybe inefficient? So today it's finally changing. We've made progress towards that single model, but just even last year or two or three years ago, the prevailing idea was that every task in natural language processing should have its own model.
6:47You have a sentiment analysis model that just classifies tweets as positive or negative. Then you have a summarization model that takes in some long input and then summarizes it in a fewer sentences. You have a translation model that just translates German to English. You have a question answering model that takes a context and says, who's the president in this Wikipedia article that's mentioned? And then you just give that. So there are all these different sub -models that people have worked on and some people have built their entire careers on just sentiment analysis models. And so the difference in something I've been very excited about for pretty much a decade now is to have a model that you keep making better, that you keep adding to it rather than restarting every new training run and so on.
7:33Imagine Wikipedia and everyone just keeps adding to Wikipedia and keeps making one dictionary better rather than everyone who wants to build a dictionary just starts their own dictionary company and builds it from scratch. It just doesn't make as much sense and makes sense when humanity and people and researchers work together and keep making a single model better and better. Richard, I think one of the reasons why this shows success was because I'm not afraid to ask the slightly more basic questions. When we're really playing it like that, it seems highly logical that you would improve and improve and improve upon a single model, versus start again.
8:08Why is that not obvious and why is it not been that way? Oh man, yeah, I have some funny rejections on that. Actually, the paper when we submitted this prompt and engineering paper, it got rejected. it and the reviewer said, oh this is such a misguided effort. There's not a single model in the world that could do any of these things together to which I thought the brain, like we don't replace our brain, we use the same brain. So there is an existence proof for a single model for all of NLP, but it was just a very different way of thinking about the world. There's just debate and people done a lot of research and were very stuck in one way of thinking about it.
8:46There's sort of academic preconceived notions about how things have been and how certain people have worked on certain problems for a long time. And the notion that you should have a single model for each specific task was prevailing for a long time because it's also really, really hard. You needed the ability to train massively large models. You needed the ability to incorporate world knowledge through a large language model to really make that happen. And you needed attention mechanisms and fast GPUs and hardware. where there's a lot of things to make it work to then eventually show that it was better.
9:19And when you had very small models 10, 20 years ago, it would have been impossible to make it work. How important is the size of the model, Richard? It is super important. You just cannot train a single model for all of these different tasks with a small model. That's exactly why and how it would have always failed in the past. I've had guests on the show before and they say the model size isn't so important, but it's the data size that is. Is that wrong? It's totally not wrong. It's just not mutually exclusive. You need a large model and you need a lot of training data for that model Either of them in isolation, like, you know, just like imagine the simplest neural network will just predict a single one -dimensional Output line, right?
10:00Let's like a regression analysis. You have some input X, some output Y, and you try to model where it goes. You can model that with a handful of neurons. And the simplest one is just like a line, right? A linear regression. And that model has even fewer parameters. A long story short, if you now give this linear regression model, billions and billions of training data, it's not going to learn magically anything, but a simple linear line. But if you give the model billions and billions of parameters, it can learn all kinds of very complex, predictive functions and abilities. So concretely, the big breakthrough on top of this idea of prompt engineering of being able to have a single model was to also use language modeling as one of those tasks.
10:44It was actually on our to -do list, but we didn't get to it before others did. And the idea of language modeling is you just predict the next word, which is very easy to get a lot of data for, because you can use anything on the internet and so on, but it's actually incredibly hard. And to do it really, really well, you have to learn so much about the world. Right? If I'm just, I have the sentence like, I'm a New York City and I'm driving North to, right? And now you want to predict what's the next word? Maybe it's Boston, maybe it's Montreal, maybe it's Yale, but you probably want to put more probability on the word being Boston than Yale because there are more people probably driving to Boston because it's a bigger city.
11:21And so not only do you learn about geography, but you also learn just to predict that one word really, really accurately from that one sentence, you need to learn everything about the geography of the Northeastern United it states. And now if you do this billions of times across the internet on all the chemistry articles and biology articles and so on, you'll learn world knowledge and few as world knowledge into that large language model. But only if you have enough parameters to learn it all. I need your help. I think they're straight away about the access to this training data. And some people say, ah, this is where incumbents really thrive.
11:54They have the consumer data. They have transactors. They have all this data. They can utilize that. And others say, no, there's so much open data today that actually access today is highly democratized. Which side would you sit on and how would you think about that question? Because I don't know the answer. It's so funny because it's another question where they're actually both right. So it's complicated in the sense that unsupervised data just raw internet text is easily accessible. But there's still a lot of data sets out there that are not out there. They're actually stored in a private data basis.
12:33And indeed, if you want to answer customer emails automatically, it's very helpful if you're Salesforce and you already have all those emails. You already have labels that people, this knowledge -based article, answered this email or answered this question, right? If you have that data, you can entrain that particular AI for that company to answer its questions from its customers automatically, you can do that much more easily than if you're a small startup and you need to just get access to that data and you need to get permissions and so on. And so that part is true. At the same time, the state before large language models and we called them foundational models because you can often build on top of them very easily, before that it was even harder.
13:16Like it would have been impossible for us as a small company to build a search engine that understands all of these different things in many different languages. It would be unthinkable, but now we can actually because of these large language models and foundational models, we can have a general sense of understanding natural language. And you can at a small start up nowadays, it built in like 80 % solution very quickly. You layer more and more specific data for your task on top as after you have an MVP, a minimum revival product, and then you make it really, really good. Now, the big incumbents, they can make it really good much more quickly, but you could get to at least an 80 % solution thanks to a large LLM site more quickly also.
13:56A lot of people suggest in the venture community, so you can poo poo this one. But like when they're denigrating kind of a lot of AI, I still have to say they say, oh, it's a thin value layer on top of foundational models, and actually really the value accrues to the foundational model layer because it's this thin line on top. Is that fair? What do you think that's total bullshit? You're asking a very good question that often have a more subtle answer than what would fit in a tweet I think there are some very thin wrapper companies out there that probably have very little moat But there are also companies that people don't realize the complexity that will make it really work And they underestimate all the sudden all the other stuff that Companies need to get right to build a viable business.
14:41So concretely, you know, you could think about Instagram Instagram is like, what's the moat of Instagram? It's certainly not their AI and their backend and the brilliance of engineering. It's just like a fairly simple photo sharing app of some fun filters back in a day. None of that was rocket science. Turns out you can have moat other than your backend AI model. It's distribution, it's partnerships, your sales funnel, processes, and so on. So there's a lot. And then there are also areas where the default large language model will not do as well. So for instance, if you want it to be more factual, more up -to -date, and have citations for the facts that it tells you, you need to have a search packet.
15:25And so at U .com, for instance, we've had to build this very complex search backend with a ton of data in it and knowing when to retrieve what facts from the internet so that your model, you can ask about Messys, Miami, Switch, or something, and it'll just talk to you about that, even though the LM itself couldn't be trained and anything, you can think of these LM's as like reasoning engines But you still need to feed them with the right facts and information so that they can reason over the right things rather than just sort of reason what they remember and their memory is a little bit like maybe your uncle who Sometimes exaggerates some idea of the stories from the past and it doesn't remember all the details exactly He often still gets it right, but not always.
16:10And so you want to infuse the fact into the LM and then reason over it, and that whole retrieval back end is also highly non -trivial, but it's something that some VCs don't appreciate and understand the complexity of, and then they say, oh, like, you don't come, it's just a thin wrapper around a large language model, which is very far from the truth. You mentioned that about your uncle, sometimes kind of being hyperbolic or exaggerating. It was Ema that's stability, who said on the show that it's like, I think it was like a crazy smart student who sometimes goes off their meds is what he has robbed a lot.
16:41And he said that hallucinations are a feature and not a bug. How do you think about hallucinations as a feature not a bug? Yeah, you know, it depends again on the context. Your questions are perfect in that they help people understand. It's neither black or white, but it's just some complexity in between. And the complexity here is that indeed the LN doesn't know your intentions and your backgrounds. And so it takes some time for them to adjust to understand and for us to train them and then to balance out like what can and cannot say. And I think some cases we even overshot a little bit. You know, there's some folks who say, oh, you cannot tell any dirty jokes anymore or like a murder mystery is like unethical to write.
17:23So we should not ever write about a murder mystery. But no one says Stephen King or Agatha Christie are like unethical people for writing murder mysteries. people just think it's great entertainment, but they're scared when in AI, it's less than in a large language model predicts the tokens given that problem to write about that kind of stuff. And so in some cases we almost overshot what we can and cannot say or hallucinate about. But hallucinations aren't a feature and are a problem for search engine, right? We do want to be most of the time when someone asks us about a fact. We need to bring in the right fact and then be very accurate.
17:59You mentioned that they don't know your context, which I think is important. We had another guest on the show that said, no models used today will be used in the years time. Do you agree with that? And how do you think about the longevity and lifespan of model usage? That's a great one that's sort of right in the sense that we're going to update all these models. Like we update our model every other week. And so it's not that exact model with those exact weights. But my hunch is we will still have a lot of large language models that are running in production that are that general model architecture.
18:32There probably still be some Transformers. Even though Transformers are mostly amazing because they can be trained on GPUs, there are lots of other models that were currently not exploring because they're not trainable on GPUs. I'm taking a collection of wisdom from other people and throwing it at you to hear your thoughts. I had Alex Nabla on the show and he said that actually, a company's ability to transition between model is what will determine their success moving forwards. How do you think about companies transitioning between models as a differentiator and a voltage mechanism. Interesting.
19:07You can almost start to in the future think about the LMS as a kin and some a ways to a database. You don't really care about which database people are using nowadays, right? If you want to go really big, you might use an Oracle database. If you just want to build the smaller thing, you some might SQL open source database and there's a bunch of of others in between. And I think it won't matter that much, which database you use, just like it won't matter that much, which LMU use. But it matters what you do with it, how you tune it, what kind of trained data you add onto it, to find you in it, how you retrieve facts into it, so it can reason well over all of those things, how you may be prompted to run multiple chain of thought steps to get to a conclusion, all of those things I think will matter more and more.
19:56And so I would argue that it's even more important than switching between models is to be able to incorporate multiple different models because you may have a predictive forecasting model and that is Important to include into an LM like an LM won't be as good in doing a financial forecast because that's not what they're trained on like LM's help a ton for all things natural language because they understand so much about natural language and have so much world knowledge, but that doesn't help you if you just have a huge sequence of numbers and you need to build a very, very accurate statistical predictive model for forecasting or something.
20:32Final one, if we do open versus close, but on the foundational model side, do you think all the foundational model companies that have been created today are the incumbent layer? Do you think that will be a new incumbent foundational model layer or have the winners been created already? I mean, certainly you can't deny that opening eye is ahead by a lot. I predicted that we'll have a GPD4 equivalent model before the end of the year that's open source. Of course, GPD4 keeps getting better and better, so my prediction was for the version we had a few months ago, but I actually think that with models like Lama 2 from Facebook, and everyone there, I do think open source will take over a lot of use cases, right?
21:15It's already getting close to GPD 3 .5. When there's this much excitement in so many careers, depending on understanding these models, imagine all the researchers in all these universities, they're all the sudden out of a job unless they have an LM that works really, really well and they can do useful things with it. They're not going to just say, oh, let's just from now on, run our entire research agenda on some closed API that we cannot analyze and understand and improve and publish papers on. So they need to have a model to exist and those are all very, very smart people that don't have as many resources usually.
21:54They can train a single model for like 20 or $50 million because they're in universities, but they're finding ways, they're collaborating and they're probably going to work on foundational models that are fully open source. And we now see this with like surprisingly Facebook being at the very forefront of it. people will layer on top of that, make it better, and then there will be open source versions you can run on your phone, and those will get better and better over time. And so I'm quite bullish on the LMs in particular, getting more and more commoditized. And yes, there will be a few foundational companies, you know, just for here in Anthropic, also behind OpenAI, working very hard to catch up with them.
22:33And they did raise a lot of money, too. And it's good to have some competition in that space, but my hunch is a lot of people will be okay with an open source model too. A duet on the show from Contextual. I said that ashy it was the lack of alignment between open ecosystems, which would lead to closed winning. Lack of alignment around bunny, goals, timelines, decisions, and that ashy to think that open would win would be naive. Can you see his point or do you think that's not right? I mean, I certainly see that there's a lot of complexity in open models, but you also have already seen a lot of open models, right?
23:11Like Lamatutus was probably like, we could together, Lamatut came out, right? And Lamatut is incredibly good. And it made even more progress than the past open model. And so we might not see for a while until some really strong smart engineering project manager comes along that can herd all these incredibly smart and ambitious cats to come together and build a single model. but that would be my dream is that we actually have a single model and just like Wikipedia, the whole world collaborates and adds to it. And then we get ultimately an AI that anyone on Earth can talk to, just like anyone on Earth can ask facts on Wikipedia.
23:53And you know, like Wikipedia takes a lot of coordination. These articles have lots of layers and hurdles you have to jump through to make an edit to that article. You have to build trust with folks and whatnot. But it has happened in the past that when there is enough interest and excitement and a strong structure that people can collaborate on OpenModel. I love that vision and I love that excitement for that view of the world. Why would it not happen? Coordination is hard, right? Someone needs to come through, needs to get funding. There's a lot of complexity. Do you think that's the role of governments to provide the funding?
24:24I think governments can help. The problem is it's very hard for governments to fund uncertain research and science projects Because when it doesn't work out, some taxpayers will foreshore complain. They're like, you guys wasted my taxpayer money. It's hard to build that. Then inevitably, I have to become even more bureaucratic than where, like, actually has some necessity for and so on. But I do think governments should fund science and especially open science. When we chatted before, you said to me that SIRCH was the highest impact of NLP technology. Why? Why help me understand that? Yeah, obviously it's a massive market, right?
25:05We have a trillion dollar plus company value in that. But more importantly intuitively, we ask search engines questions every day to learn something. They ask them their phone search engines a lot. And so it's an incredibly impactful technology to help people learn that sort of another really big foundational block for the internet. Some people even think like Google is just sort of a yellow pages, right? I mean, there's so many examples in use cases where you can do so much better than what we had in the last 15 years. Richard, how do you do this without killing those original providers? And what I mean by this is say, I type in a question into u .com, u .com scrapes it for the best information it can, answers incredibly well.
25:49I now don't go to the original site because you provided it for me instead. That site, often a news site or and information side or content side relies on clicks and traffic. How does the next generation business model of the internet work? When that attribution and that throughput is not to the provider, but it is to this chat search. Yeah, it's a great question. We need to have that search engine that will fill the open platform where you can actually contribute your apps tool and your content tool. And then if it actually brings up the context from your app and we make money with it, then you can also participate in making that money.
Read the full transcript
26:30Or you have some subscription model, right? And you have some data that you show publicly, but then you have to subscribe. You can have that subscription happening right within your search engine by virtue of this being a much more open platform than Google. Now we've launched this open platform last year, but to be honest, we haven't had a ton of really amazing apps being added to this platform because we just don't have hundreds of millions of users. And so a viewer app lets you book a kayaking trip in Indonesia or something. By the end, you only get like five people. And so then slow in an uptake on that open platform.
27:05But I think that is philosophically the right solution. And I hope people will start collaborating with us because if we don't win and others win, then they'll be actually fully left out. You mentioned slow on uptake. It just makes me think to a question that I ask myself so often in this world, which is, will the incumbent acquire innovation before the start -up requires distribution? How do you think about that in this case here? Yeah, it is the main question. The truth is, this distribution will be ever fully solved, and it's a constant uphill battle. You know, we just got the ranked in all our sites from Google and saw like a drop right away.
27:41So we have seen that drop from getting the list of it out of Google. we're working on a lot of different partnerships and yeah constantly thinking of clever ways to do it that because we have some competition small and large that you know have been copying many of our features in fairly quick succession we have to be a little more careful there and not share all the details. I mean de listing it's quite petty isn't it? Where do you think value most occurs if you were to bet on one in the next five years? Does the next way they I create more value for incumbents or more value for startups? I think it'll be a mix.
28:14I think there are companies that have been fully activated. I see Salesforce having launched a bunch of really incredible features and made incredible announcements in their AI day that will make it harder for AI for service automation, for instance, because it's so core to their business. We've also seen Bing and Google copy what we have launched late last year with you chat. They've copied us in the sense that we've launched it earlier and then they launched something very similar three to four months after. At the same time, we don't see Google change and become a chat for a search engine. They have some features somewhere else, but the main Google experience is the same.
28:51That kind of big change will be hard for Google too, because they make $500 million a day with privacy invading advertisements on that page. And so you don't just really nearly change most of that page, and you get rid of the five six ads that are on top of that page, followed by a bunch of SEO and microsites that are not as good as the ads so people click more on the ads, like you don't just replace all of that with a chat, right? Because you just lose hundreds of millions of dollars a day. And so there is still some innovator still Emma that will not make them change their main experience overnight.
29:24It's been very carefully tuned. Every shade of blue has been a be tested to death. It's very hard to make a massive change and improve that entire review overnight. So they'll slowly move carefully in their default experience. and that's sort of one opening that we have, but it's not going to be easy. You have to keep innovating and you have to keep working on partnerships too and try to stay clever and a little bit paranoid as the startup CEO and in a space where you have these huge incumbents, trillion dollars, something companies that would love to crush you. 500 million a day, my word. Okay, final one for you do one final segment on age, I wish I'm super excited for what I have to say.
30:01Which incumbent do you think's done the best and which do you think is actually lagged behind? I'm a little bit biased because I still love Salesforce and my time there. I think they have done a phenomenal job. You know, when we invented pond engineering, we actually did it at Salesforce Research when I was chief scientist there. And so they've been at the forefront for a while thanks to the research. And so they've been very, very quick also in incorporating that technology into its products. So that's been really great to see. And then of course, in my space in search, you can't deny the fact that we have invulkin and the giants, right?
30:33Like both Microsoft Bing and Google are much more actively working and innovating than they had been for the last 10 years, right? Like since we've launched UChap and had the first LN with a search backend, and citations, and web links and so on in the search context, which we launched last December before anyone else in the world since then, search has changed more than the last 10 years combined. That is exciting to see and in some ways at a very abstract level our mission has succeeded. We wanted to improve the state of the art and search. We did it, but company -wise we have a lot more work to do to actually financially benefit from those changes even more.
31:15I do want to discuss one final element though and we touched on it before the show actually, but it is around AGI because there's a lot of people who are very excited and optimistic about AGI, but you mentioned to me that people might be overly optimistic. Why do you think that people are potentially overly optimistic around AGI? I think it's very natural for people to look at a type of progress and an extrapolated further and further. And I think there's a little bit of overly strong optimism because we have again made a ton of progress, right? Like it's undeniable how much better AI has gotten, but just like with flight, for instance, in the human flight.
31:52We went from the first motorized human flight. And then literally 30, 40 years later, we could fly, loopings with machine guns and full metal airplanes, high up in the speed of sound and years of like, wow. I mean, at this rate of progress, we're gonna have vacations on the moon and we're gonna have flying cars and like everyone will just fly everywhere all the time and so on. And then in the 50s, the whole thing just stopped And we're flying slower off the now than we did before. People realize all kinds of issues. They burn a lot of gas, like airplanes and so on, the way they're structured right now.
32:30And we're slowing down. And like there hasn't been that much improvement. Like if you think about the space shuttle, it used to be that you could land coming out of space. You land like a plane in the space shuttle. Now we just drop people and a bucket with a parachute and they fall somewhere in the ocean. Doesn't look like we've made that much progress in terms of returning from space, but progress is not as linear or as exponential as a lot of people think. What do you stand to think that is science and the laws of nature, physics, gravity versus human reasons, skill, funding? It's all of the above, right?
33:07But maybe in terms of AI and the progress, right? You think like we have this keep having exponential growth, but I'll give you an example where we already see the flattening out of that exponential growth in to an S curve, And that is an image generation. It's photorealistic now. Look at some of these mid -journey images. You can just say I want this exact thing and it'll just give you a photorealistic image of the scene you're just described. Where do you go from there? You can be like hyper super duper photorealistic. Like people only can see an image, right? And that's how good it is. Now you can of course still make a ton of progress in video generation and like make the moving.
33:43But image generation itself is basically maxed out now. Well, it's hard to go. And then in some ways people think, oh, the AI has this superhuman capability. And you could argue, yes, like one AI algorithm can translate decently well into 50 different languages. No human can do that. No normal human can translate that many different languages. But it's also hard to imagine what superhuman language is because language is a human construct. And if there was superhuman, it would just be eyes be talking to each other with like 50 parallel streams or thousands of parallel streeters, but humans can only register and understand language sequentially.
34:20You can have five conversations at the same time, you can read five books at the same time. So language being a human construct and being how humans communicate thought in ideas has a limit on how superhuman it can be. It doesn't make sense to have superhuman language because humans couldn't understand it anymore to some degree. Again, you can produce it more quickly, you can generate a ton of stuff, but humans can only consume it at the rate they have over the last couple of years. It seems like there's this mismatch between the asymptotic point of development there as he said with image generation, aligned to the distribution awareness.
34:56And what I mean by that is if you go to most people on the street in a normal town, not San Francisco, and you have no idea what mid -Journey is, there's no question here. I'm just like, it's weird that we've reached plateau in technological development with zero awareness of it in 6 .9 billion people. You're 100 % right and I'll just jam on that but there's a famous saying which is the future is already here It's just not equally distributed. How long does that take? I mean, that's the interesting thing where it's like I'm so bullish and excited about the I or I have been for over a decade and I think it will meaningfully change so many jobs.
35:32The more digital they are, the more data there is about a job. The more that job is likely to go to be automated and improved, massively in its efficiency. But at the same time there are still a ton of jobs or no one collects data. No one collects data about how many data cleans and water exhalation, all the inputs and the outputs and the actions and so on. Every house is slightly different. It's hard to constrain the environment, full self -driving on off -road, dirt road, like small roads, night time, fog, like all of these things won't work for a long time, you know, but the more constrained the environments are, the more the eye can do, the more data we can collect, the more the eye can help us automate and make things more efficient.
36:12But what that means is that the tasks that are physical are getting more and more expensive and they're going to become the new bottlenecks, right? So your carpenter, people like who built your house, all of those kinds of jobs are going to be more and more expensive and then they're going to slow down overall progress so that the GDP can't like 100x because of AI, it can only increase less than we think it might when we're deep in the bubble. Richard, I'm giving you a warning that this is a shit question, okay? So warning ahead of time. Are you concerned by the job displacement question and just the awareness that when you look at industrial revolutions and technological revolutions.
36:50There's always decade -long as you transition periods. Whereas here it seems like the transition period is years, like a couple of years, not decades. Are you concerned or do you think we're overwiring about this? A hundred percent. I do think, you know, past industrial revolutions, they usually happen for an individual at a surprising rate. When you're a weaver and you get some big machine now and it's just like totally automates making clothes, you will hate that machine, right? And the Lodites tried to destroy those machines and there are people who want to slow down and destroy certain kinds of AI that change their jobs.
37:28And I feel empathy with those people. And it would be nice to have social systems that catch some of those people and help them learn new kinds of skills, help them incorporate those new technologies into their workflows so that, you know, if you're an illustrator and you used to be able to charge $1 ,000 for an illustration because it took you three days, now it takes you three minutes. If you don't use it and you still expect to get paid for three days, but now there are thousands of people who can do it in three minutes, that if you're not adapting to it, it will change your job landscape. So to try to help people in that transition isn't incredibly important.
38:03I think it's also hard to say, let's not make things more efficient. Like, let's not have more art, let's not have in healthcare, like, more healthy people cheaply, let's make more jobs. fields a little bit like you'd want to slow things down and some people feel like they want to go back to the mountains Right and just live a simple lifestyle with no technology You see that now come up sometimes, but I think overall in middle class person now lives in Along many dimensions at better life than a king did a thousand years ago or even a few hundred years ago Most of the time people appreciate the end state of that those transitions like we can have more cheaper food now There are fewer people who have to starve.
38:44You know, 150 years ago, over 90 % of people work in agriculture, just to put fruit on the table. If you told them, hey, we'll have these massive machines. If you stand in their way, they will trust you to death, but like they'll do all of this farming automatically. And now only 5 % of people need to work those kinds of jobs. They would have said, that's really scary and what are we going to do? But people always find new things to do when there are efficiencies created in existing things. I do have just one final thing before we do a quick fight. And it's just we mentioned before a GI and kind of wise potentially overhyped or over excited What are the three barriers you mentioned to me before that prevent a GI or will put a pause on it?
39:23We do have a lot of research breakthroughs we still need to make in order to achieve a GI Just like it's very hard to know when someone would have been prone to engineering or when someone would have the idea to to scale up language models and transformer networks just massively and not have other cute little models and modify those, but just scale up existing models massively, like opening eye and Ilya Satskewer and other side years. We don't know when those breakthroughs will happen. I find it hard to call something artificially super intelligent or generally intelligent. If it all it does is predict the next tokens and you say, oh, predict this next token, it'll predict that next token.
40:01I think an intelligent existence probably needs to have some of its own goals and like some of its own mind of what it might want to do. Maybe it doesn't want to generate tokens. If you just want to maximize the right token prediction, you just make sure everyone says AA only ever and then you predict the right token a hundred percent of the time. But that wouldn't be very interesting fulfilling intelligentex existence. And so because companies need to make money and governments want to have a productive economy, no one is working on AI just doing whatever it wants to do. Because that doesn't make any money, so no one's working on AI setting its own goals.
40:38And until that happens, it's going to be hard to think about general intelligence of almost life form that has some intelligence, even if it is not biological. Richard, I'm going to do a quick fire with you. So I give you a short statement and you ping me a quick thought back. It's as fast as possible. Does that sound okay? I'll try my best. As you may have noticed, I'm not really good with the short form. Not neither am I my friend, so what do others not know that you know to be true? Knowledge is similar to the future, is already here, but not equally distributed. How important retrieval augmentation is for LMS?
41:14A lot of people are not interested in that. Do AI founders need to be in the valley? It certainly helps a time. What single element would you most like to change about the AI community? To have more folks talk to each other about real risks and let a small set of folks continue to think about existential risks but not scare people so much with very interesting sci -fi general fiction scenarios that would make fun action movies but are not the real issues at hand right now. Can he has time what role does AI play in society then? And even bigger one than now. Were you against Elon Musk's petition to pause development?
41:53I did not sign it. I don't think it makes sense to pause the training of models. Just like I wouldn't want Pessa to stop updating the AI in my test car because I would assume it gets better and better and I don't assume at some point the car will just take off and watch the sunset by itself because it got too intelligent. Final one for you. 10 years time. Where's U .com then? We have this conversation in 2033. Where's U .com then? Well, I think and hope it will be the default in many hundreds or millions, if not billions, of people's computers so they can find better answers. Richard, I've loved this.
42:30We've gone off topic many times. It's been a fantastic discussion. So thank you so much for doing this, my friend. Thank you, Harry. You've been one of the most engaging podcasters I've ever talked to. I was super fine. I mean, I have to say, I think that I showed that I actually know a little bit about AI I can do more than just venture. I'm quite impressed with that knowledge. I hope you enjoyed the show there. You can find out more and see everything from that episode on YouTube. By searching for 20VC, that's 20VC. Huge science to Richard for being such a fantastic guest. But before we leave you today, you know all those mind -numbing tedious tasks that seemingly take up half your day.
43:09Well, Coda is here with their new AI powered work assistant that helps you and your team not just finish tasks, but make progress. So your product team can bring a feature to market faster by using Coda AI to tag customer feedback Draw PRDs, suggest target audiences summarizing product discussions and more. Your sales team can engage more customers by bringing in data from sources like Salesforce and then use Coda AI to suggest action items meeting agendas or lead scores and your marketing team can drive an impactful launch. with code AI summarizing user insights, creating briefs, generated from notes, and writing mock press releases to visualize taglines.
43:48With code AI, you can reimagine your stu -list and how you collaborate. So you're not only finishing tasks, but really making progress. If you want to work a system that lets you get back to work, you can get started with code AI's day for free, head over to coder .io slash 20VC, that's coder .io and get started for free. And to be here of amazing tools we cannot live without, travel in a expense and never associated with cost savings. But now you can reduce costs up to 30 % and actually reward your employees. How do you do this? Well, Navan rewards your employees with personal travel credit every time they save their company money when booking business travel under company policy.
44:29Does that sound too good to be true? Navan is so confident you'll move to their game changing all in one travel, corporate card and he spends super app. that'll give you a $250 in personal travel credit just for taking a quick demo. Check them out now at navan .com forward slash 20VC. And last but not least, we need to talk cash. As of the 29th of June, you can get 5 .5 % yield on your cash with 26 -week treasury bills. But buying treasury bills is not that easy and you have to navigate a website that looks like it was made before I was born, enter public .com. Their treasury accounts make it simple to earn a high yield on your cash, and it takes 20 seconds.
45:09Here's how it works. Sign up at public .com, easily purchased 26 -week treasury bills that automatically roll over at maturity for a compounding yield. Plus there are no minimum hold periods. You can access your cash at any time, with the flexibility of a bank account, of course, to receive the full -gown tea deal. Of course, to receive the full -gown tea deal, you do have to hold a maturity. But here's the thing, these are tea bills, which means your investment has the complete backing of the US government, making one of the safest places to park your cash. Go to public .com -20VC to lock in a historic 5 .4 % yield on your cash.
45:48As always, I so appreciate all your support and stay tuned for an incredible set of episodes coming next week.
From the publisher
Richard Socher is the founder and CEO of You.com. Richard previously served as the Chief Scientist and EVP at Salesforce. Before that, Richard was the CEO/CTO of AI startup MetaMind, acquired by Salesforce in 2016. He is widely recognized as having brought neural networks into the field of natural language processing, inventing the most widely used word vectors, contextual vectors and prompt engineering. He has over 150,000 citations and served as an adjunct professor in the computer science department at Stanford.
In Today's Episode with Richard Socher We Discuss:
1. The Decade-Long Journey to Becoming an AI OG:
- How did Richard first make his way into the world of AI over a decade ago?
- What are 1-2 of his biggest lessons from working with Marc Benioff?
- How did 5 years at Salesforce impact how he both thinks and operates?
2. Models: Does Size Matter:
- How important is model size? Is data size more important?
- What are the biggest misconceptions people have around models today?
- How does Richard respond to the suggestion that "many startups are wrappers around LLMs"?
- Are hallucinations a feature or a bug?
3. Where Does Value Accrue:
- Where does Richard believe most of the value will accrue; startup or incumbent?
- Which incumbents are best positioned to win? Which are the laggards and behind?
- What do many not see about the startup vs incumbent race in the AI war?
4. Open vs Closed: Which Wins:
- Does Richard favour Yann LeCun's open approach? Or is the world of AI more closed?
- What are the biggest challenges of an open ecosystem?
- What are the nuances that make both challenging?
5. Richard Socher: AMA:
- Why will carpenters be paid more than software engineers in 10 years?
- Why is AGI still way off? Are people too unrealistic?
- How much money does Google make off search every day? Why does that leave them vulnerable?




