The difference between early and late AI adopters

30 Apr 2025 · 50 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Azeem Azhar's Exponential View - Episode on AI Adoption

Episode Title

The difference between early and late AI adopters Guest: Steve Hsu, Physicist and Entrepreneur Podcast Host: Azeem Azhar Original Air Date: [Insert Date Here]

Episode Overview In this episode, Azeem Azhar engages physicist and entrepreneur Steve Hsu to explore the nuances of AI adoption, particularly focusing on the challenges of hallucinations in large language models (LLMs) and the impact of AI on labor markets. Hsu, co-founder of Superfocus, shares insights on overcoming the limitations of current AI technologies and discusses the broader implications for society.

---

Key Discussions

  1. AI Agents and Hallucination Challenges
  2. Definition of Hallucination: Models often produce confident but incorrect answers based on pre-training data.
  3. Superfocus's Approach: The startup aims to control AI behavior and knowledge bases to reduce hallucination occurrences.
  1. AI's Impact on Call Centers
  2. High Replacement Potential: AI could replace 80-90% of calls in call centers.
  3. Employee Reactions: Like historical workers facing technological change, current call center employees feel threatened yet uncertain about the future.
  1. Current State of AI Technology
  2. Rigorous Testing Required: A significant amount of rigorous statistical testing is necessary to understand the tail risks associated with deploying AI.
  3. Delayed Adoption: The pace of technology diffusion is slower than expected due to human decision-making and resistance to change.
  1. China's Competitive Technology Landscape
  2. Mini CEOs: Mayors in China act much like CEOs, focusing on technological advancement to improve local economies.
  3. Public Perception: The Chinese populace generally views AI and technology positively, contrasting with Western skepticism.
  1. Open Source AI and its Future
  2. Rise of Open Source Models: Open-source models allow for more tailored, efficient applications of AI in specific domains.
  3. Diversity of Applications: A shift towards specialized models rather than monolithic ones can lead to better performance in targeted tasks.
  1. Balancing AI Deployment and Human Oversight
  2. Expectations of Reliability: AI must meet higher reliability standards than humans, especially in critical applications like customer service and healthcare.
  3. Human vs. Machine Performance: The expectation gap between human agents and AI agents will likely change as organizations learn more from AI performance metrics.

---

Key Takeaways

  • AI Adoption Dynamics: There are significant barriers to AI adoption in sectors like customer service due to existing labor structures and human resistance to change.
  • Cultural Differences in AI Acceptance: Societal attitudes towards AI vary significantly, with optimism prevalent in China versus skepticism in Western nations.
  • Future of AI in Labor Markets: The introduction of AI in labor-intensive sectors could lead to massive shifts in job structures, particularly in countries heavily reliant on BPO industries.
  • Innovation from Open Source: The emergence of open-source AI technologies is likely to democratize access and encourage rapid innovation across diverse applications.

---

Potential Future Directions

  • AI in Early Childhood Education: Hsu discusses research into embedding AI into plush toys for educational purposes, indicating potential applications in language learning for children.

Conclusion This episode provides an in-depth perspective on the complexities of AI adoption and its societal implications. As AI technologies evolve, understanding the balance between human labor and AI capabilities will be crucial in navigating the future of work.

For more insights, tune in to Azeem Azhar’s Exponential View and explore the nuances of exponential technologies shaping our future.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00We have agents that are capable of replacing, you know, something like 80%, maybe 90 % of the calls that come into a call center. So what is it like when you go out there and you show this technology, particularly to those frontline operatives? I think they feel about it the way that someone who was a blacksmith or a buggy whip maker felt when they saw their first automobiles rolling down the street, right? They could tell something was happening and they don't like it, but what else can they do? I think every person who's developing software knows that getting to the first 70 % is trivially easy.

0:32It's the climb up the last 30%. 70%. You don't just want to turn this thing on and then discover like, oh, overnight it got into some bad loop and pissed off 100 ,000 customers, right? One of the things that I think the general public doesn't understand is like, how much rigorous statistical testing is required to know what are the tail risks associated with a deployment of autonomy? That could be automated customer service. That could be a driverless vehicle. Let's talk about China. There is this tournament-style competition between mayors of cities. They act a little bit like, you know, mini Y Combinator bosses or mini venture capitalists.

1:06Combined with the people, there are just more pro-technology. You say like, oh, we're going to have AI at the hospital. The average Chinese person is not thinking like, oh, what about my privacy? Or what if the AI makes a mistake? The Chinese people are more like, this is awesome. In a future that is full of AI agents, what's something that you might end up doing more of? Again, I'm going to tell you something incredibly crazy. You can break the news here. Today, we are joined by physicist and entrepreneur, Steve Hsu. The most loyal Exponential View listeners may remember. He was on my podcast back in 2019.

1:42Feels like a century ago, so much has happened. We talked about biotechnology and the promise of genomic sequencing. And I actually had to go and buy a maths book after that podcast to make sense of some of the answers. So today we're talking about something else, AI, agents, and what might come next. Steve, it's great to have you here. I want to start with what you're building. You've co-founded a company called Superfocus, which you've described as solving some of the key limitations in today's LLMs. Give us a quick tour. What are you building and what is the deeper shift that it represents?

2:19So I think everyone who's experimented with large language models knows that they hallucinate. So they will sometimes give you a very confident but completely wrong answer and that answer will always be typical of the kinds of things that it saw in its pre-training. So if you ask an LLM, is there a flight that reaches Paris from London that lands around four o 'clock? It will give you an answer, but it might not be real. It will seem like a real answer, like, oh, Air France has one that lands at 412, right? And maybe in the past it did, but maybe right now it doesn't, right? Typical large language model hallucination.

2:56And so our startup realized early on that in order to make practical applications from these AIs, one would have to solve this problem. One would have to control both the behavior of the AI agent, if you want to call it an agent, and also control its fact base, the sort of core knowledge base that it uses to answer questions or conduct operations. And that cannot be drawn from the pre-training data. There's just too much junk in the pre-training data. Or even if it's not junk, it may not be relevant to the specific problem that the AI agent is trying to solve for you. Right. So you're tackling that fundamental limitation in the way that LLMs are built.

3:38I've had that experience. The way I think about this is that they are often accurate at a conceptual level, but not necessarily accurate at the word or token level. So they'll sometimes say of me that I went to Cambridge University, which if you're an alien sitting on Neptune, I went to Oxford. Well, Cambridge is kind of the same thing. But they don't ever say, I went to West Point and I did SEAL training. I mean, I've never had as wild a hallucination as that. And that really is about the way in which they represent the world in this complex, high-dimensional space, isn't it? Yes. So there's something called the embedding space.

4:18The models are actually working in this abstract space, which is literally a space of concepts. Oxford and Cambridge are very, very close in that concept. And maybe there's no other school that's exactly in that space. But, you know, sometimes the details matter. So if I'm a fundraiser for Cambridge, I actually care whether you went to Cambridge or Oxford in deciding whether to contact you, right? Right. Yeah. Yeah. Yeah, so the details matter, but in some sense, they've still got a useful model of the world that perhaps loses some of its resolution if we push really, really hard for it to be precise.

4:54So how do you go about resolving that? Right. So we actually build systems in which we embed the language models in a larger software platform. and the language model itself, we generally are using mainly only for its language and to some extent reasoning abilities, but the knowledge base is stored separately. The fancy way we describe it is as a kind of attached memory for the AI. The AI can rely on that attached memory. And the way we program, we use sort of old style programming in this platform. We sort of force the model to only use that knowledge base in answering the fact part of questions.

5:38And so that extra constraint, it solves the hallucination problem and makes both the behavior and the knowledge base of the AI reliable. How different is that to a traditional RAG, retrieval-vented generation type of approach? So if you look at the architecture, it is actually, a piece of it is RAG. And interestingly, we actually, when we founded the startup, we actually filed a patent, the company filed a patent on our architecture And that was actually before the word RAG was in wide usage. So it is possible, who knows how the USPTO, Patent and Trademark Office, operates, but we might be issued a patent on RAG.

6:17Of course, it's not referred to RAG in the patent filing. It has a lot of similarity. Another thing you might do is you might have multiple models involved in the generation of the response in which some models are just error checking the proposed response of the big model against what the little models can see in the knowledge base. And all that, if you're doing voice, which we do, all that has to happen in a latency time of less than two seconds. So humans, if I stop speaking and I'm waiting for you to respond to me, if it goes more than a couple seconds, it's kind of strange. And so all of the stuff that I just described to you is engineered down so that the latency is between one and two seconds, so it sounds natural.

6:59One of the things that I'm curious about is whether the improvement in models really starts to tackle the hallucination question. So I've been using Gemini Pro 2.5 or Gemini 2.5 Pro for the last couple of weeks, both through the app and via the API and some applications I've built. and I've been really, really impressed with its ability to be factual, be precise and hold that over quite long context conversations. That was often a problem with LLM models, which was, they were very accurate for the first 500 words and by the time you got to 5 ,000 or 10 minutes into your discussion, they're just going all over the place.

7:41And I found that wasn't the case with Gemini 2.5 Pro. So in a sense, isn't a lot of this going to be solved by the investments in bigger and bigger models that can model more precisely this concept space? And that's the direction that the Googles and the open AIs are going? Well, it is improved. The situation is improved if you're using a model that has reasoning capabilities. What's going on in the reasoning is the model has been taught as it sort of talks to itself in trying to solve, generate a good response to your query, it has been taught to double check facts or components of the reasoning.

8:21However, if the model doesn't really have access to the actual ground truth, it can still go off the rails because it can think X is true. It doesn't have a way to independently validate X, and it can then use X in its reasoning pattern. Even though it knows it's supposed to check X, It doesn't really have a way to falsify. The other problem with reasoning models for the application that we're specifically addressing in our market segment is people want for voice, again, fast response. So if you rely on a whole bunch of tokens being generated for reasoning, that's way over the fraction of a second that you have for that generation.

9:01You know, the speech to text and text to speech, all that stuff is wrapped in there. And then the reasoning has to be, you know, even smaller than a second. like a fraction of a second. So you would need a reasoning model that generates all of its reasoning tokens that fast. And then you might be able to use it, yeah. I mean, when we talk at this level of detail, of course, it puts into context what this path to super capable AI really is. Because you're working in a very constrained environment where not only do you have to, you know, take the sound wave, you've got to break it into phonemes, you have to put it into a voice model, you then have to turn that into text.

9:39You then have to reason across it. And you've got a really limited amount of time to do all of that. You know, it's not like O1 Pro going away to think for 10 minutes. And this is sort of the reality of how these products are going to be built over the next few years. With Superfocus, I saw quite an impressive demo where essentially your AI is having like a support conversation with a real person. So where is that now? Where is that technology now? And was that a shiny TED-style demo? Or is that the reality of customers using this? It's interesting. And this gets into a very, I think, important point, even for the broader question of AGI and what's going to happen to human society in the future.

10:24So the technology for, you could call them customer support agents, is very advanced now. So we have agents that are capable of replacing, you know, something like 80%, maybe 90 % of the calls that come into a call center. So if you're ordering a pizza or you're changing a delivery address for a package, you're wondering what happened to your package, you know, all of those things, AIs can actually handle pretty well now. However, in terms of what fraction of labor that used to be done by humans, entirely by humans in these call centers, has been replaced thus far by AI agents, it's still minuscule.

11:03It's very tiny. And a lot of it's held up by human decision-making. A lot of it's held up by sunk costs in old systems and old ways of doing things. And so one of the lessons that people don't understand, like the real pain that startups have to go through isn't just inventing the technology, it's actually diffusing the technology and getting customers to understand how it works, to adapt their workflow to it, to buy it, make a bet because, you know, you're taking some risk to let some guy come in and say, hey, my black box AI is going to replace those hundred people over there in the field that are in the call center.

11:39So it's just slower than what people think. I was going to ask about that because you, I do remember you were going off to the Philippines quite regularly a couple of years ago where BPO business process outsourcing is 2 million people in the country employed there, about 10 % of GDP. It's obviously bigger in India, 6 million roughly working BPO and support by the end of this year. Those are big numbers and they tend to be above average in terms of their education, in human capital and above average in their income. So what is it like when you go out there and you show this technology, particularly to those frontline operatives who, in a sense, you are producing technology that does quite a lot of their work?

12:28I've been in many, many meetings now where in the room are usually it's managers, more senior people at a company, maybe a company that does BPO work or a company that hires outsourced BPO providers, but people who are familiar with the customer service problem and have maybe even worked, some of these people have worked their way up. So they have worked in a call center themselves or they spend time in call centers all the time. And what's amazing is how big their eyes get when, you know, we might show them a video demo, which is a recording of maybe me or somebody else talking to one of the AIs.

13:04And okay, that's impressive enough. But then like some of them are a little more suspicious. So like, well, let me talk to the AI. Can I talk to your AI? And I'm like, sure. And I turn my laptop around and, you know, hit a button and they can talk to the AI. And when they realize that the AI is actually under, you know, quote, understands what they want and can make changes in the system, generate a return label, you know, can do lots of things that we've, you know, functionality that we've equipped it with. It's kind of shocking because you see the wheels turning in their head and they say, wait, I employ a thousand people who do mostly this.

13:38Like surely there are some corner cases where you really want the AI to pass it off to a human. But 90 % of the labor that those thousand people do is being done by this AI. And if I choose to power it with DeepSeek or something now, I can do it for, you know, we used to think in terms of one-tenth the cost per hour of labor versus the human, but you could go down another order of magnitude. I mean, that sounds insane, but it's actually true. It's crazy. And at a one one hundredth of the cost, you could run this across five agents at the same time. And if four of them agree, you just move to the next step, right?

14:15You can construct a sort of a voting mechanism to reduce the number of edge cases you can't handle. That, of course, though, let's talk about what that means in those employment circumstances in those companies in that country. You said that deployment is slower than you expect because buying and decision making is slow. So that is from the bosses of the companies over in, say, the Philippines. But what's your sense for the extent to which people have understood what this might mean once the technology is being bought and deployed and ultimately copied, the deployments being copied by everyone's competitors in the various suburbs where this gets done?

14:58So we are right in the middle of one of the greatest examples of the future is already here. It's just not evenly distributed. Typically, like if you're, let's suppose you're a growth stage startup. So you found a particular niche and you're growing like gangbusters, but you're still very functional. You're a small company, relatively small, and you've got very good startup people, agentic people, etc. etc. Those companies are going to or are deploying AI very fast. And so some of those companies, when you deal with customer service there, or you have to do something in their system, you literally are dealing with an AI and they've embraced the technology.

15:39And it's what all the other companies are going to look like five or 10 years from now. On the other hand, I have to go to these like trade shows and expos, founder led sales, right? So I'm often at a meeting where Everybody at the meeting is somebody who owns like, I own the customer service function for, you know, Hilton Hotels. I own the customer service function for XYZ Bank. I own, you know, those are the people there. And those people want to deploy this. They're being told by their CEO and CFO that they should deploy it. But if you look at their incentives, okay, imagine you're someone who came up, you're 50 years old and you came up for, you know, you've been working for 25 years in a call center environment.

16:21you worked your way up, you know how to manage people, you know how to deal with, you know, contractors in the Philippines or in India. There's a bunch of like specialized knowledge that you've built up over the years. Okay. Some crazy professors, startup founder guy comes to you and says, Oh, you see this little black box. This is going to replace this little black box, which isn't even really here. It's in the cloud is going to replace all of those people that your core competency is managing and provisioning those people. My little box is going to handle that for you. My little cloud thing.

16:53Let me just give you a cryptographic key and then you can hook your system. Most of these people are incredibly threatened by that because they think ahead one year and they say, what do they need me for? They need me to manage the now reduced by an order of magnitude set of people who just handle the edge cases. And this young whippersnapper guy, his thing is now the whole thing I used to manage. One guy who's a friend of mine who owns the entire CX function for a very well-known entity, sort of similar to DoorDash or Uber Eats or what's the one you guys have, Deliveroo? We have Deliveroo, yeah.

17:31So this guy owns the whole CX function for a company like the ones I just mentioned. And he told me most of these guys you're going to deal with in my space are going to pretend they're interested in installing your system. they're going to try to learn from you as much as possible, but they're dragging their heels. They actually don't want to install it for the reasons I just, the incentives, because of the incentives that I just described to you. If you're trying to get the factory to switch from using steam to electricity, all the people who build steam and fit steam pipes are like against you.

18:06They're actually against you at some level and you have to fight through that. So, I mean, I think that that's very reasonable assessment of the incentives and the pressures. It is the difference between the long term and the short term that we see with new technologies. And in a way, it goes to answer the question of what happens to the people who are working in the call centers. But I'm quite curious just to come back to that idea, which is when you meet the frontline workers who have access to Google and YouTube and ChatGPT, I mean, they know what the trajectory of the technology is and they know what the promise is.

18:44and virtually every discussion about AI and the future of work puts BPO and call center workers right at the tip of the spear. How do they feel? How are they thinking about it? Is there a sense of, I don't know, what is the sense that is amongst those people in the conversations you've had? I think they feel about it the way that someone who was a blacksmith or a buggy whip maker felt when they saw their first, you know, their first automobiles rolling down the street, right? They They could tell something was happening and they don't like it. They're vaguely worried about the future. But what else can they do?

19:21Now, I will tell you that in some context, again, some of this is proprietary information, so I can't say it. But in some context, companies that we are working with have said to us, we cannot say anything publicly about this because it's very bad optics for us to be shown to be replacing all these humans with robots. And furthermore, we're actually negotiating a labor agreement. Like our whole industry maybe is negotiating a labor agreement. And the unions are kind of like, they're kind of aware of this. They're putting in little clauses. Like if you're displaced by technology, then you're entitled to this severance payment or there's a pool of displaced workers and the employers must hire from that pool first.

20:09if they fill any new job in the, they must hire first from that pool. So these are all standard things that a union would negotiate with the corporate employer. But the moment that like, they become more aware that this threat is imminent, then they're going to negotiate even harder, right? So all of this is very sensitive. What do you think is a workable zone of agreement to bridge this chasm that's going to emerge? I think the workable thing will be you install a super focused agent you pay us one-tenth or one-fifth of what you're paying the human and everybody's happy i'm not sure what you're asking me is you know just no what i know i mean i mean in terms of the relationship with the workers and i think in terms of you know 10 percent of philippines gdp right what what is where is there a a sort of balanced approach where or is it or is it just the law of the market and and you know we hope the labor market is flexible enough and there's enough enterprise in these countries to spring up new jobs in all of the promising areas that will emerge?

21:12The problem is particularly acute in the Philippines because it's about 10 % of their GDP. It's 2 million workers. And those workers, that's one of the best paths into the middle class. If that reduces by an order of magnitude and size, I think there are just really strong social repercussions in that country. I don't know how to fix it. I do think it's going to take some years before all of this is done. I was just giving reasons a moment ago why technology diffusion is a bit slower than you think it will be. Like even though the tech is ready, it takes some time before it's widely adopted, right?

21:45So we do have some time to react, but I just can't imagine like these jobs that can be done at one-tenth the cost using AI, they are going to be replaced. Yeah, no, I mean, it's the logic is there, the economic logic, the historical analogies are there. And what we'll also start to see is the firms who get there first will be able to compete better than the firms that don't. and I'm not sure if this will hold for AI, but it certainly held for automation in factories in the 70s and 80s and 90s, which was that while there were job losses, it tended to be in the firms that didn't automate because they got outcompeted by the companies who did.

22:28I want to pivot to the other area that I think you've spent a lot of time thinking about. You're such a polymath. You cover so many areas, but let's talk about China. I think we've all seen that There's innovation across so many sectors in China, electric vehicles, semiconductors, solar, maybe even photolithography and so much more going on. And one of the things I think that has allowed that has been that from the top down direction, there is this tournament style competition between mayors of cities. And Chinese cities are essentially nation sized. They had 20 million people, 30 million people with great universities, extreme educational competition, attainment focused to get those top schools.

23:15and the mayors themselves they act a little bit like you know mini y-combinator bosses or mini venture capitalists pushing key technologies fostering ecosystems to support research and development and competing with each other as to who can be you know the self-driving car capital the smart city capital of china and and of the world to what extent does that dynamic drive the rate of adoption and development of these types of technologies at the moment? Yeah, I think the way you described it is perfect. So I think for people who are not familiar with China, on the one hand, you might think, based on mostly Western propaganda, that China is closer to North Korea as a polity or an economic system.

24:07On the other hand, you could say, no, it's actually closer to Singapore. because Singapore also is not the freest political system, right? It's a little bit more authoritarian than maybe we're used to in the West. So is it more like North Korea or is it more like Singapore? And I would say it's much more like Singapore. So you have this economy in which the people who run a particular province or run a particular city, as you said, which could be 20 million people or even 100 million people, their incentive to rise in the party is economic and technological development. And in a way, their job is more like that of a CEO than it is like a Western politician who has to get reelected in two years.

24:53And if you think about what pressures is a CEO under, someone's gonna measure their performance, like the market or investors, or maybe the Communist Party Central Committee is gonna look at the numbers from Guangzhou from 2025. And they're going to ask you like, well, what happened to your GDP? What happened to your average level of education? What happened to the number of patents that were produced? You know, all these things they're going to get graded on. And of course, they still have a little bit of sensitivity. Like if the local labor union protests or people are very unhappy about some policy, that also gets registered, right?

25:27And they're graded on that. So they're kind of balancing the way a CEO of a company has to balance public opinion, performance of the company, internal R &D of the companies. That's kind of more like... It sounds a bit like working at Bridgewater, frankly, with all this grading going on, it was popping up in my head. So that seems to be why we often misunderstand if you're sitting in the West, that dynamism within China. Do you think that that sets up the conditions for more rapid AI deployment than we might see, particularly in the US. I mean, the US has always led on these types of technology deployments as they expand out the economy because it is a leader in developing them because there's a lot of capital.

26:14The market is big. It's a single market. So it's a great place for the entrepreneurs to sell. And of course, American bosses have an incredibly positive attitude towards technology by and large. So they're willing to move forward. Just many more of them are early adopters than you might find in the UK or Belgium or Germany. But is there anything that you've seen that suggests that maybe that calculus changes over the coming few years with AI? I think that all of those factors that you mentioned, and I'll add one more, which is that as a fraction of the total student population in the country, in China, they're about two times more likely to do really hardcore STEM majors.

Read the full transcript

26:56And so people there tend to be able to do math and do really hard STEM stuff. And so there's actually no shortage of engineers. There's just so many good, talented engineers there that they can make progress in many, many different things at the same time. And so combined with what you mentioned, that the people there are just more pro-technology. Like a lot of how you feel about technology, we don't really know how this whole thing is gonna play out, right? So it's all vibes, right? So it's like, oh, you could have a society that like is very nervous about technology, is worried about ecological disasters or AI existential risk.

27:35Or you could have another society that says like, well, we got much, much richer just in my lifetime. And a lot of that was due to just climbing the technological food chain. So we should just keep climbing it, right? And that's really where China is. China is much, much more optimistic. Like if you talk to some random Chinese person, you say like, oh, we're gonna have AI at the hospital, the AI is first going to talk to you before you talk to the doctor and it's going to like do some triage. The average Chinese person is not thinking like, oh, what about my privacy? Or what if the AI makes a mistake?

28:07Or what if Microsoft gets all this data? Those are things which a typical American might think right away. But the Chinese people are more like, this is awesome. This AI is going to like, you know, so just that vibe difference is a huge factor. Yeah, that vibe difference actually shows up in data as well. So Edelman, which is a PR company, has a long-running trust survey, the trust barometer. And for the last few years, they've been asking about people's attitudes towards AI. And this is a consistent, you know, longitudinal data set. And whether people are roughly optimistic or pessimistic about it, there are many, many more layers to it.

28:45I encourage listeners to go off and get the Edelman Trust Barometer. What I found fascinating was that if you look at the US and the UK, roughly 70 to 75 % of people are a little bit pessimistic about the impact of AI on their lives. And if you go to Asia, not just China, to Korea, South Korea, or to Singapore, it's the reverse. 70 plus percent are optimistic about it. And like you, I think it does come back to people's actual, just to use a word that's a bit unfashionable right now, but their lived experience, right? They remember living in the village. They remember, even if you're as young as your early 40s, you will have these memories and you'll have the memory of your life changing and transforming from living in 20 square meters, 240 square feet with a shared bathroom on the corridor and a block of apartments shared with other families to living in some of the most futuristic cities in the world with, you know, all of the luxuries and more in some cases.

29:53So I think that that does contribute to a sense of optimism. And I suppose it sits with your observation of the depth of technical and engineering talent. But does that, I mean, does that really matter in the current moment? And I'll give you the two thoughts about that. The first is that, you know, we are scaling human capability with AI systems. We are delivering graduate level server tools for$15 a month. And, you know, for$15 a month with Gemini or$20 a month for Claude, I can suddenly spin up a lot of third-year university students, hundreds of them. I mean, you used to be an academic. You know how good they can be and what their limitations are, but that transforms the reach of individual.

30:45So does that labor advantage in terms of high, deep STEM really hold beyond the next two or three years as AI gets better? Well, a lot of this depends on what you think about fast takeoff for AI. And not just, as I was saying earlier in our conversation, not just fast takeoff in the capabilities of the AI, but also the diffusion of it. Because you can have a situation where the AI can really do good stuff, but no one's letting it do good stuff. That's actually the situation in customer service or agentic. It's like there is this lag, right? So it's also true that I think there will be a long period of time where a skilled engineer who's trying to figure out how to do injection molding more efficiently or something to make like a car body, you know, that kind of physical stuff is still going to be done by human engineers.

31:35It's going to be a while before AI takes over that kind of very spatially oriented physical stuff. So, you know, I think some of the people in Silicon Valley who believe in fast takeoff and are very hawkish about the China AI competition, what they are thinking is, oh, all of these advantages that Steve mentions that the Chinese have, those are all going to be irrelevant once we have AGI, which is sort of the scenario that you mentioned. I think that time scale is going to be somewhat longer than what they think. So I don't think you're going to wake up and all of a sudden like, oh, that, you know, the nation of Denmark, because it has so many smart AIs, it's suddenly competing against China or the US, you know, using the AIs and not the people.

32:21But I think that's not going to happen overnight. Well, I have an essay on fast takeoff coming out tomorrow morning. So for all of those who are listening, you'll all receive exponential view anyway. So you'll see it in your inbox. I also want to touch on your perspective on this word competition or race. Is it competition in the sense of being a zero-sum game? I mean, sometimes it feels that this isn't like the scramble for Africa or the race for oil reserves in Baku and Indonesia and Pennsylvania. Is it really a zero-sum game? No, I don't think it's zero-sum at all. I mean, I think it's very healthy, all these labs competing to build better.

33:02And, you know, assuming you're not, if you're worried about existential risk, it's a slightly different question, right? Because when people are racing, maybe they're less careful about safety. But in terms of just like incremental technological progress or progress in capability of AI, I think the competition is very positive. It's great. And, you know, like, I'm super happy that there are open source models coming mostly from the Chinese side that are almost as capable or just as capable as the top proprietary models. I mean, that's incredible. The amount of like building that's being done by tech entrepreneurs using those open source models is incalculable, actually.

33:39Yeah, it's wild. I spoke to somebody in the medical space a few weeks ago, and they said they swapped over to DeepSeek. They dropped their inference costs by 97%. And it's made a real difference, actually, not just on the cost. It's about all of the other things that you can now do, because it's financially and economically viable to go off and use the AI to do it. And my favorite use of DeepSeek is when I'm on a plane. So I have DeepSeek R1, the I think it's a 7 billion parameter distilled version on my MacBook Pro. And, you know, when you're on a plane, you suddenly don't have your cognitive brain assistant with you.

34:20And now I do. Of course, it's not as good as, you know, Claude's Sonnet 3.7, but it's really better than nothing and actually quite helpful to sort of maintain some degree of productivity and co-working. because I work very collaboratively with the models throughout my day. But this open source shift is really so interesting. And I'm curious about your read on DeepSeek. I've heard a few things. I've heard this is China using open source to undercut huge investments in US industry as if China is this one thing. And I've heard other things like, you know, that Lang Wengfeng believes that open source, open weights, creates a more competitive space where one can get to AGI faster.

35:05And there are other theories running around. What is your read on it? Yeah, before I answer that, let me just mention one other thing about the benefit from open source models, which if you follow the academic research, so professors of computer science or theoretical physics or math or whatever working on AI-related stuff, we went through a period a couple years ago where like the academics were kind of locked out. But because of these open source models and the lower cost of compute, all kinds of really interesting research and papers are being written by academic researchers now. It's kind of a renaissance of, you know, new ideas, new insights into AI that are coming because these academics, instead of having to pay millions of dollars to do some kind of training run, they can use an open source model, they can tweak it, or they can do less expensive training runs.

35:53So all of this is like multiplicative for innovation. Now, in terms of DeepSeek, DeepSeek was really not in any way part of the Chinese government. So if you follow AI in China, they had a list of national champions, which were the usual suspects, you know, like ByteDance, Alibaba, Tencent, Huawei, you know, companies. The kind of things a huge bureaucracy would be able to identify. Right. Companies that were already leading Internet companies that had an AI lab and they were doing their thing. And by the way, one thing that's escaped attention in the West is a lot of the models produced by those companies that I listed are also competitive with open AI.

36:33So it's not just one DeepSeq model coming out of China. It's a handful of models that are at the frontier. But prior to that DeepSeq moment, DeepSeq, if you recall, was a quant hedge fund, which turned its attention to AI research because the founder, Liang, is a true believer. He really wants to build AGI. I have an interview on my podcast, which I'm going to release probably in the next month, with a researcher who used to be at DeepSeek. And his view, knowing all the people there, is that especially Liang is deeply committed to open source. And he just views it as the right way to advance the ecosystem.

37:17And Sam Altman, whom I also know, has also said the same thing. So remember, he says that the closed source is on the wrong side of history. We really want to get to the point where we can release open source stuff. They've now stated their intention to do so. Easy to do so after you've raised$40 billion at a$300 billion valuation. Why not? It pressures off a little bit. But I actually don't think DeepSeek has much to do with the Chinese government. You know, maybe now things will change because they've become such an icon. They're so iconic for the Chinese AI effort. But everything they did up till now really has almost nothing to do with the Chinese government.

37:53Yeah, I mean, I think that's really my main read as well, which is it's come out of left field out of this high frequency hedge fund. He's a very motivated and driven character who understands what matters to him. And who wouldn't be having so much fun doing this kind of research work with great people and now a lot of momentum behind you? People are waiting for the next DeepSeek model to launch and they want to put it to its test against all of the benchmarks. Let's think, though, about how that changes the nature of deployment of AI. The way I look at this is that with these open source, open weight models, which means that you can get access to the sort of internal mechanisms of the model, it enables a real flourishing of different types of models.

38:48and we tend to think when we're using chat gpt of these ai models being highly generalized but in business what you really want is you want a model that is really really excellent and cheap to run at the process you're putting it at that it knows the winners of every olympic marathon since the first one was run is of no use if its job is to do customer onboarding and so there is this new approach, right, where you do this reinforcement learning based fine tuning on smaller and smaller models, and you can push model performance in that narrow domain, perhaps a generation forward, right? You know, notionally, you can take a GPT-3 class model, distill it, make it small, fine tune it, and then it becomes a GPT-4 class model in the domain that you care about.

39:36And And my hypothesis is that that's going to be incredibly commonplace over the next few years. This will be the way that software companies and businesses actually deploy these models because they'll want them to be small. They'll want them to be lightweight. They'll want them to run efficiently because if they run efficiently, it's fewer flops. It's less power. It's ultimately cheaper. And what you end up seeing is rather than these single monolithic powerful models, you see this society, this mesh of models and powering agents interacting with each other. How does that align with your view of how things might develop now that we've seen open source can do what it can do?

40:19I think that that is how things are going to evolve. And our company is, in a sense, a bet on that kind of thing. Now, the previous logic that someone like Sam would have given you, say, a year ago or a year and a half ago, was we're just going to crush you because the next version of our big model is going to be so awesome. All the tweaking and stuff that you did will just be eclipsed by the increase in capability of our big model. And I think that hasn't really turned out to be the case. And I think there will be a broad diversity of models of different sizes applied to different tasks in the economy.

40:59That broad diversity, And these models will be embedded in this series of agents. You know, where we started this discussion with Superfocus were the agents that you were building for BPO and call center. But I'm curious about, you know, whether there's not another flourishing of agents now that we have cheaper open source systems. Back in 2023, I think it was, AutoGPT was this little framework that made its way onto GitHub. And it was the first sort of attempt to let people build somewhat autonomous agents, right, using what was then, I think, GPT 3.5 under ChatGPT that had 100 ,000 GitHub styles in a month.

41:44And back then that mattered a great deal. But I also see a population explosion of the use of specific tools with greater degrees of autonomy in our daily lives. You know, I have workflows, many, many different workflows that I can address directly that trigger off tasks for my research work or just my admin work. And increasingly, you start to string the workflows together. And there's an LLM that's making a decision about where it should go next. And that's starting to look a little bit more, you know, agentic because a longish process can take place. What are you seeing and how do you think that model evolves?

42:23I would say we're at an early stage for that. I don't know of too many people that are, you know, maybe you're one of the outliers in terms of how much productivity gain you're getting from, you know, using LLMs in the way you described. As a company, we're still in the situation where we have to basically exquisitely architect and test the system for the customer because you don't, you never want the AI saying the wrong thing to someone who's called in, who's a loyal customer of the company, right? So we're very worried about tail risk and mistakes that the model makes. And so one of the things that I think the general public doesn't understand is how much rigorous statistical testing is required to know what are the tail risks associated with this deployment of autonomy.

43:07That could be a driverless vehicle. That could be an automated customer service device, you need to understand what the tail risks are. You don't just want to turn this thing on and then discover like, oh, overnight, it got into some bad loop and pissed off 100 ,000 customers, right? You just can't have that, right? Now, you're more fault tolerant because if you're just like, oh, this is just my productivity. So I'll let the model make a few decisions about this and then I come back and I'm like, yeah, it didn't do a good job, just scrap that. In enterprise production, you can't, unless the purpose is research.

43:39If the purpose is like something that really matters to a customer, cannot take those risks. So there's a lot of what we do, which is almost like statistical rigor, like designing a test system, running agents through that test system, characterizing what happens, showing it to our customer. There's a lot of stuff like that, that the general public doesn't think of at all when they think about AI. Yeah, I think that comes back to one of your earlier points in discussion, which is about what deployment speed ends up looking like, because you're absolutely right. I have a particular agent that works, I would say 95 % of the time.

44:17A lot of useful information, including your podcast, Steve, are on YouTube. They run for an hour. I rarely have time to listen to them. I do try to, but I'll fire these podcasts into my agent and it comes back with a presidential daily brief style structured summary. It's incredibly useful. And I would say about once every couple of weeks, it fails for whatever reason. I can live with that. But if you've got an agent in customer service and it's failing one in a hundred times and there's no fallback, that's expensive, it's problematic. And, you know, that's actually just a soft problem, let alone a hard problem in a medical situation.

44:52I think every person who's developing software knows that getting to the first 70 % is trivially easy. It's the climb up the last 30%. And with autonomous vehicles, of course, you know, like the failure mode is like, there's a crash, there's a death, There's a multimillion-dollar lawsuit. And so, you know, that's like an extreme case. But even in customer service, you don't want cases where the AI, you know, pisses off someone. Well, it all adds to that point about the barriers to persuading people who've been used to working with, you know, a particular degree of fault tolerance. I mean, I remember just, we're circling back, but with BPO, when BPO launched as a sort of style of industry, it wasn't really even called that in the late 70s, early 80s, the completion success rate for these human outsourcers was about 30 or 40 percent and that was good enough to get started back then it's not good enough now right because they are so much better so if you're competing with with ai you can't go in and say well this is 25 success rate right you have you have to go higher we're gonna run go on jump in and then we're gonna run short of time i've got one more question for you yeah a byproduct of trying to be rigorous about how well the AI performs means you're also being rigorous about how the humans are performing in the same task.

46:04And you often learn shocking things that the managers are very shocked by, of like how often their agents give the wrong information, human agents give the wrong information to the customer. Yeah. Okay. So we've just opened that and I have to absolutely agree with that. The expectation that we have on the machines is a much, much higher level of tolerance than the expectation we have of people. And I think that that is a vibe that will start to change as we get more and more experience with this. I've used Waymo's many, many times, but I've actually had now two problems with Waymo's. One problem was it was too scared to turn left because it was rush hour traffic.

46:47So my 20-minute journey turned into an hour and a half and several calls to support because the Waymo got scared and couldn't get through the aggressive humans. and the other time was when it locked me in the car and wouldn't let me out, which in and of itself was quite amusing. Have you used the Waymos, Steve? I use the Waymos whenever I'm in the Bay Area and in San Francisco and I've had one, not really bad failure, but I've had failures. And so the failure rate's not that low. It's just, of course, we didn't drive, you know, head on into a concrete wall or anything, but, you know, I've had failures like the ones that you just described.

47:21Yeah, it's still a wonderful, wonderful experience. Okay, we're going to run out of time and I want to be respectful of your time. Super insightful going down from the kind of granularity of being at the forefront of these call centers in the Philippines through to building the technology, trying to understand what's happening with China and what open source means. I knew with you that we would cover many, many dimensions in our conversation. So thank you very much. I have just a question for you. Like in a future that is full of AI agents that are a little bit reliable and working well, What's something that you might end up doing more of and what might you do less of?

48:01Again, I'm gonna tell you something incredibly crazy. You can break the news here. We have been testing. So what do we have? We have a platform that controls the behaviors and the knowledge base of an AI and it can converse naturally with people. We have been researching putting that brain into a small plush toy. So imagine that small dinosaur bunny rabbit that your kid is carrying under her arm. Your kid is two or three years old. And this thing is teaching your child perfect French or how to count or telling its stories, singing little songs. And we find that the systems we've built, kids really enjoy interacting with.

48:41And just imagine the learning possibilities. Like there's a window of neuroplasticity where you can acquire a foreign language at native level fluency when you're quite young. If you just hear it, maybe you talk to your nanny. Well, now every kid can have that. Well, Steve, there is a Simpsons episode on exactly that, which I will look up and send to you. Everyone else, don't watch that Simpsons episode. You should listen to Steve's very rich podcast, which is called Manifold. It is really genuinely one of my favorite shows that I try to listen to, sometimes read summaries of. I am hosting these Friday Live discussions every Friday at 12 p.m.

49:19Eastern Time, 9 a.m pacific time and five o 'clock in the uk thank you everyone for tuning in thank you steve and have a great weekend

From the publisher

Physicist and entrepreneur Steve Hsu, whose startup Superfocus tackles hallucination problems in large language models, joins Azeem to discuss AI agents, hallucination challenges and what happens when technology meets labor markets. 

They discuss: 

(01:31) The deeper shift that Superfocus represents 

(07:00)  Will models overcome hallucination? 

(10:15)  AI Agents can replace 80-90% of call center calls

(12:27)  What it’s like showing customer support AI to customer support people 

(22:36)  China's mayors are like mini CEOs 

(30:05)  What will matter most in the supposed "AI race"? 

(35:58) DeepSeek was not part of the Chinese Government 

(38:23)  How open source will change the future of deployment 

(40:59)  What the public doesn't understand about AI tail risk 

(48:01) How AI plush toys can teach French to 2-year-olds 

This was originally recorded for "Friday with Azeem Azhar", a new show that takes place every Friday at 9am PT and 12pm ET. You can tune in through Exponential View on Substack. 

Produced by supermix.io and EPIIPLUS1 Ltd


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

More from Azeem Azhar's Exponential View

All 44 episodes
The difference between early and late AI adoptersAzeem Azhar's Exponential View · 50 min
Listen in VO