Microsoft CTO Kevin Scott on How Far Scaling Laws Will Extend

9 Jul 2024 · 1 h

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Training Data - Episode with Kevin Scott

Episode Overview Title: Microsoft CTO Kevin Scott on How Far Scaling Laws Will Extend Hosts: Pat Grady and Bill Coughran, Sequoia Capital Description: The episode discusses the implications of scaling laws in AI, particularly regarding large language models (LLMs) and their increasing relevance in technology and society.

Key Themes and Discussions

Introduction

  • Background on Kevin Scott
  • Kevin Scott, CTO of Microsoft, shares his journey from rural Virginia to leading Microsoft’s AI strategy.
  • Emphasizes the importance of being at the right place at the right time.

The Current Landscape of AI

  • Scaling in AI
  • Discusses how the current LLM era is driven by scaling model sizes and computing power.
  • Contrasts the better-than-Moore’s-Law price-to-performance ratios of newer Nvidia GPUs.
  • Challenges Ahead
  • Questions whether the rapid scaling will continue to yield marginal returns or if hyperscalers will end up with excess hardware and insufficient customer use cases.

Microsoft’s AI Strategy

  • AI as a Platform Company
  • Microsoft aims to create a technology platform that enables others to build on top of it.
  • Focus on making AI more accessible and robust by investing in various components of AI technology.
  • Highlights and Lowlights
  • Strengths: Partnering with OpenAI has made powerful AI more accessible to a wider audience.
  • Weaknesses: Acknowledges being late to some foundational AI efforts and investing too broadly without focus.

Future of Computing

  • Shift from Training to Inference
  • Anticipates a significant shift in focus from training large models to optimizing their inference capabilities.
  • Discusses the differences in requirements between training and inference environments.

Business Models for AI

  • Emerging Models
  • The importance of high-quality training data over sheer volume.
  • Envisions new business models for managing and monetizing data used in AI training.

Copilot Technologies

  • Microsoft Copilot
  • Overview of Microsoft's Copilot products designed to assist users in various tasks.
  • Discusses the balance between autonomy and user control in AI assistants, emphasizing the need for trust in AI decisions.

The Value Function and Reasoning Capabilities

  • Challenges in AI Reasoning
  • Discusses the complexities in constructing value functions for broader reasoning tasks beyond zero-sum games.
  • Highlights the limitations of current models and the need for better benchmarks.

Looking Ahead

Optimism for AI

  • Long-term Potential
  • Kevin Scott shares a vision of AI contributing positively to society, addressing healthcare challenges, enhancing education, and solving complex problems.
  • Emphasizes the high cost of not deploying beneficial technology effectively.

Final Thoughts

  • Admiration for Pioneers
  • Kevin admires Ray Solomonoff for his early recognition of the importance of probabilistic methods in AI, showcasing the value of long-term vision and contrarian thinking.

Key Takeaways

  • Scaling Laws: The durability of scaling laws in AI is vital for future advancements.
  • Partnerships Matter: Collaborations, like that of Microsoft and OpenAI, can accelerate access to powerful AI tools.
  • Data Quality Over Quantity: As the AI landscape evolves, the focus will increasingly shift towards the quality of training data.
  • Business Model Innovation: New economic frameworks will emerge surrounding AI data usage.
  • Cognitive Augmentation: The aim is to create assistive AI technologies that enhance human capabilities rather than replace them.

Closing Notes This episode provides insight into the strategic vision behind Microsoft's AI initiatives, the ongoing challenges in scaling AI technologies, and the optimistic outlook for AI's potential to solve pressing societal issues. Kevin Scott embodies a mindset that embraces both skepticism and hope, positioning Microsoft as a key player in shaping the future of AI.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The things that are riddled right now where you're like, oh my god, like this is a little too expensive or it's a little too fragile for me to use. Like all of that gets better. Like it'll get cheaper and like, you know, things will become less fragile. And then like more complicated things will become possible. Like that is the story of each generation in these models as we've scaled up.

0:37On any given day, Microsoft may be the most valuable company in the world, and arguably no one has been more ambitious, more strategic, or more effective in its AI strategy than Microsoft. The key architect behind that strategy? Kevin Scott, CTO of Microsoft. We've had the pleasure of knowing Kevin for a couple of decades now dating back to his time at Google when he overlapped with our partner Bill Court. Bill will join us today for a very special episode of training data. We hope you enjoy.

1:14Kevin, thank you for being here on training data. Now, glad to be here. So just to start, I know you've talked about this before, but for our listeners who might to be familiar with your story. How does a kid from rural Virginia end up becoming the CTO of Microsoft? Who knows?

1:37Certainly not a not a repeatable plan, I don't think. I don't know. It is the thing when I reflect back on my story is it's just a lot of being at the right place at the right time. So I'm 52 years old. So I was, you know, 10, 11, 12 years old when the personal computing revolution started to hit full steam. And so like right at that moment when when you're a kid trying to figure out what you're about and what you got to latch onto. I had this really convenient thing that captured my interest and was a good place for me to ground my curiosity on. And I think that's one of the object lessons in general is if you happen to be interested and really motivated to learn more and do more with something that is at the same time growing really, really quickly.

2:44Like you probably are going to end up in a reasonable place. And so, you know, I was interested in computers. I, yeah, I was the first kid in my, or the first person in my family, not that my mom, nor my dad went to university. So I was the first one to graduate with a bachelor's degree. I majored in computer science and minored in English literature. I had this moment when I was trying to decide what I was going to go do after I got my undergraduate degree where my two advisors were arguing about whether it should be PhD in computer science or PhD and literature and I was very seriously considering both but I was so broke and just so tired of being busted all the time that I picked the pragmatic path.

3:45Not that, like I still imagine what my life would have been like as person with a PhD in English literature and I think it would have been just fine, but like I chose one of my two equal interests. And then, you know, for a while, I thought it was going to be a computer science professor. And at the last minute, this is where Bill and I intersected, like I decided I was a compiler optimization and programming languages person through years and years in grad school. And I got almost all the way to the end. I was like, I don't think I want to be a professor anymore. I would work on these things where it was six months of effort to write a paper and you make some synthetic benchmark 3 % better.

4:37I was like, this doesn't feel to me like the way to have a lot of impact in the world. I don't want to do this over and over and over again for the next 25 years of my life. And so I sent my resume cold into Google in 2003 and I got an email from this guy Craig Neville Manning who had just gone off to New York to open up Google's first just remote engineering office. And like I had an amazing interview at Google. I don't know whether this was on purpose or not or like I just got luck of the draw, but like it seemed like every compiler person who was working at Google was on my interview slate.

5:28And I was like, this is amazing. Like all these people, know all of this stuff that I know and you know, we can have easy conversations I worked on nothing that was even remotely close to compilers at Google and I was confused why all of these people were there. But it was a great interview and I was super stoked and I joined Google. And Google has yet another one of those things, just like the PC and just like the internet, like it was a phenomenon that was growing crazy fast with a bunch of smart people working there and that resulted in this opportunity I had to go join this startup ad mob when it was very early on, like right at this pivotal moment when mobile was taking off and you needed things in mobile like advertising infrastructure and I helped build the seminal company I think in mobile advertising.

6:25thing. And then was back at Google and then I helped LinkedIn go public running its engineering operations team and then I was at LinkedIn when we got acquired by Microsoft. So like, none of that I think you can plant. It's just like a lot of right place at right time and yeah, trying at every point you can to do the most interesting thing you can do on the thing that's growing really fast. You know, when you talk about your personal history, you come in, I guess, you know, the focus nowadays is on AI and machine learning. A lot of the practitioners are people with PhDs. How do you think about sort of practical teams for AI?

7:17Since you're obviously doing a lot of that work at Microsoft and involved with partnerships with OpenAI and others. Yeah, I mean, I think if you are building the really complicated platform pieces of AI, so like the big distributed systems for training and inference, the big networking and silicon and you know system software components or the algorithms that you're using to do training and inference. I think a PhD is super helpful like there's just a huge amount of prior knowledge that you need to have in order to jump into the problem space and be able to like go quickly and like, you know, you need to be clever, but like, you don't, PhD is, I like, yep, I know you have a PhD, Bill and a far cleverer than I am, but like, usually folks with PhDs are clever, but like, they're not the only people in the universe who are clever.

8:26So like, I think it's mostly helpful in the sense that you've gone through like a pretty rigorous training regimen where you get a whole bunch of prior art stuff into your skull and like you, you know, demonstrably can do a very complicated project. And, you know, the PhD projects, you know, look kind of like AI platform systems projects, except the AI platform and systems projects are lots and lots of people working together, whereas, you know, when you're getting your PhD, you often are like working in relative isolation a non -particular thing. So that's one of the things people have to learn is how to get yourself docked into a group and to be able to collaborate effectively with a bunch of other people like yourself.

9:16So useful. But there's so much else in AI that needs to be done other than building the platform. For those things, PhD is helpful, but certainly not necessary. Yeah, like figuring out how do I apply this to education? How do I apply this to healthcare? How do I, like how do I build developer tools around this? How do I do all of the million things that happen when new platforms emerge that you sort of complete the whole platform into like a portfolio products and a portfolio of middleware and a portfolio of like all of the other stuff you need. Well, speaking of which Microsoft seems like it has about the most sort of far reaching or ambitious AI strategy of anybody out there.

10:03Can you just kind of say in a couple words, what is the AI strategy for Microsoft and then just for fun? If you're going to grade yourselves, what have you done particularly well? What have you done? Maybe not as well as you could. Yeah, so I mean, we've been sort of talking about the strategy. Microsoft is a platform company like we, I think have participated or like helped drive a handful of the big platform waves and computing like we were certainly one of the pillar companies in the personal computing revolution like we had a important part to play in the internet revolution although I think that one was a far more or diversely contributed to revolution, then, yeah, then personal computing.

10:51Yeah, we kind of miss the mobile computing revolution. But like each one of those things, like we have thought about, how do you go build a technology platform for this particular era of technology that allows other people to go build on top of that platform to make useful things for other people. And so that is AI strategy. It is like how do you from frontier models to small language models to highly optimized inference infrastructure, you know, like hyper scale on both training and inference, like economies of scale, like making the entire platform more accessible because it's cheaper and more powerful with every turn of the crank.

11:46And like all of the developer tools and safety infrastructure and testing and everything that has to be there in order to have robustly built AI applications, like go build that and like listen to developers and listen to people building AI is as intently as you possibly can so that you are filling in all of the gaps that you can for them as they are encountering problems deploying this technology to users. So that is that is our strategy. And so yeah, I think we're doing a reasonable job of it.

12:33I don't know, like I hate to grade myself. It seems a little bit disingenuous. Right. I like to look like. Well, so maybe before I do that, like, you know, let me describe something about my own psychology. So I, like I am an engineer, and I think most engineers are like short term pessimists, long term optimists. And so the short -term pessimism is like you come in every day and you're like, oh my god, this is like a back of crap Like I just don't like any of this and like everything's broken and like I gotta Like I got I got so much stuff to fix and I'm so frustrated But you work on all of those things anyway because you're Optimistic that all of the problems can be fixed and that they're gonna be worth fixing at the end of the day And so yeah, I mean the There's a bunch of stuff that I think we're doing really well.

13:28I think we have absolutely, along with OpenAI, made a very powerful AI, dramatically more accessible than it otherwise would have been to a larger group of people. I think you, because of that work that we've been doing alongside OpenAI, we're just seeing lots and lots of customers who otherwise wouldn't be building powerful AI applications. And so I feel like we're doing a good job and the way that we're partnering, I think we're doing a good job and having a really particular point of view. And it's not an immutable point of view, but it's a point of view about what an AI platform to look like and we're trying to like make it as complete as we can.

14:18You know, low lights is I think we were a little bit late to like some of the basic AI stuff. So it wasn't that we were not investing in AI at all. And like you can sort of look at some of the work that Microsoft research had done over the years. And like MSR was an early AI leader. And I think, you know, the, you know, bit Bill knows this just as well as I did just from his time at Google and, you know, where we overlap for a number of years. But many, you know, maybe most of the really important advancements in AI over the past 20 years have been a function of some kind of scale. And it's usually you got data scale and compute scaling combination.

15:17Let you do things that weren't possible at lower scale points. And at some point, that scaling of data and compute is so exponential that you get past the point where you can have fragmented bets, where you can literally just become economically impossible to bet on 10 different things that are all exponentially scaling or have the ambition of the need to exponentially scale simultaneously. And so I think one of the things that we were a little bit late to is like we didn't put all of our eggs into the right basket soon enough. Like we just, yeah, we were spending a lot on AI, but it was fragmented across a whole bunch of different things.

16:03And because we didn't want to hurt any of the feelings of smart people or, yeah, whatever, I don't even know what the diagnosis was because a lot of that was before I was at Microsoft. We just weren't as quick as we should have been at like saying, no scale is what matters. And like here's how we're going to focus our investments on scale in a principal way. When did you get religion that scales what matters? Was there a particular event or a moment that really crystallized that for you? So, yeah, I mean, I was, so I've been at Microsoft for about seven and a half years now and like, my, when I, when I became CTO, my job was like take a scan left to right across both Microsoft and the entire industry and try to see where, yeah, we just had holes in execution where we were not doing things at that point in time, which was, I guess, 2017 early, where, all right, what are we not doing today that we're gonna deeply regret in 2019 or 2020?

17:23So like two, three years out. And like the biggest thing on the list was like, you know, our rate of progress on AI was not fast enough. So I'd say mid 2017, like I had religion that that was gonna be a big part of my job was like helping us figure out what the strategy was gonna be. And then in 2018, like if anything, the of the publication of the birth paper from Google was like a real crystallization of that belief. So like everything that I had that was in my analysis, I was like, this is as fine an example as anything of like why we have to really, really accelerate on getting more serious here.

18:15And so like, very shortly after that, like I restructured a whole bunch of stuff inside of Microsoft to get us more focused on AI. And then about a year later, we did that first deal with OpenAI. And yeah, we have been accelerating our investments and like trying to get more focused, more crisp, more purposeful since them. When you were very early to appreciating the potential of OpenAI, what did you see in and then at that time when that first partnership was true. Well, we had, like, or at least I had, this, like, real belief that what was happening with these models as they scaled, as they actually became a basis for building a platform that, like, the big shift that was happening is it wasn't just, you know, I ran one of these teams at Google where you had a pool of data and a bunch of machines and an algorithm and you were like training a model and like the model was for a specific thing like in the other case of the thing that I was doing at Google it was like click through rate prediction for advertising and you know, like a handful of other things and like just outrageously effective, right?

19:35But most of the work before this, before GPT was about those sort of narrow use cases. Like you were purpose building models for narrow things, and it was just tough to scale. Like you couldn't, you'd invest a bunch of compute and like you couldn't amortize the cost of the compute across anything more than just the narrow thing that you were building the model for. And yeah, I'd have a lot of expertise that you know, if you wanted to replicate all of this, It's like you had to have different data and different AI PhDs and different processes every time you wanted to go build AI into an application.

20:21What was happening was, yeah, you had these big, large language models that were useful for lots of different things. So you didn't have to have a separate model for machine translation and sentiment analysis. and all of the different text things that you were doing. And I was like, okay, this is extraordinary. And they were also becoming more platform like as a function of scale. So transfer learning was working better as things scaled up. And this is still the general pattern. So like everything that we understand that large language models can do plus or minus, like we'll get better when you get to the next scale point.

21:08And on top of that, they will become slightly or maybe dramatically more general in the sense that their capability set broadens. And OpenAI had that same belief, and they also had a very principled analysis So like how the those platform characteristics emerged over time is a function of scale at a bunch of experimental validation that said that their forecast were right and so it's like it just you know You sort of like look at what the forecast says and you like this is how much money it costs to run the experiment to see if you're gonna be on forecast for the next turn of the crank and like, you know, it was felt like a big number at the time, like it was billion dollars.

21:59But like relative to what was happening, it just wasn't a large amount of money. And then, yeah, GPD -3 was on forecast and GPD -4 was on, so like it just was, like finding a partner that had the same platform belief that you did and like a track record of being able to execute through these scale points. Like it was... It didn't... Like I've done a bunch of things before that I have way more reservation about in the past, like just in terms of investments. Like this one didn't... You know, like there were a bunch of people who didn't agree with me, but like I had pretty high conviction. You touch on investment, I guess.

22:44You know, there's a lot of trade publications now speculating about the cost of doing training and so forth and you know, rumors of billions and billions of dollars being spent and so forth. And I guess based on my own background, I think training is going to get dwarfed by inference here very soon. How do you see the nope so? Well, yeah, otherwise we're building models that nobody knows what to do with. That might not be a great investment. How do you see kind of computing landscape involved and where's it going? I think people are joking that all the money is going to Nvidia at the moment. Well, look, I think Nvidia is doing a good job.

23:31So like the two interesting things that are happening with these models just in terms of the efficiency of the scale up is each hardware generation's better price performance wise, usually by an extent greater than Moore's law used to work for general purpose computing. So, A100 was about three and a half times better price performance than V100, H100, like not quite that much, but close on paper. The next generation and looks very good as well. And so like you've got hardware for a variety of reasons, like you know, part of its process technology but part of its architecture and like a lot of it is like being able to leverage narrower word size in the computation.

24:34And so like, you know, instead of needing 64 bit arithmetic, like you're, you know, you're doing arithmetic with much less precision right now. And so like that, you know, there's just an embarrassing amount of parallelism there, then like we're getting better and better at extracting that architecturally in the hardware. And, um, yeah, there's a bunch of innovative stuff happening with networking as well. Like, we're well past the point for the frontier models at least where you can do anything interesting on a single GPU. So for years and years now, like both training and inference have been multi -GPU, multi -compute node problems.

25:13And so like there's a bunch of innovation happening on the network side as well, which allows you to strap all of the compute together, like at the chassis level, the rack level, the row level, the data center level, more effectively, which is great because, you know, for the nerds listening, like we haven't had effective power scaling or denards scaling since 2012 or so. So, yeah, we're not, we're getting more transistors but like they are not, they're not getting cooler. Yeah, and like we just, we have a lot of density issues just with power dissipation that we have to go deal with. Do you see inferences driving different data center architecture?

26:11Yeah, I mean, look, we already architect our training environments and our inference environments differently. they just need different things. And like I think, you know, all the way down to, you know, Silicon and like through the network hierarchy, you need different things for inference. And like, inference is, inference is kind of easier than training, like training the way that we're doing it now is like we go build big environments that take, you know, a few years to build.

26:45And yeah, with inference, if somebody came along with a better Silicon architecture, a better network architecture, like a better cooling technology, it's a much easier experiment to go run. You just go swap some racks out. I mean, like my data center people would yell at me, like it's not quite that easy, but it is easier than having to go do a big capital project. like a training environment looks like. And so intuitively you would think that that is going to result in more diversity in the inferencing environments and more competition and more like a faster rate of improvement. Like that's, and on the software side that's certainly what we see.

27:35Like the inference stack just because it's such a large fraction of the overall compute print and it's constrained because we have more demand and supply at the moment. Like you just have very, very powerful incentives to go optimize the software stack to squeeze more performance out of it. You think we'll be in an environment anytime soon where that demand supply balance changes? Not necessarily at Microsoft, but it feels like we're seeing that at the market level as well. Yeah, I don't know. I mean, if we continue to see the platform continue to expand capability -wise, and it just becomes more useful, like I think demand increases if anything.

28:26Now, the shape of the demand is probably going to move around. Like, I think you're already seeing a little bit of that. building a frontier model is like a very, very resource intensive thing. And as long as people are building frontier models and making them accessible, and maybe they're not accessible quite the way that people want. like they're only API accessible and like there isn't an open source thing, you can go instantiate and muck around with. But it's way more accessible than it was six or seven years ago where the only way to access some of this stuff is you had to go work for two or three tech companies.

29:11And so anyway, but I think you do have to ask yourself, and somebody else should do the asking right, because I'm like all kinds of bias, right? But I don't know how many frontier models you actually need. If they're all roughly speaking in the same tier of capability. Yeah, that's an awful lot of money to spend for things that are roughly equivalent. It's sort of like, if you're starting a company right now and you believe that you have to build your very own frontier model in order to go deliver an application to someone, That's almost the same thing is saying like I gotta go build my own smartphone Hardware and operating system in order to deliver this mobile app Like maybe you need it, but like probably you don't and so I mean to build point like I think the thing that Make sense, you know for the market is like you just you you're gonna want to see lots of people doing lots of inference because that means you've got lots of products that have found product market fit and those things are scaling.

30:32But like lots of speculative dollars flowing into infrastructure R &D, like probably ends the same way that the, you know, many speculative infrastructure blooms who ended. on this, excuse me, on the scaling front, Microsoft published a paper sometime ago pointing out that the quality of training data is maybe at least as important as volume. And I think one of the speculations you see now in the industry is that we're running out of sources of high quality training data and you're reading at least some articles claiming that various partnerships are being struck to get access to training data and they might be behind paywalls and so forth.

31:30How do you see that evolving? Because it feels like we have more and more computation but we may not have more and more training data. Yeah. Yeah, I mean, I think that was almost inevitable. It is, in my opinion, a good thing that quality of data matters more than quantity of data because of give or you an economic framework to go do the partnerships that you need to go do to make sure that you're feeding your AI training algorithm a curriculum that is going to result in smarter models. And like, honestly, not wasting a whole bunch of compute feeding at a bunch of things that are not. And I think from an infrastructure perspective, one of the things people have been very confused about is like a large language model is not a database, it's not a repository of facts.

32:28It's like important for it to quote unquote, well, knows some factual things, but it is like the world's crappiest database, honestly, like if you need it to be your retrieval engine. And so like you just shouldn't think about it as, like, hey, I got this thing and like it has to have everything baked into it, like into the model weights themselves so that you can recall a bunch of stuff. We got, you know, like as you've seen, like the recall is imprecise in the same way that human recall is imprecise. So they're just much more efficient ways to do recall. Yeah, so look, I think you are, I mean, the way that we see things developing is like you have data that is valuable for training models and then you have data that you need to have access us to for an application in order for the model reason over.

33:31And like those are two different things. And I think they're probably two different business models around those things. So at the end of the day, this is all just, this is about business models, right? Like people who produce data want to be compensated for use of that data. And so, yeah, we have all of this data sitting inside a search engines right now, like not in randomized weights, but quite explicitly it's like sitting in indices and being in Google and whatnot, just waiting to be retrieved. And plus or minus, everybody's okay with that because there's a business model there that makes sense.

34:22So you enter a query and you're either sending traffic or there's SEO and advertising a whole bunch of business model that surrounds that. I think we'll figure out a business model for that referral data so that when an agent or an AI application needs to retrieve some information from someone so that it can reason over it and give the user, an answer, like we will figure out the business model for that. Like it'll either be subscription revshare, it'll be licensing, it'll be like some new flavor of advertising. Yeah, I was just telling someone the other day, like if I was in my 20s right now, like, you know, for all of your entrepreneur, like somebody ought to be out right now figuring out what the new ad unit is for agents and like just building the company.

Read the full transcript

35:18because it will have the same characteristics and qualities as previous ad units. You have people with information and products and services who are going to want to get to the attention of someone who might want those data and products and services and quality is going to matter and relevance is going to matter and a bunch of other things. and I would be shocked if there is an auction model for, you know, that's going to be the right way to value everything. And, you know, like, yeah, maybe there's referrals and like referrals will have some economic value. I think retraining is just a little bit different because it's very, very hard when you're building that model at the time that you're doing in the building to really be able to ascribe a monetary value to a particular token of input.

36:23Just because it's contribution to the model like the same way that a word from Mobey Dick has like a very diffuse contribution to like your own human intelligence, even though you definitely read it at some point in your career or your life, like how valuable that is to like forming you know Bill Koran or Kevin Scott's you know useful intelligence like who knows. Speaking of which this one of the things that we hear a lot of times is the value function isn't some ways the bottleneck to broader reasoning capabilities. It's easy enough to construct a value function when you're playing a game with a you know, with a known winner and loser like Go or Chess or poker or diplomacy, but it becomes a lot harder to construct a value function when you're going into broader domains and it's things like assigning the value of Moby Dick to Kevin Scott's life, you know, that sort of thing.

37:25Are there practical solutions to this? Are there practical implications to this? I guess the broader question would be where do you see the overall field of reasoning and LLM's going? Well, look, I think people are trying to get at this. So you've got a bunch of benchmarks like GPUA and MMLU and, you know, like we're just sort of rolling through a bunch of benchmarking paradigms that, like, try to come up with scores of performance for these models, like whether it's reasoning capability or, yeah, And yeah, I think we have one of the interesting things we've seen over the past handful of years, is like we just are very quickly saturating these benchmarks, like where you like one emerges and then, yeah, within a model generation, like you'll completely, or like get very close to saturating the particular benchmark, and then you gotta go find something else to help be your guiding light.

38:24And so, but like let's just sort of assume that like, you know, you'll have some interesting benchmarks that are correlated with the reasoning capabilities you want models to have. Then the question is just an expensive experiment to run. Like you can run an experiment where it's like, okay, I'm gonna train a model with this information in and out and like does it get better or worse at performance on these reasoning benchmarks. And, you know, like I think all of us have done different versions of those experiments. Like they're just extraordinarily expensive to run at the most granular scale.

39:07You can imagine running them. Part of that paper that Bill referenced, like textbooks are all you need is like a... It's not the full story, but it's like part of a story that is like, you know, just sort of evaluating like token contribution quality to a model performance. So yeah, I think it's, and everybody's got every incentive in the world right now to try to figure out what that is. If for no other reason than you, like in a world where you're synthesizing data, you're literally spending compute to generate synthetic tokens for training, you really want to make sure that the tokens that you're generating are actually useful or not?

39:57Where do you think the models are at the moment? I think Microsoft has introduced a whole bunch of co -pilelets to try to help end users with your products and so forth. On the other hand, I see lots and lots of companies trying to build agents that can be kind of autonomous actors now that that's a wide spectrum of kind of expected performance of what the models can do. Where do you think we are? Where do you think we'll be in a couple of years? Well, yeah, so I think there's a super good question and there's a, you know, there's even a philosophical thing there about, you know, what it is we should want.

40:38And whether is the the spectre of everybody's job getting replaced, right? So yeah. Yeah, you know, somebody, like that we, we chose the, the name co pilot for the things that we were doing relatively deliberately because we, we want to, at the very least, encourage everyone who's building these things inside of Microsoft to think about how can I help augment someone who's doing some form of cognitive work. So like we want to build, you know, assistive, not substitutive tech. And, you know, the good news is also easier to think about how to go from, you know, sort of rough frontier model capability to useful tool when you're narrowing it down to a domain.

41:36And so like I think that's been a reasonable deployment path. And like we've got a handful of our co -pilots that have real market traction right now and are like in daily use by like a lot of people doing real non -trivial cognitive work. And I think that will expand over time. What are just on that? What are some examples of co -pilot set of really like hit the bulls eye already versus maybe co -pilot for the technology is not quite ready? In terms of like jobs to be done. Yeah, like I think you know get a get up co -pilot is like probably the you know the thing we've talked most about and like there's you know the most public conversation around.

42:26It's been a hit. It is genuinely useful. We've got some other co -pilots that are like that, that are super useful. But I think the thing that Bill was getting at, the more general the co -pilot is, like the, you know, the harder it is to have it actually take a very high precision action on your behalf autonomously. You know, particularly if, you know, it's doing something where it's representing you, where, you know, there are stakes and, you know, consequences and accountability back to you if this agent makes a mistake. And we're trying to be very deliberate there because I think one of the things that you don't want to do is introduce a thing that's going to make a whole bunch of these sort of errors where the user's first reaction is like this doesn't work and like I'm not going to try it again for a good long while.

43:40So we'd rather have it be very good before we introduce it, which again, you have means you're optimizing for use cases, not for super, super broad things. I mean, there's a good, we did a partnership recently with Devon, which I think is another one of these very interesting, use case specific things where it's frontier model, plus a whole bunch of other stuff that is optimized for giving humans high quality recommendations for actions that they can take. Then when you click accept on the action, you have reasonably high confidence that it's going to work and you haven't made another set of problems for yourself.

44:35So I'm guessing, too, you all see this in your portfolio companies. There seem to be a bunch of companies out there right now that are doing exactly this and that it's useful and working. Well, it's interesting because we hear one of the things that we hear fairly consistently from the companies that are further on in their AI journey. You know, everybody kind of starts in the same way where they start playing with OpenAI and then maybe they start using some of the other proprietary foundation models that incorporate some open source models and maybe they have some of their own stuff. There's a vector database in there somewhere.

45:07Yep. From an architectural standpoint, it feels like people tend to go on a not quite the same journey, but journeys that sort of rhyme. But then what we hear from them when they're 12 or 18 months down the road is there's kind of this massive 80 -20 rule at play and maybe it's a 92 -rule, but you can automate most of a task pretty quickly and pretty effectively, but getting it to the point where it's actually end to end running autonomously in a way that is compelling and consistent enough that you can actually trust it. Kind of that last mile, that last couple percent that makes you really trust it.

45:44Yeah, it seems like that's been pretty elusive for a lot of tasks. And so one of the things that we're really curious about is, okay, well, when do the foundation models themselves, you know, get good enough to knock out that last 2 % or is that a domain specific thing? And that's really the job of the software vendor who lives on top of the platform to figure out that last 2%. I look, I think it's going to be both for a while. Like the two things that I think you can trust. Yeah, I know you guys are probably going to ask this question at some point, but

46:21We're not diminishing marginal returns on scale up. I try to help people understand, there is an exponential here. The unfortunate thing is you only get to sample it every couple of years because it just takes a while to build supercomputers and then to train models on top of them. And so the next sample is coming. And like, I can't tell you when and I can't predict exactly how good it's going to be, but it will almost certainly be better at like the things that are brittle right now where you're like, oh my god, like this is a little too expensive or it's a little too fragile for me to use. Like all of that gets better.

47:07It will get cheaper and things will become less fragile. And then more complicated things will become possible. That is the story of each generation in these models as we've scaled up. And so we even think about this inside of Microsoft and one of the category errors that our own developers who are building these AI products can make is get two convinced that the The only way to solve my problem is I have to go take the current frontier and supplement it with a whole bunch of things. But what you do have to do, but you want to be very careful architecturally when you're doing that, that it doesn't prevent you from taking the next sample when it arrives.

47:50So you just want to architect these applications where when the new goodness comes, you can go plug it in. And you'll have to go optimize that as well. like I think that's just sort of the grind that we're all on. But like you just like the thing that was killing us internally is I would have teams inside of the company who would look at a frontier model and say, oh my god, there's no way that we can ever deploy a product on top of this because like this is fragile and this is too expensive. And so, you know, like please give me like giant pools of GPUs and like let me go spin up a big team doing like a very tailored version of this and we're going to build a specific model.

48:33Yeah, they would go off and spend a whole bunch of money and the thing would be a little bit better cost wise at the same level of performance as the current frontier. And then the frontier would snap to the new point and it would be just doomed. and so like you just architecturally don't want to get trapped by that, I don't think. I mean like that'd be my advice I give to everyone is like just give yourself the flexibility to snap to the new, to the new frontier when it emerges. And like that lets you preserve all of your skepticism. You can believe all you want that the new frontier's not coming.

49:15and like go, you know, like, you know, read your favorite Twitter troll that says it's all over on a sham and like, but, but just give yourself the option that, you know, it may, maybe what's been happening for six years now is going to continue. Well, hearing, hearing that, hearing that we are not at diminishing returns to scale, I'm going to count that as good news. And so let's, you stay on the theme of good news. I know So that your short -term pessimist, long -term optimist, can you give us some of the optimistic point of view for where we're heading in this world of AI? What are some of the things that you're most excited to see in the world in five or ten or fifteen years or whatever you count as a longer term horizon?

49:58Well, look, I think the thing that everybody ought to spend some time thinking about is where are the gnarliest zero -sum problems that we have in society? like where are the things where like we just are fighting with one another or you know like we are emisorating people because Whatever it is that people need there doesn't appear to be enough of it Uh, and I think for a good number of those things like what you have to have to to turn them into non -zero -sum games to Yeah to create abundance and to relax some of these constraints is you have to have technological breakthroughs. It's like the only thing reliably that's ever turned zero some to non -zero some in human history is like some tech has to come along that lets us have more.

50:56You know, whenever tech comes along and creates more, it doesn't mean that the more gets equitably and uniformly distributed. And like I think they're real conversations to go have about that. But like what you do want is the more. And you want it to be directed at things where like we're just having a tough time right now. Yeah, I'll tell this story that I've told a couple of other times recently. But you know, my mom, like I grew up in rural central Virginia, my mom 74 years old, and like she's suffered from this thyroid condition called Graves disease for 26 years. And so, you know, when you have Graves disease, your thyroid is hyperactive, you're generating too much thyroid hormone.

51:47And so they go in and like irradiate your thyroid gland to reduce its activity level. And then you take hormone replacement therapy to like upregulate your hormones for the rest of your life. And so she was having some blood pressure issues and her doctor, you know, dorked around with her dosage of this hormone medicine. And then like she just had like some serious health issues as a consequence of that that landed her in the ER and you know, rural central Virginia, like six times, like in a pretty, you know, short number of weeks. And yeah, the interesting thing there was the first time she went to the ER like she was presenting all of these cardiac symptoms and it was pretty clear they hadn't even read her chart.

52:34Like it hadn't registered on the that she had graves disease and like the thing that they needed to do was like go water a TSH panel to see what her like thyroid hormone levels were. Like if they'd done that right away, like they would have said, okay, like we gotta go, you know, adjust your medicine. Yeah, and like I'm not even ascribing ill intent, like, you know, this is a healthcare system that is egregiously overburdened. Like this is not a place where, you know, like you've got this influx of, you know, talent, and coming into this part of the country. And it's a lovely part of the country.

53:19I love it. So I'm not criticizing anything. It's just they have an aging population and they don't have enough young people coming in to be things like doctors to help this healthcare system keep up with all of the challenges that they have. If some of those doctors had access to GPT -4 and like it was an approved product use. All they would have needed to do was put the symptoms that she was presenting in and her medical record. And it would have said, hey, she needs a TSH test. And if you would put the TSH test resulted, the recommendation it would have was like, look at the dosage of her Harma replacement therapy that she's on.

54:11And like if like this isn't theoretical like I did this Like it could have helped alleviate a massive amount of her suffering right now and like I think the only reason like that she Got out of this tough situation that she was in was I Had to intervene like I I've sent her to a specialist that was you know 400 miles away that she couldn't have gotten into without, you know, special, you know, and like, it's ridiculous. Like, yeah, there are so many 74 -year -old, old Southern ladies or old Midwestern ladies or like folks who are going through similar sorts of things who do not have someone who's going to go in and intervene on their behalf who are suffering unnecessarily because we're not even deploying the technology that we've got right now and like it's just going to get better.

55:07And so like that's the thing that I'm excited about. It's like let's let's go let's go give kids a leg up in education. Let's go fix some of these crazy problems. We've got in a healthcare system that is just you know absent technological intervention is just going to get more strained over time. Um you know let's equip our scientists with better tools so that they can find better carbon capture catalyst so that we can design safer modes of transportation so we can more quickly get to a post -carbon economy. It's just so many things we can go do with this stuff. I'm super, super optimistic about it.

55:54And so, you know, like it just kills me, like,

56:01what we don't wanna do is get distracted on, you know, like, things that are, just, you know, effectively noise in the ecosystem right now. And, you know, we're getting so sideways sometimes with this model said something that hurt my feelings. And I don't want to disment people's feelings matter. And I'm not trying to be a jerk here. But I do want to make sure that as we're thinking about how we develop and deploy the technology that we are always remembering like what the cost of not deploying the good is, because that is a high, high cost. Mm -hmm. Very well put. Yeah. Yeah, I'm here, I'm here.

57:01Son, probably a good note to end on, I think. Well, we have one other question that we like to ask people, and it's a quick one. Some of them asked it, and I'm gonna ask you this one. Who do you admire most in the world of AI? Yeah, you know, I was thinking about this. So I think it's race alumina off, who was one of the folks who was at that Dartmouth workshop in the 50s where Marvin Minsky and Simon and a whole bunch of folks convened that summer and like they were all interested in machine intelligence and they coined the term artificial intelligence at that workshop. And like the reason that Salamunov is so interesting and like not many people I think know who he is outside of computer science is like he was the one from the very beginning who was pushing on this whole notion that probabilistic methods were going to be very important for or the development of AI.

58:14And when I was in grad school in the 90s,

58:24prevailing academic theories about how we were gonna get to AI were all about like, okay, well, like there's some magic, minimalist calculus about human intelligence and like we've just got to figure it out and it's got to be, you know, rule -based systems and ontologies and, you know, symbolic reasoning and like a bunch of like stuff where, you know, we like we do in physics, we were going to have to like, you know, divine, you know, the inherent simplicity in the system, like figure out what the rules are and as soon as we understand the rules, like we'd be able to make software emulate human intelligence.

59:02And Slominoff is like, no, no, no, like this, this is just a intelligence is an extraordinarily complicated phenomenon and like the only way that we're ever going to really get there is modeling it with probabilistic methods and he was right and like he was, he was judge wrong for a very long time and so I really admire his contrariness. Yeah, all the way back in the 1950s, he's in like, he stuck with his beliefs, his entire career. And I don't know whether Ray actually like lived to see how right he actually was. Hmm. That's a great answer and a great story. Thank you, Gavin. Yep, you're very welcome.

59:55Thank you.

From the publisher

The current LLM era is the result of scaling the size of models in successive waves (and the compute to train them). It is also the result of better-than-Moore’s-Law price vs performance ratios in each new generation of Nvidia GPUs. The largest platform companies are continuing to invest in scaling as the prime driver of AI innovation.

Are they right, or will marginal returns level off soon, leaving hyperscalers with too much hardware and too few customer use cases? To find out, we talk to Microsoft CTO Kevin Scott who has led their AI strategy for the past seven years. Scott describes himself as a “short-term pessimist, long-term optimist” and he sees the scaling trend as durable for the industry and critical for the establishment of Microsoft’s AI platform.

Scott believes there will be a shift across the compute ecosystem from training to inference as the frontier models continue to improve, serving wider and more reliable use cases. He also discusses the coming business models for training data, and even what ad units might look like for autonomous agents.

Hosted by: Pat Grady and Bill Coughran, Sequoia Capital

Mentioned:
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, the 2018 Google paper that convinced Kevin that Microsoft wasn’t moving fast enough on AI. 
Dennard scaling: The scaling law that describes the proportional relationship between transistor size and power use; has not held since 2012 and is often confused with Moore’s Law.
Textbooks Are All You Need: Microsoft paper that introduces a new large language model for code, phi-1, that achieves smaller size by using higher quality “textbook” data.
GPQA and MMLU: Benchmarks for reasoning
Copilot: Microsoft product line of GPT consumer assistants from general productivity to design, vacation planning, cooking and fitness.
Devin: Autonomous AI code agent from Cognition Labs that Microsoft recently announced a partnership with.
Ray Solomonoff: Participant in the 1956 Dartmouth Summer Research Project on Artificial Intelligence that named the field; Kevin admires his prescience about the importance of probabilistic methods decades before anyone else.

00:00 - Introduction
01:20 - Kevin’s backstory
06:56 - The role of PhDs in AI engineering
09:56 - Microsoft’s AI strategy
12:40 - Highlights and lowlights
16:28 - Accelerating investments
18:38 - The OpenAI partnership
22:46 - Soon inference will dwarf training
27:56 - Will the demand/supply balance change?
30:51 - Business models for data
36:54 - The value function
39:58 - Copilots
44:47 - The 98/2 rule
49:34 - Solving zero-sum games
57:13 - Lightning round

More from Training Data

All 110 episodes
Microsoft CTO Kevin Scott on How Far Scaling Laws Will ExtendTraining Data · 1 h
Listen in VO