With Groq, Jonathan Ross is taking AI inference to new speeds

9 Apr 2025 · 35 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: With Groq, Jonathan Ross is Taking AI Inference to New Speeds

Podcast Overview Title: Pioneers of AI Host: Rana el Kaliouby Description: The podcast explores how artificial intelligence (AI) is transforming our lives, featuring discussions with leading figures in technology.

Episode Description In this episode, Jonathan Ross, founder of Groq, discusses the development of AI chips focused on fast inference. The conversation emphasizes the benefits of faster processing, including reduced energy costs, increased job creation, and the democratization of AI access.

---

Key Themes and Concepts

  1. The Importance of Fast Inference
  2. Definition of Inference: The process by which AI models use learned information to make decisions.
  3. Speed's Impact: Faster inference allows for rapid responses from AI systems, analogous to the transition from dial-up to broadband internet.
  1. Grok's Innovations
  2. Low Latency Chips: Groq designs AI chips (LPUs) that focus on reducing latency, thus improving inference speed.
  3. Unique Architecture: Groq's architecture avoids external memory, using an assembly line approach with multiple chips to enhance efficiency and speed.
  1. Market Dynamics
  2. AI Chip Competition: The AI chip market is competitive, with Groq positioning itself not as a direct competitor to NVIDIA but rather focusing on inference rather than training.
  3. Large-scale Deployments: Groq has made significant strides, including a $1.5 billion deal with Saudi Arabia, enabling rapid deployment of chips.
  1. Future of Work and AI
  2. Job Creation: Ross proposes that AI will create new job opportunities, mentioning the emerging role of "prompt engineers."
  3. Democratization of AI: Emphasizes the need for equitable access to AI technology, aiming to provide it to a broader audience, including startups.
  1. Energy Efficiency
  2. Sustainability Concerns: Groq's LPUs are designed to be more energy-efficient than traditional GPUs, reducing overall energy consumption in AI deployments.
  1. Open Source vs. Proprietary Models
  2. Shift Towards Open Source: The release of models like DeepSeek has accelerated the trend towards open sourcing AI, fostering innovation and broadening accessibility.
  3. Data Privacy: Concerns regarding where data is processed and the implications of using models developed by foreign companies, particularly in relation to privacy and governance.

---

Key Takeaways

  • Speed Matters: Faster inference directly correlates with better user experiences and lower operational costs.
  • Grok's Strategic Positioning: Groq emphasizes a focus on inference, leaving training to companies like NVIDIA, thereby creating a niche in the market.
  • Access to AI: Ross argues for making AI technology accessible to everyone to prevent a monopolization of compute power.
  • Changing Landscape of Employment: Automation through AI is expected to redefine job roles and create new opportunities.
  • Sustainability in AI: Energy consumption in AI is a critical issue, and Groq aims to address this through innovative chip design.

---

Conclusion This episode of *Pioneers of AI* offers deep insights into the rapidly evolving AI chip landscape, the critical importance of inference speed, and the broader implications of AI technology on society and the economy. Through Groq's innovations, Jonathan Ross exemplifies how AI can be made more accessible and efficient, promoting both technological advancement and societal benefits.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Starting a business comes with its share of ups and downs, which is why staying true to your vision is essential. a non-negotiable for Romeo and Milka Bregali, Capital One business customers and co-owners of Ra's plant-based restaurant in New York. Romeo and Milka took a leap of faith when starting their own restaurant, gutting an empty space and building it from the ground up, every pipe, every wall, every detail. But building from scratch came with a heavy financial burden, which is when they turned to their Capital One business card. With the flexibility of the card's no preset spending limit, they were able to spend more and earn more rewards while bringing their vision to life.

0:36Today, Roz's success is proof that with passion and the right support, it's possible to make your dreams a reality. Learn more at CapitalOne.com slash business cards. I want you to imagine walking up to a computer and typing in a Google search and waiting a minute for an answer. You would just give up and go on and do other things, right? Once you get speed, you can never go back. It's sort of like getting broadband. You never want to go back to dial up. That's Jonathan Ross, founder of Grok, one of the rising AI chip makers out there. And Jonathan says that the reason why people want fast internet is the same reason why people want fast AI.

1:16And so right now, everyone is fighting for market share in AI. And if you have a high latency version of a product, you are competing with one hand tied behind your back. Grok, with a Q not a K, so not to be confused with X's AI chatbot, is not one of those companies. They make low latency a priority. Their AI chips, called LPUs, focus on inference. This is when AI takes all the information it's been trained on to make a decision. So basically, that answer ChatGPT just gave you on what to cook for dinner, that's inference. Grok is taking their claim on fast inference, and as a multi-billion dollar company, their bets are paying off.

2:01Today I'm talking with Jonathan about how his chips provide faster inference while saving energy, why AI will create more jobs, and the importance of democratizing access to AI.

2:17I'm Rana El-Khalyubi, and this is Pioneers of AI. A podcast taking you behind the scenes of the AI revolution.

2:33Jonathan, welcome back to Pioneers of AI. We last saw each other in December at the Fortune Brainstorm AI conference, which was an awesome conversation. So I am so glad to have you back on. Thanks for having me. I appreciate it. So to kick us off, I want to take us all the way back to when you started Grok. Before that, you were at Google. And I'm just always curious to hear people's kind of founding stories. And what was the tipping point that led you to start Grok? So you want to know what my radioactive spider bite is? Yeah. So I started the Google TPU. That's the AI chip that Google uses. And I did it as a side project.

3:12It started to become known that Google had made an AI chip. We actually had it in production for about a year before it was ever announced. And so I started getting calls from other companies that also wanted to build AI chips. And one of them made a pitch to me, which was, come join them because Google had a tremendous advantage in having their own AI chip. At the time, it was 10 times faster than any GPU. And make sure that there isn't a concentration of AI in one place. Help us make sure that there are two players. And I was thinking, gosh, that's a pretty good argument. But one versus two, that's not a huge improvement.

3:48I want to make sure that everyone gets access. So that was in the back of my mind, but I still didn't intend to do a hardware startup because, frankly, it's hard. And then I realized, well, gosh, people are giving away the models for free. They're giving away the frameworks like TensorFlow for free. I didn't feel that software was going to be a great moat. So as I was going out and talking to investors, one of them asked, well, what would you do differently? And I said, two things. Number one, I would start with the software. The software is a mess on all of these AI chips, even TPU. It's very difficult.

4:25As in it's like not optimized for these chips? It's optimized for it, but you have to optimize it by hand. It takes a long time. But the other realization was everyone knows about Moore's law and that the number of transistors doubles every 18 to 24 months. What we realized was actually the number of chips was also doubling every 18 to 24 months. And so if you were doubling the number of chips, that meant that effectively you could act as if the number of chips was infinite. And if the number of chips was infinite, how would you design differently? And so we came up with a very different architecture that doesn't even use external memory.

5:05We just use a larger number of chips, and we lay out the models and the computational problems across those chips. And it reduces the cost. It improves the speed. It's a very different architecture than anything else that exists. Yeah, so the AI chip market looks a lot different today than it did when you started Grok. And so much has happened, including, of course, the generative AI explosion. And there are giants like NVIDIA, but there's also competitors like you and other companies. We actually had Gavin Oberti, the CEO of Etched, on our show. They do specialized chips that are very focused on transformer models.

5:44I just think the whole AI chip space is very ripe for disruption. And I'm curious, like, what do you make of the current kind of AI chip space? And where does Grok fit into all of that? Well, history doesn't repeat itself, but it rhymes. And our inspiration came from AlphaGo. AlphaGo is an AI model that was developed at Google DeepMind to play the game Go. At the time, Jonathan was still at Google, and he says that the AlphaGo team reached out to him during a critical stage in its development, a five-game match with top Go player Lee Sedol. The test game against Lee Sedol didn't go very well. And so we got an email saying, is your chip as fast as we've heard?

6:33This was the TPU at Google. And the answer was, yes, it is. We didn't know how fast they had heard, but that's the appropriate answer, right? So we spent the next 30 days furiously recompiling their model to run on the TPU. Even though the model didn't change, it performed dramatically better. It went from losing to winning dramatically. And that was the realization that compute was going to influence the quality. Now we're seeing this with the reasoning models. Why would that be the case, by the way? Why would it change its answer? Well, let's think of it this way. Imagine that you are writing an essay.

7:09And I tell you that when you write that essay, you have to do it without hitting the backspace or delete key even once. How good is that essay going to be? Right. Not so good. Now imagine you ask 10 different people to write essays in parallel, not seeing what the others are writing, without the backspace or delete key. You pick the best one. It's better, but it's not going to be great. Now you let one person iterate 10 times. There's a good chance that it's going to be a lot better. Right. That's especially true with coding and technical problems. So by being able to iterate, you actually improve the answer.

7:45This is one of the reasons we focus so much on latency, because as long as it takes to get these answers from LLMs today, as you add reasoning, it gets worse. The AlphaGo trials unfolded in 2016, but that realization of the importance of compute speed eventually led to Grok and its chip specifically designed for generative AI. Now, nearly a decade later, Grok is a multi-billion dollar company. They actually just landed a$1.5 billion deal with Saudi Arabia. So the way that we deploy around the world is we partner up. We've been working with Saudi Arabia. We've been talking to them for about a year, but this moved very quickly.

8:28We decided to do this about four months before we completed it. We signed the contract, and 51 days later, we had almost 20 ,000 chips up and running and serving traffic in the country. Yeah, amazing. What are they doing with all of these chips? What are the applications? It's mostly commercial, and it's not all Saudi Arabia, actually. So we're serving traffic from all around the world. And when you think about it, training is a bit of a local game. You do it in a data center somewhere in the world. Inference is a global game. So this is one of the reasons why you see so many people trying to build out quickly, build all their infrastructure.

9:05But one of the advantages of being fast is even from halfway around the world, we can complete a query faster than anyone else. We did those 20 ,000 chips, and that worked out so well, we've signed this deal to expand it quite significantly this year. Yes, Grok is going full speed ahead with their LPUs. But Jonathan doesn't see his company in competition with the biggest AI chipmaker. NVIDIA. We'll get to why after a short break. Stay with us.

9:45If you've spent any time building AI products or leading technical teams, you know this. Transformation doesn't fail because of ideas. It fails because teams can't move together. Enter Atlassian's Teamwork Collection. It has planning in JIRA, documentation in Confluence, video updates in Loom, and now AI agents in Rovo, which connects the dots across your work so nothing gets lost. It's one AI-powered teamwork platform designed for how modern teams actually build. Learn more at Atlassian.com slash TeamChanger. That's A-T-L-A-S-S-I-A-N dot com slash teamchanger. Do you see Grok as an NVIDIA challenger?

10:35Not really. I think when you look at what NVIDIA does well, they do training amazingly well. And with training, all you need is brute force compute. It doesn't matter how long it takes. It doesn't even matter if it's particularly expensive because you're going to amortize it across all of the usage. When you're talking about inference, typically you're going to have about a 10 to 20 times larger inference deployment than training. That's where it starts to become very cost sensitive. And that's also where you just need more scale. And also it's latency sensitive. And so our expectation is that NVIDIA is going to continue doing very well in training.

11:13In fact, we've had customers ask us when we show them the speed of our demo, should we just not buy NVIDIA GPUs? Should we just buy LPUs? And my answer to that is always, no, get every GPU you can. First of all, it's hard to get them. Second of all, you need them for training. And third of all, the more inference you do, the more you're going to want to train your models to optimize them more because then you get more out of each inference. So the more inference, the more training, more training, the higher the quality, the more the demand goes up. Every time a new model's released, our usage spikes.

11:45So it's a virtuous cycle. Now, of course, every large incumbent wants every part of the market. But success is usually based on an element of focus. And GPUs are not designed for inference. They're expensive, and they're just not low latency. Yeah. What is your competitive moat? Like, what is stopping an NVIDIA from, like, designing chips that are optimized for inference? One of our moats is that not only did we start early, not only did we build something general that seems to work no matter where the market is moving on different model architectures. But also, if we need to adapt to what's happening, we're going to be able to do that faster than almost anyone else because we just don't have to rewrite the software.

12:29We're actually moving to a much faster chip cadence. We've hinted that maybe there's a new chip coming, but there might be another one coming soon after that. And if you can get into a very fast cadence, because most of it's the software, if your software is automatic, then you can iterate. And the speed of iteration is the speed of innovation. Okay. Let's talk about energy consumption and just sustainability of AI, both on the training side and the inference side. LPUs are a lot more energy efficient than GPUs. So why is this important? Well, you're only going to be able to have as much AI as you have energy to power it.

13:07There's a stack up in civilization, right? So first you have materials, then you have energy, then you have information. and then you have compute. And compute is different. It's about creating something contextual in the moment creatively. That requires compute, but it requires everything else down that stack. The materials to build the chips, then you can't have the chips. If you don't have the energy, you can't power the chips. If you don't have the information, you can't train the models. And if you don't have the compute, you can't run them. And so it's absolutely fundamental. And then what we did that was quite unique, rather than retrieving data from external memory we actually just have a large number of chips and each chip does a little bit of the computation and then passes it on to the next set of chips to sort of like an assembly line right and the reason that that's so efficient is when you're reading from external memory there's a wire and that wire gets charged and discharged and the longer that wire is the more energy the wider that wire is the more energy that memory is not inside the chip It's actually quite far away.

14:12And so all of that data that you're reading is requiring a large amount of energy to move from that memory over into this chip. And so when we look at the amount of energy used by GPUs, just the memory reads alone are more energy than we use. Yeah, that's amazing. I love how you explained it to one. I love the stack explanation too. I had not heard that before. It's very powerful. So many people think of Grok primarily as an AI chip company. but you're a vertically integrated solution. You know, again, more like an NVIDIA from that perspective. And the Grok cloud platform, I don't know the latest number, but I don't know, it's probably over 500 ,000 developers at this point.

14:53So last week we crossed over a million. That is amazing, congratulations. So why was it so important to be a vertically integrated platform, both hardware and software? And I would love to hear some of what are some of the developers doing on the Grok cloud? So we always intended to do this. But when we started, capital wasn't as easy to get. So our initial plan was, let's sell hardware. And eventually, we can use that money to build our own cloud. What ended up happening was, it turns out, when you develop your own chip, and you try and get others to adopt it, you're actually not solving a problem.

15:34You're creating a problem. The problem is, they had something that worked. and now you're trying to give them something that requires a lot of work to get working. When we were trying to explain what we could do to people who were serving their own models and that it would be much faster, the feedback that we got was, why would you ever need an LLM to be faster than you could read? Like just like a teletype thing, like dial up, whatever. And we're like, but that's like not how the internet works. No one wants web pages to load that slowly. Use your intuition. And they couldn't. once we put an LLM on our website and showed how fast it was, all of a sudden we went viral.

16:13And we immediately started making it available as an API. We had to develop the software for that so that people could build on it. So if you've got a real amazing product, rather than trying to sell it to people on how, you know, what they're doing will get better, just put it out there yourself and go viral yourself. Like bring the possibility to life and then people see it. Now, as for what people are doing on it, you name it. I mean, a million developers, we've seen it all. Everything from people searching documents to people doing legal work. What we're seeing is most customers are either in the startup bucket or actually in the very large, like Fortune 50 bucket.

16:54There's not a lot in the middle. So we've analyzed that. And what we've come to a conclusion on. We could be wrong. But startups, when you're going from having no customer support to I've got a chatbot that has access to a little bit of information and it sometimes gives wrong answers, but it works. It's better than nothing. So it's an improvement. It solves a problem. So they're not looking for perfect. They're looking for solving a problem. When it's the larger companies, what we've seen is they're willing to try a very large number of initiatives. Now, most medium-sized companies can't afford to do that.

17:28Yeah, that's super fascinating. I am very passionate about democratizing access to AI. And, you know, in our previous conversations, I know you are too. And with my investor hat on, I'm always thinking about how these startups get access to compute, right? Because that's going to unlock a lot of innovation. And I sometimes worry that, you know, a lot of the compute is hogged by the bigger players. Who do you think is currently being left out? And I guess, how do we fix that? As always, the way that you're left out is either it's just too expensive and you can't launch what you're doing, right?

18:09Or you can't get enough to launch what you're doing. Like what we hear from customers all the time is they have requested a large number of GPUs and the wait time is just too long. Others who have paid for GPUs are waiting sometimes as long as a year to get them. But what it comes down to is, can you launch the application that you've developed? So very often people will develop on the most capable model and then they will attempt to deploy it. It's too expensive. So then they look for open source models, smaller models. Can they port their application? It's not like people are typically saying, oh, I just don't have enough compute.

18:50They're typically saying it as it's too expensive, but they mean the same thing. Can you be very specific around what we mean when we say too expensive? Yeah. So it used to be that your largest line item in running a business was your employees. and nowadays your cloud budgets are getting pretty close to to your people budgets and that's only going to get worse if you replace your customer service with an llm you're effectively replacing what some people did right now people are growing sort of hybrid they hire people but they also hire LPUs and GPUs, right? Now, it's not too surprising that people like to do stuff with AI because it's repeatable.

19:40Once they get it to a certain capability, they can just scale that up. It's not like I have to train another person to do that thing, or then they want to scale up really quickly, and they can't. They can't get access to enough compute. And so what you're seeing is we're getting to a point where we're almost able cost-wise to get the same capabilities out of the hardware as hiring sort of minimum wage workers and so on. And then you're just not gonna be able to scale if you can't get that. Yeah, as we think about the cost equation, it's back to this cost of every call you make. That factors in too, right?

20:19So if you have a customer service bot and this bot serves, I don't know, a thousand customers, right? and these customers are prompting the bot all sorts of questions, you got to kind of put into consideration the cost of each of these calls to the LLM or to the underlying model. Is that also how it gets expensive? Well, there's also another thing. I actually believe that AI is going to cause labor shortages. I don't think we're going to have enough workers. I'll give an example. I was talking to someone recently. they were explaining why they thought all customer support jobs were going to go to AI.

20:57And then what they explained was they just had had this amazing experience calling into a call center. And it was all AI up until the final step where it went to a human being to verify the correctness of the actions. So I asked them, how often do you typically call support? He's like, never. It's a terrible experience. I'm like, but that was a good experience, right? Yeah. Would you call more often if you got that same experience? Absolutely. So now that you're making that experience good, people are going to want it. If you know that you can call an airline and change your flight easily, if you know that you can call a hotel and get any issues sorted, you're going to do it more.

21:43And so the demand for that service is going to increase. And even if the proportion of time that's spent with a human being on the other side doing it, is lower, overall, the value to that business because of the better experience is going to mean that they're going to spend more on it. But there's other reasons why I think it's going to cause labor shortages. The other one is, think about jobs that exist today that didn't exist 100 years ago. Software engineer, right? It's not like those people just sat at home and did nothing. we invented new jobs for them to do. And then the last part is if we end up in this sort of deflationary economy because things are so inexpensive to build, a lot of people are going to opt out of the workforce because they're going to be able to work either part-time, they're going to be able to work fewer days a week, or they'll be able to work fewer years of their life before they retire and then go on and do whatever they want to do.

22:42And so that's also going to lead to labor shortages. So the fact that businesses can exist with fewer employees means there'll be more businesses doing things that weren't even profitable before, that many people will opt out sooner, and that there will be greater demand for services because the experiences will be better. And I think all of those will come together to cause enormous labor shortages. Yeah. What are some new jobs that you're seeing AI create? The most obvious is prompt engineer, right? Yeah, right. We're actually looking to hire our first head of prompt engineering. How do you hire for that position?

23:18So we're trying to figure that out. But here's some of the initial intuition that we have. Prompt engineering is more about communication and leadership than it is about doing sort of, you know, grindy work where you're just like iterating and tweaking. And so can you communicate clearly? Can you visualize what you're to get at the end. The way that we look at it is it used to be that you would have hardware engineers that would wire things up. And then this wacky thing called software engineers, which at first people were like, that's not a real job, right? You just sit there and you type stuff like what?

23:54And then eventually people became full-time software engineers. But at first they had to do a little bit of hardware engineering in order to make the computers work. And then And what we see happening next is right now, prompt engineering involves a little bit of software writing. But eventually, it's going to be purely through language. Now imagine anyone who can speak can now create their own business. It's going to unlock billions of people. There's, what, 1.4 billion people in Africa, about the same number in India, right? All of a sudden, they will be unlocked. Yeah, it's very exciting. AI can unlock human potential.

24:35But with more access, how can we make sure that our data is safe? Jonathan thinks the answer is open source. That and more after a short break.

24:56Meet Nicole Nicholas, Capital One business customer and co-owner of Ansett Uncles, a plant-based restaurant and community space in Brooklyn, New York, that got its start from a need for unity. The inspiration, it was born from the desire to create a space that felt like home, where we can connect community culture, good food, and come together. with family and friends. That's how we birthed aunts and uncles. Nicole and her husband, Mike, were fulfilling their dream of bringing people together out of their home kitchen, but they soon learned that the demand for community was greater than they knew.

25:27It became overwhelming and we were like, we need home, but not in our actual home. We realized that there was also a need in our community for something bigger and in our neighborhood, so we had to find a place. Moving from a home operation into a storefront was a huge next step, but Nicole and Mike were able to take it on with the help of Capital One Business. It's not for the weak. As a small business, finding resources is super important because that's the way you'll be able to manage and scale. We would have never done that without having Capital One to be able to help us along the way. The cashback rewards are very helpful.

Read the full transcript

26:03You know, it just gave us that runway to be able to breathe a little bit. Then you get to focus on the cooking of the food and making the experience great. To learn more, go to CapitalOne.com slash business cards. So let's talk about DeepSeek. You know, obviously rocked the AI world, primarily because they were able to build a model that has very similar capabilities to OpenAI and other foundation models out there for a fraction of the cost. And of course, they open source the model. I personally think this is great news because it makes AI more efficient and more accessible, which only increases demand.

26:39But you guys called this a game changer moment. And I'm curious, yeah, curious about your thoughts around DeepSeek. You also, I think, offer it through GrokCloud. So yeah, tell us your thoughts about all of this. So for about six months before DeepSeek released their R1 model in the distilled versions, there was a lot of murmuring in the venture capital community of, are the LLMs going to be commoditized? And the general sentiment was yes. Meaning that it doesn't really matter anymore what LLM you're using. They essentially all do the same thing. What DeepSeek really did was it ended any questioning about whether or not these models had been fully commoditized.

27:26And so at this point, no one is seriously entertaining, as far as I know, building a model company where it's purely based on the quality of the model. right if there was any harboring of a notion that that was a good business that's over now since deep seek has been released now everyone's rushing to open source these models and if you haven't open sourced your model yet at this point since it was never really a moat to begin with it's kind of what's what's wrong like why are you holding it back and what deep seek also did was they published a lot of the details behind even how they ran the models.

28:06And they even shared their revenue numbers and their profit margins. So it showed, hey, you can make money if you do this and you do these things and that. So like, it's really changed the game from being a little more closed to being a little more open. Yeah. You know, there's always been this tension between open source and proprietary, right? Like in software. Do you think it's important for these models to be open source? And how does that help, I guess, the company that's open sourcing its models, but also, yeah, how does it accelerate innovation or not? Well, if I look at the word important, important to whom, right?

28:41And instead of important, I would focus on, can you compete with a closed source model if there's a bunch of open source models. And open always wins, always. Now, people have shifted from thinking that open is full of vulnerabilities and lacks quality to open is always better, safer, and so on. So now the onus is on those who are doing closed models to prove that their models are as good. And there's just this ground swell of support for open. Yeah. Let's talk about safety for a second and trust, right? So, So, for example, you know, I've been playing with the DeepSeq model. And at the end of the day, I don't know if I trust, right, Chinese-based company.

29:27I don't know where the data is going. I don't know what they're doing with the data. And so I think as we continue to see more and more applications of AI, I am very interested in the question of safety and governance and just responsible AI. Curious what you think about that. Yeah. Yeah, so safety is a complex topic, and in particular around the Chinese models. So there's two different things to worry about. The first is, where's your data going? And if your data is going into China, it's not like these companies have any way to say no to the CCP. If the CCP asks for the data, they have to give it.

30:02So you have to assume that any query that you do to the DeepSeek service itself, you might as well be sending that straight to the CCP and putting your name right on it. And so David Sachs, the AI czar for the U.S., advocated for a couple of different U.S. companies that are running the DeepSeq model to be used. So that way the data is not going to China. And Grok, he said that we're one of those. So you can use us. We actually delete all queries after they're done. So we don't retain your data. Interesting. But there's another concern as well in the Chinese models, which is they've also been trained to answer certain questions like, tell me about Tiananmen Square.

30:41Or how do you treat the Uyghurs? And if you ask these queries, they will give answers like the CCP treats all people with equal respect. But the bigger concern there isn't the censorship. That's not great. But worse is what if they intentionally bias the model? So imagine that the CCP says, we want this person to win the next election. And then all the answers are, well, you should vote for this person for this reason. And do we want to give that kind of control to the CCP? And I would say no. Yeah, absolutely. How do you think we should be building guardrails into these systems and these models?

31:21Guardrails are a little bit different than dealing with the biases. On the guardrail side, one of the things about these LLMs is remember that they've been initially trained on the Internet. And the internet doesn't necessarily have the most eloquent, most subtle, most nuanced content, right? But interestingly, the more you train the models, and if you give them just a little bit of fine-tuning, they start to act more like rational, intelligent entities. So it's probably not too surprising that as the models get more intelligent, they actually start acting more intelligent. They start treating things with more subtlety and nuance.

32:02They have more creativity, more understanding. They can even help people resolve issues. And so I suspect that while it does take work, the work actually gets easier the smarter the models get, not harder. Okay, two more questions. One is you've said in the past that down the line, at some point in the future, you would love to offer free access to Grok around the world. who would you want to offer this technology to and why for free? Well, first of all, it's not down the line. We already do. And our mission is to drive the cost of compute to zero because we want everyone to have access, right?

32:43Not just, you know, the wealthiest in society, right? We're worried about a concentration of compute power. What we want is for everyone to have accesses. Just imagine if you had the ability to, on demand, get access to 1 ,000 PhD students to help you solve a problem. But the other person that you're competing against doesn't have that. That's going to cause an imbalance. So our goal is to make sure that everyone gets equivalent access to AI. Now, of course, it has a cost. We can't give all of it away for free. But once you start getting into the sort of token rates that are required to run a business, then we start charging.

33:27It's a little bit like electricity, right? If you plug your phone in at a hotel or a restaurant, no one's going to charge you for that. On the other hand, if you're trying to run a data center, you're going to get charged. Yeah, that's awesome. Okay, final question. And we were starting to kind of allude to it a bit. I really believe that AI ought to be applied to unlock human potential. But I also think a lot about, OK, so if AI is going to be smarter and even nicer, nicer, more creative, more empathetic than humans, what does it mean to be human in the age of AI? And I think that's something we're going to have to figure out.

34:05And so one of our core missions internally is to preserve human agency in the age of AI. Think about it this way. Think about someone who's very affluent and has had children and how that affluence is a benefit, but sometimes it's also a hurdle, right? There's a lack of mission. There's a lack of sense of purpose because you could literally sit on a couch all day. You have a trust fund, right? And as a society, I am less worried about AI taking over. I'm less worried about AI taking jobs. I'm more worried about what happens when we don't have to work and how we're going to find our own purpose, how we're going to also preserve our own agency.

34:48Because I'm worried that we will give our decision-making authority over to AI because decisions are hard. And we will probably hand over simple decisions. That will be a wise thing to do because you can make a finite number of decisions in a day. right? But the important decisions, we should continue making those ourselves. And I think as a society, we're going to have to figure this out. Yeah. Super fascinating, right? Like, how do we, yeah, how do we continue to have that agency, but also the motivation and a sense of purpose? Great way to end our interview today. Thank you, Jonathan, for joining us.

35:23Super fascinating. Thanks for having me. purpose and profit not all multi-billion dollar companies are thinking about marrying the two but the ones that do are going to have the biggest impact as ai becomes more efficient there will be more demand for it demand will mean more users which will mean we'll need more ai chips for training and for inference and we actually also need more energy to run these chips But that's for a different conversation. So this means two things. One, there's lots of room for innovation and disruption across the entire AI tech stack. And two, unless we're proactive, compute will only be available to a select few.

36:08This is where Grok comes in. It's committed to ensuring that that doesn't happen. Sure, Grok is partnering with entities like Saudi Arabia and others to scale their AI platform. But they also want to bring AI compute to the masses, and they want to do that in part for free. On Pioneers of AI, I'm so excited to continue featuring entrepreneurs passionate about democratizing access to AI so that we can all reap the benefits. If you like what you heard on this episode, take a moment to rate and review us wherever you're listening. Your feedback is so important to us.

37:14Thank you. on LinkedIn, Instagram, TikTok, YouTube, and X. Just search for at Pioneers of AI. Thanks so much for listening.

From the publisher

Jonathan Ross and his company Groq are building AI chips focused on fast inference, increasing the speed with which AI models can process new information. But it’s not just about speed. Faster inference means that AI can deliver answers more quickly, and in turn save on energy and compute costs. In this episode of Pioneers of AI, we explore why Groq’s chips are making waves, how AI will create more jobs, and the importance of democratizing access to AI.

Learn more about Pioneers of AI: http://pioneersof.ai/

Follow Pioneers of AI on all channels: https://linktr.ee/pioneersofai

At the center of AI is people, so we want to hear from you! Share your experiences with AI — or ask us a burning question — by leaving a voicemail at 601-633-2424. Your voice could be featured in a future episode!

See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

More from Pioneers of AI

All 125 episodes
With Groq, Jonathan Ross is taking AI inference to new speedsPioneers of AI · 35 min
Listen in VO