In short
The AI Daily Brief - Episode Summary
Episode Title
Groq is 10x Faster than ChatGPT and Gemini
Overview In this episode of The AI Daily Brief, host NLW delves into the implications of Groq's impressive speed in the realm of AI language models, particularly in comparison to well-known competitors such as ChatGPT and Gemini. The discussion also touches on significant industry developments involving AI chips, data usage agreements, critiques from prominent figures, and the potential future of AI applications.
---
Key Topics
- The Emergence of Groq
- Groq claims to be the world's fastest LLM (Language Model), achieving response times around 10x quicker than ChatGPT and 18x faster than Gemini.
- Groq utilizes a foundational model, Llama 270B, and runs on its proprietary Language Processing Unit (LPU) inference engine.
- Its technology is designed to overcome bottlenecks faced by GPUs, particularly in terms of compute density and memory bandwidth.
- AI Chip Developments
- OpenAI's Ambitious Plans: Sam Altman is exploring a $5 to $7 trillion initiative to develop a network of chip fabrication plants globally.
- SoftBank's AI Chip Project: Reports reveal SoftBank is investigating a $100 billion project named Izanagi, focused on AI chips, potentially utilizing $30 billion from their liquid assets.
- Data Usage and Legal Agreements
- Reddit signed a $60 million annual deal with an unnamed AI company for data usage, marking a shift in their earlier stance against data scraping.
- Such agreements are significant amid ongoing legal battles regarding copyright and data usage for AI training.
- Critiques of OpenAI's Sora
- Jan LeCun from Meta raised concerns about the effectiveness of generative models like Sora for simulating the real world, suggesting limitations in handling complex sensory inputs.
- The Video Joint Embedding Predictive Architecture (VJPA) introduced by Meta aims to improve predictions in video processing.
- Industry Responses
- Notable figures (including Elon Musk) have voiced opinions on the effectiveness of OpenAI's Sora versus alternatives like Tesla's video generation capabilities, emphasizing the need for accurate physics predictions.
---
Implications of Groq's Speed
Revolutionary Potential
- Groq's speed opens new possibilities for real-time AI applications, potentially transforming user experiences across various sectors.
- The rapid response times could enable instant human-AI conversations, redefining interaction dynamics.
Economic Considerations
- Questions arise regarding the economic viability of Groq's technology, especially in relation to its high upfront costs.
- Industry experts believe that while initial costs may seem steep, the efficiency and speed improvements will justify the investment over time.
Future Use Cases
- The advancement of Groq could facilitate innovative applications, such as seamless text-to-speech integration for real-time interactions.
- The potential transformative impact of ultra-fast LLMs may lead to unprecedented AI functionalities and user experiences.
---
Conclusion This episode of The AI Daily Brief highlights the rapid advancements in AI technology, particularly focused on Groq's exceptional performance. The ongoing evolution in AI chips and data usage agreements signifies a dynamic landscape where speed and efficiency are becoming paramount. As the industry progresses, the possibilities for novel use cases and applications are expanding rapidly, prompting discussions about the future of AI and its practical implications.
For more insights, check out [The AI Breakdown](http://breakdown.network/) and subscribe to their newsletter for the latest updates in AI.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:27Today on the AI Breakdown, we're talking about Grok. news you need in around five minutes. There has been a lot of scuttlebutt in the news around AI chip efforts. OpenAI's Sam Altman, of course, has recently redefined the term ambition with his interest in exploring a$5 to$7 trillion initiative to build a network of chip fabrication plants around the world, but also with a focus on the United States. Perhaps in the realm of a little bit more realistic, if still intensely ambitious, Bloomberg reports that SoftBank's Masayoshi-san is exploring a$100 billion AI chip project. Now, the ups and downs of SoftBank are of course at this point legendary, but they are certainly currently on an upswing, which has been largely driven by their 90 % share of ARM holdings, the AI chip designer, which has seen enormous stock price increases this year.
1:14Bloomberg writes this new$100 billion chip venture would complement ARM and is codenamed Izanagi. Apparently, one of the scenarios that SoftBank is imagining would see them putting in about$30 billion, with another$70 billion coming from institutions in the Middle East. At the moment, SoftBank has around$41 billion in cash and cash equivalents on hand, which would mean that this would represent a major part of their liquid assets being invested into this new project. Over the last 10 trading days, as ARM has increased by more than 80 % in the markets, SoftBank shares have gained about 30%, and after this news broke, SoftBank's stock was up another 3%.
1:47Now at the moment, $100 billion chip project would represent a little less than a fifth of the global semiconductor market. But I don't think anyone on the planet, especially not people who are paying attention to AI, thinks that that market is going to stay this size for very long. Now, moving on from chips into the realm of data, last year in April, Reddit started to posture like it was going to take a more harsh stance towards companies that were scraping its data to train their AI models. Now, around 10 months later, they have actually signed a$60 million annual deal with an as-yet-unnamed AI company to allow that company to train on Reddit content.
2:21One thing we don't know is whether the deal is exclusive, but the contact who gave Bloomberg the news speculated that more likely was that the contract would serve as a model for future agreements with other AI companies as well. Now, from a Reddit standpoint, this is something they wanted to get done before a potential IPO, which could happen as soon as next month. And obviously, from the larger AI industry standpoint, any of these types of agreements right now, as legal battles around copyright and fair use are being fought in the courts, are going to be influential in shaping how the industry moves next when it comes to proprietary sources of data.
2:50Next, we have a little bit of follow-up around Sora. If you've been listening to my shows recently, you will know that Sora has been the major topic of conversation over the last few days, really ever since OpenAI announced it, but some people aren't quite as impressed. Jan LeCun, who is of course the head of Meta's AI department, said that if OpenAI's goal is really to simulate the world, that Sora's approach is ill-suited for that. He said, Modeling the world for action by generating pixels is wasteful and doomed to failure. Writes the decoder, There has been a historic debate about the merits of generative versus discriminative classification methods, with generative methods considered more difficult and less effective.
3:25Lacoon believes that generative models for sensory inputs will fail because it is too difficult to deal with the prediction uncertainty of high-dimensional continuous sensory inputs. Basically, he's saying that while generative models work for text, because there are a finite number of symbols, uncertainty is easier dealt with in that context. Sensory input, on the other hand, just has a huge additional level of complexity. It will perhaps surprise you not at all to know that Lacoon and Meta have their own approach to this problem, which they are calling the Video Joint Embedding Predictive Architecture, or VJPA.
3:53Again from the decoder, the model predicts complex interactions and interprets them by adding hidden parts of video to convey the dynamics of objects and interactions to the AI. VJPA focuses on predictions in a broader conceptual space, similar to human cognitive image processing. This architecture allows Vijaypa to adapt to different tasks by adding a small task-specific layer rather than retraining the entire model. Elon Musk also had some words for OpenAI Sora, saying on Twitter, where Tesla video generation exceeds OpenAI is that it predicts extremely accurate physics. That is essential for self-driving.
4:23So again, what we have here is another critique, not of Sora's ability necessarily to make incredible-looking videos, but instead how accurate it truly is as a representation of the real world and its ability to be a world simulator. which seems of course to be the ultimate goal in OpenAI's pursuit of AGI. Now staying on the theme of Elon or at least electric vehicles for just a minute, a Chinese EV maker, Xpeng, has announced that it would hire 4 ,000 people and invest millions in AI in a strategy that is contrasting with other Chinese and global EV makers who are, instead of investing more, currently racing to cut their own costs.
4:56Microsoft continues its streak of investing in European countries, announcing an AI infrastructure bid in Spain, coming along with a$2.1 billion investment. This, of course, follows our recent announcement of a$3.45 billion AI-focused investment in Germany, and will be centered around AI and cloud infrastructure. Lastly today, one that I'm only just starting to see talked about on Twitter a little bit, the University of Pennsylvania's Penn Engineering Today wrote a blog post at the end of last week called New Chip Opens Door to AI Computing at Light Speed. The piece begins, Penn engineers have developed a new chip that uses light waves rather than electricity to perform the complex math essential to training AI.
5:31The chip has the potential to radically accelerate the processing speed of computers while also reducing their energy consumption. They call this a silicon photonic or SIPH chip, and it's based on recent research around manipulating materials at the nanoscale to perform mathematical computations using light. And so we close this brief basically where we began with the continued focus on AI chips. In many ways, these two bookending stories represent the spectrum of what we're seeing right now. On the one hand, people trying to solve the compute access problem by throwing money at it and just building out more infrastructure, versus on the other hand, thinking in fundamentally new ways about the actual underlying technology itself.
6:06It is almost for certain that as the AI revolution continues, we will see immense developments in both ends of the spectrum. For now though, that is going to do it for today's AI Breakdown Brief. Up next, the main AI breakdown. Hello AI friends. Quick note before we get back into the show. we have just opened up registration for the March edition of the AI Education Beta Program. The whole philosophy of this program is to get you learning by doing. So we have short tutorials, think three minutes, five minutes, seven minutes, around specific features and use cases in AI, followed by challenges that are step-by-step instructions that get you actually using the most interesting and relevant tools.
6:42We have now built out a library of more than a hundred of these lessons and step-by-step companion instructions, and we'll be dropping more each week. For the first time, we'll also be moving beta users this month to a new dedicated platform where you can access that library of content, build lists of lessons you want to learn from later, and other features that we hope will help make this the single best AI learning experience available. If you want to check it out, go to bit.ly slash AI beta. That's bit.ly slash AI beta. Registration is only open this week until next Monday, so go check it out.
7:16A quick message before we get back to the episode today. At this point, you guys know that Notion is one of the major tools that I use day in and day out across the Breakdown Network, the AI Education Beta Project. Basically, anything that I'm doing in any sort of professional or entrepreneurial endeavor is going to be anchored by Notion. Now, you also know that one of the big themes that I keep talking about for 2024 when it comes to artificial intelligence is the integration of AI into our workflows. I think in many ways it's not just about which third-party AI tool is best for any given use case, but how they actually fit into what we're doing in ways that are actually time-saving.
7:53And that's why I love that Notion now has AI so deeply integrated across its entire suite of tools, which means that it's everywhere in your entire workspace. Now, for those of you who don't know, Notion combines your notes, documents, and projects into one space that is simple and beautifully designed. It's your one place to connect teams, tools, and knowledge, so you can do your most meaningful work. Unlike other solutions, it doesn't have you bouncing between six different apps. It is seamlessly integrated, infinitely flexible, and incredibly easy to use. Now, with the new, fully integrated Notion AI, you can work faster, write better, think bigger, and take care of tons of tasks that might normally take you minutes or hours in just seconds.
8:29One of my favorite use cases is to use Notion for brainstorming. So, for example, what would a great launch strategy be for some project? Use Notion AI to help you think through all of the different dimensions of how you could tell that story. Now, the proof is in the pudding, and Notion is used by over half of Fortune 500 companies. And most importantly, probably for you guys, the teams that do use Notion, send less email, cancel more meetings, save time searching for work, and reduce spending on tools. Right now, you can try Notion for free when you go to notion.com slash AI breakdown. That's all lowercase letters, notion.com slash AI breakdown, to try the powerful, easy-to-use Notion AI today.
9:06And of course, when you use our link, you're supporting the show. One more time, that's notion.com slash AI breakdown. On a recent video about Sora that I published, YouTube commenter Coldly Analytical wrote, I regard February 15th, 2024 as AI's day zero. Sora and Gemini 1.5 both announced on the same day, and both pushing us from the beta test phase into the AI is a real usable technology zone. I think it's an astute comment, and quietly, there is another leg of this next phase of AI stool that came over the weekend. It started for most with a tweet from Matt Schumer, the CEO of HyperWrite, who said, Wild tech you have to try.
9:45Grok, G-R-O-Q dot com. They are serving Mixtrel at nearly 500 tokens a second. Answers are pretty much instantaneous. Opens up new use cases and completely changes the UX possibilities of existing ones. Matt was the first to notice a live demo from what claims to be the world's fastest LLM. When you go to grok.com, it says, We'd suggest asking about a piece of history, requesting a guide on how to achieve your New Year resolution, or copy and pasting in some text to be translated by prompting Make It French. This alpha demo lets you experience ultra-low latency performance using the foundational LLM, Llama 270B created by Meta AI, running on the Grok LPU inference engine.
10:22Now, they actually give you two options for model. You can use either Mixtral or the Llama 270B, but suffice it to say that the speed at which generate responses has people's minds scrambling. A little later over the weekend, Matt again writes, The first public demo using Grok, a lightning-fast AI answers engine. It writes factual, cited answers with hundreds of words in less than a second. More than three-quarters of the time is spent searching, not generating. The LLM runs in a fraction of a second. So what is going on? Well, Grok, on its website, on its Why Grok section, says, Grok is on a mission to set the standard for Gen AI inference speed, helping real-time AI applications come to life.
10:57In its FAQ section, Grok writes, What is the LPU inference engine? An LPU inference engine with LPU standing for Language Processing Unit is a new type of end-to-end processing unit system that provides the fastest inference for computationally intensive applications with a sequential component to them, such as AI language applications or LLMs. On the question of why it is so much faster than GPUs for LLMs and Gen.AI, Grock writes, The LPU is designed to overcome the two LLM bottlenecks, compute density and memory bandwidth. An LPU has greater compute capacity than a GPU and CPU in regards to LLMs.
11:29This reduces the amount of time per word calculated, allowing sequences of text to be generated much faster. Additionally, eliminating external memory bottlenecks enables the LPU inference engine to deliver orders of magnitude better performance on LLMs compared to GPUs. Jay Scrambler on Twitter wrote a slightly more comprehensive explanation which I found useful. Jay writes,
12:16the SIMD, Single Instruction Multiple Data Model used by GPUs, and favor a more streamlined approach that eliminates the need for complex scheduling hardware. This design allows every clock cycle to be utilized effectively, ensuring consistent latency and throughput. For developers, this means that performance can be precisely predicted and optimized, which is critical in real-time AI applications. Energy efficiency is another area where LPUs shine. By reducing the overhead of managing multiple threads and avoiding the underutilization of cores, LPUs can deliver more computations per watt. Grok's innovative chip design allows multiple TSPs to be linked together without the traditional bottlenecks found in GPU clusters making them extremely scalable.
12:51This enables linear scaling of performance as more LPUs are added, simplifying the hardware requirements for large-scale AI models, and making it easier for developers to scale their applications without re-architecting their systems. So what does this all mean? LPUs could provide a massive improvement compared to GPUs for serving AI applications in the future. If anything, it will be great to have alternative high-performing hardware since A100s and H100s are so in demand. So even if all of that doesn't make it necessarily too much clearer, what you should be taking away is that there is a new hardware approach underlying this.
13:20This is not a different model akin to GPT-4 or Gemini or anything like that. This is a new approach to processing that lies underneath. Carlos Perez at Intuit Machine explains further. He writes, Grok is a radically different kind of AI architecture. Among the new crop of AI chip startups, Grok stands out with a radically different approach centered around its compiler architecture for optimizing a minimalist yet high-performance architecture. Grok's secret sauce is this compiler-first method that shuns complexity in favor of tailored efficiency. At the heart of Grok's architecture is an almost surprisingly bare-bones design that does away with unnecessary logic in favor of raw parallel throughput.
13:52The hardware itself is comparable to an ASIC, an application-specific integrated circuit finely tuned for machine learning. However, unlike a fixed-function ASIC, Grok leverages a custom compiler that can adapt and optimize across different models. It is this combination of a streamlined architecture and an intelligent compiler that sets Grok apart. The key insight is that many AI chip stack components, like GPUs, bring extraneous hardware and bloat. Grok returns to first principles, recognizing that machine learning workloads are about massive parallelism over simple data types and operations.
14:20By eliminating generic hardware and even concepts like locality, the design maximizes throughput and efficiency. This is enabled by Grok's compiler that sits between software frameworks like TensorFlow and the hardware. The compiler analyzes and optimizes neural network graphs, tailoring and mapping them to the underlying architecture for accelerated execution. It breaks computations into the smallest operations to unlock parallelism. The compiler also enables capabilities like batch size 1 inference that ensures all hardware is usefully leveraged. Critically, Grok built its compiler before even finalizing the hardware design.
14:48The software insights directly inform the architecture. This co-design process allowed inference-specific optimization without legacy limitations. The innovative compiler-first methodology allows custom optimization that balances flexibility with performance. So basically the idea here is that whereas, for example, NVIDIA GPUs do lots of different things, they're used to run gaming. They were, for a time, used for crypto mining. The Grok chip is totally optimized for the generative AI world. Carlos used that phrase, first principles, and that's really what this seems like, a design from the ground up based on this particular use case, which is admittedly a set of different use cases.
15:21Still, for most people, this is really just all about speed. Dini Yerlin writes, side-by-side Grok versus GPT 3.5, completely different user experience, a game changer for products that require low latency. Ethan Mollick writes, quote, quote, GBT 3.5 class LLMs are too slow. Sure, that was true last week. Here is Grok running Llama 2. My favorite comment on this when I posted the same video on LinkedIn, it is too fast. It shouldn't be this fast. Tom Osmond writes, Love seeing all the Grok demos on the feed, but is it only good for working with LLMs? Answer, nope, it's insane at other stuff too.
15:52Watch this clip from Grok Labs that shows it running style clip on an image to create eight different styles in 1024 pixels in just 0.185 seconds. Basically, this is showing an image generation capacity that is equally insanely impressively fast. Gabor Sell writes, Comparing time to complete answer for a simple code debugging question. Grok wins on speed 10x faster than Gemini, 18x faster than ChatGPT. Although Gabor did say that Gemini wins on quality of answer. Some people think this is so disruptive that we're going to see Grok get scooped up by one of the big AI labs. Farron Mather writes, Prediction, Grok will get a$10 billion acquisition offer within this month.
16:27Grokker Tom Ellis responded and said, add two zeros and we'll think about it. Just kidding, we're not for sale. But we are building out more and more infra every day to serve customers with the lowest latency LLMs available. Now, to the extent that there has been any critique of this or skepticism, it's been around the potential price. FelixRedPanda writes, How does Grok make economic sense? One of these cards costs 20k and has 0.23 gigabytes of memory. So do people buy 320 of these cards and fill two full racks of them to serve a single Llama 70B for 10 million including servers? That can't be how this works, right?
16:57Bindu Reddy from Abacus, however, bites back, saying, I'm seeing lots of takes claiming that Grok does not make economic sense. Barely anything makes economic sense at the beginning. LLM companies still lose billions. Vision Pro is too expensive. Even the much-loved Cybertruck is too expensive. The cool thing about Grok is the blazing fast inference. The economics will make sense in time. Now, others are just thinking about what opportunities this opens up. Responding to a question, what are some novel use cases now possible because of this speed, AI solopreneur Levels.io says, So I thought about this.
17:26If you hook up a just-as-fast text-to-speech model and fast whisper-speech-to-text model to Grok, you can have instant conversations from human to AI and back without any delays, like ChatGPT's TTS 5-second delay. Arvid Kahl responded to that, AI could respond while you're still speaking the last syllables of your word. I think we have to rethink what this could be used for. Honestly, I feel the humans are too slow to be even interesting for this kind of speed. Can you imagine how fast this thing could code as an autonomous agent? Sully Omar writes, imagine the possibilities with a model that has 1 million contexts like Gemini Pro 1.5, instant and cheap inference like Grok, GPT-5 like reasoning.
18:00We'll be building insane things. And Andrew V responds, imagine how dated this post will be in a year, and what our imagination will be dreaming of then. Now obviously people are only just starting to dig into Grok and there'll be a lot more to talk about in the coming days and weeks, but you can go check it out right now at grok, G-R-O-Q dot com. This is definitely one that is benefited by a live demo. So I hope you go check it out. For now, though, that is going to do it for today's AI breakdown. Until next time, peace.
From the publisher
Alongside Gemini 1.5's massive new context window, and Sora's mindblowing video generation, Groq has come along to redefine how fast we think LLMs can be. NLW explores people's reactions and the implications for new use cases.
INTERESTED IN THE AI EDUCATION BETA?
Learn more and sign up https://bit.ly/aibeta
Today's Sponsors:
Notion - Notion AI. Knowledge, answers, ideas. One click away. - https://notion.com/aibreakdown
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI.
Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe
Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown
Join the community: bit.ly/aibreakdown
Learn more: http://breakdown.network/
