In short
Podcast Summary: The AI Daily Brief - Episode on Google Gemini 1.5
Podcast Title
The AI Daily Brief (Formerly The AI Breakdown)
Episode Title
Google Launches Gemini 1.5 with ONE MILLION Token Context Window
Episode Description
Google surprises everyone with a massive context window on their new Gemini 1.5 model. Additionally, Magic.dev raises over $100M for a model that can handle multiple million token context windows.
---
Key Topics Discussed
Introduction
- Google announced the launch of Gemini 1.5 Pro, featuring a groundbreaking 1 million token context window.
- Magic.dev secures $100 million funding to develop a model capable of handling multiple million token context windows.
Importance of Context Windows
- The term context window refers to the number of tokens a model can engage with at a single time.
- A larger context window allows for more coherent reasoning and processing of vast amounts of information.
Magic.dev's Funding and Goals
- Magic.dev aims to create an AI programmer that can reason over entire codebases due to the expanded context window.
- Key beliefs include:
- Code generation is a pathway to Artificial General Intelligence (AGI).
- AGI safety is both significant and solvable.
- Bigger models are necessary for creating advanced AI products.
Meta's Shift in Coding Focus
- Mark Zuckerberg emphasized coding's importance for Llama 3, highlighting its role in improving an LLM's understanding of knowledge and logic.
- The shift indicates a strategic move to compete in the AI landscape, aiming for state-of-the-art models.
Opinions on AI's Impact on Programming
- Debate exists on whether AI will replace programmers.
- Ahmad, CEO at Mercury, argues that misunderstanding the role of programmers fuels this misconception.
- Programming is seen as an art that involves understanding customer needs, integrating codes, and scalability—tasks that AI can assist with but not fully replace.
Additional AI News
- Slack integrates AI for personalized search results and summaries.
- Amazon develops a large text-to-speech model, Base TTS, with emergent qualities, improving its natural speech capabilities.
- NVIDIA’s investments lead to stock rises in various companies, indicating a bullish sentiment in AI technologies.
Google’s Gemini 1.5 Announcements
- Sundar Pichai, Google CEO, introduced Gemini 1.5 Pro as capable of processing a million tokens—equivalent to the content of entire films.
- MOE (Mixture of Experts) architecture allows for more efficient training and high-quality responses while supporting multimodal inputs.
Competitive Landscape
- The launch of Gemini 1.5 raises questions about OpenAI's position in the market.
- OpenAI is reportedly developing a web search product to compete with Google, although details remain unclear.
Conclusion
- The episode highlights the rapid advancements in AI technology and the ongoing competition between major players like Google and OpenAI.
- The future of programming and AI integration raises important debates, warranting close observation of industry developments.
---
Key Takeaways
- Gemini 1.5 Pro’s 1 million token context window represents a significant leap in LLM capabilities.
- Companies like Magic.dev are pioneering AI tools that enhance software development processes.
- The conversation around AI's role in programming reveals nuanced views about job displacement and the evolving nature of work in tech.
- Ongoing developments in AI, especially from Google, suggest a competitive landscape that could shift the current dynamics of the industry.
---
Further Resources
- Subscribe to The AI Breakdown newsletter: [Link](https://theaibreakdown.beehiiv.com/subscribe)
- YouTube Channel: [The AI Breakdown](https://www.youtube.com/@TheAIBreakdown)
- Community: [Join Here](bit.ly/aibreakdown)
- More about AI news: [Breakdown Network](http://breakdown.network/)
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Today on the AI Breakdown, Google has announced Gemini 1.5 Pro with an unbelievable million token context window. Before that on the brief, another very long context window project, specifically focused on coding, has raised a fresh$100 million. The AI Breakdown is a daily podcast and video about the most important news and discussions in AI. Go to breakdown.network for more information about our YouTube, our Discord, and our newsletter. Welcome back to the AI Breakdown Brief, all the AI headline news you need in around five minutes. One of the big pursuits for AI developers right now is building software that can build software.
0:38This is important on a number of different levels. As we will discuss, it is not only about building a new type of tool, but also about giving AI's more broad reasoning power. Well, the specific context that we're starting with today is a new nine-figure funding round for Magic. Former GitHub CEO and now AI investor extraordinaire, also the guy behind the Vesuvius challenge that we talked about last week, Nat Friedman tweeted,
1:14If this sounds like magic, well, you get it. Daniel and I were so impressed. That's Daniel Gross, his often investing partner. We are investing$100 million into the company today. The team is intensely smart and hardworking. Building an AI programmer is both self-evidently valuable and intrinsically self-improving. If this sounds interesting to you, consider joining them. Now, one thing I want to note is this many millions of tokens of context, which is something that we're going to get deeper into in the main part of the episode. But here, obviously, part of the big innovation is that with that big of a context window, as Nat puts it, the AI programmer can reason over the entire codebase.
1:48The ability to ingest an entire codebase all at once obviously seems like it could give magic a very different set of capabilities. Now, in their announcement post, Magic writes that they're working on frontier-scale code models to build a co-worker, not just a co-pilot. They write, Things we believe. Code generation is both a product and a path to AGI. AGI safety matters and is solvable. To build a great AI product, we need to train our own frontier-scale model. Transformers aren't the final architecture. We have something with a multi-million token context window. So there is a lot in just a few bullets there.
2:19Code generation has both product and a path to AGI. We're about to discuss Mark Zuckerberg's comments recently, where Meta has come around to this thinking that code generation is indeed a central capability for the state of the art, not just from a feature parity standpoint, but because it improves general reasoning. The other bombastic note here is this idea that transformers, which is of course what GBT is built on, the T stands for transformer, aren't the final architecture, and that they have a multi-million token context window. Now given all this, it's not surprising that Magic is very buzzy.
2:47However, they're not the only project thinking in similar ways. Anton Osika, who has been building GPT Engineer, which we've covered on the show before, yesterday tweeted, Happy Valentine's Day, everyone! Launching a new AI startup out of Europe today. Lovable. We're building software that builds software. How we are approaching it at a high level? Step one, a tool for anyone to prototype web apps and collaborate via natural language on the same code base as developers. Step two, make fully autonomous code generation work on one single set of technology choices. Now, if you go to their website, lovable.dev, The claim they make is that this is the last piece of software.
3:20They write, we're building the software that builds other software. Why? The world, they write, is full of ambitious people who want to solve important problems. For three decades, software has been the most significant tool to unleash the world's ambition. Still, less than 1 % of the world has the skills required to create software. If we succeed, they write, everyone will have the same capabilities that entire product development teams at Stellar Tech Companies have today, right at their fingertips. We will unlock a new era of innovation, empowering dreamers everywhere to shape the world. We're reducing the barriers to build and staying committed to one goal, to unleash human creativity on an unprecedented scale.
3:53Now, about a month ago, Mark Zuckerberg talked a bit about what was coming with Llama 3, and some of these same themes were on display. He discussed how the company had changed its position on coding and AI, saying, one hypothesis was that coding isn't that important because it's not like a lot of people are going to ask coding questions in WhatsApp, which is, of course, a big part of the MetaPortfolio. Zuckerberg went on, though, it turns out that coding is actually really important structurally for having the LLMs be able to understand the rigor and hierarchical structure of knowledge, and just generally have more of an intuitive sense of logic.
4:23You remember there were a lot of headlines around this time last month about Zuckerberg's goal to build an open source AGI. Now the consequence of this shift to focus on coding inside Llama 3 is not just to compete on that open source access, to compete at the state of the art in general. Zuckerberg said, Llama 2 wasn't an industry-leading model, but it was the best open source model. With Llama 3 and beyond, our ambition is to build things that are at the state of the art and eventually the leading models in the industry. Now, of course, there are big debates around what the net impact of this type of self-generating software can do.
4:55Holding aside all the arguments about how this is where AI goes off the rails and all the safety risks, which is worthy of an entire episode on its own, even the debate around whether AI is going to replace programmers has a lot of interesting dimensions. Ahmad, the CEO at Mercury, tweeted about this yesterday writing, why AI is not going to replace programmers. When I was in college studying computer science in 2005, I was told that outsourcing to India will remove the demand for programmers. It was a real fear. AI replacing coders, I think, is based on a similar misconception. He goes on, this misconception is rooted in a misunderstanding of what programmers do that non-programmers and even many engineers have.
5:31Programming is sometimes seen as a science where you have a very specific spec and you convert that through a series of repetitive work to code. People perceive it similarly to high school level math. You get an equation and you solve it and get an answer. And to be fair, that is what learning to program and entry-level programming is like. So it's easy to see why people have this misconception. In reality, a lot of programming is 1. Understanding customers, internal or external. 2. Converting that understanding into potential implementation. 3. Thinking about integrating seamlessly into existing code.
6:004. Thinking about building in a scalable way for future recs. 5. Etc. A lot of it involves having deep taste and applying that taste against the customer need. This part is much more art than science even. So what he says AI will actually do? One, make programmers more efficient at the repetitive part. Two, enable more people to be programmers. Three, turn more of the world into bespoke software. Four, increase the demand for great programming artists, elevating the craft even more. Ultimately, he concludes, if you're thinking of learning to code, don't be put off by the AI will replace programmers meme, and just do it.
6:31Another take on this that we've often heard from Sam Altman is that AI won't replace programmers because the demand for things that programmers build is just going to go up consequently with our capacity to actually deliver against that demand. Anyways, it's a super interesting discussion and honestly could have been a full episode rather than just a brief, but before we get to the main part of this episode, let's actually cover a few additional stories. Slack becomes the latest Web 2.0 software to deeply integrate AI, announcing that Slack AI has arrived. There are a bunch of different parts of this.
6:58Personalized search results, channel recaps and thread summaries, so for example, if you've been gone for a while, Slack AI can summarize what you've missed, which for anyone working with me on Slack would be a very useful thing given how many messages I send. And of course, Slack and its owner Salesforce are promising that users will control their data, that they'll have more granular ability to implement tools, etc, etc. Interesting news out of Amazon. Researchers at Amazon have apparently trained the largest ever text-to-speech model, which of course matters given that the company is so invested in things like Alexa, and they claim that they're seeing what they call emergent qualities, which is improving that model's ability to speak naturally, even in the context of complex sentences.
7:35The new model is called Big Adaptive Steamable TTS with Emergent Abilities, or Base TTS. TechCrunch writes, at 980 million parameters, Base Large appears to be the biggest model in this category. The largest version was trained on 100 ,000 hours of public domain speech, 90 % of which was in English, with the remainder in German, Dutch, and Spanish. Areas where this large model improved from previous models, they said, include compound nouns, emotions, foreign words, paralinguistics aka readable non-words like shh, punctuation, questions, and syntactic complexities. The authors of a paper wrote, These sentences are designed to contain challenging tasks, parsing garden path sentences, placing phrasal stress on long-winded compound nouns, producing emotional or whispered speech, or producing the correct phonemes for foreign words like qi or punctuations like at, none of which Bayes TTS is explicitly trained to perform.
8:25TechCrunch writes, Such features normally trip up text-to-speech engines, which will mispronounce, skip words, use odd intonation, or make some other blunder. But while Base TTS still had some trouble, it apparently did far better than its contemporaries. So once again, we have an advanced model that is working in ways we just don't quite understand. Over on Wall Street, a number of companies saw their stock rise after NVIDIA disclosed that it was an investor. Those investments included Recursion Pharmaceuticals, conversational voice assistant developer SoundHound AI, and fellow AI chip designer, ARM.
8:56So friends, AI's Whitehawk streak continues, but for us, that will end the AI Breakdown Brief. Thanks for listening or watching as always, and up next, the main AI breakdown. Welcome back to the AI Breakdown. Today, we had some unexpected and exciting news out of Google. CEO Sundar Pichai tweets, In December, we launched Gemini 1.0 Pro. Today, we're introducing Gemini 1.5 Pro. This next-gen model uses a mixture of experts MOE approach for more efficient training and higher quality responses. Gemini 1.5 Pro, our mid-sized model, will soon come standard with a 128k token context window. But starting today, developers and customers can sign up for the limited private preview to try out 1.5 Pro with a groundbreaking and experimental 1 million token context window.
9:46The 1 million tokens feature unlocks huge possibilities for devs. Upload hundreds of pages of text, entire code repos, and long videos, and let Gemini reason across them. It's still experimental and early, and we'd love your feedback. Now, context windows are something that people have been talking about ever since ChatGPT launched. It refers to the number of tokens that any given model can engage with at a particular time. The larger that window, the more coherent an LLM can reason across a bigger volume of text. To get a sense of how much of a difference it is to be talking about 128k in million token context windows, people were incredibly excited to see ChatGPT move from 8k up to 32k.
10:24So obviously we're talking about significantly longer than that. Anthropics Cloud has also used longer context windows to try to compete, although of course one of the things that people watch out for is whether performance starts to degrade when you're actually using longer context windows. Still, the initial response has been incredibly excited. Lior at AlphaSignal writes, Just in, Google releases Gemini 1.5, a powerful MOE model. It's a huge breakthrough. The model has the longest context window ever seen, 1 million tokens. It can process 1 hour of video, 11 hours of audio, 30 ,000 lines of code, or 700 ,000 words in a single prompt.
10:57When tested on text, code, image, audio, and video evaluations, 1.5 Pro outperforms 1.0 Pro on 87 % of the benchmarks used for developing LLMs. Jeff Dean, the chief scientist at Google DeepMind and Google Research, says, One of the key differentiators of this model is its incredibly long context capabilities, supporting millions of tokens of multimodal input. The multimodal capabilities of the model mean you can interact in sophisticated ways with entire books, very long document collections, code bases of hundreds of thousands of lines across hundreds of files, full movies, entire podcast series, and more.
11:28Now, given that Google has kind of a history of announcing things before they make them available, another thing that people were very excited about is that early testers were actually allowed to start using this 1 million token context window at no cost during the testing period. Writing about the technical approach to this model on their announcement post, Google says, Gemini 1.5 is built upon our leading research on transformer and MOE architecture. While a traditional transformer functions as one large neural network, MOE models are divided into smaller expert neural networks. Depending on the type of input given, MOE models learn to selectively activate only the most relevant expert pathways in its neural network.
12:04This specialization massively enhances the model's efficiency. Google has been an early adopter and pioneer of the MOE technique for deep learning. Our latest innovations in model architecture allow Gemini 1.5 to learn complex tasks more quickly and maintain quality, while being more efficient to train and serve. These efficiencies are helping our teams iterate, train, and deliver more advanced versions of Gemini faster than ever before. Another performance piece that many people noticed was this one. Quote, Gemini 1.5 Pro maintains high levels of performance even as its context window increases.
12:32In the needle in a haystack evaluation, where a small piece of text containing a particular factor statement is purposely placed within a long block of text, 1.0 Pro found the embedded text 99 % of the time, in blocks of data as long as 1 million tokens. Jeff Dean again tweeted about this a little bit, saying, Needle in a haystack tests out of 10 million tokens. First, let's take a quick glance at a needle in a haystack test across many different modalities to exercise Gemini 1.5 Pro's ability to retrieve information from its very long contexts. He points out that the results are almost entirely good, meaning 99.7 % recall, even out to 10 million tokens.
13:05That's in audio, video, and text. Avi Schiffman, who's working on the AI Wearable tab, writes, Perfect recall with 10 million tokens? The Messiah has arrived, and its name is Google. Developer Nick Dobos responded, What? That's crazy. Okay, Google, you might have a chance. And indeed, this is the sentiment that I am seeing all over Twitter slash X today. In many ways, this Gemini 1.5 Pro announcement has developers and people who are deep in the technical part of this space more excited than Gemini Advanced even did. There is so much sentiment like, wow, Google is really shipping, which is of course a complete switch from the discussion around them last year.
13:42The Verge's piece about this is Gemini 1.5 is Google's next gen AI model and it's already almost ready. The next version of Google's model is better and faster, sure, but it also has one pretty remarkable new party trick. Now, the Verge does a good job of contextualizing what 1 million tokens means in this multimodal model. Apparently, Sundar Pichai told The Verge, that means you can fit the entire Lord of the Rings film trilogy into that context window. As you might imagine, this has a lot of people saying, is OpenAI starting to fall behind? Where is GPT-5? Is this going to put increased pressure on them?
14:13Well, interestingly, we did get some new leaks about some things going on inside OpenAI that do in some ways relate to Google as well. Last week we heard that they were working on AI agent startups, but now, according to people with knowledge of the project, OpenAI is also working to develop a web search product that would bring them into direct competition with Google and, of course, other AI search tools like Perplexity. Writes the information, OpenAI has been developing a web search product that would bring the Microsoft-backed startup into more direct competition with Google. Now, details are scant right now.
14:42The information source said that the search service would be partly powered by Bing, but it wasn't clear whether it would be a separate product from ChatGPT or in some ways embedded into ChatGPT. Now, this is a story that I could see going in a bunch of different ways. It could be as simple as a slightly different interface and a set of expanded features around what they already have, which is effectively what Perplexity has done, although of course, done in such a valuable way that many people have shifted their research behavior entirely to Perplexity. Or it could be a fundamentally different and expanded product.
15:12In any case, it feels to me like there is a lot brewing and bubbling behind the scenes at OpenAI, and it'll be interesting to see what actually pans out. Now, staying on the topic of OpenAI for just a minute, they just published a piece called No, Sam Altman Isn't Raising Trillions of Dollars for Chips. And basically what they're saying is that while the Wall Street Journal report from last week implied that he was actively seeking that much capital for some sort of joint venture, instead, that's Altman's calculation of the total cost of everything from real estate to power for data centers that it would take to actually achieve the type of objective that he wanted to see.
15:44So it's not just some specific company that he's out trying to raise money for. To me, it's pretty interesting to see Google putting the pedal to the metal so hard and finally starting to reclaim some ground against OpenAI, which had been so lost throughout the course of 2023. Will this mean that we'll see an increased urgency from OpenAI to release more advanced models or new types of products? Or are they comfortable just continuing to go at their own pace. That is certainly what I will be watching and I will share it with you as I get any clarity on it. For now though, that is going to do it for today's AI breakdown.
16:13Until next time, peace.
From the publisher
Google surprises everyone with a massive context window on their new Gemini 1.5 model. Speaking of huge context windows, Magic.dev raises $100M+ for a model that apparently can handle multiple million token context windows.
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI.
Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe
Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown
Join the community: bit.ly/aibreakdown
Learn more: http://breakdown.network/
