In short
AI Daily Brief: Episode Summary
Podcast Information Title: The AI Daily Brief (Formerly The AI Breakdown) Description: A daily news analysis show on artificial intelligence, examining its impact on creativity, work, industries, and ethical questions.
Episode Details Episode Title: GPT-4o Mini and the Rise of Smaller, Low-Cost AI Models Episode Description: OpenAI released GPT-4o Mini this week, which is more powerful and 60% cheaper than GPT-3.5. The episode discusses the competition among AI models shifting towards smaller, state-of-the-art options.
---
Key Highlights
Opening Remarks
- Introduction of the new GPT-4o Mini model by OpenAI.
- Overview of AI competition shifting towards smaller models with enhanced capabilities.
Headlines Segment
- Google at the Olympics
- Google will be the official AI sponsor for Team USA at the upcoming Olympics.
- Plans include using immersive views from Google Maps for broadcasts and integrating AI features for commentators.
- OpenAI and Chip Development
- OpenAI is in talks with Broadcom to develop its own chips due to reliance on expensive GPUs.
- Hiring former Google employees to bolster chip development capabilities.
- Samsung's AI Features
- Introduction of generative AI features on the Galaxy Z Fold 6, including a tool called Sketch to Image.
- Concerns raised about the implications of AI-generated content and its potential to deceive viewers.
- Google's AI Industry Forum
- Launch of the Coalition for Secure AI (COSAI) to focus on AI security.
- Participation of major companies including Amazon, Microsoft, and NVIDIA.
---
Main Discussion
GPT-4o Mini
Overview of GPT-4o Mini
- Performance: Scores 82% on MMLU, outperforming GPT-4 on chat preferences.
- Pricing: Costs $0.15 per million input tokens and $0.60 per million output tokens, making it significantly cheaper than previous models.
Implications of Cost Reduction
- OpenAI's focus on making AI more affordable to expand its applications.
- Lower costs enable businesses to implement AI solutions in customer service, data extraction, and other areas.
Competitive Landscape
- Discussion on the shift towards smaller, low-cost AI models as a response to market demand.
- Notable advancements in smaller models by other tech companies, focusing on functionality over sheer size.
Community Reactions
- Mixed reviews from industry experts, with a recognition of GPT-4o Mini's capabilities but also a note that it may not replace more powerful models in all scenarios.
- Emphasis on cost-effectiveness and affordability as a primary driver for development.
---
Future Directions in AI
- Predictions of a trend towards the development of smaller, efficient models that require less computational power while still delivering robust performance.
- Excitement around potential new applications as AI becomes more accessible due to decreased costs.
---
Conclusion
- The announcement of GPT-4o Mini reflects a significant change in the AI landscape, with a clear trend towards cost reduction and increased accessibility.
- The potential for smaller models to enhance various applications underscores the evolving nature of artificial intelligence and its integration into everyday technology.
---
Additional Resources
- Discounts on AI Tools: Listeners can access a 20% discount on Venice Pro and 50% off tutorials at Superintelligent using specific codes.
- Community Engagement: Join the conversation on Discord and subscribe to the newsletter for updates on AI developments.
---
For continuous updates and insights, listeners are encouraged to subscribe to The AI Daily Brief on their preferred podcast platform.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Today on the AI Daily Brief, we're talking about OpenAI's new GPT-4O mini model and why it says a lot about where AI development and competition is happening. Before that, in the headlines, Google is bringing AI to the Olympics. The AI Daily Brief is a daily podcast and video about the most important news and stories in AI. To join the conversation, follow the Discord link in our show notes. Welcome back to the AI Daily Brief Headlines Edition. All the AI Daily News you need in around five minutes. We have the Olympics coming up pretty soon. They are kicking off on July 26th in Paris. and Google will apparently be the official AI sponsor for Team USA.
0:41So what does that actually mean? Well, it means a few things. First of all, the broadcast will pull from the immersive views that were added to Google Maps over the last few years, adding additional information and giving 3D views for the venues where the competition is happening. Announcers and commentators will use Google Search AI overviews in broadcast segments, for example, answering trivia questions about the Olympics. In one part of it, comedian Leslie Jones will ask Gemini to help her learn a new sport, and so on and so forth. This is a straight-up advertising partnership. This is a chance for Google to show off what it's got, which could go either very well for the industry or quite poorly.
1:20For now, we will just have to wait and see. Over in chip land, the information is reporting that OpenAI is holding talks with Broadcom about developing new chips. Summarizes Yahoo Finance, OpenAI is exploring the idea of making AI chips on its own to overcome the shortage of expensive GPUs that it relies on to develop models. The report says that OpenAI is hiring former Google employees who worked on Google's Tensor Processing Unit, and that they've decided to develop an AI server chip. An OpenAI spokesperson was circumspect, saying, OpenAI is having ongoing conversations with industry and government stakeholders about increasing access to the infrastructure needed to ensure AI's benefits are widely accessible.
1:56Which basically reads to me like them saying, yeah, of course we're having conversations about building our own chips. Everyone is having conversations about building their own chips. I think for now, it's not clear exactly what will come out of those conversations or where any actual partnerships would come down. Next up, there has been a lot of discussion this week of Samsung's AI features on their new Galaxy Z Fold 6. The Verge, for example, writes, The Galaxy Z Fold 6 sketch-to-image tool is ridiculous, fun, and slightly worrying. They write, Samsung would very much like us and its shareholders to know that its new phones are the AI-est phones that ever AI'd.
2:30And the Fold 6 that I'm testing comes with a new tool called Sketch to Image. Draw a rough sketch on a photo or an empty note page, and it will use generative AI to turn that into an image. I shrugged it off as just another AI thing when Samsung announced it on stage, but it's really good. So good that it worries me a little. Using the Sketch to Image tool in a note is pretty harmless. You draw something, highlight it, and choose from a variety of styles like 3D cartoon and illustration, and turn your doodle into something more detailed. The author says that they worked with their two-year-old to draw things like goofy-looking dump trucks and school buses.
2:59But then they say using sketch to image on a photo was where things got weird. The problem, they said, is something they called the bee problem. I took a photo off a dock just south of downtown Seattle with some flowers in the foreground. Because they're close to the camera and my focus was in the distance, they're slightly blurred. I drew the world's worst sketch of a bee on one of those flowers, figuring AI would insert an in-focus image of a bee, giving it away easily as a fake. Wrong. The issue, the author writes, If I didn't know the bee's origin story, there's no way I'd think twice about it if I scrolled past that on Instagram.
3:27I'd assume the photographer snapped the picture at just the right time and hung around waiting for a bee to fly into the frame. To be fair to Samsung, there is a little watermark that says AI-generated content down at the bottom, but it would be very, very easy to miss. The author did note that the results aren't always that good. For example, a giant pirate ship in the harbor is probably not tricking anyone. Ultimately, for this author, there are questions but no good answers, and ultimately just a recognition that the world we're moving into is very new and unpredictable in how it will all play out.
3:56Analysts also pointed out that Apple's AI strategy and Samsung's AI strategy could not be more different. Yahoo Finance writes, Samsung is angling to quickly build a large user base for its generative AI services, thus incentivizing developers to build apps for its Galaxy AI platform. To do that, they're making those AI features available on phones from the last three generations. Overall, they write, Samsung says it expects to put Galaxy AI on some 200 million devices by the end of the year. Apple, meanwhile, will basically only give the very most recent generations of the iPhone access to Apple intelligence.
4:25That's because, of course, Apple is trying to use this to boost iPhone sales. They want to use it as a way to get people to upgrade rather than holding on to their quote-unquote good-enough iPhones. Whether that works will likely have a big impact on how Apple's stock performs at the end of this year. Lastly today, Google has announced a new AI industry forum called the Coalition for Secure AI or COSAI. They write the new industry forum will invest in AI security and leverage Google's secure AI framework. Their first three work streams are software supply chain security for AI systems, preparing defenders for a changing cybersecurity landscape, and AI security governance.
5:01Founding members of the organization include Amazon, Anthropic, Chain Guard, Cisco, Coher, GenLab, IBM, Intel, Microsoft, NVIDIA, OpenAI, PayPal, and Wiz. And all of this will be housed under Oasis Open, the international standards and open source consortium. On the one hand, it feels to me like there are so many of these announcements that it's never exactly clear which ones are actually significant or not. But the fact that all of the big players are here, with the notable exception of Meta, suggests that this might be a fairly significant effort and one that we will keep an eye on. For now though, that is going to do it for today's AI Daily Brief Headlines Edition.
5:30Next up, the main episode. Today's episode is brought to you by Superintelligent, the platform for fun, fast AI learning. Super has a ton of new things going on. We recently announced our partnership with Spotify, through which users of that app can now access super intelligent content directly from their mobile apps. We've also just launched the AI learning feed. In addition to seeing the tutorials that we're dropping, there are polls, news items with related lessons, and a chance for people to show off the projects and use cases that are making AI come alive for them. We've also just kicked off the Super Summer Challenge, where each week we'll share a new challenge that you can use to discover new AI tools and use cases.
6:08Go to bsuper.ai and use code SUPERFUN for 50 % off your first two months, that's besuper.ai. Today's episode is brought to you by Venice. The leading AI companies store your entire conversation history and attach it to your identity forever. Every question you ask, every answer you receive, every image you generate, every thought you share with the machine, it's all being spied on. If you trust all the companies, hackers, and NSA board members that will ever have access to your AI conversations, then rejoice, for you are well served. For the rest of us, Venice is an alternative. Venice is a powerful AI app for text, image, and co-generation that respects you as a sovereign individual and believes privacy and free speech are not only human rights, but are necessary for civilizational advancement.
6:50Private, permissionless, and uncensored. You can try it for free without an account at venice.ai. Welcome back to the AI Daily Brief. Today we are talking about GPT-40 Mini, and it is both more powerful and cheaper than GPT-3.5 Turbo. Now, in addition to discussing GPT-40 mini specifically. We're also going to be talking about the changing nature of competition among models, price decreases among models, and what it all means for the state of AI. But first of all, let's talk about this announcement. The way that Sam Altman, CEO of OpenAI, teed it off was saying, towards intelligence, too cheap to meter.
7:27And cost, as we will see, really is the theme. In their announcement post, they write, OpenAI is committed to making intelligence as broadly accessible as possible. We expect GPT-40 mini will significantly expand the range of applications built with AI by making intelligence much more affordable. GPT-40 mini scores 82 % on MMLU and currently outperforms GPT-4 on chat preferences in LIMSYS leaderboard. It's priced at$0.15 per million input tokens and$0.60 per million output tokens, an order of magnitude more affordable than previous Frontier models, and more than 60 % cheaper than GPT-3.5 Turbo.
8:00OpenAI also gives a sense of what types of new use cases this low-cost opens up, including applications that chain or parallelize multiple model calls, applications that involve a full codebase or conversation history, and applications that interact with customers through what they call fast real-time text responses like customer support chatbots. Currently, the model supports text and vision, with full multimodal support for text, image, video, and audio coming. It's got a 128K token context window, as well as supporting 16 ,000 output tokens per request. And basically, they try to point out is the best-performing small model.
8:32It scores higher on the MMLU than Gemini Flash or Claude Haiku. Same on the MGSM, which measures math reasoning. On that, it scores much higher. They also point out that it scores much higher on coding performance, which seems to be one of the use cases that they're most focused on. Before releasing this broadly, they say they partnered with companies like Ramp and Superhuman, who they say found GPT-40 Mini to perform significantly better than GPT-3.5 Turbo for tasks including extracting structured data from receipt files or generating high-quality email responses when provided with thread history.
9:01Some of the response was lukewarm. Professor Ethan Malek writes, First impressions with GPT-40 mini is that it is impressive for a small model but no replacement for a frontier model. When given complex education prompts, it can't follow instructions as well and misses nuanced GPT-40 nails. At AI for Success writes, Point of view. When you are waiting for GPT-5 but GPT-40 mini is coming. And it shows a guy punching a laptop. They continue, OpenAI is on a mission to piss everyone off. Everart's Pietro Charano writes, never seen such a lukewarm reaction to an OpenAI model release. Sean Ralston, who does API dev support at OpenAI, writes, Pietro, this release is about price and performance.
9:40GPT 3.5 was no longer competitive against Gemini Flash, Cloud Haiku, etc. Many business customers using LLM APIs for routine functions, like company chatbots and standard work, need and want low token costs. GPT-40 mini sets a new standard of affordability with close to frontier model performance. Of course, we all want smarter models forthcoming, but those tend to require more compute and therefore are more expensive. Pietro responds, I agree on all these fronts, just sharing the overall sentiment. From my standpoint, however, the response has been exactly focused on what it seems like the important thing was, which is cost.
10:12Ben's Byte summed this up by saying, OpenAI's 4.0 mini brings big brains on a budget. It's a small model that beats the early version of GPT-4 and it costs 30x less than its big bro GPT-4.0. They sum up, you can now say bye to GPT 3.5 Turbo as 4.0 Mini will take its place in ChatGPT's free tier. If you were using GPT 3.5 Turbo API in your applications, you should switch to 4.0 Mini. That leads to two reasons to care. People are getting increasingly smarter AIs for free directly via ChatGPT and app developers can build powerful AI tools without breaking the bank. And really, I do think that cost is the story here.
10:46The end of OpenAI's blog post reads, The cost per token of GPT-40 mini has dropped by 99 % since Text DaVinci 003, a less capable model introduced in 2022. Vittorio says 99 % reduction in cost in two years is crazy. And that cost reduction has big implications. Back to that blog post again, they write, We envision a future where models become seamlessly integrated in every app and on every website. GPT-40 mini is paving the way for developers to build and scale powerful AI applications more efficiently and affordably. The future of AI is becoming more accessible, reliable, and embedded in our daily digital experiences.
11:21Vittorio even jokes, in 2026, OpenAI will give you money to use its models. Professor Ethan Mollick came back again and said, The big story with GPT-40 mini is its cost, and how quickly the price of intelligence of a sort has dropped. Now this is the point where I want to return to that now infamous Goldman Sachs report, Gen AI, too much spend, too little benefit. This is from June 25th, and I've already done a full takedown of this, but one of my biggest critiques is the argument from Goldman Sachs head of global equity research Jim Covello that he thought it was unlikely that the cost of AI was going to come down.
11:54The interviewer asks, even if AI technology is expensive today, isn't it often the case that technology costs decline dramatically as the technology evolves? Part of Jim's response was to say, the tech world is too complacent in its assumption that AI costs will decline substantially over time. Now, the point that he was making in this section is that he didn't think there was going to be enough competition for state-of-the-art GPUs to actually bring those costs down. What he seems to have missed is that we're officially in a time where the AI that we have right now is performant enough that there are many applications that don't require the absolute state-of-the-art, but instead just require near-state-of-the-art.
12:29And those costs are coming down and coming down dramatically. What's more, it's not even clear that costs of state-of-the-art aren't coming down dramatically as well. Claude 3.5 Sonnet, which is Anthropics' most intelligent model, and which many people have switched entirely, including me, basically, from GPT-4-2, does also represent a cost reduction. Claude 3.5 Sonnet, for example, costs one-fifth of what Claude 3 Opus does. Andre Carpathy wrote, LLM model size competition is intensifying backwards. My bet is we'll see models that think very well and reliably that are very, very small. Two parameters for which most people will consider GPT-2 smart.
13:07The reason current models are so large is because we're still being very wasteful during training. We're asking them to memorize the internet, and remarkably they do and can, e.g., recite SHA hashes of common numbers or recall really esoteric facts. Actually, LLMs are really good at memorization, qualitatively a lot better than humans, sometimes needing just a single update to remember a lot of detail for a long time. But imagine if you were going to be tested closed book on reciting arbitrary passages of the internet given the first few words. This is the standard pre-training objective for models today.
13:35The reason doing better is hard is because demonstrations of thinking are quote-unquote entangled with knowledge in the training data. Therefore, the models have to get larger before they can get smaller, because we need their automated help to refactor and mold the training data into ideal synthetic formats. It's a staircase of improvement, of one model helping to generate the training data for the next, until we're left with a perfect training set. When you train GPT-2 on it, it will be a really strong and smart model by today's standards. Just maybe the MMLU will be a bit lower because it won't remember all of its chemistry perfectly.
14:04Maybe it needs to look something up once in a while to make sure. This certainly reflects what I'm seeing. And what's more, I think it's responding to a business demand. There have been numerous posts recently about how enterprises are getting more sophisticated in their needs with AI. They realize that, yes, there are some cases that they need the absolute state-of-the-art for, and they need it immediately when they need it, and are willing to pay for the fastest-best answers. But there are tons of functions that don't need to be immediate. and that don't need the absolute state of the art. The Information recently published a piece called Why Smaller Could Be Better.
14:35They write, While some developers are racing towards super-intelligent AI technology, others are focused on building cheaper, more practical models. Over the past six months, several big tech companies, including Google and Microsoft, have released small language models in a bid to stake their ground in a burgeoning area of AI research. These models are lightweight enough to run on phones instead of on the cloud. They generally have fewer than 3 billion parameters, a tiny fraction of the more than 1 trillion parameters believed to support OpenAI's GPT-4. And this is a whole additional case that we haven't talked about, the fact that everyone is racing for on-device AI.
15:05The reality here is that the lower the cost, the more different types of applications can be built with AI in it. And that really seems to have been the motivation. OpenAI president and co-founder Greg Brockman wrote, We built GPT-40 mini due to popular demand from developers. We love developers and aim to provide them the best tools to convert machine intelligence into positive applications across every domain. It's hard to overstate how big a difference it makes to reduce cost by this much. Latenspace's Swix, for example, writes, If you're interested in LLMs for summarization, my evaluation of GPT-40 Mini is out.
15:35TLDR, Mini is the same or mildly worse in some cases, but because it's 3.5 % the cost of 4.0, I can run 10 versions of the Mini and use Mini 4.0 to judge and still have money left over to donate. Point being that we are just at the beginning of seeing how many more applications AI will come to as costs come down, and I think that's incredibly exciting. For now, though, that will do it for today's AI Daily Brief. Thanks for listening or watching as always, and until next time, peace.
From the publisher
OpenAI released GPT-4o Mini this week. It's both more powerful and 60% cheaper than GPT-3.5. In this episode, NLW discusses how model competition is moving towards small at the same time it's proceeding towards state of the art.
Concerned about being spied on? Tired of censored responses? AI Daily Brief listeners receive a 20% discount on Venice Pro. Visit https://venice.ai/nlw and enter the discount code NLWDAILYBRIEF.
Learn how to use AI with the world's biggest library of fun and useful tutorials: https://besuper.ai/ Use code 'podcast' for 50% off your first month.
The AI Daily Brief helps you understand the most important news and discussions in AI.
Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614
Subscribe to the newsletter: https://aidailybrief.beehiiv.com/
Join our Discord: https://bit.ly/aibreakdown
