How Big a Deal is Llama 4's 10M Token Context Window?

8 Apr 2025 · 24 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The AI Daily Brief: Episode Summary Episode Title: How Big a Deal is Llama 4's 10M Token Context Window? Podcast Title: The AI Daily Brief Date: [Insert Date Here]

Episode Overview This episode of "The AI Daily Brief" dives into Meta's recent launch of the Llama 4 models, which feature a significant 10 million token context window and a new architecture based on a mixture of experts. Despite impressive benchmark scores, real-world performance has been a point of contention among users.

---

Key Topics Discussed

  1. Meta's Llama 4 Launch
  2. Model Features:
  3. 10 Million Token Context Window: A significant increase compared to previous models.
  4. Mixture of Experts Architecture: Allows the model to use a subset of parameters for more efficient inference.
  5. Model Variants: Includes Llama 4 Scout, Maverick, and the upcoming Behemoth.
  • Performance Claims:
  • Scout is touted as the best multimodal model in its class.
  • Maverick reportedly beats competitors like GPT-40 and Gemini 2.0 Flash on various benchmarks.
  • Concerns Raised:
  • Many users report disappointing real-world performance.
  • Allegations surfaced that Meta may have manipulated benchmark results to showcase superior performance.
  1. Response from the Community
  2. Mixed Reactions:
  3. Some users felt Llama 4 failed to meet the hype, calling it "garbage" in practical applications.
  4. Others praised it for being fast and cost-effective, noting improved visual capabilities over previous versions.
  1. The Context Window Debate
  2. Importance of the 10M Token Window:
  3. Enables the model to handle larger tasks without losing coherence, particularly beneficial for coding assistants.
  4. Potentially shifts the conversation away from retrieval-augmented generation (RAG) models.
  • Diverse Opinions:
  • Some industry voices heralded it as a game changer, while others questioned its practical applications.
  • Concerns about slow performance and the feasibility of utilizing long context windows were prevalent.
  1. Meta's Strategic Positioning
  2. Market Context:
  3. The launch comes amid competitive pressures from other AI firms.
  4. Meta aims to commoditize foundation models and create an open ecosystem to leverage its social graph advantage.
  • Future Implications:
  • The 10 million token context window is viewed as a signal of future directions in AI capabilities rather than an immediate game changer.
  1. Summary of Llama 4's Performance
  2. Benchmark Controversy:
  3. Significant discrepancies were reported between benchmark scores and real-world user experiences.
  4. Critiques highlight a gap between lab performance and practical utility, raising questions about benchmark integrity.
  • User Feedback:
  • Many users reported issues such as freezing when run locally and poor coding capabilities.
  • The overwhelming sentiment is that while Llama 4 boasts impressive specs, the actual performance may not live up to expectations.

---

Conclusion The episode wraps up with reflections on the implications of Llama 4's launch. While the model introduces exciting features, the community's mixed reactions and reported performance issues suggest that the true potential of these advancements may take time to unveil. Meta’s strategy appears to be focused on long-term positioning in the AI landscape, signaling a shift in how AI capabilities are evaluated and utilized.

Key Takeaways

  • New Milestones: Llama 4’s 10 million token context window is a groundbreaking feature, potentially shifting the landscape of AI applications.
  • Market Reactions: Despite positive benchmarks, real-world performance has left much to be desired, leading to widespread skepticism.
  • Future Directions: The ongoing discussions about context windows and retrieval mechanisms indicate a dynamic and rapidly evolving AI environment.

---

For further insights and discussions, follow the podcast or subscribe to the newsletter for more in-depth analyses on AI developments.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Today on the AI Daily Brief, Meta launches Llama 4 with a massive new context window. Before that in the headlines, MidJourney launches v7. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. To join the conversation, follow the Discord link in our show notes.

0:23Welcome back to the AI Daily Brief Headlines Edition, all the daily AI news you need in around five minutes. Big model release day here on the AI Daily Brief. Our main episode is all about Meta's Llama 4, and Mid-Journey also has released their first new model in almost a year. Called V7, the model is obviously incredibly gorgeous. It's hugely capable in both photorealism as well as stylized modes. It features things like voice prompting, image personalization based on your own preferences, and multiple speed settings. The new model doesn't really introduce any novel features, it's just a better version of the Mid-Journey experience that has been so popular so far.

0:59Now, of course, the context for this is very different in the wake of OpenAI's ImageGen release. Indeed, when that came out and was using a different approach than Diffusion, many wondered if the AI approach that underpins MidJourney was going away, to be supplanted by natively multimodal image generation. Swix even tagged in David Holds, the founder of MidJourney, and said, I try not to drink hyperbole, but will MidJourney go the same way? Now that we have this in Gemini 4.0, I don't see how I ever go back to anything else. David's one-word answer was nah. On the release of V7, he wrote, This is an entirely new model with unique strengths and probably a few weaknesses.

1:38We want to learn from you what it's good and bad at, but definitely keep in mind it may require different styles of prompting. So play around a bit. The reaction was honestly a little mixed. Content creator Freebatar wrote, Two-year mid-journey power user. Gotta say it. Kind of disappointed. OpenAI set the bar sky high. Talk to your image gen like it's your bro. Mine equals blown. MJ7 looks more realistic, but did we really need that? Midjourney plus Magnific already nailed it. Might pause my sub to be honest. Yavi Lopez, the founder of Magnific added, yep, I mostly agree. But Midjourney still has that super cool artistic aesthetic and richness of styles.

2:13Though to be fair, they already had that in version 6.1 and previous. The problem is V7 doesn't really feel like V7. It feels more like version 6.2. Jamie Ortega spotlighted, they knew it wasn't ready yet, which is why it wasn't released. Then 4.0 ImageGen took off, and they were forced to put out whatever they had to gain momentum. It's just a retaliatory move, and a poor one at that. Professor Ethan Malek wrote, I'm a big fan of MidJourney if you want to do visually interesting AI images and have been using it since 2022. I like their new release, but the problem with the new v7 release today is that v6 was already really good.

2:44Experimenting will be fun, though. Tatiana Sigileva, however, the creative ambassador at Perplexity, doesn't know what everyone's complaining about, posting, my mind is blown exploring MidJourney V7, huge jump in quality, planning to have a lot of fun this weekend. Now, I've been thinking a lot about MidJourney ever since the release of the new ImageGen. On the one hand, the style and aesthetic and quality of MidJourney continues to just be tops for me relative to the other image generation models out there, and when I have any sort of deep creative or artistic project, it's still my go-to. However, even before ImageGen, I had found myself switching almost entirely to Ideogram when it came to day in and day out type of usage.

3:20Now, a big part of that was that Ideagram had better text adherence, so I could actually create cover images for podcasts and YouTube videos and things like that, which is a primary use case for me. But part of it is that it just had much better adherence to my prompt. It always felt to me like I was fighting with MidJourney, like it thought it knew better than I did how to make something cool. And so when I was willing to let the AI kind of do its thing, it was still great for that. It's getting harder and harder though now with ImageGen coming out, because not only now do those other tools have better prompt adherence, you also have the ability for inline editing where you can just talk at it and it actually makes the changes you want.

3:55That is so transformative and so different across so many different use cases. It makes the band of things that I want Midjourney for more and more narrow. I've generated tens of thousands of images on Midjourney, absolutely love it, and I'm a person who has no problem spending money on multiple different subscriptions. But even I myself am wondering at what point I'm going to pause Midjourney because I'm just not using it enough anymore. Midjourney is also interesting just as a case study. The company is bootstrapped and highly profitable, so it's not clear that they actually need to do anything other than maintain being a useful model.

4:26Back in December, David Holtz wrote, By VC standards, we should either conquer the world or die in a fire, and neither of these are spiritually compelling to me. I never wanted a company, I just wanted a home. At this point, we have a large and loyal paid community, we build tons of features for them, and they're pretty happy. We have enough revenue to fund tons of crazy R &D, and our models are still the best by the metrics we care about, which is how the images look and how fun it is to make things. We have a huge backlog of exciting things to make our models way better. Zero risk. We did this all with no investors.

4:52Honestly, it feels like we are successful. The next metric of success I think about most is a big, well-funded R &D lab with cool people free to work on whatever they want. Can we now build something that would make baby David proud? Can we now tell bold stories about a human future that people want to be part of? I think we can. And I think it's awesome that we have a company that's really pushing the boundary of doing it this way. At the same time, I will be interested to see how durable that large and loyal paid community is as all these things change around them. Next up, at an event celebrating their 50th anniversary, Microsoft has rolled out the agents.

5:24Copilot is now able to handle internet use with agentic features allowing the model to book tickets, reserve restaurants, and more. The feature is configured to work in the background so you can keep working on other tasks while having an agent run digital errands. The agentic assistant also now has memory so can remember the user preferences between sessions. Microsoft has also introduced a podcast generation feature similar to Google's audio overviews, and Copilot now finally also has a deep research feature. Now, of course, none of these features are pushing the limits. Basically, they just represent Microsoft keeping up with their AI rivals.

5:54Then again, as we see Apple and Amazon and their struggles, being able to offer feature-complete agentic AI if these things really work would still be a great deal better than some of the other big tech firms are doing. What's more, Microsoft AI CEO Mustafa Suleiman clearly sees the importance of agents commenting, I think that this will completely change the way we use computers forever. And speaking of Microsoft, the company is getting back in the game with an AI-generated version of Quake 2. They released a tech demo of the classic 90s shooter powered by its Muse generative game model. The demo is very basic, has blurry enemies, low resolution, and only replicates a single level.

6:27But it's still a playable game generated frame-by-frame using AI. Back in February when Muse was first unveiled, Microsoft gaming CEO Phil Spencer said,

6:58Researchers wrote in a blog post,

7:04the camera, jump, crouch, shoot, and even blow a barrel similar to the original game. Additionally, since it features in our data, we can also discover some of the secrets hidden in this level of Quake 2. At the same time, they were careful not to oversell the technology. They wrote that they didn't intend to fully replicate the original game, and that the demo should be thought of as, quote, playing the model as opposed to playing the game. Derek Strickland wrote, I'm trying the Quake thing and Gen.AI is just so weird. It's like playing a dream where you turn around and things are different, never ending hallways, etc.

7:30I know it's experimental, but it's still freaky. And indeed, whereas most of the conversation around this is just how useful it is for game preservation, I think it's much more interesting as an early example of what it's going to look like to generate games on the fly. A big part of my thesis for how the world evolves is more and more custom experiences, and part of that is going to be, I think, live generation of those experiences based on the user that's actually experiencing it. This is a very small step in that direction, but a step nonetheless. That, however, is going to do it for today's AI Daily Brief Headlines edition.

7:59Next up, the main episode. A quick note before we get into today's ads. For those of you who are looking for an ad-free experience, we are now up on Patreon. You can go to patreon.com slash ai daily brief. Right now, the benefits of this are that you will get the episodes without ads, and they will also come out a little bit earlier. We'll be exploring things like community and additional content in the weeks to come. But for now, I've heard from lots of you that you are looking for an ad-free experience. And so again, it's patreon.com slash ai daily brief. Thanks as always for supporting the show.

8:27Today's episode is brought to you by Vanta. Trust isn't just earned, it's demanded. Whether you're a startup founder navigating your first audit or a seasoned security professional scaling your GRC program, proving your commitment to security has never been more critical or more complex. That's where Vanta comes in. Businesses use Vanta to establish trust by automating compliance needs across over 35 frameworks like SOC 2 and ISO 27001. Centralized security workflows, complete questionnaires up to 5x faster, and proactively manage vendor risk. Vanta can help you start or scale up your security program by connecting you with auditors and experts to conduct your audit and set up your security program quickly.

9:08Plus, with automation and AI throughout the platform, Vanta gives you time back so you can focus on building your company. Join over 9 ,000 global companies like Atlassian, Quora, and Factory who use Vanta to manage risk, improve security in real time. For a limited time, this audience gets$1 ,000 off Vanta at vanta.com slash nlw. That's v-a-n-t-a.com slash nlw for$1 ,000 off. Today's episode is brought to you by Super Intelligent and more specifically, Super's Agent Readiness Audits. If you've been listening for a while, you have probably heard me talk about this, but basically the idea of the Agent Readiness Audit is that this is a system that we've created to help you benchmark and map opportunities in your organizations where agents could specifically help you solve your problems, create new opportunities in a way that, again, is completely customized to you.

10:00When you do one of these audits, what you're going to do is a voice-based agent interview where we work with some number of your leadership and employees to map what's going on inside the organization and to figure out where you are in your agent journey. That's going to produce an agent readiness score that comes with a deep set of explanations, strength, weaknesses, key findings, and of course a set of very specific recommendations that then we have the ability to help you go find the right partners to actually fulfill. So if you are looking for a way to jumpstart your agent strategy, send us an email at agent at bsuper.ai and let's get you plugged into the agentic era.

10:39Welcome back to the AI Daily Brief. Some exciting new model announcements to close out the end of last week. On Friday, Meta revealed their new Llama 4 family of models. As is the case every time Meta announces a new set of models, there is a lot to dig into here. These models feature all-new architecture, including multimodal functionality for the first time. The models are the first to utilize the mixture of experts architecture that most recently has been seen in DeepSeek. It's an architecture that allows the models to access a subset of parameters within a larger model, making inference more efficient.

11:09The Lama 4 family includes three different models. Lama 4 Scout is a 17 billion parameter model with 16 experts, which Meta claims is the quote best multimodal model in the world in its class, and is more powerful than all previous generation LAMA models, while fitting in a single NVIDIA H100 GPU. LAMA 4 Maverick has the same 17 billion active parameters, but includes 128 experts, basically meaning it's a total of 400 billion parameters. Meta states that the model is the, quote, best multimodal model in its class, beating GPT-40 and Gemini-20 Flash across a broad range of widely reported benchmarks, while achieving comparable results to the new DeepSeq v3 on reasoning and coding, at less than half the active parameters.

11:50Llama 4 Behemoth is still in training. It's set to feature 288 billion active parameters with 16 experts for a total of 2 trillion parameters. So Meta here is taking the same strategy that they did with Llama 3, which is release a couple of the smaller models early to get people excited, and then release the biggest version of the model a couple months later. Now, when it comes to Llama 4 Behemoth, this will be the first time a model has reached into the trillions of parameters that we know for sure, and the first mixture of experts model of this size, so we don't really know how model performance will be affected.

12:18Looking at costs, Lama Force seems to be pretty competitive. Inference service provider Grok has the hosted model available already. Scout costs$0.11 per million input tokens and$0.34 per million output tokens, while Mavericks prices are$0.50 and$0.77 per million for input and output respectively. In that, both models undercut DeepSeek Gemini 2.0 Flash and Quen's QWQ32B. When it comes to benchmarks, the new models look comparable to their peers. Scout outperforms models like Mistral 3.1, Gemini 2.0 Flashlight, and Gemma 3 on some benchmarks, while Maverick beats out GPT-40 and Gemini 2.0 Flash on most multimodal reasoning benchmarks.

12:57Notably, neither of those models are a true reasoning model utilizing chain of thought or test time compute. Now, one thing that's really important to note is the context into which Llama 4 is entering. A couple months ago, we got this leak from inside the company, which claimed that the meta Gen.AI organization was in panic mode. The leaker wrote, It started with DeepSeek V3, which rendered the Lama 4 already behind in benchmarks. Adding insult to injury was the unknown Chinese company with 5.5 million training budget. Engineers are moving frantically to dissect DeepSeek and copying anything and everything we can from it.

13:27I'm not even exaggerating. Management is worried about justifying the massive cost of the Gen.AI org. How would they face the leadership when every single leader of Gen.AI org is making more than what it costs to train DeepSeek V3 entirely, and we have dozens of such leaders? DeepSeek R1 made things even scarier. I can't reveal confidential info, but it'll soon be public anyways. It should have been an engineering-focused small organization, but since a bunch of people wanted to join the impact, grab, and artificially inflate hiring in the org, everyone loses. So this was the type of report that we were getting behind the scenes.

13:56And in the wake of these announcements, there was a lot of discussion about the feeling that maybe this release was rushed, and that there might even be something more nefarious than that going on. Min Choi writes, Yikes, Llama 4 benchmarks looked insane, but something feels off. Reddit leak claims Meta cooked it. In the 24 hours following the announcement, as people started to dig in, they seemed to be finding a fairly big difference in output between what Meta was claiming and what seemed to be the reality. TechCrunch writes, Researchers on X have observed stark differences in the behavior of the publicly downloadable Maverick compared to the model hosted on LMArena.

14:29The LMArena version seems to use a lot of emojis and give incredibly long-winded answers. Even more concerning was a Reddit post from someone who claimed that they were a meta-engineer. The post they shared said this, Despite repeated training efforts, the internal model's performance still falls short of open-source state-of-the-art benchmarks, lagging significantly behind. Company leadership suggested blending test sets from various benchmarks during the post-training process, aiming to meet the targets across various metrics and produce a presentable result. Presentable in air quotes. Failure to achieve this goal by the end of April deadline would lead to dire consequences.

15:02Following yesterday's release of Llama 4, many users on X and Reddit have already reported extremely poor real-world test results. As someone currently in academia, I find this approach utterly unacceptable. Consequently, I have submitted my resignation and explicitly requested that my name be excluded from the technical report of Llama 4. Notably, the VP of AI at Meta also resigned for similar reasons. There have been a lot of people referencing this post without a ton of verification yet. Bernie Tech wrote, Lama 4 gamed benchmarks so hard LMAO, completely out of touch with reality and practice.

15:33Andrew Allen summed it up this way. He wrote, Meta just dropped Lama 4 and scored number two on LMA Arena, beating GPT-40 and Grok, but users are calling it garbage and vaporware. Let's unpack the biggest benchmark controversy of 2025 so far. The numbers look incredible on paper. 10 million token context window, 1417 ELO score on LMA Arena, the second highest, beating many top-closed models. but something doesn't add up when users actually try it. He pointed to a tweet from Harshvardhan that writes, tried out Meta's Llama 4 for coding-related tasks, found it super basic and almost useless. Didi Das from Menlo Ventures writes, Llama 4 seems to be actually a poor model for coding.

16:09ELO maxing on LM Arena doesn't create the best models. Back to Andrew, he continues, the disconnect is stark, on paper second highest on LM Arena leaderboard, in practice super basic and almost useless for coding. Marketed revolutionary capabilities, reality struggling with basic instruction following. The most serious allegation? Meta may have submitted a different model for benchmarks than what's publicly available. This raises major questions about benchmark integrity. User reports highlight specific failures, freezing when run locally on Macs, poor coding capabilities compared to Claude and GPT, inability to follow instructions consistently, declining quality with longer contexts.

16:44Many users are calling the 10 million token context window marketing fluff that doesn't translate to better performance. And we'll be coming back to that 10 million token context window in just a minute. But Andrew also points out there are bright spots. It's fast, 512 tokens per second on Grok, cost-effective, improved vision capabilities over Llama 3, and open-source, enabling community innovation. Ultimately, he writes, what this reveals about AI development, benchmark scores do not equal real-world utility, the gap between lab performance and practical use is widening, and users increasingly value reliability over raw specs.

17:17Obviously, we'll get a lot more information in the days to come, and even if there hasn't been nefarious behavior here, there's still some pretty big gaps between the marketing promise and what people are actually finding in practice. Outside of all of that dubiousness, the big point of discussion and the thing that has everyone's mind racing is that Llama 4 Scout theoretically features a 10 million token context window. Until now, Google's development of a functional million token context window for their Gemini models was state-of-the-art. It was five times as large as the same class of models from OpenAI and Anthropic.

17:48Now, ultra-long context windows are a really big deal for a variety of use cases, for example, for coding assistants. The longer the context window, the more a coding assistant is able to ingest an entire code base to be understood all at once. For agents, long context allows for much longer tasks to be completed before losing coherence. Meta demonstrated the performance with a retrieval needle and a haystack test across 10 million lines of code. Scout didn't have a single failure across their testing. Now, independent benchmarks weren't anywhere near as impressive. And yet still in this case, most of the conversation wasn't so much about Meta and Llama 4 specifically, but about what the implications are as the tech improves.

18:29Representing around a million variants on this take, Marvin Aziz, the community manager at Lindy, wrote, RAG is dead. Why bother with a knowledge base when you can shove 10 million tokens into a context window and call it a day. RAG, of course, refers to retrieval augmented generation, which is the process of hooking up an LLM to a database or knowledge source to search up any information it might need. Then again, the opposite take was just as prolific. With AI evaluations designer Hamil Hussain writing, RAG is dead post R annoying AF. R is retrieval and AG is the LLM. This means you think retrieval is dead.

18:59Seriously, you think retrieval is dead? Keyword search, metadata filtering like dates and users, grep and other filtering are retrieval. Good luck without retrieval. Charles Fry writes, RAG is dead is also the sort of thing only said by someone who has never run LLM inference themselves, let alone been on the hook for cost and latency. Enigmatically, Swix writes, unpopular opinion right now, but LLAMA4's 10 million token window will finally actually end the long context versus RAG debate, but not the way the other guy is thinking. For those trying to toe a more middle of the line, they basically point out that we just don't know enough yet to declare the end of RAG or really understand how well long context windows are going to work.

19:32Nir Sian wrote, I haven't played with the Llama 4 series, but needle in a haystack is woefully insufficient to know the strength of a context window. If you want needle in a haystack, we have grep for that. Grep is a Linux command for searching databases. OpenAI co-founder Andre Karpathy falls into the category of wanting to believe, adding, My reaction to when reading all the rag is dead tweets earlier today. Huge amount of optimism that the context window is also usable in practice for real problem solving and not just in theory. Could very well be true, I just don't super know. Now, the community with the most enthusiasm about an ultra-long context window was the Vivecoders.

20:05Plain Game creator Peter Levels wrote, This is insane and makes it finally possible to Vivecode up to giant code sizes. The limit just weeks ago was context window. AI would get lost once your Vivecoded game or app became too big. Imagine an AI with memory loss that starts breaking stuff. With 10 million tokens, there's practically no limit. Really quite big for Vivecoding and another big hit for the perpetual naysayers. AI consultant Sasha Lecti added,

20:36Still,

20:40there were many trying to harsh the vibes with practical issues of using a 10 million token context window. They assumed that loading that many tokens would be painfully slow, and questioned whether a Gemini 2.0 flashlight class model would be up to the task of generating functional code. Developer Nick Dobos rebutted, Lazy take. Use it to ask questions and plan. Use the high-tier models to write the actual code. Not hard. His point being that even if the model isn't really up to writing code or even developing a plan, simply creating an outline of a large code base for use in another LLM is a new feature that hasn't previously been accessible.

Read the full transcript

21:10LinkedIn co-founder Reid Hoffman had a less combative take, posting, Spending the day playing with Llama 4. One of the many interesting things, the massive context window is a game changer. I don't think it's the end of RAG, but for a surprising number of workflows, the long context alone is enough. And I think this is an important point. Ultra-long context doesn't have to be perfect or completely replace RAG to be a really big deal. To the extent that it holds up at all, this feature could unlock a huge range of functionality that wasn't possible before. Orchestration platform Obelix commented that this is just one tool in future workflow design, writing, Long context doesn't replace RAG, but it absolutely shifts the trade-offs.

21:45For structured, contained workflows like contracts, single docs, or chat history, context alone is simpler, faster, and good enough. RAG still shines when you need external dynamic or filtered retrieval. The future probably blends both. Long context for memory, RAG for knowledge access, orchestrators for choosing the best tool in real time. Matthew Berman zoomed out even more. While noting a ton of Llamaforce shortcomings, he added, Here's the strategic insight that everyone's missing. Meta's 10 million token context window isn't about today's performance. It's about signaling tomorrow's direction.

22:14They're showing us a future where AI doesn't just retrieve knowledge, but transforms your entire knowledge base into manipulable working memory. Zuckerberg understands the truth Google accidentally leaked. Closed-source AI has no moat. Foundation models are becoming commodities faster than anyone predicted, and meta is accelerating this transformation. Meta strategy becomes clear when you connect the dots. Commoditize foundation models through open source, make context the new competitive battleground, force innovation up the application layer, leverage their massive social graph advantage, and ultimately create an open ecosystem where social and application data become the true moats.

22:46Still, at the end of the day, as much as they are helping shape the conversation, it's hard not to view this release so far as a disappointment. Professor Ethan Malek even commented that their flagship model doesn't stack up, writing, Looks like even Llama Behemoth doesn't come that close to Gemini 2.5, so no open model parity with the state-of-the-art enclosed models. We will see what happens when people slap a reasoner on Llama, though. Doesn't seem like they're launching with one. And indeed, this was another common take. Andrei Burkoff writes, If today's disappointing release of Llama 4 tells us something, it's that even 30 trillion training tokens and 2 trillion parameters doesn't make your non-reasoning model better than small reasoning models.

23:21Model and data size scaling are over. And so as we wrap up here, I'm not yet exactly sure what to make of this. On the one hand, it feels a little rushed. It does seem like the deep seek pressure is getting to meta. At the same time, given that they are taking an open strategy, the consequences of releasing early are a little bit less severe for them than perhaps for other companies. If it's cost effective, better than some of the things that people had access to before, there's still going to be a lot of developers building on it. Indeed, holding aside wanting every single model to break the mold every single time, ultimately for developers, this just represents another set of choices, which in a very fast-moving environment is nothing but a good thing.

23:59For now though, that is going to do it for today's AI Daily Brief. Appreciate you listening or watching as always, and until next time, peace!

24:13Thank you.

From the publisher

Meta’s new Llama 4 models have a massive 10 million token context window and a fresh architecture using a mixture of experts. Scout and Maverick are out now, and Behemoth is still training. Despite strong benchmark scores, many users report underwhelming real-world performance.
Get Ad Free AI Daily Brief: ⁠⁠https://patreon.com/AIDailyBrief⁠⁠


Brought to you by:

KPMG – Go to ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://kpmg.com/ai⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ to learn more about how KPMG can help you drive value with our AI solutions.

Vanta - Simplify compliance - ⁠⁠⁠⁠⁠⁠⁠https://vanta.com/nlw

The Agent Readiness Audit from Superintelligent - Go to https://besuper.ai/ to request your company's agent readiness score.

The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614Subscribe to the newsletter: https://aidailybrief.beehiiv.com/Join our Discord: https://bit.ly/aibreakdown

More from The AI Daily Brief: Artificial Intelligence News and Analysis

All 1,099 episodes
How Big a Deal is Llama 4's 10M Token Context Window?The AI Daily Brief: Artificial Intelligence News and Analysis · 24 min
Listen in VO