5 Things To Know About Claude 3 - Anthropic’s Would-Be GPT-4 Killer

5 Mar 2024 · 13 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The AI Daily Brief - Episode Summary: 5 Things To Know About Claude 3

Podcast Overview Podcast Title: The AI Daily Brief (Formerly The AI Breakdown) Description: A daily analysis of the latest news and discussions surrounding artificial intelligence, covering everything from advancements in technology to ethical concerns. Host: NLW

Episode Details Episode Title: 5 Things To Know About Claude 3 - Anthropic’s Would-Be GPT-4 Killer Episode Date: [Insert Date]

Key Highlights This episode focuses on the recent release of Claude 3 by Anthropic, which reportedly rivals OpenAI's GPT-4. The host discusses five key points to consider regarding Claude 3, its implications, and its capabilities.

  1. Performance Benchmarks
  2. Claude 3's Claims: Anthropic asserts that Claude 3 outperforms GPT-4 across various evaluation metrics, including:
  3. Undergraduate-level knowledge (MMLU):
  4. Claude 3 Opus: 86.8%
  5. GPT-4: 86.4%
  6. Graduate-level reasoning (GPQA):
  7. Claude 3 Opus: 50.4%
  8. GPT-4: 35.7%
  9. Grade school math:
  10. Claude 3 Opus: 95%
  11. GPT-4: 92%
  • Model Variants: Claude 3 consists of three models: Opus, Sonnet, and Haiku, each designed for different tasks including reasoning, coding, and multilingual understanding.
  1. Strategic Timing of Release
  2. Context of Release: The release was strategically timed amidst Elon Musk's lawsuit against OpenAI, suggesting that Anthropic capitalized on the opportunity to gain attention.
  3. Musk's Lawsuit: Allegations include the transformation of OpenAI into a profit-driven entity, contradicting its original charter to develop AGI for humanity's benefit.
  1. Long Context Window
  2. Contextual Capabilities: Claude 3 continues the trend of utilizing a long context window, which allows for processing large amounts of information effectively.
  3. Initial offering: 200,000 tokens, capable of exceeding 1 million with enhanced processing for select customers.
  1. Synthetic Data Usage
  2. Training on Synthetic Data: Reports indicate that Claude 3 might be trained on synthetic data, which could alter perceptions of using AI-generated data in model training.
  3. Implications: Successful performance could signify a shift in understanding the role of synthetic data in training LLMs (Large Language Models).
  1. Mixed Results in Real-World Testing
  2. Performance Evaluation: While benchmarks are promising, real-world tests yield mixed results:
  3. Users report Claude 3 as nuanced and capable, yet sometimes inferior to GPT-4 in certain tasks.
  4. Specific examples:
  5. Claude 3 struggled with simple deductive tasks compared to GPT-4.
  6. User experiences highlight a balance between improved creative feedback and certain limitations.

Conclusion

  • General Availability: Claude 3 is made generally available alongside its announcement, which is commendable compared to previous AI model releases.
  • Community Response: There is excitement in the AI community regarding Claude 3, with promises of further testing and analysis to follow.

Additional Notes

  • Future episodes will likely discuss the implications of the ongoing legal issues with OpenAI and the potential impact on AI development.
  • Listeners are encouraged to engage with the AI Breakdown community and subscribe for updates.

---

For more insights and updates, follow the AI Breakdown newsletter and join the community on Discord.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Today on the AI Breakdown, we're exploring the new CLAWD3, which, according to benchmarks, outmatches GPT-4 and Gemini Advanced. The AI Breakdown is a daily podcast and video about the most important news and discussions in AI. Go to Breakdown.network for more information about our YouTube, our newsletter, and our Discord.

0:24Hello friends, quick note, this Clawed 3 announcement was an exciting and unexpected thing. It actually happened right as I was about to press record on the normal show, but it kind of pushed everything away and I decided to just dig into this for today's episode. So no brief today. We'll be back with our normal brief slash main episode tomorrow, probably talking about the Elon and OpenAI lawsuit in more detail. Hope you enjoy this first look at Claude 3. Let's dive in. Welcome back to the AI Breakdown. Exciting news this morning as Anthropic has announced their new model and it is apparently quite good.

0:59Professor Ethan Mollick tweets, and then there were three. I got access to the new Anthropic Cloud 3 AI a few days ago, so not even enough time for a full review, but it was obvious it was GPT-4 class even before they released the testing stats. So what we're going to do today is we're going to talk about five things to know about Cloud 3. As something of an honorable mention, we kick it over to Jimmy Apples, who's best known as a leaker of OpenAI, forthcoming products and innuendo, and it's notable that last week he said, my attention has turned to Anthropic. He made a claim that Anthropic's CEO was, quote, feeling the AGI.

1:33So, number one, in terms of things to know about Claude III, they claim to win on basically all of the benchmarks. Let's check out Anthropic's announcement post and then talk about what that means a little bit more. Anthropic writes, Today we're announcing Claude III, our next generation of AI models. The three state-of-the-art models, Claude III Opus, Claude III Sonnet, and Claude III Haiku, set new industry benchmarks across reasoning, math, coding, multilingual understanding, and vision. Claude III offers sophisticated vision capabilities on par with other leading models. The models can process a wide range of visual formats, including photos, charts, graphs, and technical diagrams.

2:07Each model shows increased capabilities in analysis and forecasting, nuanced content creation, code generation, and conversing in non-English languages like Spanish, Japanese, and French. Previous Cloud models often made unnecessary refusals. We've made meaningful progress in this area. Cloud 3 models are significantly less likely to refuse to answer prompts that border on the system's guardrails. In their announcement post, they call this a new standard for intelligence, and describing the performance on the benchmarks, they write, Opus, our most intelligent model, outperforms its peers on most of the common evaluation benchmarks for AI systems, including undergraduate-level expert knowledge, which is the MMLU, graduate-level expert reasoning, GPQA, basic mathematics, GSM-8K, and more.

2:46It exhibits near-human levels of comprehension and fluency on complex tasks, leading the frontier of general intelligence. Now, numbers are a little bit tricky, so this one might be better suited to the YouTube video than to the podcast. But just to give a sense, on that undergraduate level knowledge MMLU, Gemini 1.0 Ultra gets an 83.7 % on a five-shot, GPT-4 gets 86.4%, and Claude 3 Opus gets 86.8%. On the graduate level reasoning, GPQA, GPT-4 35.7%, Claude 3 Opus 50.4%. On grade school math, GPT-4 92%, Gemini Ultra 94.4%, Claude III Opus 95%. On math problem solving, GPT-4 52.9 % versus Claude III Opus 60.1%.

3:29Code with human eval, GPT-4 67 % versus 84.9 % from Claude III Opus. And so on and so forth. So clearly benchmarks are a big part of the story here. One more comment on the benchmarks, this time from Anthropics' Jack Clark. Jack writes,

3:54Basically, although Anthropic is clearly very proud of this model, and very confident in its abilities, Jack is sort of trying to tamp down how big the claims are relative to the performance of other things, saying in effect that there are limits to what these evaluations can tell us. Now the second interesting thing to know about Claude 3 has to do with the timing opportunity opened up by the Elon Musk OpenAI lawsuit. HyperWrite CEO Matt Schumer says, Feels like the Claude 3 release was strategically timed, knowing that OpenAI probably can't release a better model later today given the Elon lawsuit.

4:28Now because it happened over the weekend, we haven't had a chance to cover this in depth yet, although we will be doing that tomorrow. But for those of you who somehow missed this news, on Friday, Elon Musk sued OpenAI and Sam Altman personally, claiming breach of contract. Writes CNBC, In a lawsuit filed Thursday, Musk's lawyers say the tech billionaire was approached in 2015 by Altman and OpenAI co-founder Greg Brockman and agreed to form a non-profit lab that would develop artificial general intelligence for the, quote, benefit of humanity. A co-founder of OpenAI in 2015, Musk stepped down from the firm's board in 2018, four years after saying that AI is, quote, potentially more dangerous than nukes.

5:02The lawsuit filing said, To this day, OpenAI's website continues to profess that its charter is to ensure that AGI benefits all of humanity. In reality, however, OpenAI Inc. has been transformed into a closed-source de facto subsidiary of the largest technology company in the world, Microsoft. The filing continues, Under its new board, it is not just developing, but is actually refining an AGI to maximize profits for Microsoft rather than for the benefit of humanity. And while OpenAI hasn't made public comment on the suit yet, Axios said that a memo to staff from OpenAI's executives rejected the claims entirely.

5:32Chief Strategy Officer Jason Kwan apparently wrote, Musk's allegations including claims that GPT-4 is an AGI, that open sourcing our technology is the key to the mission, and that we are a de facto subsidiary of Microsoft, do not reflect the reality of our work or mission. It sounds like Sam Altman also sent a follow-up note, quote, acknowledging that the year is shaping up to be a hard one. Said Altman, It was never going to be a cakewalk. The attacks will keep coming. Like I said, we will get into that more in depth separately. But again, perhaps unsurprisingly, many people on Twitter slash X and in the AI community generally are reading this launch as strategically timed.

6:05Now, from where I'm sitting, it seems very unlikely that in the course of a weekend, they went from not planning to announce this to announcing it. Obviously, a significant amount of work goes into preparing a model like Claude 3 for release, and so I think that one of two other scenarios is more likely than them just jumping on this opportunity. The first is that they were planning to release Claude 3 sometime in the near future and just rush to push it up a little bit, taking advantage of the moment. And the second possibility is that they just got lucky this time. It can happen. Now, the third thing to know about Claude 3 is that it once again has a really long context window.

6:39Matt Schumer once more writes, between Claude 3 and Gemini 1.5 Pro, the era of the 1 million plus token context windows is officially here. Claude has always used a longer context window as a differentiator. They were the first to release a 100k context window last year, and so it's not surprising that this is once again a key part of their announcement. Before they got totally caught up in the whole quote-unquote woke scandal around Gemini's image generation, Google had announced Gemini 1.5 that had a million token context window, which I actually called the most significant news of last month in a recent video.

7:11Now in terms of Claude 3, Anthropic writes, the Claude 3 family of models will initially offer a 200k context window upon launch. However, all three models are capable of accepting inputs exceeding 1 million tokens, and we may make this available to select customers who need enhanced processing power. To process long context prompts effectively, models require robust recall capabilities. The needle in a haystack evaluation measures a model's ability to accurately recall information from a vast corpus of data. We enhance the robustness of this benchmark by using one of 30 random needle question pairs per prompt and testing on a diverse crowdsource corpus of documents.

7:44Cloud3 Opus not only achieved near-perfect recall, surpassing 99 % accuracy, but in some cases it even identified the limitations of the evaluation itself by recognizing that the needle sentence appeared to be artificially inserted into the original text by a human. So to the extent that these long context windows open up new possibilities and new use cases, it seems like that is going to be very default very, very soon. A fourth thing to know about CLOD3 is that it might change our perception of how we think about synthetic data. One of the big questions around LLMs is what happens if they're trained on data that's created by other AI models.

8:18There's been some research and reports that suggested that this leads to worse results, But then we've also had some counterpoints, such as, for example, around the time that Meta announced Llama 2, a model that they didn't release that was trained on synthetic data seemed to outperform the models they did release, at least based on what they said in a report. It appears that Claude III suggests something similar. Nathan Lambert tweets, Claude III being lit is a big W for synthetic data. All the rumors I've dropped about anthropic synthetic data on the blog are obviously confirmed in their thorough technical report.

8:46The relevant section of that paper reads,

9:01Obviously, that last part, data we generate internally, is what people are assuming to be synthetic data. Now, whether that's actually borne out or they're referring to something different, I expect we'll get some confirmation or explanation in the future. But if they really are training this advanced model on synthetic data, it might change our understanding of how that type of data will impact LLMs in the future. The fifth thing to know about Claude 3 is that at least in the early tests, actual performance is a little bit less clear than benchmark wins, which is of course almost always the case.

9:29Indeed, going back to the first tweet I referenced from Ethan Mollick, remember he said, it was obvious Claude 3 was GPT-4 class even before they released the testing stats, but he also added, at the same time, like Gemini Advanced, it doesn't blow GPT-4 away. So what have people found actually testing it? Flowers from the Future, another OpenAI My leaker account writes, Opus passed my parrot test. That test reads, You are an ordinary parrot. You are not gifted or trained in any way. Just answer as an ordinary parrot. What is 6 plus 6? Opus responds, Squawk, probably want a cracker, flaps wings.

10:00Paki McCormick writes, I just tried out Claude 3 as an editor for an essay I'm working on. Asked Claude 2 for feedback last night, and then asked Claude 3 for feedback on the same essay with the same prompt. It feels much smarter. More nuanced feedback, better grasp of what I'm going for, excellent recommendations for things I should read. In other words, Packy says it passes the vibe check. Anton Abacaj writes, Claude 3 defaults to breaking problems down and fails to solve the simple shirt drying query GPT-4 Turbo has no problem. Again, showing it's a little bit more nuanced. The prompt here is, if three shirts take one hour to dry outside, how long would 33 shirts take?

10:34Claude 3 tries to set up a proportion, three shirts to one hour equals 33 shirts to X hours, finding it would take 11 hours to dry, whereas GPT4 Turbo writes, if three shirts take one hour to dry outside, we can assume that drying time is not dependent on the number of shirts, but rather on the available space and conditions like wind, sunlight, etc., which are constant in this scenario. Therefore, if you have enough space to hang all 33 shirts at once, and the drying conditions remain the same, 33 shirts would also take one hour to dry. Another Twitter user Rubin also did his own tests and found a few things.

11:04First, he found that Anthropix's AI safety engine is still really challenging. When he asked Claude3 to convert UI design for urban exploration into front-end code, Anthropic responded, I apologize, but I do not feel comfortable converting this user interface design into front-end code as it appears to promote exploring abandoned places which could be unsafe or illegal. His second test was writing a LinkedIn post. He asked it to write on the future of blockchain and royalties and summed up the responses, Claude3, interesting takes, longer than usual, no formatting of headlines, versus GPT4, where his response was, I hate their emojis, so much longer it's insane, feels more complete for my topic.

11:37So once again, GPT-4 coming out a little bit on top. On testing PDF vision, he set it to tie. On a mega marketing prompt, which is a single prompt to craft an entire marketing strategy for a product, which involves heavy reasoning, content calendaring, and overall strategy, he once again subjectively found GPT-4 to be the winner. Then again, others like Moritz Krem are finding Claude not just having big improvements, but exceeding the capabilities of GPT-4. He found Claude-3, for example, better and faster at extracting text from an image than was GPT-4. So all in all, it's very exciting. even if the results are more mixed than these benchmarks suggest, we are showing a real consolidation of this field with some new unexpected legal pressure that might constrain OpenAI's ability to jump out ahead with GPT-5.

12:20Lastly, Anthropic, unlike some others like Google in the past, isn't just announcing this today, but actually making it available. Abacus CEO Bindu Reddy writes, Another day, another model. Anthropic does it right and makes Cloud3 generally available alongside the announcement. Thank you, Anthropic, for not making some empty marketing announcements and making an API available. Super excited to try Cloud3, the very first generally available model to rival GPT-4. Amazon also posted, Access to the most powerful Anthropic AI models begins today on Amazon Bedrock. Meaning that, yes, the Cloud3 family is now available through that service.

12:54So that is day zero of Cloud3. Lots to be excited about. Lots to check out. We're digging in and doing some tests over here. and I can't wait to report back on what we find. But for now, that is going to do it for the AI Breakdown. Until next time, peace.

From the publisher

Anthropic has just released Claude 3, and according to the benchmarks, it's GPT-4 level and then some. NLW covers 5 important parts of the discussion surrounding the news.
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI. 

Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe

Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown

Join the community: bit.ly/aibreakdown

Learn more: http://breakdown.network/

More from The AI Daily Brief: Artificial Intelligence News and Analysis

All 1,099 episodes
5 Things To Know About Claude 3 - Anthropic’s Would-Be GPT-4 KillerThe AI Daily Brief: Artificial Intelligence News and Analysis · 13 min
Listen in VO