Early Uses for Anthropic's Claude 3.5 and Artifacts

21 Jun 2024 · 15 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The AI Daily Brief: Episode Summary

Podcast Information

  • Podcast Title: The AI Daily Brief (Formerly The AI Breakdown)
  • Description: A daily news analysis show on all things artificial intelligence, exploring creativity, industry disruptions, and philosophical questions surrounding AI.

Episode Title

  • Title: Early Uses for Anthropic's Claude 3.5 and Artifacts
  • Description: This episode discusses the launch of Anthropic's Claude 3.5 Sonnet, highlighting its performance metrics, new features, and early use cases.

---

Episode Summary

Introduction

  • The episode opens with discussion surrounding the release of Anthropic's latest model, Claude 3.5 Sonnet.
  • Brief commentary on AI becoming a political topic, referencing former President Donald Trump's remarks on AI and energy needs.

Key Highlights

  • Trump's Comments on AI
  • Emphasizes the increasing energy demands of AI technologies.
  • Relates AI competitiveness with global energy solutions, particularly in the context of competition with China.
  • Government Engagement with AI
  • OpenAI’s Mira Mirati mentions providing governments early access to new AI models for education and risk mitigation.

Funding Announcements in AI

  • HeyGen: Video avatar company raised $60 million, reaching over $35 million in annual recurring revenue (ARR).
  • Poolside: A French AI-focused company is raising $400 million at a $2 billion valuation, emphasizing the growth of AI companies in France.
  • Cerebrus: Chip manufacturer filed for IPO, benefiting from the surge in AI hardware investments.

Introduction to Claude 3.5 Sonnet

  • Performance Metrics:
  • Outperforms GPT-4 in various benchmarks including coding evaluations and reasoning tasks.
  • Improved capabilities in visual understanding, including interpreting charts and transcribing text from images.
  • Safety Evaluation: Claude 3.5 Sonnet was provided to the UK's AI Safety Institute for pre-deployment testing.

User Experience Enhancements

  • Artifacts Feature:
  • Introduction of a new interface allowing users to generate documents, code, diagrams, and games.
  • Improved user interaction by separating input instructions from output previews.

Early Use Cases

  • Practical Applications:
  • Examples of creating SOP documents, visual graphs, and coding simple games.
  • Highlighted user experiences shared on platforms like Twitter, showcasing the creative and functional potential of Claude 3.5 Sonnet.

Industry Reception and Critique

  • General excitement around the advancements in AI, particularly the enhancements offered by Claude 3.5 Sonnet.
  • Some skepticism regarding whether these advancements represent a significant leap in AI capabilities.
  • Discussion around Dario's claims about not pushing the AI frontier and contrasting it with the actual improvements seen in Claude 3.5 Sonnet.

Conclusion

  • The episode wraps up by encouraging listeners to explore the new Claude model and share their creations.
  • Reminder to check out additional tutorials and community engagement through the Super platform.

---

Key Concepts and Takeaways

  • AI and Energy: A critical relationship where the growing demands of AI technologies are noted alongside the need for innovative energy solutions.
  • Funding Landscape: Robust funding activity in the AI sector signifies ongoing growth and interest.
  • Advancements in AI Models: Claude 3.5 Sonnet illustrates tangible improvements in performance, with a strong emphasis on user experience and interface design.
  • Engagement with Government: The evolving relationship between AI developers and government bodies may signal a new phase in AI governance and safety considerations.
  • Incremental vs. Revolutionary Advances: Discussions reflect on the nature of progress in AI, recognizing both small improvements and their substantial impact on productivity.

Call to Action

  • Listeners are encouraged to experiment with the new Claude model and share their experiences, fostering community interaction and innovation.

---

This summary captures the essence of the podcast episode while outlining significant discussions and insights pertaining to AI developments, particularly focusing on Anthropic's latest model.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Today on the AI Daily Brief, Anthropic releases its latest model. Before that in the headlines, is AI becoming a presidential issue? The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. To join the conversation, follow the Discord link in our show notes.

0:22Welcome back to the AI Daily Brief Headlines Edition, all the AI daily news you need in around five minutes. We kick off today with some comments from former President Donald Trump, who is of course also running for president right now. He recently appeared on the All In podcast, part of a courting of the technology industry that he's doing in his campaign right now, and spent a bit of time talking about artificial intelligence as it relates to energy. Let's quickly listen to this one minute clip. We have a phenomena coming up right now, and I was talking about it the other day to David, it. And that's AI, little things, simple, two little simple letters, but it's big.

0:59And I realized the other day, more than anything, when we were at David's house and talking to a lot of geniuses from Silicon Valley and other places, they need electricity at levels that nobody's ever experienced before to be successful, to be a leader in AI. The amount of electricity that And it's like double what we have right now and even triple what we have right now. It's incredible how much they need to be the leader. And we're going to have to be able to do that. And a windmill turning with its blade, knocking out the birds and everything else is not going to be able to make us competitive.

1:39You'll have China. What about nuclear, Mr. President? Yeah. So let me just give you a statistic on this. Nuclear is okay. So what's interesting to me about this is not only that like him or loath him, when Trump speaks, it does set the tone of the conversation to come, but more that we're not just seeing a superficial surface level conversation around AI competitiveness and AI as it relates to a global geostrategic battle with China, but this specific awareness of the relationship between AI and energy. This is something that Sam Altman has talked about a lot before as well. When he's been asked about whether he's concerned about the environmental impact of artificial intelligence, he's basically said that we're going to have to innovate new energy solutions for the sake of AI because we simply can't get enough with what we have.

2:22I'll be watching closely to see if AI starts appearing in this type of political discourse more. But for now, speaking of the government and AI, also in a recent interview, OpenAI's Mira Mirati said that they give the government early access to new AI models. How do you minimize risk and providing people the tools to do that? And in the case of government, for example, it's very important to bring them along and give them early access to things, educate them. As you'll hear more in the main episode today, as part of its announcement of its new model Claude 3.5 sonnet, Anthropic shared more details about how they gave access to the UK AI Safety Institute.

3:07So again, maybe a bit of a shifting of the tide in terms of how these frontier labs think about their relationship with the government, even in advance of any sort of comprehensive legislation. Next up today, we've got a set of different funding announcements. HeyGen, the video avatar and visual storytelling company has announced a$60 million Series A. They write that in just over a year, they've grown from$1 million in ARR to over$35 million and have been profitable since quarter two of last year. What's more, they say more than 40 ,000 companies are currently using the tools. These numbers are pretty impressive to me, especially from the standpoint that we're only just beginning to scratch the surface on avatar use cases in the workplace.

3:46I happen to think that this is going to be an incredibly rich area of experimentation where we're going to see everything from company onboarding to change sales processes and 60 million more in the war chest will definitely give HeyGen the ability to compete at the front of that pack. Another French company has also raised a big round. Poolside is, according to TechCrunch, raising$400 million at a$2 billion valuation. This is not yet an announced deal. TechCrunch writes that Bain Capital Ventures and DST are in talks to co-lead the round, with BCV being a previous investor and DST being a new investor.

4:17You might remember this company from when they raised a$126 million seed round last year. Poolside is focused on using AI to accelerate and augment human software developers. TechCrunch points out that not only is France pumping out some very high-profile AI companies, not one but two others have also raised nine-figure seed rounds. Mistral raised$113 million, and H raised$220 million. TechCrunch writes the City of Light might need to be renamed the City of AI at this rate. AI language tutor Speak announced that it had reached a half-billion-dollar valuation, raising$20 million in a Series B3 financing.

4:55The company writes,

5:06An impressive statistic from this press release, learners speak a thousand times on average in their first week. Finally today, on the other end of the financing spectrum, the information reports that chip company Cerebrus has quietly filed for an IPO. The information writes, The IPO plans show the eight-year-old company wants to ride a wave of investor enthusiasm over AI hardware sales that have made NVIDIA the world's most valuable company and boosted scores of other stocks. The startup's financial results couldn't be learned. It said in a blog post in December that it had recently reached cash flow break-even without elaborating.

5:39Still, the information writes that a new share authorization suggests Cerebrus is valuing itself at around$2.5 billion. So friends, that is the news from here, and that's going to do it for today's headlines. Next up, the main episode. A quick note before we get back to the show, today's episode is brought to you by Super Intelligent. Super is, of course, the platform that we built and released a couple months ago to help people learn how to use AI. It's built around fun, fast tutorials that get you actually using AI in minutes, not hours and certainly not days. Their learning all happens in the context of an engaged community and we've got a bunch of exciting features rolling out in the weeks to come.

6:16One of those is a new Teams version of the platform, which includes a custom curated playlist, as well as a showcase where people on your team can share their use cases and projects in AI across the organization. If you are interested in being a part of the Super for Teams beta, go to bsuper.ai slash partner so you can learn about the program. Welcome back to the AI Daily Brief. Today we have one of my favorite types of stories, which is when a very cool new tool gets released and people get all excited to go try it out. Of course, this tool coming from one of the biggest labs in the space, Anthropic, and it beating out GPT-4.0 on a number of different metrics, means even more excitement.

6:53First, let's talk about what was actually released. It is called Claude 3.5 Sonnet. Anthropic writes, we're launching Claude 3.5 Sonnet, our first release in the forthcoming Claude 3.5 model family. So basically, this is an update to their mid-tier model Sonnet, which was previously behind their top-tier model Opus. Claude 3.5 Sonnet is more intelligent based on benchmark scores, but also still costs less in terms of price per million tokens. The 200k token context window has remained the same, but the benchmarks appear very impressive. A highlight, for example, that they call out, in an internal agentic coding evaluation, Claude 3.5 Sonnet solved 64 % of problems, outperforming Claude 3 Opus, which solved 38%.

7:32On graduate-level reasoning, 3.5 Sonnet scored a 59.4 % compared to GPT-40's 53.6%. On the MMLU, its five-shot score of 88.7 % was the same as GPT-40's zero-shot score. The Claude zero-shot score was just a little bit behind at 88.3%. And so on and so forth. You get the idea. Basically, on most of these tests, with the exception of one particular math benchmark, Claude 3.5 Sonnet was outperforming not only Claude 3 Opus, but GPT-40 as well. Anthropic also calls out Claude for having state-of-the-art vision. They write, Claude 3.5 Sonnet is our strongest vision model yet, surpassing Claude 3 Opus on standard vision benchmarks.

8:11These step change improvements are most notable for tasks that require visual reasoning, like interpreting charts and graphs. Cloud 3.5 Sonnet can also accurately transcribe text from imperfect images, a core capability for retail, logistics, and financial services. Another interesting note on the safety side of things is they say that they provided Cloud 3.5 Sonnet to the UK's Artificial Intelligence Safety Institute for pre-deployment safety evaluation. So basically, the model testing wasn't just internal. Okay, so basically, part one of this is that we have a model upgrade. The Verge writes, Claude 3.5 Sonnet is apparently Anthropic's smartest, fastest, and most personable model yet.

8:46The Verge sums up, For now, the model is the big news, and the pace of improvement here is wild to watch. Anthropic launched Claude 3 Opus in March, proudly saying it was as good as GPT-4 and Gemini 1.0, before OpenAI and Google released better versions of their models. Now, Anthropic has made its next move, and it surely won't be long before its competition does so too. Claude doesn't get talked about as much as Gemini or ChatGPT, but it's very much in the race. Another reason for it to remain in the race is that this model update also comes with a pretty meaningful interface update. Mike Krieger, the company's chief product officer, who was previously the co-founder of Instagram, wrote, we're also launching a preview of artifacts on Claude.ai.

9:25You can ask Claude to generate docs, code, mermaid diagrams, vector graphics, or even simple games. Artifacts appear next to your chat, letting you see, iterate, and build on your creations in real time. I've used it to work on writing, coding, and diagram projects. Yesterday for Superintelligent, I did a tutorial about four early use cases for Claude and artifacts. The first was writing an SOP document for a remote tech company. If you're watching this, you can see that what an artifact represents is basically a right side of the chatbot that's all about previewing the output of your instructions, which are going on on the left side.

9:56This is in many ways primarily a user experience update. By separating the section where you are providing instructions and interacting with the chatbot from the output is arguably just a better way to set up the interface. For example, after I got that SOP document, I asked it to add a tongue-in-cheek section about taking advantage of the fun parts of work from home, including dress code and family time, which it did as a second document preview that I could see alongside the first version. A couple other use cases I shared. I did a screenshot of growth statistics from our super intelligent YouTube channel that I got from Social Blade and asked Claude to turn it into a line chart, which it dutifully did.

10:32I had it write the copy and the code for an AI consultancy to help small businesses. One interesting note here is after I got the first version, I asked it to update the name to Mighty AI Consulting and explained that I wanted the name to reference the idea of these businesses as small but mighty, which it then took as a prompt to update the tagline, not just the name as well. If you've spent any time on Twitter or X since Claw 3.5 came out, you might have seen people coding up complete games. Allie Miller, for example, provided a screenshot of the instructions for the classic game Mancala and asked Claude 3.5 to read the instructions, code the game, and then preview it so she could test and play.

11:10It was able to successfully do this in a matter of seconds. There are literally dozens of examples as well of this type of game popping up as an example of what Claude 3.5 Sonnet can do. Outside of the capacities, people took note of the speed. Perplexity writes Claude 3.5 Sonnet is now available on Perplexity. With 2x faster speed than Opus, Quad 3.5 Sonnet unlocks new possibilities for complex AI applications across reasoning, knowledge, and coding tasks. Dan Shipper, who just put out a new product called Spiral, writes, The really fun thing about building Spiral is the product just got way better and cheaper to run, and we did nothing.

11:42Anthropic just dropped a new model that's smarter and lower cost. All we have to do is change one line of code to take advantage. Pretty crazy world. While a lot of the focus was on the Artifacts interface, some like Maratka and Koilen, also tested it for its reasoning capabilities and found it much improved as well. So what about any critiques or skepticisms? One thing I did see a little bit from the safety crowd was summed up by Mikhail Samin who wrote, as a reminder, Dario, the CEO of Anthropic, told multiple people Anthropic won't release models that push the frontier of AI capabilities. He then shared a couple of screenshots, seemingly referencing that point.

12:16Given that Claude 3.5 Sonnet does seem to be ahead of GBT-40, is this a betrayal of that? The take that I've seen most often in response to this, is that while 3.5 Sonnet might be slightly better across a number of different dimensions, it's not some phase-shift revolutionary jump up, and so the spirit of what Dario is saying is not necessarily undermined by releasing a slightly more advanced model than the state-of-the-art. Also, the information about him telling people that they won't release models that push the frontier is all second-hand, so it may not even be really true. Wired continued their pattern of getting more and more negative about AI, with a piece titled, We're Still Waiting for the next big leap in AI.

12:52Anthropik's latest clawed AI model pulls ahead of rivals from OpenAI and Google, but advances in machine learning have lately been more incremental than revolutionary. This is functionally true. It's exactly the argument that I was just giving about Dario saying that they wouldn't push the state of the art. At the same time, I think this sort of critique from media rings a little hollow to me. One, because as I said, I think Wired is now in a category of publications that is just looking for things to not like about artificial intelligence. And second, I think that it inherently fails to recognize how much additional value in terms of productivity, time saved, new opportunities, even quote-unquote incremental advances really represent.

13:28One of my bully pulpits here on this show is the idea that functionally the way that AI is impacting the world right now is not that we wake up and entire categories of jobs are gone, but that people are winning back their lives 20 minutes at a time. Ultimately, I think the big advance here, even though the updated model is great, is really the Artifacts interface. Professor Ethan Malek writes, the thing Anthropic is nailing is making their systems fun to use. ChatGPT has a lot of key features Claude is missing. Web access, full code interpreter, voice, GPTs, but it requires some trial and error to figure out since it isn't obvious.

13:58Even Gemini feels more complicated. For us, the consumers, it's basically nothing but upside. So if you haven't tried it out yet, go check out the latest Claude. It's available right now at Claude.ai. And if you create anything really cool, please tag me on Twitter slash X so we can see what you did. Like I said, we're dropping a new tutorial about this on Superintelligent today, so if you are a member, you will get access to that. Alright guys, thanks as always for listening or watching the show, and until next time, peace!

From the publisher

Anthropic has launched the latest model, Claude 3.5 Sonnet, and a new feature called artifacts. Claude 3.5 Sonnet outperforms GPT-4 in several metrics and introduces a new interface for generating and interacting with documents, code, diagrams, and more. Discover the early use cases, performance improvements, and the exciting possibilities this new release brings to the AI landscape. Learn how to use AI with the world's biggest library of fun and useful tutorials: https://besuper.ai/ Use code 'youtube' for 50% off your first month. The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614 Subscribe to the newsletter: https://aidailybrief.beehiiv.com/ Join our Discord: https://bit.ly/aibreakdown

More from The AI Daily Brief: Artificial Intelligence News and Analysis

All 1,099 episodes
Early Uses for Anthropic's Claude 3.5 and ArtifactsThe AI Daily Brief: Artificial Intelligence News and Analysis · 15 min
Listen in VO