In short
The AI Daily Brief: Episode Summary - OpenAI DevDay: Everything You Need To Know
Podcast Overview
- Podcast Title: The AI Daily Brief (Formerly The AI Breakdown)
- Description: A daily news analysis on artificial intelligence, covering creativity, potential disruptions in industries, and philosophical questions surrounding AI.
Episode Highlights
- Date of Episode: November 8, 2023
- Focus: Coverage of significant announcements made during OpenAI's Dev Day, including new models and features.
Key Announcements from OpenAI DevDay
- Introduction of 128k GPT-4 Turbo
- Key Features:
- Expanded context window from 32k to 128k tokens, enabling more comprehensive interactions.
- Price reductions:
- Input tokens cost reduced by 3x.
- Output tokens cost reduced by 2x.
- Additional Benefits:
- Integration of vision capabilities and DALL-E 3.
- Enhanced control features, including valid JSON response formats and built-in retrieval functions.
- Whisper 3 and New Text-to-Speech Model
- Whisper 3:
- OpenAI's speech-to-text model is being updated and will be released as open-source.
- Text-to-Speech Model:
- Features six realistic voices and allows for professional voice cloning.
- Competitive pricing: 10-20x cheaper than existing solutions like 11 Labs.
- Assistants API and Custom GPTs
- Assistants API:
- Allows developers to create AI assistants with a set of tools including threading, code interpretation, and function calling.
- Aimed at democratizing AI assistant creation for smaller developers.
- Custom GPTs:
- Users can create tailored versions of ChatGPT for specific tasks using natural language.
- Expected to revolutionize user interaction with AI by simplifying workflows.
- Legal Support for Developers
- Introduction of a Copyright Shield Program protecting developers from legal issues related to copyright infringement when using OpenAI’s tools.
Reactions and Implications
- Industry Response:
- Positive reception at the event, especially regarding pricing and capabilities.
- Noted impact on competitors, with comments about "killing it" in the AI startup space.
- Future Considerations:
- Potential for significant advancements in AI capabilities and developer engagement.
- Encouragement for developers to innovate using the newly available tools.
Conclusion This episode centers around the major advancements presented at OpenAI's Dev Day, highlighting the implications for the AI landscape, particularly around accessibility, cost reductions, and the potential for new applications through custom AI solutions. The episode concludes with expectations for ongoing discussions and developments in the weeks to come.
Additional Resources
- Subscribe to the AI Breakdown Newsletter: [AI Breakdown Newsletter](https://theaibreakdown.beehiiv.com/subscribe)
- Join the community: [AI Breakdown Community](bit.ly/aibreakdown)
- YouTube Channel: [AI Breakdown on YouTube](https://www.youtube.com/@TheAIBreakdown)
Sponsors
- Check out the podcast "Web3 with a16z crypto" for insights into the intersection of AI and cryptocurrency.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:01Today on the AI Breakdown, we're talking about all the biggest announcements from OpenAI's Dev day which happened yesterday. The AI Breakdown is a daily podcast and video about the most important news and discussions in AI. Go to breakdown.network for more information about our YouTube, our Discord, and our newsletter.
0:24Welcome back to the AI Breakdown. We are back in studio finally after a few days of travel, but today will still be a slightly different episode than normal. There is so much to cover from yesterday's OpenAI Dev Day, as well as a little follow-up from the weekend, that instead of having the normal convention of a brief followed by a main episode, we will just be doing the main episode today. Our normal format will be back, I anticipate, tomorrow. Today we are getting into everything you need to know from OpenAI's Dev Day, which was held yesterday in San Francisco, but I would be remiss if we didn't at least mention the announcement from over the weekend, which was, of course, Elon Musk and XAI announcing their new Grok.
1:03Now you might have caught the episode that I did entirely with AI, including a HeiGen video avatar, that I posted as a bonus episode over the weekend, so I won't get into this much, I just wanted to touch on it briefly as myself. The way that the XAI team described Grok was as an AI modeled after the Hitchhiker's Guide to the Galaxy. So they say, intended to answer almost anything and far harder, even suggest what questions to ask. Grok, they say, is designed to answer questions with a bit of wit and has a rebellious streak, so please don't use it if you hate humor. A unique and fundamental advantage of Grok is that it has a real-time knowledge of the world via the X platform.
1:37It will also answer spicy questions that are rejected by most other AI systems. Grok is still a very early beta product, the best we could do with two months of training, so expect it to improve rapidly with each passing week with your help. So even just in that announcement, they hit on a number of the main things. The first is the idea of it having humor in its responses. I mentioned before the post that Elon shared where someone asked Grok how to make cocaine step by step, and Grok said, oh sure, just a moment while I pull up the recipe for homemade cocaine. You know, because I'm totally going to help you with that.
2:07The announcement also points out its access to real-time data, which Elon showed off with a set of questions about Elon's most recent interview on Joe Rogan, which happened just a couple weeks ago. Now, of course, much has also been made of Elon's anti-woke positioning of the chatbot. That's something that I noted in that previous piece about just how much that seems to be a clear emphasis in the way that they're trying to differentiate from ChatGPT and all the others. Now, one thing that got a little bit lost in the announcement was the prompt IDE. XAI describes it like this. The XAI prompt IDE is an integrated development environment for prompt engineering and interpretability research.
2:42It accelerates prompt engineering through an SDK that allows implementing complex prompting techniques and rich analytics that visualize the network's outputs. We originally created prompt IDE to accelerate our development of Grok, it has helped us iterate quickly over different prompts and prompting techniques. We're now making the IDE available to members of our Grok early access program. Now, I think this is relevant in the context of today's OpenAI Dev Day focus, given that what that event represents and what we're seeing more broadly is competition for developer affiliation as these AI models get more advanced, and especially as closed models face more competition from their open source peers.
3:17Now, there is going to be a ton more to talk about when it comes to Grok in the coming weeks, I anticipate. But that is really not the focus of today's show. The focus of today's show is, of course, OpenAI's first developer conference, Dev Day, happened in San Francisco yesterday, and there was so much to cover. In fact, basically the entirety of the AI Breakdown First Five this morning was about announcements from that event, and there could have been a lot more. So what we're going to actually do today is break the announcement into four categories. And even with this, there are still some things that I'm missing, but this should give you the high level of the most important parts of the announcements and the reactions from the developer and the larger AI community.
3:53So first, let's talk about Whisper and text-to-speech. Whisper is OpenAI's speech-to-text model. If you've ever used the ChatGPT mobile app, you'll know that Whisper is frankly what you expect Siri to be. It's an unbelievably accurate model that is unbelievably fast and can even handle environments where there's a lot of background noise. Whisper is getting a third update, which will be released to open source. although it was one of the only things that wasn't actually available upon the announcement of it, and so should be something that we get in the weeks to come. Now, in addition to Whisper's speech-to-text model, we also got a new text-to-speech model that can convert text into spoken audio.
4:31Now, out-of-the-gate text-to-speech comes with six highly realistic voices, but they're also apparently going to allow people to have professional voice cloning of their own voices, as well as creating up to 30 custom voices. At first glance, AI developer Nate Chan wrote that OpenAI's text-to-speech look to be around 10 to 20x cheaper than 11 Labs, which, if that's true, would obviously have huge implications for that company, which, as you guys know, is one that I use very frequently. Now, some people have already begun the comparison. Justin de Guzman writes, Just tried OpenAI text-to-speech.
5:00Non-HD about the same speed as 11 Labs, HD is slower. Audio quality, even HD, is lower. 11 Labs' expressiveness of voice is still a lot better. No voice cloning yet makes prototyping less fun, but 10x cheaper is compelling. It's interesting because this is the first time that we've seen OpenAI really, really compete on price first versus just on quality. Now, we will in just a moment get to another part of the announcement where again, OpenAI is being at least conscious of cost, but that's different than just focusing on it as a main competitive advantage. Just benchmark the new OpenAI text-to-speech, coming in at around one second for TTS1.
5:37However, they've massively undercut others on price. Daniel Mange summed it up, OpenAI came out of left field and today revealed an 11 labs level TTS that costs 1.6 of the price or 1.12 if you use the standard model. That's kind of crazy. Now, this was a big theme of the discourse surrounding the event, summed up quite well by Stability AIs and Mod Mustoc, who wrote, OpenAI killing it, where it is all those dead AI startups. Now, these whisper and text-to-speech announcements would be huge on any other day, but were totally overshadowed by some of the other announcements at the event. The next one that I'll mention is the new GPT-4 Turbo.
6:13Now, this was a huge update for the GPT-4 API. First of all, we are clearly heading towards multimodal, given that the Vision API is integrated, as is Dolly 3 and this new text-to-speech model. But the biggest news comes in terms of the context window and in terms of the price. GPT-4's API context window used to be$32K, but it is now$128K. That's effectively as long as most any book that you might read. holding aside, I don't know, Brandon Sanderson novels or something like that. The cost for input tokens is down 3x and the cost for output tokens is down 2x. And this is what I was referring to when I said that OpenAI is clearly conscious of cost questions when it comes to developer affiliation.
6:52In fact, one of the loudest and most raucous sections of applause at the event was when Sam Altman announced those price cuts. Every co-founder, Dan Shipper, pointed out another set of benefits from the new GPT-4 Turbo, including, quote, more control. GPT-4 will respond with valid JSON and can call multiple functions. Better knowledge. The API now comes with retrieval built in. And knowledge cutoff is now April 2023. He also points out the multimodal that we just talked about with DALI, text-to-speech, and vision all being in the API as well. He also points out that fine-tuning for GPT-4 is coming out today in experimental access.
7:24Now, another announcement that once again went a little bit under the radar is that OpenAI is following the model of a number of companies in the space, including Microsoft and Adobe in that they're guaranteeing that they will pay legal fees for developers who build on top of their platform and ultimately get sued for copyright infringement. Said Sam Altman, We can defend our customers and pay the costs incurred if you face legal claims around copyright infringement and this applies both to ChatGPT Enterprise and the API. This Copyright Shield program, which is what they're calling it, applies to, as Sam pointed out, the Enterprise users and developers using the API, but not to free ChatGPT or ChatGPT Plus users.
8:00People are, of course, already racing to figure out how they might use the new expanded context window. And to get a sense of one of those use cases, let's listen to this short video from AI content creator and educator, Riley Brown. The new ChatGPT is here. Let's break it down. This video is how to write, and it has 8 million views. With one click, I can download the transcript. I can go back to ChatGPT. I can pull up the text file, and the text file is just a giant wall of text. You see this? This is an hour and 20 minutes of a lecture. And we're going to hit control A, control C, copy everything, go back to chat GBT, and we are going to type summarize every section of this video.
8:43Break down every important point. Place that entire transcript in here and press enter. Now let's see how this bad boy does. So here is critique of traditional writing, challenges faced by expert writers, interference in reader comprehension, and it's basically going through every single part of the video, creating value in writing, differences in academic writing, and standardized testing. Now, every piece of text that you own now has so much more value with this chat GPT, and this doesn't even include the feature of creating GPTs where you can create a chatbot on hundreds of these documents right here.
9:21Then you can ask it to expand on any part, right? So effective problem construction. We're going to copy this right here. And please expand on this part based on the text I provided earlier. And we're going to hit enter. And this is very, very just great information for writing scripts for any form of social media or for any storytelling in any manner. And now a word from today's sponsor. Are you interested in how two top of mind trends, AI and crypto, can work together? If so, I have the perfect podcast recommendation for you. Web3 with A16Z Crypto, the chart-topping show brought to you by venture firm Andreessen Horowitz.
9:59Web3 with A16Z Crypto is your definitive resource for the future of the internet, whether you're already building in these spaces or simply curious about what's next. If you need a place to start, they recently released an excellent episode with Stanford cryptography professor Dan Bonet and former Google Xer Ali Yahya in conversation with hosts on Al Choksi about the intersection of AI and crypto. From fighting deepfakes and proving humanity to large language models like ChatGPT, they cover it all. I highly recommend checking it out, especially if you'd like to learn more about how AI and crypto will impact our everyday lives.
10:31Beyond Crypto and AI, this show is for creators seeking more ways to truly own their work, for business leaders trying to prepare for the future today, and for innovators exploring trending tech topics. So go ahead, listen to Web3 with A16Z Crypto wherever you get your podcasts. Now, somehow, we still haven't really gotten to the part of the presentation that the most people were excited about, but that's where we're headed next. In his keynote presentation, OpenAI CEO Sam Altman talked a lot about the future of AI agents. This has been one of the hottest areas of AI development throughout the entire year.
11:08You might have heard an episode or seen a video about AutoGPT or Baby AGI or any of the numerous other projects that are trying to build autonomous agents that can actually go out and solve problems on their own and complete tasks on behalf of their creators and users. Albin shared the company's belief that these agents were going to play an increasingly important role in society and the economy, as well as the company's belief that when it comes to these types of disruptive changes, the best way to figure out how to adapt to them is to slowly walk down the path towards them and figure out how to adapt as we get real-life evidence of what happens.
11:41With that in mind, two of the biggest parts of the announcements from Dev Day were the Assistance API and custom GPTs, both of which they viewed as very first steps towards that AI agent future. Let's start with the Assistance API, and in terms of a summary, I'll share a tweet from Sully Omar that reads, OpenAI literally flipped the entire AI landscape on its head. Most AI code is likely tech debt. Companies have spent hundreds of millions of dollars building their own Assistance API, and now it's available to everyone. This is huge for the little guys, brutal for the big guys. So what is the Assistance API?
12:16Well, it's an API for building agent interfaces. It has a set of core primitives including threading, retrieval, code interpreter, and function calling. Here's how Conrad Nat sums it up. Threads and messages. For each user, create a stateful thread, add messages. Improved function calling allows AI to respond with actions on your front end. Function calling also includes guaranteed JSON output. and retrieval allows for the upload of related documentation. Finally, Code Interpreter figures out if there is a need to write custom code and then can actually go generate those files. OpenAI writes, The Assistance API allows you to build AI assistants within your own applications.
12:53An assistant has instructions and can leverage models, tools, and knowledge to respond to user queries. The Assistance API currently supports three types of tools, Code Interpreter, Retrieval, and Function Calling. In the future, we plan to release more OpenAI-built tools and allow you to provide your own tools on our platform. At a high level, a typical integration of the Assistant's API has the following flow. One, create an Assistant in the API by defining it custom instructions and picking a model. If helpful, enable tools like Code Interpreter, Retrieval, and Function Calling. Two, create a thread when a user starts a conversation.
13:26Three, add messages to the thread as the user asks questions. Four, run the Assistant on the thread to trigger responses. This automatically calls the relevant tools. So almost immediately, people started hacking. Yohei, who's built things like Baby AGI, writes,
14:07Let's set sail on a treacherous journey to explore this dire threat to our world. Mermaid speaking. Like, oh my gosh, global warming is like seriously the worst. Anyway, kind of a cheesy application, but just something that was thrown together to demonstrate what this could do. Another application that looks a little bit more like something that someone might use comes from Brian Sunter, who writes, Using the OpenAI Assistance API to make an AI chatbot from my blog post in less than five minutes. We see a video in the Assistance Builder Playground where Brian writes, Instructions, you are the public-facing AI assistant for Brian Sunter, answering questions in the style of the uploaded documents.
14:40Answer questions about the uploaded documents. The documents are writings from his personal blog. He then selects the model GPT-4-1106 preview and uploads a set of a dozen files. Then in the preview section, when a user asks what is Brian Sunter's blog about, the new assistance API chatbot responds with information that comes from the post that he uploaded. It doesn't take much imagination to see how this could be applied to any content creator on the web, or frankly, any company with products or services that people might want to know more about. Now, what about limitations? Well, Zhao Aguiam gets into some of those.
15:13He writes, a maximum of 20 file uploads per assistant. Each file can be up to 512 megabytes with a 100 gigabyte max at the organization level. Function calling has a maximum wait time of 10 minutes for execution. There's no support for streaming output. Image generation is not supported. You need to call the DALI 3 API separately. Image analysis is not supported. You need to call the Vision API separately. Retrieval capabilities do not extend to XLS or CSV files. But still, and this is something that Bennett's strategically pointed out, as opposed to the way that tech presentations have trended recently, where products are announced weeks before they're actually available, OpenAI was, by and large, with the one exception of the Whisper 3 API, actually putting all these tools out into the world as soon as they had announced them.
15:54Given that, there's going to be, I imagine, a lot more patience for some of those limitations that we just listed. Still, I think for me and for many people, the most potentially game-changing aspect of the announcement was the announcement of custom GPTs, which are basically customized specific purpose versions of ChatGPT that anyone can create with natural language. Here's how Sam Altman described them. GPTs are tailored version of ChatGPT for a specific purpose. You can build a GPT, a customized version of ChatGPT, for almost anything, with instructions, expanded knowledge and actions, and then you can publish it for others to use.
16:27And because they combine instructions, expanded knowledge, and actions, they can be more helpful to you. They can work better in many contexts, and they can give you better control. They'll make it easier for you to accomplish all sorts of tasks or just have more fun. And you'll be able to use them right within ChatGPT. You can, in effect, program a ChatGPT with language just by talking to it. OpenAI CTO Mira Mirati writes, GPTs are not omniscient. They're custom versions of ChatGPT tuned for specific tasks. Smart tools that I'm certain we won't be able to live without. OpenAI describes GPTs as a new way for anyone to create a tailored version of ChatGPT to be more helpful in their daily life, at specific tasks, at work, or at home, and then share that creation with others, no code required.
17:08So what are some examples of this? Well, one that Sam Altman himself built live was a startup mentor that used previous speeches of his from his time-leading Y Combinator to help give founders of startups advice that he might give them, but automatically through a chatbot. To do this, Sam goes to the GPT builder, where it asks him to describe what he wants to build. Sam writes that he wants to build a chatbot that gives advice to founders, and in a preview window on the side of the builder, you can actually see how the GPT builder is interpreting his instructions and starting to turn it into what will eventually be the custom GPT.
17:41The GPT builder comes up with a name and a suggested icon, both of which Sam accepts, calling it Startup Mentor, and then he goes into the Configure menu, where, among other things, he can upload a file, in this case, the transcript from a speech he's previously given about startups. From that same configure menu, he can also add or subtract capabilities, including web browsing, DALI image generation, and code interpreter. And within just a few minutes, he's got a working version of one of these custom GPTs. Professor Ethan Mollick gave another example. He writes, Here's a little GPT, the name for the new agent-like thing released by OpenAI, that I threw together in less than a minute.
18:16It looks up the latest trends for a product category on the web, and then creates prototype images for it. takes less than 90 seconds end to end. So in the demo video, Ethan shares, the tool, which is called Trend Analyzer, asks the user what type of product they're interested in. It says, are you thinking about technology, fashion, home goods, or something else entirely? Once I know the product category, I can look into the latest trends for you. Ethan writes sneakers, and the Trend Analyzer goes off and sources some trends around 2023 that could be built into the product design. It comes back with six different trends, all featuring sources, and then asks which, if any, the user wants to proceed with.
18:50When Ethan writes, you decide, Trend Analyzer says, we'll go with a high top silhouette with chunky elements and bold and vibrant colorways. From there, it says, next I'll create realistic photo shoots of our futuristic sneaker concept incorporating these trends. That, of course, is where Dali 3 comes in. And boom, all of a sudden you have this futuristic shoe concept. Nick Dobos had another example. He writes, playing with the new ChatGPT custom GPTs, introducing GIF PT, automatically turn Dali images into GIFs. Now, as Nick points out, quote,
19:28And this is something that I think people are initially perhaps missing a little bit, at least those who are inclined to be contrarian about all of the excitement surrounding this. For example, if you go back to Ethan's post, someone writes, But as Ethan points out, it uses the same tools. It just makes it easier to share and to work with large prompts. It also includes a lot of features that Bing hasn't implemented yet. Connection to outside systems, CI, and working with files. I think it would be incorrect to underestimate how much it matters to simplify workflows in the way that these custom GPTs do.
20:01My general belief is that every simplification of a workflow leads to a massive increase in the number and variety of use cases for the thing that's underlying that workflow. In other words, OpenAI making this a lot easier means a lot more people are going to use it and for a lot different purposes than they might have before. Now, to the extent that there was any skepticism around custom GPTs, it wasn't around their usefulness, but about whether there will actually be demand to buy these. As part of their announcement, OpenAI said that within a month, we would have a custom GPT store. Ex-user and AI trend watcher Boris writes, The idea of building a private library of custom GPTs is a great and very useful one.
20:37I'm a little skeptical of the store, but it might turn out to be great as well. AI entrepreneur Bindu Reddy, as part of a much larger critique post, one of frankly the only ones that I saw, was also skeptical of this. She writes,
21:09Robert Scoble, however, disagrees. He writes,
21:15AI quite a bit of lock-in even after competitors arrive. Think about it this way. If you're getting paid$1 ,000 a month for building a useful GPT, will you leave just because Elon Musk has a better AI? Nope. So the race is on to get developers addicted to your ecosystem today. I could see buying quite a few GPTs to help me run my business and life. That revenue, even if expected to be small for a while, like Bindu says, will provide lock-in. Now, speaking of Google and Gemini, NVIDIA's Dr. Jim Phan writes, expectation for Google Gemini is now ridiculously high. Gemini has to check off at least one of the following.
21:48120 % IQ of textual GPT-4, or 100 % of GPT-4 but at half the cost or 2x speed of turbo, or 100 % of visual GPT-4, or natively support long videos, and ship the API in Q1 of 2024. It's about time that DeepMind recovers the glory of AlphaGo in 2016. I'm looking forward to it. However, they have their work cut out for them. Investor Ali Miller writes, I'm at OpenAI Dev Day in San Francisco. I was front and center for the keynote. I've tweeted so many tweets, but there is one big takeaway that you're not going to see over the live stream. One feeling that you only get in person. And that is, compared to every other big tech event I've been to, OpenAI Dev Day is the highest, okay, I have to go build something with this new release immediately score.
22:35I'm talking 11 out of 10 builder activation score. It's incredible. Putting an exclamation point on that, AI entrepreneur and developer Sam Whitmore writes, when you run away from Dev Day to go integrate everything as soon as humanly possible, dream day for people building AI-powered products. Thank you, OpenAI. This is certainly what I felt when I was watching these announcements, which I was doing in an airport and an airplane on my way back from Mexico. Right now, OpenAI has builder imagination in a huge way. They are moving quickly. They're introducing new things. They are to the point of the video that I released over the weekend.
Read the full transcript
23:09Walking down a path to a very different type of relationship with computing. And frankly, it's exciting to watch. I'm sure that over the course of this week, we will come back to these topics and see what people are hacking on and what early experiments are showing the possibility of things like the assistance API and custom GPTs. But for now, after this long episode, we are going to wrap it there. I appreciate you guys listening or watching as always. I'm very excited to be back with you all. And until next time, peace. Thank you.
From the publisher
Yesterday OpenAI announces 128k GPT-4 Turbo at 1/3rd the price; a new Text-to-Speech model; Whisper 3; and proto-agent features like the Assistants API and Custom GPTs.
Today's Sponsors:
Listen to the chart-topping podcast 'web3 with a16z crypto' wherever you get your podcasts or here: https://link.chtbl.com/xz5kFVEK?sid=AIBreakdown
Interested in the opportunity mentioned in today's show? jobs@breakdown.network
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI.
Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe
Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown
Join the community: bit.ly/aibreakdown
Learn more: http://breakdown.network/
