In short
The AI Daily Brief: Episode Summary
Podcast Title The AI Daily Brief (Formerly The AI Breakdown)
Episode Title OpenAI's DevDay And A Preview Of Our Agentic Future
Episode Overview This episode focuses on the highlights from OpenAI's second annual Dev Day, showcasing significant updates and announcements for developers. The discussion elaborates on how these innovations signal the impending evolution of artificial intelligence (AI), particularly with respect to "agentic" AI.
Key Announcements from OpenAI's Dev Day
- Real-Time API
- Offers near real-time speech-to-speech interactions.
- Key Features:
- Native speech-to-speech without a text intermediary.
- Low-latency responses, crucial for applications like customer service.
- Currently provides six voice options but no third-party voice integrations.
- Public beta is now open.
- Cost Considerations:
- $0.06 per minute for audio input and $0.24 for output, which may not be cost-effective compared to human labor initially.
- Vision Fine-Tuning
- Enables developers to use images along with text to enhance the performance of GPT-4.0 applications.
- Potential applications include autonomous driving features like traffic sign detection.
- Safety policies prevent the upload of copyrighted or violent images.
- Prompt Caching
- Allows developers to save frequently used context between API calls, cutting costs and reducing latency by up to 50%.
- Raises questions about its implications for partnerships, particularly with Microsoft.
- Model Distillation
- Permits the fine-tuning of smaller models for better performance using larger models like O1 Preview and GPT-4.0.
- Accompanied by a beta evaluation tool for performance comparison.
Noteworthy Absences
- No updates on the GPT store, Sora, or GPT-5.
- O1 is positioned as a new category of reasoning models.
Key Themes from the Event
- Speed of Progress in AI Development:
- Developers can now achieve what once required extensive teams and resources.
- Divergence in AI Models:
- OpenAI is distinguishing between general-purpose LLMs (like GPT-4) and reasoning models (like O1).
- The emergence of a specialized family of models is anticipated to cater to different use cases.
Insights from Sam Altman and Kevin Wheel
- Discussion about the proximity of achieving Artificial General Intelligence (AGI):
- Altman noted that advancements are fast-tracking the definition of AGI, suggesting that O1 approaches level 2 AGI.
- Alignment concerns:
- OpenAI seeks to develop capable models that enhance safety over time, emphasizing iterative deployment as a safety mechanism.
- Future of AI agents:
- O1 is viewed as pivotal for developing autonomous agents, capable of performing tasks rapidly and effectively.
Takeaways
- OpenAI's Leadership:
- The company is positioning itself ahead of competitors by focusing on reasoning AI, rather than merely enhancing existing LLMs.
- AI Agent Evolution:
- The capabilities showcased indicate a shift towards creating more autonomous, intelligent agents, potentially revolutionizing how tasks are executed.
- Future Implications:
- The delineation between LLMs and reasoning AIs is becoming clearer, suggesting a transformative phase in AI development.
Conclusion The developments announced during OpenAI's Dev Day represent crucial strides towards a future where AI capabilities are not just expanded but fundamentally redefined. OpenAI's focus on reasoning models may signal a significant shift in the landscape of AI technologies, enticing both developers and industries to rethink their strategies and applications.
For more updates and discussions, listeners are encouraged to join the AI Daily Brief community on their Discord and subscribe to their newsletter.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Today on the AI Daily Brief, what we learned about our agentic future at OpenAI's Dev Day. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. To join the conversation, follow the Discord link in our show notes.
0:21Hey, hello, friends. Quick note. There ended up being so much to talk about with OpenAI Dev Day that today's episode is just that. It is just a main episode. We are not doing the headlines. And frankly, part of the other reason for that is that the headlines are, in fact, other really cool product announcements. Pika, Microsoft, and 11 Labs all announced really interesting products yesterday, and there was also an XAI meetup. So there was a lot going on outside of just OpenAI's Dev Day, and we will get to all of that in a future show this week. But for today, it is just going to be all about OpenAI, and it's pretty interesting.
0:52Welcome back to the AI Daily Brief. Today, of course, we are talking about OpenAI's Dev Day. This is their second annual Dev Day. It's a chance for them to give a bunch of updates, share a bunch of new products. And as you'll see, one of the big themes from this particular Dev Day was that this one really was for developers. This did not feel like an event that was focused on consumers. There was, as you'll see, very little talk of GPT-5. And instead, what we got was something of a preview of the future. There is a picture that is starting to become clear that's not necessarily new, but is really rounding out in its vision of the agentic future that is hurtling down the pipeline at ever-increasing speed.
1:32So let's do a little bit of what was actually announced, and then let's talk about how different people reacted to it and what some of the big themes were. To start, as much as they would have liked to, there was no way to avoid the recent departures of CTO Mira Mirati and Chief Research Officer Bob McGrew. In a briefing with reporters ahead of the event, new Chief Product Officer Kevin Wheel said, I'll start with saying Bob and Mira have been awesome leaders. I've learned a lot from them, and they are a huge part of getting us to where we are today. And also, we're not going to slow down. Like I said, a lot of the discussion was about developers in this session.
2:03OpenAI said they currently have 3 million developers building on top of their models, which is triple the number of active apps on the platform compared to last year. One of the things that came up in discussions a lot was the decreasing cost of intelligence. OpenAI said that they had cut costs for developers to access their APIs by 99 % over the past two years, and obviously this was driven by incredible competition from Google and Meta who are also pushing the price down. Now, in terms of what was announced, there were four big things. The real-time API, which probably generated the most discussion on places like Twitter, vision fine-tuning, prompt catching, and distillation.
2:36Let's talk about the real-time API first. The real-time API gives developers nearly real-time speech-to-speech experiences in their apps. The really important thing here, which seems like a small detail but really isn't, is that this is native speech-to-speech. As they put it, there is no text intermediary, which means low-latency nuanced output. So of course, when you're thinking about applications like customer service, this is a total game changer for that sort of experience. In terms of other details, there is a choice of six voices provided by OpenAI. These are different than the ones used by ChatGPT.
3:06And currently, there's no ability to integrate third-party voices, probably in order to avoid copyright issues. Head of developer experience, Romain Hewitt, demoed a trip planning app using the real-time API. He verbally discussed a trip to London with the AI assistant, got low-latency responses, and the app was also able to annotate a map with restaurant locations as it answered questions. A second demo showed how the real-time API could speak on the phone to order food for an event. Unlike Google's Duo, the API can't make calls natively but can integrate with APIs like Twilio to do so. OpenAI is not currently forcing the AI models to identify themselves as such on phone calls, something which could be illegal under new California laws.
3:39The real-time API returns recorded responses, transcript, and function calls in real time. Now, once again, this is not a, hey, this is coming in the future product. The public beta is now open. In terms of costs, one of the big conversations was about how at$0.06 per minute of audio input and$0.24 per minute of audio output for an evenly mixed use case of about$0.15 per minute, this was not only more expensive than 11 lab speech-to-speech products, which cost$0.11 per minute, but might not actually, in the short term at least, be cost savings relative to human labor. Many people pointed out that the cost for one hour of this would be around$18, which is more than what companies would pay for many international call centers.
4:17Now, others pointed out that that would assume constant talking, which isn't necessarily going to be the case. And I also think more broadly than that, trying to understand the prices now as some sort of static thing is probably not really going to give you a good picture of what the future is actually going to look like. Today's episode is brought to you by Plum. Generative AI promises to supercharge your productivity and give you superpowers. But if you're not an engineer, trying to harness AI can be incredibly frustrating. Hours wasted wrestling with complex tools only to give up when they don't work.
4:47We all have tedious tasks we'd love to automate and challenges AI could solve, but few of us have the skills to fully leverage these game-changing technologies. That's where Plum comes in. The mission? To make automating your work feel like magic. Imagine typing out, AI, read my Gmail and ping me in Slack when something critical comes in, and watching it come to life before your eyes. No coding required. Whether you're a marketer, salesperson, or founder, Plum enables you to create custom AI workflows in minutes, not hours. Check out useplum.com, that's Plum with a B, for early access to the future of workflow automation.
5:42other AI apps, Venice won't tell you what's okay to say or not. Venice won't patronize you. It simply provides direct access to machine intelligence. No topics are off limits. No ideas are taboo. With Venice, you're in control of the AI as you should be. Pro subscriptions are available for$49 a year or$8 per month. AI Daily Brief listeners receive a 20 % discount on Venice Pro. Visit venice.ai slash nlw and enter the discount code nlwdailybrief. That's nlwdailybrief, all one word. Next up was vision fine-tuning. This is a really cool one that I think is going to open up a lot of new use cases as well.
6:17Developers can now use images as well as text to fine-tune applications of GPT-4.0 using the API. This should help massively improve performance of tasks that require visual understanding. Some pointed out that this could be significant for autonomous driving, think traffic sign detection, and the OpenAI team said that in general this was the top feature request to the fine-tuning team. Now, when it comes to safety, head of product API Olivier Godman told TechCrunch that safety policies are in place to prevent uploading copyrighted and violent images. The third notable feature was prompt caching.
6:47This is a similar feature to the one that was introduced by Anthropic several months ago, but basically it allows developers to cache frequently used context between API calls, which reduces costs and latency. OpenAI say developers can save 50 % using the feature for repeated API calls. For For what it's worth, Anthropic promoted theirs as capable of reducing costs by up to 90%. This is clearly part of the deflation of API costs, but some wonder if this is going to cause issues with Microsoft. Dan Shipper from Every writes, It's great for developers, but it also creates an interesting dynamic with its biggest partner, Microsoft.
7:15I heard from DevDay attendees that Microsoft has been pushing large enterprises to commit up front to buy a certain amount of GPT-4 API calls in order to guarantee capacity. One wonders how Microsoft and its customers who have already committed feel about these price reductions. Last big update was model distillation. developers can now use larger models like O1 Preview and GPT-4.0 to fine-tune smaller models like GPT-4.0 Mini using the API. The idea here is basically to allow developers to get better performance out of smaller models. OpenAI are also launching this feature alongside a beta of an evaluation tool so developers can compare performance.
7:46Now, in terms of what wasn't announced, there was no update on the GPT store, there was nothing about Sora, there was no full size O1, and there certainly wasn't mention of GPT-5. That said, the buzz around O1 was still palpable. And what this session really reinforced was that this is a different category of product than GPT-4O. That's exactly how head of product API Olivier Godman described it, basically saying that this was a new family of models distinct from GPT-4O. You're starting to see then a cleave between, on the one hand, general purpose LLM models, and on the other hand, reasoning models, and a sense that there's going to be different use cases that fit each of these different categories.
8:27Head of Developer Relations, Roman Hewitt, again, did a live demo of O1 where he used it to build an iPhone app with a single prompted 30 seconds. He also prompted a web app to control a drone that was present on stage and then used the app to pilot the drone. As Shipper again from EveryPointsout, it would have been possible to do these demos with previous GPT models, but they would have taken much longer to build and probably wouldn't have been suitable for a live audience. Now, one of the big culminations of the event was a fireside conversation between Sam Altman and Chief Product Officer Kevin Wheel.
8:54One thing that's notable is that this was not live-streamed, and so what we got was basically summaries from people like Greg Camerat, who was there, in the room. A few of the highlights. One of the big topics of discussion was how close we are to AGI. Sam said that they would finish a system and ask, in what way is this not an AGI? Basically arguing that the word is overloaded, and that O1 is level 2 AGI. For Sam, though, the rate of scientific discovery is the benchmark. He said that the fact that definitions matter this much means we're getting close. He also said we're in this period where it's going to feel blurry for a while.
9:23If we can make an AI system that is better at AI research than OpenAI is, then that feels like a real milestone. Another part of the discussion was alignment. Altman said it's true we have a different take on alignment than whatever that internet forum is, presumably talking about LessWrong. We want to figure out how to build capable models that get safer and safer over time. We have an approach to figure out where the capabilities are going to work, then make it safe. O1 is our most capable model, and it's our most aligned model, too. Sam reinforced the idea that iterative deployment is the best safety system they have, because as Kevin Wheel put it, no matter how many smart people you have inside your walls, there are way more smart people outside your walls.
9:56There was more discussion of O1. Before the end of the year, they said that O1 would support function calling along with system prompts and structured output. Altman said, The model is going to get so much better so fast. It's at GPT-2 level. We know how to get it to GPT-4 level. Plan for the model to get rapidly smarter. At one point, Sam asked the crowd if they thought they were smarter than O1. A few hands went up, and he said, do you think you'll still think this by O2? No one wants to take that bet. This was yet again, the company really reinforcing that they're going to be the best at reasoning models, which I think is one of the biggest takeaways of this event.
10:26There were a bunch of other little things that were interesting too. Sam, for example, said that they thought that infinite context length would happen within the decade. And finally, there was the discussion of agents. Basically, O1 was presented as the path to making agents actually happen. Altman said, people get used to any new tech quickly, but agents will be a big deal. People will ask an agent to do something that would have taken them a month, then it'll take an hour. Then they'll have 10x agents, then they'll have 1000x agents. And because of 01, for the first time, the agent conversation doesn't feel like it's some super far in the future type of thing.
10:55And so let's talk about takeaways. There were two kind of big themes that stood out to me. One is just the speed of progress. As Dan Schipper put it, it's easy to forget that just a few years ago, none of the things that were shown today were possible or even on anyone's radar. Today, a single developer making an app in their spare time can build things that entire teams of developers wouldn't have been capable of previously. However, I think in many ways, the even bigger theme was summed up by developer Nick Dobos who wrote, OpenAI is challenging what a computer can do. Everyone else is playing with LLMs.
11:24It really does feel like there's this fork that's starting to happen between LLMs and reasoning AIs, and that O1 is where these two genetic lines diverge. They still feel really close right now, but it seems like they will be fundamentally different in the future. Ethan Mollick put it this way, OpenAI put a lot of interesting pieces on the board today. With the API, you can get good planning, real-time two-way voice, and a few ways to get cheaper specialized answer when you don't need an expensive generalist. A lot of the key components of AI agents coming into view. Shipper again reinforced that OpenAI believes that O1 is an important step towards agents.
11:56Agents have long been one of the sexiest AI applications, but previous GPT models were likely to get off track if they were left to figure out a task by themselves. O1, because of its ability to reflect on its own thought processes and plan next steps, is a key pillar in making agents that are actually autonomous. That leads to another conclusion for Dan. OpenAI is leading into the race to build different kinds of models for different use cases. The company believes that the most effective applications are going to string together multiple models rather than use one for everything. Now, this might seem like a small thing, but I actually think it might be the beginning of a fundamentally different positioning.
12:26Everything in AI since ChatGPT launched was just about benchmarking the most state-of-the-art models or doing a different version of that, constrained by size or constrained by the fact that a model was open source. Now we're talking about a new family, a new category even of models. And OpenAI is clearly planting a flag that says, leading the AI space isn't just about GPT-5. It's about how far you are in this fundamentally new branch on the AI family tree. I think it's going to take a little bit of time for us to really grok the implications of that. And to the extent that they're right, OpenAI will have once again pushed the GenAI space into its next evolution.
13:04Really interesting stuff. Appreciate everyone who is tweeting and sharing their thoughts from the actual event so we could share it here as well. Appreciate you guys listening or watching as always. And until next time, peace.
From the publisher
OpenAI’s second annual DevDay showcased a glimpse into the agentic future with powerful new tools for developers. This video covers all major announcements, including the real-time API for speech-to-speech interactions, vision fine-tuning, prompt caching, and model distillation. Discover how these updates set the stage for the next evolution in AI, and why reasoning models like OpenAI’s O1 might be the foundation for autonomous AI agents in the near future.
Concerned about being spied on? Tired of censored responses? AI Daily Brief listeners receive a 20% discount on Venice Pro. Visit https://venice.ai/nlw and enter the discount code NLWDAILYBRIEF.
The AI Daily Brief helps you understand the most important news and discussions in AI.
Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614
Subscribe to the newsletter: https://aidailybrief.beehiiv.com/
Join our Discord: https://bit.ly/aibreakdown
