147 | OpenAI’s 12 Days of Releases, Google’s Game-Changing Video Generation Leap, and Amazon’s New Models (Nova) and more AI news for the week ending on Dec 7, 2024

7 Dec 2024 · 42 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: Leveraging AI - Episode 147

Overview The episode discusses significant developments in the AI landscape for the week ending December 7, 2024, focusing on innovations from OpenAI, Google, Amazon, and others. It raises questions about the rapid pace of AI advancements and their implications for businesses and creators.

Key Highlights

OpenAI's 12 Days of Releases

  • 12 Days of OpenAI Campaign: A series of daily releases showcasing new capabilities.
  • O1 Pro Model:
  • Achieved an 83% success rate on the International Mathematics Olympiad qualifying exam.
  • Considerable improvements in problem-solving and reasoning, surpassing GPT-4.
  • Pricing Tiers:
  • Introduction of a new ChatGPT Pro subscription at $200/month, significantly higher than the regular $20/month plan.
  • Expanded access to advanced features for heavy users.
  • Reinforcement Fine-Tuning: Researchers and enterprises can now apply for alpha access.

Google's Innovations

  • Veo Video Generation Model:
  • Generates 1080p videos from text or images.
  • Incorporates content style consistency and safeguards against harmful content.
  • Gemini Update: Enhanced control over essential apps on Pixel devices.

Amazon's Nova Series

  • Launched a family of AI models:
  • Nova Micro: Fast, cost-effective text model.
  • Nova Lite: Affordable multimodal model for text, image, and video.
  • Future Releases: Planned models for advanced reasoning and video generation in 2025.

Multi-Agent Systems

  • Microsoft's Magnetic One: A multi-agent system with specialized agents for browsing, coding, and file management.
  • Anthropic's MCP Protocol: A new standard enabling seamless connections between AI assistants and data sources.

Other Notable Developments

  • Runway's Frames: New image generator emphasizing stylistic control for video creation.
  • Eleven Labs' Conversational AI: Introduction of a customizable agent builder and a new podcasting capability.
  • Surgical Robots: Training systems using videos to improve surgical procedures without human intervention.

Industry Trends and Predictions

  • A shift from experimental AI applications to focused implementations in enterprises is noted.
  • Predictions of AI tools capable of producing indistinguishable video content by 2025.
  • Discussions around OpenAI's transition from a non-profit to a for-profit model, with implications for its future.

Community Engagement

  • Viewers encouraged to share the podcast and provide reviews to help others benefit from the insights shared.

Conclusion The episode encapsulates the rapid advancements in AI technology and their potential transformative impact on various sectors, while also acknowledging the ethical considerations that accompany these innovations. It serves as a call to action for business professionals to leverage these tools responsibly.

---

For more resources and information about the podcast, visit

  • [The Ultimate AI Course for Business People](https://multiplai.ai/ai-course/)
  • [YouTube Full Episodes](https://www.youtube.com/@Multiplai_AI/)
  • [Connect with Isar Meitis](https://www.linkedin.com/in/isarmeitis/)
  • [Join Live Sessions and Newsletter](https://services.multiplai.ai/events)

If you found this episode helpful, consider leaving a review on your favorite podcast platform!

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Hello and welcome to another weekend news episode of the Leveraging AI podcast, the podcast that shares practical, ethical ways to leverage AI to improve efficiency, grow your business, and advance your career. This is Isar Maitis, your host, and we have a jam-packed week. It seems that there has been more new feature and model releases this past week than any week in history. So while we usually dive into several large topics and then rapid fire, this is going to feel more like rapid fire beginning to end, but the entire first segment of the show is going to talk about new models and new features that have been released or just about to be released by more or less anybody who has anything to do with AI on the planet.

0:46So we have a lot to talk about. So let's get this started.

0:56If you've been listening to the show for a while, you know, there have been huge anticipation to see what OpenAI is going to do towards the end of the year, or more specifically towards the two-year anniversary or birthday of ChatGPT. Well, nothing actually happened on the ChatGPT day itself, but what they have done is they just announced 12 days of OpenAI, which basically are 12 days in which every single day they're going to release a new capability. Some are big and some are small, or how they labeled it, big ones and some stocking stuffers. So the first one that we got is the full O1 model.

1:31So we had access to O1 Preview and O1 Mini. Both were groundbreaking new kind of family of models that can think and analyze things in a much deeper way than before. This took the world by craze and now basically everybody's chasing the same thing. And we saw multiple releases from Chinese companies, which you're going to talk about as well in the show that are trying to do the same thing. Well, now we finally got the full O1 model, and that was the first release in the first days of the 12 days of OpenAI. Just as a quick reminder, we talked about this when this was announced, but O1 is achieving incredible solutions in problem solving, and it scored an 83 % success rate on the International Mathematics Olympiad Qualifying Exam, which is a very big spread than GPT-4-0, who scored 13%.

2:20So that's a huge spike. Also big, huge decline in error, a significant improvement in error reduction and a very high success on anything that has to do that requires deeper reasoning as well as STEM capabilities. Now, in addition, they announced a new ChatGPT Pro subscription that is going to be$200 per month. So that's a pretty big spike from the$20 a month most of us are paying of the$25 or$30 for the Teams versions, depending exactly which plan you pick. And then there's obviously the enterprise versions that are priced very differently. So right now there is going to be one more tier. There's the free version.

3:00There's the plus version, which most of us use if you're just a regular user, which is$20 a month. and then it's going to be$200 a month for the pro version. And I'm reading for their website. It's going to have everything in plus and unlimited access to GPT-40 and O1 and unlimited access to advanced voice and access to O1 Pro mode, which uses more compute for the best answers to the hardest questions, which basically tells us that there is going to be limited access to O1 if you're not on that mode and standard and advanced voice mode accessible, but with some limits if you're just on the 20 bucks a month plan.

3:41Basically what this is, it's targeting people who are heavy users who are going to use this all the time and needs more bandwidth. And for them, it will probably make sense to pay$200 a month. It will be very interesting to see if that actually works because that's a very, very big spike. That being said, people who actually use this at that capacity, probably see the value in it and probably will pay the$200. This will open the door for many other providers to do similar significantly higherly, to follow a similar approach with significantly higher monthly rates to use various advanced services.

4:14On day two, which was on Friday, December 6th, they released the capability to allow people to do their own reinforcement fine-tuning. So you can now apply, and now I'm reading from their website again, apply to the reinforcement fine-tuning research program. We're expending alpha access to reinforcement fine-tuning and inviting researchers, universities, and enterprises with complex tasks to apply. Spots are limited. So that was the announcement of day two. On the day this is released, it's going to be day three and so on. So if you're listening, They're releasing one of those at 10 a.m. Pacific time every single day, and you can chime in and see exactly what the release is.

4:57But there's going to be a lot of exciting stuff. What might be some of that exciting stuff? I really hope and many other people hope that Sora will be released as part of those capabilities, which is the really incredible, presumably, video generation model that they have announced and showed off in February of this year and haven't released to the public yet. and now a lot of the other companies are catching up to it. And so that's a highly anticipated one and they may or may not release that and maybe glimpse or segments or previews of GPT-5 or whatever they're going to end up calling it. We may get that.

5:32Maybe more advanced voice capabilities and stuff like that. So there's a lot that they may release, but we don't really know. And like I said, some of them announced that are going to be big, some are going to be small, but definitely pay attention because every single day in the next few days, there's going to be another announcement from OpenAI. Now, staying on the topic of interesting releases, a Chinese AI firm called O1.ai has achieved a really interesting release. So they released an AI model that is at the level or close to the level of GPT-4. The company is called Yi Lightning, and their model is currently ranked number six on the LMSYS chatbot arena.

6:11and they have trained the model with an investment of$3 million. To put things in perspective, ChatGPT 4.0 was trained at an$80 to$100 million investment. They've done this with only 2 ,000 GPUs, despite the US restrictions. That's what they were able to get their hands on. And they were able to achieve that by just solving for problems and being more innovative than what the big companies that just have the money and the brute force to put into this. And they solved things like reduced computational bottlenecks and multilayer caching implementation, specialized in inference engine that they developed specifically for their needs and optimized memory usage.

6:56So when we hear people like Ilya Saskover saying that it's not about just old scaling laws, but we need to think out of the box, this is a very good example in my eyes. So this model is now available. It's open source. You can use it. And it's significantly cheaper on inference as well than the big models, while it's generating results that, as I said, place it six out of hundreds of models in the world right now. Another big announcement came from Google this week. So they're rolling out a major update to Gemini on Pixel phones and Pixel watches. And what it allows you to do is that allows you to use multiple extensions to connect and control and work with multiple apps on your phone.

7:37The biggest ones that they just released are Spotify, Messaging, Calling, and Smart Home. The biggest gripe that people had with the new Gemini that they're trying to replace just a Google Assistant is that it didn't control most of the important apps on your phone. And now this feature has been released that allows these extensions to talk to the other apps. As I mentioned, a few has already been released, and I'm sure we're going to see more and more of them, which will most likely make Gemini the new tool for most people use Android phones like myself to engage with your phone and with different applications.

8:10But that's not the biggest announcement from Google. Google has launched Vio, which is its AI video generation model that was in private release for a very long time and now is available for preview via Vertex AI platform, which anybody can have access into if you're just creating an account and logging into that. And it is incredible. So it generates 1080p resolution videos from either text or images. It is very good at following content in various visual cinematic styles. So you can pick a specific style and it sticks with that style through the entire video generation. It can produce videos beyond a minute per them.

8:51I haven't tested it yet. And it includes built-in safeguards against harmful content so people cannot generate violence and otherwise problematic content. The other additional interesting thing is that it also includes DeepMind's Synth ID watermarking technology, meaning it's nothing you can see in the video, but it allows Google to detect that these videos were generated. Synth ID was released as open source. Technically, now anybody will be able to find out which videos are created with Vio. This should be a very serious catalyst for OpenAI to release Sora because this is now Sora domain capabilities.

9:281080p, one minute long is not something we have from any other supplier. It's not surprising to me that Google are able to make that jump because they can train on every piece of video from YouTube, which is the largest video repository on the planet. And they obviously have the compute and the human resources to develop this kind of tool. But I'm personally very curious to start playing with this and see what it can produce. Now, in addition to that, they are planning to also release a new version of Imagine 3, their text to image generator to all Google Cloud customers. And the combination obviously will allow you to create very specific images, high resolution and detail to present exactly what you want, and then use that image as a feed into the new VO engine to create videos.

10:14Again, very exciting to any creator who wants an additional tool at their fingertips. I want to share with you some exciting news from Multiply, my company. We just opened the registration for our January AI Business Transformation course. The current course was sold out. The current course that started in November and is ending this week was sold out. We're not going to launch one in December because of the holidays, but there's another cohort that's opening on January 20th of 2025. This course has transformed hundreds of companies and business leaders that have taken the course since its inception in April of 2023.

10:54So if you're looking for ways to start 2025 with the right foot forward, whether for your personal career or your knowledge, or for the sake of the success of your company, business, team, organization, et cetera, come join us. There will be a link in the show notes so you can open your phone right now or your computer and go and see exactly all the details over there. And now back to the episode. Now, staying on the same topic, but from a different company, Runway, which is one of the leading providers of video capabilities, just launched Frames. So Frames is an image generator built into Runway.

11:29They had one before, but this is a completely new model that they've developed from the ground up, and it prioritizes stylistic control and visual fidelity. What it basically means, it means that you can generate multiple images with consistent artistic styles which can be used as the starting point to create videos, which is really important when you're trying to create a consistent video that is generated from multiple images. So right now, the videos that Runway can generate are significantly shorter and they are a few seconds each. So in order to create a longer video, you need to start with multiple images and you need these images to look consistent.

12:07Otherwise, your video will not look consistent. And this is exactly the capability they have just provided. Now, it's already rolling out gradually to anybody who has access to Gen 3 Alpha, basically any paying user of Runway, and it will be available through the API as well. Staying on the topic of video generation, Kling has released Motion Brush. So it allows you to select up to six elements in a single image, paint them with a little brush and point an arrow where you want their motion to go. They also provided camera movements with six types of cameras with camera motion, horizontal pan, vertical pan, zoom, tilt, and roll, which means you can now control both the motion of the camera as well as motion of elements within the scene.

12:56And they released two modes, standard mode with 720p video iteration that is faster and more cost effective and professional mode that allows you to release 1080p HD capabilities. capabilities. Now, I've got to go back to a prediction I made in Q1 this year, and you can go back and check those episodes. But I was saying that by the end of 2025, we will, by the end of 2024, we will be able to create videos that will be indistinguishable between professionally generated videos of real life or cartoons or any other style we would want. And by 2025, we will have full control over the cameras, the scenes, the view, the story, and everything else in the video.

13:34And the first half of the prediction is already correct. Now, these tools are not perfect yet, and they still have some morphing, and there's still some issues in consistency, but it's getting better and better every day. And now with the release of Vio and potentially Sora as well before the end of this year, we will have even more advanced capabilities with more consistency. And I think the consistency issues are going to be mostly resolved in the near future. And then the full control over camera and scene and action will happen in 2025. The other thing that I think will happen in 2025 is AI editing, meaning you'll be able to go to an existing video, whether a real one that was shot with a camera or an AI generated one, and go back and request edits.

14:15Like take this section out, change the camera pan from this to this, and do whatever you want in order to actually edit the video that you already have created to have a lot more control. And if I have to bet, we will have those capabilities before next year is over, which will allow and we'll probably start seeing full videos created completely with AI. Right now, people are kind of like playing with it, creating short commercials and 30 seconds videos. But I assume and I bet that by the end of next year, there's going to be a full featured film created with AI or at least an episode of like 30 minutes of something that would be like a series.

14:52What I just said obviously have profound implications on TV and Hollywood, and it will be very interesting to see where that goes. Another company that made some big announcements in the past few days is Microsoft. Microsoft launched Magentic One, which is an open source multi-agent system that is featuring five specialized agents working in concerts. So the idea is to create a multi-layer approach where you have an orchestrator, which is the lead agent that's coordinating all the other operations, a web surfer that handles browsing to navigate the web and get information from the web, file surfer that manages document and file operations, coder, which writes and analyzes code solutions, and computer terminal, which executes the codes and provides system operations.

15:38So with all of these working together, you can create magical, really advanced, complex processes that are auto run and executed through the orchestrator agent that is going to run everything. Now, they're not the first ones to release something like this. AWS has already released something like this. IBM released B agent. OpenAI has Swarm. So this multi-layer agent control system is something that we're going to see more and more of, and it will allow real advanced development of sophisticated processes that will become easier and easier as these systems will learn our needs. And with simple prompts and instructions, we'll be able to complete very complex tasks.

16:19Now, Microsoft also is launching its screen reading AI that they're calling Copilot Vision. It's something that they demoed before and it can analyze text and images on web pages in real time. And then you can ask it to provide summaries and translations and it can help in product discovery in online catalogs and it offers gaming assistant if you're in the middle of playing a game while you're on a browser, et cetera, et cetera. Basically, it's going to be AI layer integrated into the Edge browser that can see everything in the browser and provide information and assistant to everything that's on that web page.

16:55Right now, it's going to be released in the US only, and it requires you to be a paid member of Copilot with a$20 pro subscription. Now, the good news is that it only works with pre-approved popular websites, and it cannot do the same thing with like your bank account or stuff like that. So there's some guardrails that has been put in place out of the box for us to use it. The other safety guard is that it deletes all the data that it captured at the end of each session, and there is no storage or model training using the process content that was generated in each of those sessions. So in theory, you can use it safely to get assistance on every web page you're on if you are using the Edge browser.

17:35Google has already shared that they're releasing something like this in the very near future, probably in the beginning of Q1 of next year for Chrome. This will become a standard thing, regardless of which browser you're using, to be able to ask AI about what's on the screen right now. The next big and interesting feature comes from Anthropic. They just added Google Docs support, so you can connect multiple documents in a single chat. But the coolest feature in this is that there's an automatic sync, meaning if you are changing the version of the document, Claude will already know that. That, to me, was one of the biggest benefits of Gemini.

18:10So Gemini allows you to basically look into a folder and use that as the data set for a Gemini conversation without having to actually connect the files or upload them into the chat. this is still the case, meaning it's still a unique feature of Gemini that allows you to look into a folder, but this new functionality by Claude closes the gap somewhat that allows you to at least see updates when there's an update to a document and re-ask questions and get updates without having to re-upload the file into Claude. The disadvantage of this tool is that it cannot process images or comments in Google Docs, so it doesn't know how to process that.

18:48I would say it doesn't know how to process that yet. Another big news that came from Anthropic this week is a game-changing MCP protocol. And MCP stands for Model Context Protocol, which is a technical standard that enables seamless connection between AI assistants and data sources and tools. So if you think about to create successful working agents, which is the future of everything AI, you need three things. You need the model itself, the thinking entity, you need access to data, so whatever the sources the data are, and you need tools, you need access to the internet, a browser, an application, and so on to actually execute the things you need to execute in order to complete the task.

19:30And this MCP protocol that Anthropic just open source is supposed to be an infrastructure that will enable all of that. And because it's now open source, anybody can use it and more and more people can contribute. And I actually see that as a very solid step in the right direction of sharing this kind of information across different companies and organizations. What it allows you to do is it allows you to connect AI applications with databases and knowledge bases and business intelligence graph sources and then link chatbots to development tools and development environments and so on. So it basically allows you to connect the three main components seamlessly through this interface.

20:09Now, they already have several connectors built into it, like Google Drive and Brave Search and Slack, and they're planning to connect this to any other known organization data sources. The idea is obviously to break the silos of information. So far, to do stuff like that, you needed to deploy data lakes and data lake houses, which is in many cases a very complex and expensive process. And this may or may not replace DataLakes, but it definitely provides a smaller footprint, significantly cheaper and less complicated solution for at least some of the company needs that can be deployed very fast using now an open source architecture.

20:48Now, we both know that Anthropic just recently got another$4 billion investment from Amazon and AWS for a total of$8 billion so far. But Amazon are not just counting on Anthropic. They're actually developing their own models. And they just released a whole family of new models called Nova. And that series of models is going to be available on AWS Bedrock, which is their infrastructure for everything. And now they currently release three models. Nova Micro, which is fast, cost-effective text model. Then there's Nova Lite, which is a low-cost multimodal model for image, video, and text. And then there's Nova Pro, which is advanced multimodal model.

21:26And in 2025, they're planning to release Nova Premiere, which is advanced reasoning model, basically like GPT-01, Nova Canvas, which is image generation tool, and Nova Real, which will be video generation model with watermarking. And they're planning speech-to-speech and native multimodal models in 2025 as well. So as if we didn't have enough models in the mix, now we have a whole set of new models available directly from Amazon that are available on their Bedrock platform. I assume they will make it highly competitive rates for anybody on their platform to give some benefit to people to use their models over other models on Bedrock.

22:06Now, staying on Amazon and AWS, they released three big tools in the AWS environment. One is automated reasoning check. So it's a new tool that allows you to check the data to reduce or maybe eliminate hallucinations. It can cross-reference both public data as well as customer supplied information that resides in AWS. And it's similar to such offerings that already exist from Microsoft and Google. They also added model distillation, which basically allows to take information that was put into one large model and move it into smaller models. The idea is obviously to get higher efficiency and lower cost and higher speed on tasks that do not require the large language model.

22:52there's a lot of discussions about this and you heard me talk about this. That's the future where everything is going, right? And we talked about this before with delegation and multi-agent collaboration, which is the third thing that they have released. So they released a multi-agent collaboration tool that enables AI tasks to be distributed with supervisor agents and allows parallel processing of complex tasks by multiple smaller agents. So this idea of multi-layered, multi-tiered approach is something that's going to become the norm across all the different platforms and will allow the main tool to understand what the task and what the needs are and then orchestrate it across multiple smaller agents that will work in parallel to complete the task faster, but also with more specifically oriented models that will be tailored for specific tasks.

23:37Another company that we talked about last week, which is Cohere. Cohere is focusing on enterprise solutions. They released Rerank 3.5, which sets new benchmarks on enterprise search capabilities, its accuracy and speed. So it sees 23.4 % improvement over existing hybrid search systems. It's doing 30.8 % better than traditional search algorithms on financial data sets, etc. And it supports 100 languages on the data sets, including Arabic, Japanese, and Korean, which were considered to be more complex to query through. So significant benchmark achieved by Cohere in this particular solution, Cohere has selected to not compete in the crazy race for the most advanced models and instead are developing tools that are geared and tailored for enterprise benefits.

24:27And this new release is just going to push them even to be more competitive in that particular field. Another company that has made a release, and I told you lots of releases this past few days, Eleven Labs, which is a company that has been around for a while doing really advanced voice models, has made two releases. One of them is a customizable conversational AI agent builder. So what they released is a complete conversational agent building platform. It has multiple LLM options behind the scenes. So you can use Gemini, GPT, and Claude. And you can customize the voice parameters and response characteristics.

25:03And you connect it to multiple databases that they already have integrations with. So it comes with a full SDK for Python and JavaScript and React and Swift. So whichever platform you're using, you can connect it to that. And it has the capability to have multiple customizable variables for tone and response length. So you can adjust parameters such as the language selection, the response temperature, the token usage limit. So you can limit the length of each answer, the latency and stability of the model. So you can control how fast it's going to respond and take that into consideration when creating the answers, conversation length, et cetera, et cetera.

25:41So multiple controllable parameters in there. And they're obviously coming to compete with OpenAI real-time conversation API that has been released in the past few months, basically jumping into their field. So they've been the leader in the voice field for a while. And now both OpenAI and Google has stepped into their field. So this is just them upping their game and providing a complete SDK and controllable capability for agent development. From an availability perspective, this is awesome for anyone who wants to develop any kind of voice agents, whether for internal or external usage, whether for customer service, employee support, whatever you want, these tools will become more and more capable.

Read the full transcript

26:19And there's going to be a whole universe of applications that are going to be developed on top of these capabilities. In addition, Eleven Labs introduced an app that literally just competes with the podcast capability of Notebook.lm that we showed you several times in the past. So you can now upload PDFs and articles and eBooks and other types of data into it. and it will generate a conversational style podcast between two AI generated hosts. Again, literally directly competing with Notebook LM's podcast capabilities. And the last big release for this week, which is not a release, but more an announcement, Apple is planning to do a major overhaul for Siri to be LLM driven, which will allow it to be a lot more conversational.

27:04So that's the exciting news for anybody in the Apple universe. The not so exciting news is that they're planning to release it in 2026. So that's at least a year out, more likely more. And that's really disappointing from Apple's perspective. They have been very late to the AI game. What they released so far has been nothing impressive. And even that has been released later than they have suggested initially. I don't really know what's happening with Apple and their AI capabilities, but so far what they're doing is far from impressive and very late behind everybody else. And this is just another example of that.

27:41On the other hand, their competitors, Google, are releasing more and more advanced Gemini capabilities into Android already. Now, we're going to stay on rapid fire and now talk about stuff that is not new releases. And the biggest one for this week is that a few very, very interesting individuals are starting a company to create an operating system for agents. So the company has the weirdest name ever, and it's basically called forward slash dev forward slash agents. And they secured a$56 million seed round at a 500 million valuation led by some of the biggest VCs in the world. So you're asking, how is that possible?

28:20How can a company that just got founded gets this level of funding? The reason is the level of people and their background. So the CEO is David Singleton. He is the former Stripe CTO and was the Android Wear lead. The CPO is Hugo Barra, the former Android VP and Meta Oculus leader. And the team also includes former executives from Google, Meta, Dropbox, and Figma. Now, the goal is to create an operating system that will enable AI programs to collaborate on complex multi-step tasks. Based on what I mentioned earlier in this episode, you understand that's the holy grail. And if they can do this across multiple tools, multiple platforms, and not within a single platform, so it's not just on AWS, it's not just on Google's tools, et cetera, it can work across all of them.

29:07It's obviously very powerful and will provide an infrastructure for basically the future. So when you have people who have created Android that became the operating system for most of the phones in the world, saying that they're going to create a new operating system for AI agents, you understand the excitement. Now, even the investors have some really big names like Andrei Garpathy, which was one of OpenAI founders and worked for many years in Tesla, and ScaleAI CEO, Alexander Wang, and Palo Alto Networks CEO, Nikesh Arora. So really big names are in the investor list as well. Definitely a company that we need to pay attention to.

29:46And I will keep updating you as I learn new stuff. Right now, there's very little to know. Their website basically doesn't say anything. It's very generic. And they're in quasi stealth mode other than saying what they're going to develop. The interesting news is that they're going to deploy it, they're saying early 2025. So the first versions of this are coming very quickly. There's been some big poaching of talent this past week. So OpenAI just recruited three senior engineers from Google DeepMind, and all three are going to work on multimodal development at OpenAI Zurich's office. As I mentioned earlier, OpenAI is in a crazy race right now between Sora and Dali and image generation and voice and audio that is coming and that is being attacked from all different angles.

30:28And definitely having three leading engineers from DeepMind will help them stay ahead of the curve. Staying on the topic of OpenAI, Elon Musk seeks injunctions to block OpenAI's for-profit transition. So we talked about this a lot in the past few weeks. OpenAI are in the process of switching from a non-profit structure to a for-profit structure. Elon Musk has donated the first$44 million to get OpenAI off the ground in its early days. And there's a whole history in there that I'm not going to dive into in this episode because it will take another 30 minutes. But what Elon is targeting is OpenAI, Sam Brockman, Microsoft, Reid Hoffman, and Dee Templeton.

31:06And he basically trying to prevent this move. Now, this may have very significant implications on OpenAI because they just raised a huge amount of money and a huge amount of debt for a total of over$10 billion, all tied into a promise that they will be able to make that transition. Now, that transition is not easy, period, regardless of being sued by somebody, because the whole point behind collecting money and building value in a nonprofit is keeping it in the nonprofit realm. And I talked about the different implications two weeks ago on the episodes. If you want to dive into that, you can go and check it out.

31:40But right now, this will definitely put another roadblock in their path to success. Combine that with the fact that Elon Musk is now a buddy with elect President Trump, and he's going to play a role and have the president's ear. And there's a lot of conversations talking about potentially nominating an AI czar. And I assume that it will have at least some influence on the decision on who that is going to be and the decision that the new president and the new administration are going to put in place. So this definitely doesn't sound like promising news for OpenAI. Now, obviously, Sam Altman did not stay quiet on this, and he called the potential political interference of OpenAI's growth as profoundly un-American.

32:23Now, I tend to agree with him, but that being said, when you pick a fight with somebody like Elon Musk and you are growing the most influential startup in the past decade, but maybe in history, and you want to change it from a non-profit to a for-profit, who should expect that this was coming? It will be very interesting to see how that evolves. While Elon is a big name, and now again, he has the new elect president's ear, there are a lot of people in the OpenAI corner as well that are big and influential and have very deep pockets. So it'll be very interesting to see how this evolves, and I'll obviously keep you updated as this moves forward.

33:02Now, there's some new rumors that OpenAI may start adding advertising as another way to generate revenue. While this was very clearly denied by Sam Altman in the past and still officially is the stance of the company that no active plans to pursue advertising, CFO Sarah Fryer confirms that OpenAI are exploring advertising possibilities. So that might show up on whatever way in the new AI universe of OpenAI. OpenAI also had a new partnership deal this week, this time with Future PLC, which is the company behind some of the biggest publications in the world, like Marie Claire, PC Gamer, TechRadar, Tom's Guide, and many other names.

33:43There's 200 plus media brands under that umbrella. And there's been a relationship between these two companies before where Future has been using more and more of AI's technology, but now it's going to be a two-way relationship where OpenAI's GPT platforms will be able to use the data and the information coming from these platforms and to show it as part of the results in OpenAI. That's just another licensing deal that they have in place, like many others that they put in place before. And this is good news for anybody who's using ChatGPT to get real-time information across multiple domains. An interesting piece of news that is relevant to anyone who is trying to implement AI in their company actually comes from an interview with AWS CEO that is talking about the shift in enterprise and companies implementation of AI on AWS.

34:31So Matt German, the CEO of AWS have shared in the interview that he's seeing a big shift of companies from broad experimentation to focused implementation. and he sees a dramatic change from companies trying hundreds of proof of concepts to picking five or fewer e-applications that drive the highest ROI and just focusing on those as a step one. When I work with my clients, we actually try to do kind of like both in parallel, but we try to guess what are going to be the things that are going to drive the highest ROI to begin with and focus on those while allowing employees to experiment and develop small proof of concepts within very well-defined guardrails.

35:12And this actually proves to be very helpful and very useful to most companies where they can benefit from just a few bigger projects, but also benefit from a lot of small efficiencies that are generated by employees after they get properly trained and provided the right tools and the right guardrails in order to improve their own day-to-day work. We spoke earlier about competition to Notebook LM that now comes from several different directions. Well, Notebook LM leadership team has left Google to start their own stealth startup. Now, it's unclear what the startup is, but the Notebook LM team lead, the Notebook LM designer, and one of the top engineers, all three of them have left to have a stealth startup.

35:53They haven't said anything yet about what it is, other than the plans to focus on consumer-facing AI products. That's obviously riding on the success on the work that they've done in Google. Now, we mentioned Elon Musk in several different segments in this episode, but we told you before that XAI raised a lot of money and that they built the largest single location training computer on the planet, which is called Colossus, which is in Memphis, Tennessee, and it has 100 ,000 GPUs. We also told you that there's a short-term plan to make it into 200 ,000 GPUs from the latest version as well. So they're not just adding 100 ,000 GPUs, they're going to be the stronger, newer version of GPUs from NVIDIA.

36:33But apparently there's a plan to bring it to 1 million GPUs in total. And that's what they're working on right now as a longer term project. That's not something that's easy to achieve from any perspective, both in means of power supply, as well as many things don't scale linearly. So the fact that we're able to build 100 ,000 doesn't mean they can be 200 ,000. It definitely doesn't mean you can grow it up to a million, but that's the plan they have set themselves to pursue. And just to put things in perspective, when they built the current Colossus computer, which has 100 ,000 GPUs, as I mentioned, it's the largest GPU cluster in the world.

37:12They've done this in 122 days. Similar smaller projects have taken 9 to 15 months. And so if anybody can pull this off, it's Elon Musk and XAI and the people that they've been working with to create this new, incredibly powerful and incredibly power consuming new computer that will definitely overshadow any other computer on the planet from an AI perspective. Staying on the Elon universe, Tesla just unveils their Gen 3 Tesla bot with advanced dexterity capabilities. So they significantly improved the dexterity and the eye-hand coordination capabilities. And they've demonstrated some very incredible things.

37:53The most incredible one was that the robot can catch a tennis ball in midair. So catching something that's moving fast in midair is a very, very complex task to do. And now the robots can do that. But obviously, the idea is not that, is the fact that it has very high dexterity, which allows it to move stuff around, grab different objects in different ways and perform much more delicate tasks than most robots can perform today, which will allow it to be more useful in a lot more tasks, definitely in industrial manufacturing, but also in healthcare, etc. etc. And later on, also in household activities.

38:30And with a very interesting final happy and positive piece of news that stays on the robot realm, John Hopkins and Stanford researchers were able to build a training system for a surgical robot using videos. So they're using the DaVinci Surgical System Platform, which has been around for many years, but it was always operated by humans to do multiple types of surgery faster. And what they were able to do is they were able to build a training platform that will allow it to watch videos from hundreds of other surgeries. And the robots basically learn from watching these videos. They're now performing very successfully and very accurately three different critical surgical procedures.

39:16They do needle manipulation, tissue lifting and suttering all on their own and all without any human intervention and all just by watching videos. So this will obviously allow surgical robots to do a lot of surgeries in a highly accurate way. And they're claiming that they can train new procedures in just a couple of days by allowing the system to watch these videos. This could be transformative to the surgical world if this can be done at scale. And I don't see a reason why it couldn't. This could dramatically reduce the cost of surgery, allows to do more surgeries per day with doctors that don't get tired and can literally do this 24-7 without issues across multiple types of surgeries.

40:01And over time, this will be probably any kind of surgery or at least most, which I find as a really important piece of news. That's it for today. As I mentioned, drinking from a fire hose with new capabilities and new features from more or less everybody in the AI industry. Stay tuned for the daily release on 10 a.m. Pacific time for everything OpenAI is going to release this week. If not, just come back next week and I will share with you everything that they have released. Until then, we will be back on Tuesday with a really fascinating episode on how to. We're going to compare different leading large language models, but more importantly, we're going to show you how to do this with Hunch.

40:41Hunch is an incredible AI platform that is flexible and allows you to do magical things, connecting multiple AI platforms together, but in a very easy to use user interface. And we're doing this with the help of Hunch's CEO. Don't miss this Tuesday's episode. And my last thing, as I request every week, if you're a regular listener to this podcast and you find value in it, please open your phone right now, share the podcast with people that you know that can benefit from it. And rate and rank us on your favorite platform, whether it's Apple Podcast or Spotify. I would really appreciate that. And until next time, have an awesome weekend.

41:20Enjoy your time with family and friends. Keep on experimenting with AI. Share what you learn with the world, because we all have more to learn about it so we can have the AI revolution help us while reducing the risk as much as possible.

From the publisher

Are we living through the fastest AI innovation sprint in history?

This week, the AI world was set ablaze with groundbreaking releases from OpenAI, Google, and Amazon, among others. From OpenAI’s ambitious “12 Days of OpenAI” campaign to Google’s jaw-dropping Veo video generation capabilities, and Amazon unveiling the Nova series, innovation is moving at lightning speed.

What do these advancements mean for businesses, creators, and technology leaders? How will the introduction of multimodal agents, advanced reasoning tools, and video AI redefine the future of work and creativity? This episode dives into the transformative announcements and their implications.

If you’re looking to stay ahead of the curve, you can't miss this episode: 

In this session, you’ll discover:

  • OpenAI’s O1 Pro model: what it is, why it matters, and how it’s setting new benchmarks in STEM reasoning.
  • Google’s Veo: a 1080p video generation model designed to disrupt content creation.
  • Amazon’s Nova family: affordable and advanced AI models for text, images, and video.
  • The rapid evolution of multi-agent systems and their potential to redefine enterprise operations.
  • How OpenAI’s new pricing tiers could shape access to advanced AI tools.
  • Why video AI could soon surpass human-generated content in quality and speed.

About Leveraging AI

If you’ve enjoyed or benefited from some of the insights of this episode, leave us a five-star review on your favorite podcast platform, and let us know what you learned, found helpful, or liked most about this show!

More from Leveraging AI

All 330 episodes
147 | OpenAI’s 12 Days of Releases, Google’s Game-Changing Video Generation Leap, and Amazon’s New Models (Nova) and more AI news for the week ending on Dec 7, 2024Leveraging AI · 42 min
Listen in VO