In short
Podcast Summary: Leveraging AI - Episode 73
Episode Overview In this episode of Leveraging AI, host Isar Meitis discusses significant advancements in the field of artificial intelligence, focusing on their implications for businesses and personal lives. A variety of topics are covered, including NVIDIA's groundbreaking Project Groot, insights from leading AI figures like Sam Altman, and emerging technologies from companies such as Microsoft, Apple, and Google.
Important Announcement
- AI Business Transformation Course: Starting April 1. Aimed at helping individuals understand AI's impact on business and career advancement.
---
Key Topics Discussed
- NVIDIA's Project Groot
- Overview: A general-purpose foundation model for humanoid robots.
- Capabilities: Includes understanding natural language, human movements, and learning new skills.
- Significance: Positions NVIDIA at the forefront of humanoid robot infrastructure, expanding AI's reach into physical tasks and blue-collar jobs.
- Sam Altman's Insights
- Interview Highlights: Discussed AI governance, personal ambitions for improving human lives, and reluctance to specify future releases such as GPT-5.
- Critique of GPT-4: Altman describes it as "sucks," indicating significant advancements beyond it are forthcoming.
- Elon Musk and XAI
- Grok Release: An open-source AI model released by Musk's company, XAI.
- Philosophy: Musk advocates for open-source AI to avoid concentration of power in a few corporations.
- Microsoft's Strategy
- Satya Nadella's Statements: Microsoft’s robust relationship with OpenAI ensures continuity of AI services even if OpenAI were to cease operations.
- Hiring Mustafa Suleiman: Former DeepMind co-founder joins Microsoft to lead AI integration across various products.
- Emerging AI Technologies
- Google's Vlogger: An AI system capable of generating realistic videos of people from a single photo.
- Stability AI's SV3D: Allows for the creation of 3D videos, enhancing e-commerce capabilities.
- Apple's MM1 Research: Demonstrates the importance of data quality over quantity in training AI models.
- Devin.ai: Introduced as the first AI software engineer, capable of generating complex source code.
- Open Interpreter: A voice-controlled interface for personal computers, emphasizing future interactions with technology.
---
Key Takeaways
- AI's Rapid Evolution: The episode highlights the rapid advancements in AI technologies and their broad impact across industries.
- Ethical Considerations: Discussions around governance and ethical AI use are pivotal as capabilities grow.
- Business Applications: The introduction of advanced AI tools promises to democratize tech development and enhance business processes.
- Future of Human-Computer Interaction: The trend is moving towards more intuitive and natural interactions with technology, emphasizing voice communication.
---
Conclusion The episode encapsulates critical developments in AI, emphasizing the importance of ethical considerations and practical applications in business. The ongoing evolution of technology suggests a future where AI plays an integral role in both professional and personal realms.
Additional Resources
- AI Business Transformation Course: [Enroll here](https://multiplai.ai/ai-course/)
- YouTube Full Episodes: [Watch here](https://www.youtube.com/@Multiplai_AI/)
- Connect with Isar Meitis: [LinkedIn Profile](https://www.linkedin.com/in/isarmeitis/)
- Join Live Sessions and Newsletter: [Event Information](https://services.multiplai.ai/events)
---
*If you found this episode insightful, consider leaving a five-star review on your favorite podcast platform!*
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Hello, and welcome to a weekend news edition of Leveraging AI, the podcast that shares practical ethical ways to improve efficiency, grow your business, grow your business, and advance your career. This is Isar Maitis, your host, and we have a very exciting week of news from the AI world, some really interesting news from the biggest personas in the AI world, as well as very interesting news from different companies and new developments. Before we get started, I would like to remind you that our AI Business Transformation Course, Open Public Cohort, is starting on April 1st, so roughly a week from the time this podcast gets released.
0:35if you are interested in learning how AI can impact your business, or if you're looking for a way to advance your career with the developments of AI, this course is an incredible kickstart to that process. We teach everything from the basics all the way to frameworks on how to implement AI strategically in your business across different departments and everything that comes with it. We have been teaching these courses since April of last year. Most of the courses we teach are private sessions for companies and organizations who book it in advance. So we have open to the public courses roughly once a quarter.
1:11We did one in the beginning of January. We're doing one in the beginning of April. And I'm not sure when the next one is going to be. So if this is something that's interesting to you, and it should be, check out the course. There's going to be a link in the show notes. Don't miss it out because again, I'm not sure when the next one is going to be. We're fully booked for the next three months with private closed session courses. Speaking of which, If you are a head of an organization or association or company and you're looking for a way to dramatically advance the AI knowledge of people in your organization, reach out to me on LinkedIn and I will gladly let you know how we can set this up.
1:46And now for this week's news.
1:55We're going to start with the biggest news of this week. NVIDIA came out with a huge announcement this week. They have announced what they call Project Groot, which is actually not said Groot. It's just spelled G-R-0-0-T. And the reason they're saying G-R-0-0-T is so they don't get in trouble with Marvel and Disney. But if you try to read G-R-0-0-T, it spells Groot. So Groot is a new general purpose foundation model for humanoid robots that is built on NVIDIA, both software and hardware infrastructure. And what they're developing is basically an operating system and infrastructure for humanoid robots.
2:34The goal of this platform is to create an infrastructure for all the capabilities of humanoid robots, including understanding of natural language, human movements, learning new skills, etc. All this infrastructure that is going to be required for any new development of humanoid robots is coming from this NVIDIA platform. And as I mentioned, it could include more hardware and software, or what NVIDIA calls system on a chip or SOC, and that is optimized for performance and power specifically geared towards humanoid robots. Now, in addition, they've announced multiple upgrades to their ISAC platform that they've announced previously.
3:13That is a set of foundation models and a set of tools that will allow other companies to develop new models and new AI capabilities on top of their architecture. They're expecting the new Isaac platform capabilities to be released in the next quarter. So this is not sometime in the very far future. This is coming in the immediate future. If you have been following everything that's happening in the humanoid robots world, you know that this is accelerating very fast with some of the biggest players in the world jumping in, but also smaller startup are coming in. And this new architecture will enable new startups and companies and even existing established companies to start with a very solid starting point.
3:52It also places NVIDIA in a very interesting junction, not just for training and running models, which drove their huge growth so far, but also as the next infrastructure for humanoid robots, which is the next frontier that will allow AI to not just address knowledge work and knowledge capabilities, but also do things in the physical world. And if you want, address blue collar jobs as well. As I told you, there have been some very interesting interviews and announcements on the personal level of people at the top of the totem pole when it comes to AI development. And we'll start with Sam Altman, the CEO and the founder of OpenAI.
4:31He was interviewed by Lex Friedman this week. I highly recommend listening to the podcast because there's a lot of new answers, but they talk more or less about everything. They talk about the issue of Sam being fired and then put back as the CEO of OpenAI. They're talking about the relationships and the lawsuit with Elon Musk, and they're obviously talking about new and existing developments of open AI. It's a fascinating interview if you want to understand what's going through the head of the person that's leading the charge of AI development in the world today. I want to share two important aspects the way I see it in this interview.
5:06The first one is that Sam is really driven by trying to make our world and human lives better with AI. After listening to this interview and other interviews he's done in the past, I think it's very clear he's very sincere in his wish to make humanity better using AI tools. I think it's also clear that he believes that it will require some serious governance that may or may not exist right now. And he admits that in order to control what these systems can do, whether through a company, an organization, governments, international cooperation. But I think it's very clear listening to him that he's still not sure that there's a clear solution for that.
5:47And yet they're moving forward very fast in the development. So that's on the conceptual side. On the very practical side, he wouldn't release any clear information or when or what they're going to release. And he wouldn't name exactly what's coming in GPT-5. But he did say when Lex asked him about what he thinks about GPT-4, he said, and I'm quoting, it sucks. which really tells you that what they have right now is probably significantly better or the things they're seeing for the future are way better than what we see in GPT-4 which is still most likely the most capable model we have out there today.
6:23Some would say that Claude 3 is better but it's still top of the line model and if Sam Altman is saying that it sucks it means what they have right now is so much better that it will make a very big difference. He said that the jump GPT-3.5 to GPT-4 is a jump that is not even as significant as we're going to see going from GPT-4 forward. But as I mentioned, he wouldn't name exactly when things are going to be released, whether asked about Sora or when asked about GPT-5. But he did say that they're going to release something significant that is going to be much smarter, and you can decide what that means still this year.
7:05Specifically about Sora, he's saying that it still has a lot of issues, and despite the fact that it looks amazing and everybody's really excited, there are still a lot of issues that they're still fighting through, and it's still not good enough to be a product that can be released to the public. So overall, very interesting interview, giving a glimpse into how OpenAI thinks and approaches the development of new models, how much they feel the responsibility of doing it right, as well as the fact that we're most likely going to get a really advanced model from them sometime this year. As I mentioned, as part of the interview, Sam Altman refers to his interesting and complex relationship with Elon Musk and everything that happened with Elon leaving and now Elon suing the company.
7:45But in a very interesting timing, when Sam is being interviewed by Lex about this topic, Elon Musk's company, XAI, has released the source code and everything you need of their AI model, Grok. They literally took everything they have other than the training data itself. So everything they have, meaning the actual source code, the base model weights, the network architecture, 314 billion parameter expert model, and have released it to the public to be used however people want. That's as part of, obviously, Elon Musk's statement that these really powerful models have to be open source in order to ensure the future of humanity, the way he refers to it, but at least the benefit for all humanity versus closed models that will benefit a few giant companies, which is the reason why Elon Musk jumped in and financed the beginning of OpenAI is to be a counterweight to Google after they bought DeepMind and made it into a closed source, controlled environment by Google.
8:51He wanted to have an open source alternative, and that's why he joined the team to start OpenAI. And that's part of the reason for the beef between Elon and Sam Altman right now. But as I mentioned, in a very interesting timing, they've just released Grok. Now, Grok hasn't proven yet to be a very powerful model. They've developed it extremely fast, which is interesting. but they haven't shown anything that is as impressive as some of the leading models right now. That being said, I will never bet against Elon Musk because everything he has done so far has been extremely successful after he pushed it far enough and hard enough, and he's definitely pushing hard on this very particular angle.
9:30As of right now, Grok is a fully open source model, one of the few that could be very powerful player in the future. Another very interesting person in this AI race is obviously Satya Nadella, the CEO of Microsoft. And he was quoted saying something very interesting about their relationship with OpenAI. Most of what Microsoft is releasing right now as part of their AI capabilities is based on OpenAI's chat GPT, architecture, infrastructure, models, and so on. And they're including them as part of everything Microsoft from Copilot to future operating systems to Office 365, AI capabilities, and so on and so forth.
10:12There's been a lot of negative feedback from within Microsoft to the deep reliance on OpenAI. And obviously, there was all the questions about OpenAI themselves with the issue of firing and returning and the control of the board over everything happening in OpenAI back at the end of last year, as well as the current lawsuit by Elon Musk and so on. So Satya Nadella was just recently asked about it. And he said the following, we have all the IP rights and all the capability. I mean, look, if tomorrow OpenAI disappears, I don't want any customers of ours to be worried about it. Quite honestly, because we have all the rights to continue the innovation, not just to serve the product.
10:54But we can go and just do what we were doing in partnership ourselves. And so we have the people, we have the compute, we have the data, we have everything. While this is not surprising, it's at least putting it out in the open that while OpenAI has benefited, obviously, from this partnership with Microsoft, Microsoft made sure they are fully covered in this relationship. And as Satya Nadella says, if OpenAI disappears tomorrow, Microsoft can continue from that point. And in a very interesting development on the same topic, Mustafa Suleiman, who was one of the co-founders of Google DeepMind, who has left DeepMind to found Inflection AI, the company who gave us Pi AI.
11:36You can go and visit their website at heypi.ai. Inflection was developing a new kind of AI model that will be more personal companion that you can have real deep conversations with. A lot of people really apply for that reason, because you can really have a deep conversation with it, including over voice, which is very interesting and different. If you haven't tried it, I suggest doing it because we'll give you a glimpse on what these systems will probably do in the very near future. But Mustafa Suleiman just announced that he's leaving his position as the co-founder and CEO in Inflection and moving to Microsoft to serve as the executive vice president of Microsoft.
12:13and he's going to be joined by several other senior leadership team that are leaving inflection together with him. He's going to report directly to Satya Nadella, which shows you the importance of this role. He is going to be in charge of everything AI Microsoft, including developing future consumer products and AI efforts that puts Under Y Umbrella, CoPilot and Bing and Edge and everything else that they're doing with AI right now. Now, it was very clear that Microsoft is all in on AI and everything that they're doing is going to be AI based moving forward. It had made them the most valuable company in the world right now.
12:47So that definitely tells you that at least from a shareholder perspective, they're moving in the right direction. And putting somebody with that level of experience in both the research side from DeepMind as well as creating products from inflection is a very logical move. move, and I'm sure it will yield very interesting results for integrating AI into everything Microsoft. And from these big tectonic moves or announcement for big people, there were a lot of small research-related news this past week. Many of them are video-related. So Google researchers just shared that they have developed what they called Vlogger, which is an AI system developed to generate lifelike videos of speaking people, so an avatar of people, including the gestures and the movement and so on from a single photo and using that photo together with text to generate a complete video clip of that person speaking.
13:39They have trained that model with over 800 ,000 diverse identities and over 2 ,200 hours of video, which per them, enabling it to generate videos of people from varied ethnicities, ages, clothing, poses, surrounding environments, and so on without bias. Those of you who have been using these kind of tools like Synthesia and HeyGen know that these tools are already very capable. And if Google now develops the capability to do this at even a better quality, it is very promising and at the same time very disturbing. The reason it's very promising is it allows to create training videos, marketing videos, product explanation videos, or any other video you want to create of any person for any target audience very easily without the need for camera and lighting and editing and so on.
14:31So the creation of videos for specific needs that we have today will become significantly easier. That being said, it also generates very significant concerns because deepfake will also become very easy because you'll be able to take a single image of any person and make them say whatever you want them to say in whatever scenario you want to create. And so that's obviously very troubling because right now, and I don't see that changing in the near future, there's no real way to detect these videos. Now, right now, they're not fully realistic, and you can tell that they're not real, but sometime in the next 6 to 18 months, we will not be able to distinguish between a real video and a fake video.
15:11And those of you who have seen Sora know what I'm talking about. So combine that with the research capabilities from Google, and you have the perfect storm. On one hand, amazing efficiencies. On the other hand, very serious risks of deep fakes. Staying in the field of video generation, Stability AI just released what they are calling SV3D, which stands for Stable Diffusion 3D, which is an AI model that enables to generate 3D videos building on top of their previous model that was called Stable Video Diffusion. So what you can do right now is do one of two things with SV3D. You can either create a 360 orbital video of e-commerce object.
15:51So the trick here is obviously keeping the consistency of the object from 360 degrees without actually having the object. So that was the core of the development and the innovation that they are sharing. You can actually do a 360 degree video of a thing that you want to sell without having the thing and without actually having the capability to shoot a 360 orbiter video. The other part of this model allows you to actually define the 3D path around the object and have the camera move however you want it to move versus just a 360 orbit and still shoot the 3D object while having it staying consistent and while having the background move in a way that looks realistic.
16:32So we are moving forward to a situation where the consistency of object, which is the biggest issue with creating videos with AI is getting resolved from multiple angles, whether it's Google or Stability AI or MidJourney that announced something similar. So MidJourney announced that they are working on a new 3D video and real-time creation of models that will allow to simulate the entire world. So those of you who have seen the Sora examples from OpenAI and have been in the discussion, the discussion basically said that what Sora is doing is not just generating video, it's actually simulating the real world, hence allowing it to generate these highly realistic videos.
17:17And this is exactly the direction that Me Journey announced that they're going. Those of you who don't know, Me Journey are currently the top image generation model as far as creating realistic images. And they have shared this news in part of their office hours on Discord where people are asking them questions and they're sharing news about what they are working on. So this new development is focused on creating a real-world model that will allow people to create video games and shoot movies and do everything they want in like a sandbox of the real world. They also shared that the jumps from version 6 to version 7 in the capabilities of Me Journey is going to be significantly better than the jump from version 5 to 6.
17:59And so big, new, interesting developments coming from Me Journey. They also shared that the capability to create 3D models most likely is going to arrive before the capability to create videos. And this comes shortly after they released another capability into version 6, which is the ability to create consistent objects and people across various images. So as I just mentioned, the biggest thing here is the ability to create consistent people and consistent items and consistent background across the video and across various angles, because that is the key to creating a highly realistic video and looking at Sora and listening to these kind of news on what's in development.
18:39This is the direction that we're going, and we're most likely going to have this capability still this year from probably multiple companies. And if speaking about research, a company that we don't talk about a lot, but is definitely doing a lot, at least behind the scenes, is Apple. So Apple just released an interesting research where its researchers have developed what they called MM1. It's a set of large language models. But the main point of the research was to show that the data that is being used to train the model is as important as the amount of compute or the amount of parameters you're going to use.
19:12And what they've done, they were able to achieve state-of-the-art performance by using couples of data such as images and text and voice and text and so on, and by doing so, achieving very successful results with significantly less parameters and significantly less compute. In the past few weeks, we shared several different companies and several different processes on the research side that are showing multiple ways of how to achieve these amazing results that these models are achieving or that needs to achieve in the future to achieve AGI, and while investing less resources, which is very important because right now the process of training the models and running the models requires a huge amount of compute and a huge amount of resources and hence have a very bad impact on the world.
19:59So learning how to do this more effectively, such as these methods that were just recently shared by Apple's researchers, is a very important step in the right direction. And two small but very interesting releases that happened this week. One is Devon AI was announced, and Devon is what they call the first AI software engineer. It can generate complete source code from a single prompt. It can generate hundreds of lines of codes and perform the debugging and handle the deployment of the code. So this is a very step forward from the code generators that we had before that could generate snippets of codes that you had to put together and rig.
20:38The other very interesting thing about this tool is that it's basically an agent that is geared towards creating code. So the tool can search the internet and go through tutorials to learn how to accomplish different tasks and troubleshoot issues with its existing code or existing code that you give it. So it is a software engineer in a box. It can really do all the steps of the process, including developing new algorithms based on existing algorithms by learning how they work, understanding the needs, understanding the gaps, searching the internet, and then creating new models based on all of that information.
21:16The direction is very clear. Creating code, creating software will be something extremely different than it is right now in the future. I don't know how far into the future, and I don't know how much of software development will change, but most likely almost every person that wants to will be able to create new applications and new software using natural language and everything else that comes with a software will be done by the AI. Will this actually replace software architecture on a large scale? I don't know. This new tool definitely hints that this is the direction this is going. I think that's going to be the biggest difference between major big software and small applications.
22:03Everybody will be able to develop. Major software will probably still require some major architecture and planning beyond what AI can do at least now, but that doesn't mean it won't be able to learn that as well in the next few years. And the last thing, a company or group called Open Interpreter has released an open source code to a new, very interesting tool, which is a voice interface that controls your home computer. It's a little device that you can wear on you or hold in your pocket, however you want, that can understand your voice when you speak to it and can take actions from your home computer with everything in it.
22:40So that includes access to the internet, to search anything you want, the ability to connect to your emails, to your calendar, and basically any piece of data that you have on your computer and take actions on your behalf while you are not next to your computer, or in theory, while you are next to the computer, it's just an easier user interface. So two very interesting things about this. I think the direction this is going is very clear. We will be able to do a lot more by just talking to computers. Exactly how the interface is going to look like, whether it's going to be a wearable or our phones or something that will replace our phones, or whether it's going to talk just to the cloud and the world and the universe or to our personalized computers, or maybe all of the above, is something that still takes time and will evolve into something that I think will become common, just like cell phones are now common, but the direction is very clear.
23:29Voice communication, natural voice communication will control everything in our future lives. And it's just a matter of time until this becomes a product that we use every single day. That's it for this news edition. We are coming back on Tuesday with a fascinating episode, interview, deep diving into a specific topic. So don't miss that. And a final reminder, the AI Business Transformation course is starting in about a week. Don't miss out on that because the next one may be months out. So check out the link in the show notes. And until then, have an amazing rest of you.
From the publisher
In this episode of Leveraging AI, Isar Meitis discusses groundbreaking developments and the potential future of artificial intelligence, focusing on the implications for businesses and personal lives.
IMPORTANT: Don't miss this 2nd open publicly available cohort of the AI Business Transformation starting on April 1. We won't know when the next open-to-public will be offered so check it out here: https://multiplai.ai/ai-course/
Topics we discussed:
- NVIDIA's Project Groot and its potential to revolutionize the field of humanoid robots.
- Insights from Sam Altman's interview with Lex Friedman, highlighting OpenAI's future directions.
- Elon Musk's company, XAI, releasing the AI model Grok as open source.
- Satya Nadella's perspective on Microsoft's independence from OpenAI.
- Mustafa Suleiman's new role at Microsoft and its significance for AI integration.
- Advanced video generation technologies from Google and Stability AI that could reshape media and raise deepfake concerns.
- Apple's research on efficient AI training methods, promising less resource-intensive model development.
- Devin.ai, a groundbreaking tool that could democratize software development.
- Open Interpreter's new voice interface, hinting at the future of human-computer interaction.
Stay ahead of the curve in the AI revolution.
About Leveraging AI
- The Ultimate AI Course for Business People: https://multiplai.ai/ai-course/
- YouTube Full Episodes: https://www.youtube.com/@Multiplai_AI/
- Connect with Isar Meitis: https://www.linkedin.com/in/isarmeitis/
- Join our Live Sessions, AI Hangouts and newsletter: https://services.multiplai.ai/events
If you’ve enjoyed or benefited from some of the insights of this episode, leave us a five-star review on your favorite podcast platform, and let us know what you learned, found helpful, or liked most about this show!



