In short
Leveraging AI Podcast Episode 176 Notes
Episode Overview Title: ChatGPT’s Change Everything (Again) with a New Image Generator Description: The episode discusses significant advancements in AI, focusing on new capabilities that enhance productivity and the impact of AI in various sectors. It highlights the necessity for businesses to adapt to these rapid changes.
Key Topics Covered
- Impact of AI on Productivity
- Research Findings:
- A Harvard study revealed that individuals using AI tools can match the output of two-person teams, representing a 37% productivity boost.
- Solo employees using AI tools saved 16% of task time while teams using AI saved 12%.
- Teams using AI produced three times more high-quality ideas compared to those working without AI.
- Emotional well-being improved for those using AI, contradicting the traditional trade-off of "faster, better, cheaper."
- Rapid Advancement of AI Capabilities
- Doubling Speed of AI Abilities:
- AI capabilities are doubling every 3–7 months.
- Current models can autonomously handle tasks that previously required human input, showing significant cost-effective advantages (e.g., AI performing coding tasks at 10% of the cost of human coders).
- New Developments in AI Tools
- OpenAI's GPT-4o Image Generation:
- GPT-4o includes advanced image generation capabilities that outperform previous tools like DALL-E.
- Users can create high-quality images directly through the ChatGPT interface.
- This feature has led to high demand, causing service slowdowns.
- Future Predictions and Concerns
- Bill Gates' Insights:
- Gates predicts that AI will make roles such as doctors and tutors as common as smartphones within the next 10 years.
- Potential Job Displacement:
- The rapid implementation of AI could lead to significant workforce redundancies, necessitating societal adjustments including new educational frameworks.
- Legal and Ethical Considerations
- Copyright Issues:
- A federal judge allowed lawsuits against OpenAI regarding copyright use of material for training AI.
- Contradictory rulings highlight the ongoing debates around AI's use of copyrighted material.
- Advancements in Robotics
- Humanoid and Functional Robots:
- Emergence of robots capable of manual tasks and working alongside humans, emphasizing efficiency in logistics and workforce management.
- Companies like Dexterity and Figure are leading developments in humanoid robotics with improved mobility and functionality.
Key Takeaways
- AI as an MVP: AI is transitioning from a mere assistant to a crucial team member in workplace efficiency.
- Broader Workforce Implications: The enhancements in productivity from AI tools could lead to significant shifts in labor demand, requiring proactive strategies from businesses and policymakers.
- Legal Precedents: The evolving legal landscape regarding copyright and AI's use of creative content will continue to affect how companies utilize AI technologies.
- AI in Everyday Life: The integration of powerful AI tools into business functions will soon make them commonplace, impacting various sectors beyond tech.
Upcoming Events
- AI Business Transformation Course: Registration is open for a course aimed at helping leaders implement AI across their organizations effectively.
- Promo Code for Discount: LEVERAGINGAI100 for $100 off.
Closing Remarks Listeners are encouraged to share feedback and engage with the episode's content. Insights gained from the discussions can empower individuals and businesses to navigate the rapidly evolving AI landscape effectively.
---
This structured summary highlights the key discussions, findings, and implications surrounding AI within this episode, providing a comprehensive overview for business professionals and enthusiasts interested in the ethical and practical applications of AI technology.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Hello, and welcome to a weekend news episode of the Leveraging AI podcast, the podcast that shares practical ethical ways to improve efficiency, grow your business, grow your business, and advance your career. This is Isar Maitis, your host, and we have a packed show today. There are so many big things that we need to talk about that I could have probably done two episodes this weekend, but we're going to try to make it as efficient as possible for all of you. So as a first step, we are going to talk our three main topics today, even though they could have been easily six, but I still had to pick three, are going to be two research papers, one talking about how AI improves the efficiency of individual and teams on a research done by Harvard together with Ethan Mollick.
0:42The second is going to be a research that is showing how fast AI is doubling in its ability to perform business tasks. And the third is an incredible new release by ChatGPT. But then there are a lot of other releases that are still really important and capable and many other things to discuss, including some copyright laws, improvement in robotics, and many other things to talk about. So let's get started.
1:15As I mentioned, our first topic is going to be about the impact that AI can have and is having on people who know how to use it at work. In a study that was done by Harvard together with Procter & Gamble, they revealed that individuals using AI can match the performance of a two-people team. And if you want the specifics, that's a 37 % boost to a solo individual performer. The way this research was done is it tested 776 Procter & Gamble professionals in several different categories. There were people working alone, there were people working in teams, and then there were individuals working with AI and then teams working with AI.
1:53And they compare the quality and the quantity of their work as well as their emotional well-being during the work performed. The tools that we're using are GPT-4 and GPT-4-0, and they were providing clear training for these people on how to use AI properly for these tasks, which if you've been listening to this podcast, I told you many, many times before is the number one key for success in AI implementation. Now, what they found is, as I mentioned, very, very interesting. They found that a single individual that knows how to use AI properly matches the results of a team of two people. Basically, you can save 50 % of the workforce and get the same results.
2:33They also found that teams working with AI dramatically outperformed teams working without AI. And the best output came from teams working with AI that worked better than teams working without AI, outperforming solo individuals working with AI by 40%, and creating three times more top percent ideas and solutions than teams working without AI. Now, in addition to generating better results, they also worked faster. So individuals shaved 16 % of their task time working with AI and teams cut off 12 % of their time working with AI. I can tell you from the things that I'm doing with clients and myself that once you deploy this company wide and not on very specific use cases, these numbers could be significantly higher, shaving 50 % of some tasks and 90 % of others on specific tasks.
3:25Now, in addition, there were a few other benefits that maybe weren't expected. One of them is that in the teams that working without AI got stuck on their area of expertise. So the tech teams came up with tech ideas and the salespeople came up with more commercial ideas, but the groups that worked with AI came up with a more balanced solution. So the tech team working with AI also came up with commercial ideas and vice versa, which tells you that the versatility of bringing in a quote-unquote AI teammate that is more well-rounded in its approach can provide significantly more benefits and generates value to the company way beyond the team working without AI.
4:06In addition, the other interesting thing is that participants that worked with AI reported a more positively emotional experience to the work compared to people who worked without AI. So what does that tell us? It tells us what we already knew, but it's a scientific proof that working with AI breaks the old concept of faster, better, cheaper, pick two. That has been the common state in every work in the company, in every task in the world. You can either, if you want to do something faster, it's going to cost you more money. If you want to do it better, it's going to cost you more money or going to take you more time, et cetera.
4:39And now for the first time in history, literally anybody or any group or any company can work faster and better and cheaper and apparently have a better work experience from an emotional perspective by working with AI. Now, what do those improvements to capacity and capability mean for the broader workforce? That's not something that this particular research has tackled. But if you think about it, even if they're talking about 12 to 16 % improvement in efficiency, you can use that to grow the company by 12 to 16%. But if your market or your competition will do the same thing and will not allow it, that means that 12 % to 16 % of the capacity of the company will not be necessary anymore, which can have really bad implications on the global workforce.
5:23There is not endless demand, and hence, every time I hear that people saying, oh, so you can grow a business, it's true right now. But when everybody does it, and everybody will have that extra capacity and extra capability, or when these tools become significantly better and are not providing 12 % improvement by 30, 40, 50 % improvement, then we have a serious problem on our hands when it comes to our society. Now to pour some more gasoline to what I just said, in a new study by a organization called METR, M-E-T-R, they research the length of tasks that AI can handle autonomously. And what they have found is that this capability for the past six years doubled every seven months.
6:06And they also found that in recent months, it's been doubling every three months. So not only that it's doubling very fast, way faster than Moore's law, it is now accelerating in the speed of what it can handle. But let's dive a little deeper to how this research was done and what exactly they're trying to show. So first of all, what they were trying to see is the length of time of a task that AI can perform at a 50 % success rate. Basically, they took humans and tried to see how long it takes the average human or great humans to perform a specific task. That by itself is an interesting benchmark because does that compare it to an average human or a very good human in that task?
6:45But it was mostly coding related tasks. AI back in 2019 could perform tasks at a 50 % success rate for only four tasks that were just a few minutes. So anything longer than that, it would fail miserably. And now, Claude 3.5 Sonnet has hit the 50-minute mark, so almost a full hour of coding on its own without any human inputs and still performing tasks at a 50 % completion success rate. Now, they shared another interesting statistics that over 80 % of successful AI runs has cost less than 10 % of the human software engineer wage for the same exact task. Going back to what I said before, that is a 90 % savings if you let the AI code it versus the human code it.
7:34Now, this wasn't just one task. It was all across the software development universe with 66 different software tasks, including coding, debugging, and other software-related tasks. So while it's still within the software development domain, it's not just one type of task, like writing a specific kind of code. Now, what they are stating is that if this trend holds, and we're going to talk about there's a lot of ifs in the background of this, if the trend holds by 2027 to 2029, somewhere in that timeframe, AI might be able to manage eight hour days or maybe even full weeks or in the end of that, maybe months of work unsupervised.
8:13That provides a huge amount of potential and a huge amount of risk and fear, if you ask me, but it is where this is going if we follow the existing trajectory. Now, they are themselves are clearly stating that they're not sure of any of this, including the fact that they looked at tasks that are at 50 % success rate, which is obviously completely unacceptable in any workforce. Think about your employees failing in everything that they do 50 % of the time. So that's obviously not a good enough benchmark, but what they were trying to show is not AI's ability to solve today's business problems, but its ability to improve in doing this over time.
8:52And I think it will be very interesting to see what happens if they try to do similar tests or somebody else picks up this concept and does similar tests for the same success rate as humans, right? It doesn't need to reach 100%. It just needs to be as good as the average coder, not even the best coder because you have a few great code or any other people in any other domain, and then most of the people are average or around the average. So that will be very interesting to see where that goes. The other things that they're saying that it's not clear that the scaling laws will continue because of multiple reasons.
9:25We ran out of data, there's not enough chips, and many other reasons that can slow down the acceleration of AI. The flip side of that, not from the research, we are seeing many new innovations that drive acceleration beyond the amount of chips, beyond the amount of capacity in GPUs and power, beyond the amount of data through through reasoning, through better algorithms, through things like that, like all the latest updates that we talked about from companies like DeepSeek, et cetera. So we're seeing these things happening. So it may actually go the other way. And they're stating that as well.
9:58So that timeframe of within a few years, these tools might be able to handle full days or full weeks or even months of independent work may actually happen faster and not just slower. Why is this important? And especially when you tie it back to the previous topic of humans working with AI delivering significantly better results than humans working without AI, the world we know is changing. And it's changing pretty fast. And they're even stating themselves that even if they're wrong by an order of magnitude, this thing will change everything we know. It will just take a little longer. One thing that I can tell you is that nobody is ready for that because this may mean that we need significantly less people working.
10:39This may mean we need a completely different education system. This may mean we need a way to make people find stuff to do and get paid so they can buy things so the economy continues to work. And this is just another wake-up call to the trajectory of this thing that is potentially accelerating. And if you combine all the points and if you combine what all the leading labs and the leadership of the leading labs are saying that AGI is coming, Some of them are saying within this year, some of them are saying within the next two years, but in the very, very near future, we have to start taking action as a society and figuring out what that means on a broader scale.
11:16And from research to very practical news that we're going to dive into, which is the latest release by OpenAI. OpenAI released a new version of GPT-40 this week. Now, it does two things remarkably well, and I'm going to dive into one of them and then mention the other. The thing that it does incredibly well is that includes the capability to create images that is replacing DALI. So you heard me say many times before, DALI in the past year was an embarrassment. All the other tools out there, whether it's open source like Flux or Me Journey or other professional tools, were running circles around it.
11:50And DALI was, eh, it could generate images, but none of them were great. This new release that is just baked into GPT-4.0, there is no tool. You just use it within the regular chat GPT. Just choose GPT 4.0 in the dropdown menu and you can create images. It's nothing short of magical and incredible. I've tested it across dozens of business and fun use cases in the past three days since it came out. And it's blowing my mind every single time. Anything from its ability to understand context, from its ability to follow instructions, to keeping consistency in characters and objects and rooms and views from its ability to combine things it understands from the chat and apply it into the images, its ability to iterate on images while keeping consistency, its ability to understand style guides and so on, literally everything I threw at it, text very accurately, including the capability to generate entire infographics from scratch, just based on whatever data or topic you wanted to create the infographics 64.
12:49It just does things that were not possible before. I tried combining several different images into one image and it does that very, very well. It creates text on objects accurately. So it understands the volumetric nature of things. It can change the point of view. I took an existing image from the internet and asked it to create a God's eye view of that room. And it did with the sofa and the cushions and the pictures on the wall and the carpet and the design of all the different things. Now, it wasn't 100 % accurate, but it was good enough to fool a person or be one of those tests where you need to find the differences between two different images.
13:22It's really good across a huge variety of capabilities that were not possible before unless you were a professional designer. Now, this broke the internet because literally everybody went in and started generating images. Most people started creating images of themselves, their family, and their co-workers in Ghibli-style anime, which is really cute, but not very helpful from a business perspective, but all the other things that I said, and there were a lot of people who tried like me to do actual business use cases. And that drove Sam Altman to tweet the following, it is super fun seeing people love images in ChatGPT, but our GPUs are melting.
13:59So creating these images is a high compute demand. And the fact that more or less every ChatGPT user went in and started creating images drove the demand to a very, very high limit from the first time I tried it, which was shortly after they announced it, it was a lot faster generating images. Now they're generating relatively slow. Sometimes it gets stuck in the middle. Sometimes you have to refresh the page and so on. So it's bringing ChatGPT to a halt just because of demand. I think this will die soon in a week or two or three weeks after everybody's exhausted their need to just play with it and we'll start using it for actual business use cases.
14:36But in the meanwhile, they stopped access to free users. So initially it was available to free users as well. Now you have to be a paid member, either the$20 a month or the$200 a month. But they're saying that they will provide it to free users eventually. Now, in addition, it allows you to create images that are questionable. And now I'm going to quote Sam's tweet on X. This represents a new high watermark for us in allowing creative freedom. People are going to create some really amazing stuff and some stuff that may offend people. What we'd like to aim for is a tool that doesn't create offensive stuff unless you want it to, in which case, within the reason, it does.
15:14Now I'm jumping a little bit forward. We think respecting the very wide bound of society we'll eventually choose to set for AI is the right thing to do, and increasingly important as we get closer to AGI. Thanks in advance for understanding as we work through this. So what does that tell us? It tells us that Sam Altman is leaning more towards his nemesis, Elon Musk's approach with Grok. If you think about why people really like Grok, and I started using Grok all the time, is that it's more edgy and punchy and cares less than the average AI about what is politically correct. And this move by OpenAI is moving in that direction.
15:50It's basically saying, we're not going to be the people who decide the red line in the sand, what's acceptable and what's not and what people can and cannot generate with this, at least in most cases. So you will be able to create more or less anything you want with this image generator. And it will be very interesting to see. I haven't seen that yet blowing up on the internet, but I'm sure that it will because it will now allow to do things that other closed source models will not allow you to do. Now, they're also asking for feedback on that. So it will be very interesting to see how this evolves and how they correct.
16:23Now, in addition to this amazing image generation capability, and I'm going to record an entire episode about this together with a similar capability that was released by Gemini on their experimental platform on Google's AI Studio, which they actually released before ChatGPT that has similar but not as advanced capabilities. So I'm going to do a complete episode on a Tuesday, diving into how to use these tools and what you can do with them. In addition to that capability, it is a very powerful coding tool. So artificial analysts, that is an independent analysis group that measures AI models to choose the best models on API providers for different use cases.
17:01So they just evaluate AI APIs on different tasks. They're saying, and I'm quoting from their tweet, today's GPT-4.0 update is actually big. It leapfrogs Cloud 3.7 Sonnet non-reasoning and Gemini 2.0 flash in our intelligence index and is now the leading non-reasoning model for coding. This makes GPT-4.0 the second highest scoring non-reasoning model coming just behind DeepSeek version 303.24 released earlier this week. So what does that mean? It means that in addition to all the stuff that GPT-4.0 did before, it now is one of the top coders in the world, second only to, surprise, surprise, the latest Chinese version from DeepSeek, but it passed the leading tools before that.
17:45It is also now ranked number two on the LM Chatbot Arena leaderboard that we talked about many times in the past that is based on votes of people not knowing which tool they are using. And it surpassed GPT 4.5 and many other models that were just recently released by both Chachapiti and other companies. And it's now second only to, well, Google's latest release, which is Gemini 2.5 Pro that was launched on the same day. So if this doesn't make you feel like everything is accelerating, I don't know what will. So think about it. In the same week, we have a new model from Quinn, a new model from DeepSeek, a new model from Gemini, and a new model from Chachapiti.
18:26All of them are dramatically better than the models before, which is shuffling who's on top of the leaderboard, but all of them are ahead of the models we had just a week ago. Now, beyond the acceleration, this new image generation capability is dramatic change from everything we've seen before. From my perspective, it's a new chat GPT moment. And the reason I'm saying that is it's now good enough and actually a lot better than good enough to replace many SaaS types of software that were designed for image generation and editing and replace a lot of people who are now taking a lot of time doing this work.
19:02So the implication of this is not just a cool image generation capability. It is the fact that it can do a lot more professional work than could be done before with these tools, all baked into ChatGPT as a feature. Now, I think the next evolution of this, and it's practically already there, just not without the right user interface, the tool itself identifies every aspect of the image. I did some incredible things in the past few days. So it knows every component of what it sees because it understands the images, it understands the depth, it understands what you're asking from it. And you can make small nuance changes to the image and get the same image again, just with small changes.
19:39To tell you how crazy this is, I decided to test a campaign for a sunscreen lotion. I took a low quality image of the sunscreen lotion from the internet, asked it to create a female hand that's holding it with the background of the beach. It did that. The product looks perfect. The hand looks perfect. The beach looks perfect. It looks like a professional thing. I asked it to extend the image, I asked it to add text to it. I asked it to paint the nails of the woman in the colors of the US flag and turn this into a 4th of July campaign. And it did. So it knows how to grab fingernails and paint them in the US flag.
20:12Like the nuances I was able to get to across multiple use cases, this is just one example, is insane. So I think the next evolution will be that these tools will allow us to also visually pick specific components. So think about doing what you can do in Photoshop, but literally just by asking for it. I want you to allow me to grab this part of the text. I want you to allow me to manipulate this component and make it brighter, make it darker, change the direction of it, and so on, just by grabbing components out of an image and being able to move them around, manipulate them in any way we want, either through user interface or through words, literally just by talking to it.
20:48And that will be the end of professional design tools and definitely tools like Canva. And I'll be really surprised if that doesn't happen very, very shortly. Now there's the issue of it will stop you every now and then from doing things you want because it's not aligned with its policy. Well, as we've seen before, every time one of these tools comes out, the open source universe catches up. So I'll be really surprised if within a few months from now, we don't have this capability on open source, which means you can do whatever you want. We talk a lot in this podcast about the importance of AI training.
21:23Multiple research from leading companies has shown that this is the number one factor in success of AI deployment in businesses, large and small. I'm excited to announce that we just opened the registration for our spring cohort of the AI Business Transformation course. I've been teaching this course for two years, starting in April of 2023, and hundreds or maybe thousands of business leaders has went through the course. We had people in the recent cohort that ended in February from India, the Emirates, several different countries in Europe, South Africa, many places in the US, Canada, and even Hawaii.
21:56So regardless of where you are in the world, this could be a great opportunity for you. In previous courses, we had people as far as Australia and New Zealand. So weird hours of the day, but still getting a lot of value from this course. The course is four sessions of two hours each spread over four weeks on Mondays, noon Eastern time, starting on May 12th. If you are looking for ways to accelerate your personal knowledge and career or to change the trajectory of your team or your entire business, this is the right course for you. It is really a game changer. And within four weeks and only eight hours with some homework, you will dramatically change your understanding of how to use AI in a business.
22:34We give multiple hands-on examples and use cases and teach you the tools and the processes on how to use them. And we end up with a detailed blueprint on how to actually implement AI successfully business-wide. So if this is interesting to you, go and check the link in the show notes. You can open your phone and click on it right now and go and check all the information about the course. And because you are a listener of this podcast, you can get a hundred hours off of the price of the course with promo code leveraging AI100. I would love to see you join our course in May. And now back to the episode.
23:07And now we're going to change to rapid fire items. There's a lot to talk about. And the first one relates to everything that I'm saying about the world and how it's going to look like. Bill Gates, Microsoft co-founder, was interviewed on the NBC Tonight show, and he stated that within 10 years, AI will render humans unnecessary for most things. He calls it the era of free intelligence. And a few specific examples that Gates shared with Jimmy Fallon is things like that top doctors and teachers will become commonplace. Basically, AI will deliver great medical advice and great tutoring at scale, potentially slashing the costs and boosting access of these really important capabilities.
23:47But what does that mean to human doctors and human tutors? That's not clear to anybody. And the interesting thing that he's painting a future where AI tutors, AI doctors, and if you relate to that and broaden that AI everything will be as normal as smartphones. So if you think about 20 years ago, nobody had smartphones. It was a concept we couldn't even grasp. And now everybody has smartphones and we probably cannot imagine our day-to-day without them. Well, he's stating that within 10 years, AI will be that. But while we're just talking about these examples, if you generalize, this means manufacturing, logistics, farming, managing, writing code, literally everything humans can do, AI will be able to do.
24:28And the only thing Gates admits that will probably survive is things like sports, so basketball players, baseball players, and people who do theater and things like that, which people would still want to watch, but more or less everything we know as a professional job, not in sports or arts, will be done by AI, probably a lot more broadly than by humans. Now to a few quick news from OpenAI. On March 24th, OpenAI revealed a major leadership reshuffle, so Sam Altman will pivot to focus more on technical research and product development, while Brad Lightcup, who is their COO, will step up to run the day-to-day operations.
25:03In addition, Mark Chen rises to chief research officer to steer the scientific breakthroughs, and Julia Villagra takes chief people officer position to manage all their talent. So a lot of shuffle in the top leadership in OpenAI. That obviously follows departures of a few big names in the past 12 months, people like Mira Morati, who was the CTO, and Ilya Suskover, who was one of the leading or maybe the leading scientist and many others. So new shuffle in OpenAI. And it makes sense. The company has grown dramatically in the past 12 months. They have different needs right now. They have 400 million clients that are active all the time.
Read the full transcript
25:41They're delivering new products and they're in the process of converting from a nonprofit to a for-profit. So it makes sense that it's going to be reshuffling. It will be interesting to watch to see what impact does it have on the performance of the company. Another big piece of news from OpenAI and to the AI world in general is that OpenAI is adopting Anthropic's model context protocol called MCP, which is an open source standard that connects AI models and AI agents to data and tools. The protocol was released by Anthropic a few months ago and has become extremely popular with developers of agents.
26:17And now OpenAI is basically saying, this is a great model. We love it. And we're going to join and use the protocol as well. So as of now, it's already available as part of OpenAI's agent SDK, the desktop app, and the API support coming soon, per Sam Altman's tweet, which said, people love MCP and we're excited to add support across our products. Anthropic chief product officer, Mike Prager, obviously was very happy. And he said, excited to see that MCP love spread to OpenAI, welcome. What does that mean? It means that MCP's approach is successful. It is open source. It is open to everybody and it allows to create a unifying environment and solution for agent development.
26:54I think the more standardization we're going to see across this industry, the better because it will allow for collaboration. It will allow for companies to choose, switch and replace the agents that they're using because the underlying infrastructure and connectors between them are going to be standardized. So overall, I'm myself very excited about this as well. OpenAI are not the first company block Apollo, Replit, Codium, Sourcegraph, and other companies has already adopted MCP. So it seems that it's going to be the infrastructure to allow agents to connect with data, connect with themselves, and connect with tools moving forward.
27:28Now on a different topic from OpenAI, it seems that they are talking to several different providers to purchase billions of dollars worth of data storage, hardware, and software aiming to build its first ever data center for themselves. That's per the information article from March 26. Now, they're planning based on this article to purchase five exabytes of storage. That means absolutely nothing to me and probably to you, but this rivals Apple's iCloud capacity just a few years back. So this is a huge amount of data storage that OpenAI wants to purchase and own. And this will obviously work in tandem with their Stargate project, so their GPU project that they are working together with SoftBank, that is a$500 billion investment over the next few years to build data centers.
28:17So these two will probably work hand in hand. On a different topic, OpenAI, in collaboration with MIT Media Lab, has released a research exploring ChatGPT's effect on emotional well-being of its users based on its huge amount of, as I mentioned, 400 million people using it every single week. What they found is only a small fraction of users were emotionally connected to ChatGPT, despite the very large volume. Now, they've done two studies in parallel that were measuring different things and different emotional connection and impact on people. But what they did find is that users who trusted and bonded with Chachipiti were more prone to loneliness and dependence, hinting that there is a way to develop emotional dependency on these tools, which is something we all need to be aware of, especially those of us who have kids who are already potentially addicted to social media.
29:10this might be a lot worse if it spreads. And it's, again, something that we need to be aware of and at least educate people about. Now, for those of you who are enjoying deep research on ChatUPT, I'm definitely one of those people. As if you are a regular paid users who pay 20 bucks a month, you only have 10 searches in deep research a month. And if you're paying the$200 a month, you're limited to 120 deep research queries, but it was impossible to know where you are in the account that was solved. Right now, it creates a pop-up that tells you how many you have left every time you run it. But something even cooler, if you hover your cursor above the deep research button, it will tell you where you are on your current account and how many you have until the reset date and what the reset date is.
29:53So this is actually very helpful. Just hover your mouse over it and you will know how many deep research queries you have left. On a not so exciting news for Open AI. On March 23rd, Kai-Fu Lee, who's the ex-Google China head and the founder of ZeroOne.ai, which is a really successful AI startup, he's claiming that they're ditching their proprietary models, mostly ChatGPT, to adopt DeepSeek's open source tech. He's claiming that running DeepSeek's free and efficient models is costing him 2 % of what it cost him to run the OpenAI APIs in the backend to perform the same tasks that his company is performing.
30:30This is a huge push towards the really advanced capabilities of open source AI. And I think this will continue pushing the pricing of AI down, going back to what Bill Gates dubbed the age of free intelligence. Now we're going to shift gears and talk about contradicting signs of where the AI copyright battle is going. So the first one will stay with open AI. So on March 27, 2025, a federal judge in New York ruled that the New York Times and other newspapers can move forward with their copyright lawsuit against OpenAI and Microsoft. Those of you remember, there's multiple lawsuits against OpenAI and other AI-leading labs on using copyrighted material to train their chatbots.
31:11They claim that this is fair use, and obviously the owners of the copyrighted material think otherwise. And so the judge allows you to move forward, which may go to a jury trial. Now, as you know, a jury trial can go either way because the jury will decide what's going to be the fate of this lawsuit. That being said, that will end up most likely the Supreme Court either way. So whatever the jury decides will stand for a very short while, but there will be questioned in the Supreme Court and whatever they decide will be the way this moves forward. In the current political environment, it's pushing for AI innovation regardless of the costs and the consequences.
31:46And so I think in the current political system and with the way the Supreme Court is leaning since the previous Trump administration. I think the outcome will be in favor of AI innovation, but I might be wrong. But to tell you how confusing this is, Anthropic on March 25th actually got a really big win. And in a case in California, a federal judge rejected a bid by the music giants like Universal Music Group and others to stop the AI firm from using copyrighted song lyrics to train its chatbot Claude. So the judges were calling the claims vague and noting that they couldn't prove irreparable harm to their business in the lawsuit.
32:24Now, the interesting thing here is you have two states that are on the same side of the map, both are hardcore democratic states, and yet we found two opposite rulings by a court, which aligns with the call from OpenAI last week to the federal government to step in and define rules and regulations that will be on the federal level instead of allowing a state-level patchwork that is already in the making across multiple states and multiple regulations. How that evolves, I don't know, but it will be very interesting to track and I will keep you posted. Now, staying on Anthropic, they released a, I don't know if final, but a new update on their ability to look into the brain of Claude and see how it works, introducing new tools and methodologies on how to, if you want, read the mind of how AI works.
33:10So historically, AI was a black box. We know how it's getting trained and we know the outputs, but it was very, very hard to track how and why it's actually doing what it's doing. And Anthropic has been at work at trying to understand how the AI brain works. They now came up with a tool called Microscope that allows them to trace Claude internal steps and how it thinks and solves different problems. It's worth reading their release. It's written in non-technical terms, and it actually reveals that Claude thinks in very universal concepts before translating it into specific language. So it may think and analyze a question in several different languages like English, French, and Chinese to a simple question like what is the opposite of small, and only then it comes up with the language to describe it to the user.
33:55They also found that when given a simple math question, it actually computes and goes step by step, but when it gets a very complex mathematical question, it tries to wing it and if you want, bullshit the user with an answer without actually trying to solve the problem. The importance of this research, it goes way beyond that because the goal is the ability to see what AI is doing as it's doing it, to be able to prevent it from doing things we don't want it to do, either in the current scenario, but way more important when we hit AGI or ASI and beyond, where AI will be significantly smarter than us.
34:25Now, is that really doable at that point? I don't know, but I think it's a very, very important step by Anthropic in the right direction. And I really hope that these tools will also be able to run on other AI models and that other AI companies, such as OpenAI and open source companies, will adopt these tools in a similar way to the way they adopted MCP to allow us to see how these models work and stop them before they do something that we do not want them to do. Still, Anthropic joins forces with Databricks. Those of you who don't know, Databricks is one of the more successful enterprise data management platforms.
34:58and this collaboration will allow the 10 ,000 corporate customers of Databricks to run Claude's models, including the latest Claude 3.7, to get insights from their data. This is obviously the big promise of AI when it comes to corporations and then eventually smaller companies as well, is to be able to look across the entire company data and get insights that so far were either very hard or on the verge of impossible to get. And that will become obvious to anyone just by asking the questions or even just the AI by itself, suggesting insights based on what it's seeing. Anthropic also released version two of their economic index, which is their view on how people are using Claude for which tasks we shared with you when they shared the first one a few months ago.
35:42This one now has Claude 3.7 Sonnet in it that wasn't available when they released the previous survey. And it's very, very clear that Claude is the king of coding because it's now taking 37 % of usage of Claude 3.7 Sonnet. If you go to the review, you'll be able to see bar chart that shows the different tasks that people are using Claude for, and the first 10 or so tasks are computer related. Beyond that, the next things are education, science, and healthcare, but these are very far behind in volume compared to coding capabilities. I use Claude 3.7 for coding small things, and it works very well for me.
36:18I'm not a coder, so it's very hard for me to compare to other tools, but I did find the most amount of success building small applications with Claude. Now, from Anthropic to Google, as I mentioned earlier, Google launched Gemini 2.5 Pro Experimental, which is what they're claiming is their most intelligent AI yet, with built-in reasoning, and it jumped to the top of the LM arena, as I mentioned earlier. Now, if you remember a few weeks ago, we talked about humanity's last test, which is trying to come up with the hardest questions on earth to test AI's capabilities. So previously, OpenAI O3 Mini scored 14 % and DeepSeek R1 scored 8.6%.
36:57Well, Gemini 2.5 Pro Experimental now scores almost 20 % with 18.8. But beyond that, as I mentioned, it's on day-to-day tasks that people are evaluating it. It's scoring better than any other tool on the planet as of right now. It also has the same huge context window as the other Gemini tools. So right now, 1 million token context window soon to be switched to a 2 million context window, and it handles text, images, audio, video, and code, making it probably the most versatile tool out there. It is still the only tool that knows how to quote unquote, watch video and understand what's happening in the actual frames, not just listen to the audio track and can understand what people are saying.
37:36So if you need to analyze video, this is a very powerful capability that did not exist before this tool was released. In the demo by Demis Asabis, he showed significant improvement in coding, and they showed a few demos of creating mini web apps and interactive games with just a few prompts. It is currently available at the Google AI Studio, and if you're a paying user of Gemini, of the$20 a month, basically Gemini Advance, you have access to it as well. Staying on the topic of new models, as I mentioned, DeepSeek released a new version of its V3 model. This is dubbed V3-0324 for the date it was released.
38:12It has an MIT open license, so it's an open source model that you can take and use. And it has a significant upgrade compared to its predecessor, mostly in reasoning and code generation and user intent understanding. So it scores very high on all these things. As I mentioned, it's right now on several different benchmarks, the best AI coder out there, period. What does that mean? It means that there's two races that are very active and aggressive. One is the race between China and the US. The other is the race between open source and closed sourced AI. And it is very clear right now that both these races are very close and are going to be very close moving forward with both open source and China pushing models that are as good and in some cases better than the leading labs in the US.
38:57And if that's not enough from China, Alibaba just released QEN 2.5 Omni 7B, which is a 7 billion parameter multimodal AI model that is designed for end-to-end processing of text, images, audio, and video. Again, something that will compete with the US models across the board in a real true multimodal AI solution. Now, despite the fact it is a relatively small model with only 7 billion parameters, it rivals many other larger single modality models and multimodal models as well from the leading labs around the world. And like its previous models, this is also an open source model that everybody can go and use and it's available on both their website as well as other solutions like Hugging Face.
39:39Continuing on this path of this insane week as far as releases, Mistral just launched Mistral 3.1. 3.1 has 24 billion parameters and it is scoring better than some of the other open source models like Google Gamma 3 and even GPT 4.0 mini across several different benchmarks. It supports 21 languages, including European and East Asian languages. It's still lacking on the Middle Eastern side of capabilities when it comes to languages, and it has 128 ,000 tokens context window. And it's an open source under Apache 2 license, which means it's free to tweak, deploy, and allows developers to do basically whatever they want with it.
40:19So what does that mean? It means there are more and more capabilities that are available as open source with very powerful tools coming from everywhere around the world for free or almost free. And they're going to keep on pushing the pricing of AI down while the capabilities are going to keep on moving upwards. Staying on the open source topic, on March 19th, Hugging Face submitted a open source AI blueprint to the White House AI action plan. And what they're claiming is that open source models can outperform closed source models. One of the examples they gave is Olympic Coder, which is a 7 billion parameter coding parameter, outcodes Claude 3.7, which was considered the top-of-the-line coder until that point, again, now was taken by the new DeepSeq model.
41:02Now, Hugging Face is definitely a powerhouse when it comes to open source. They're hosting 1.5 million public models that are available for everybody to use and co-develop, and they're claiming that these models can compete with a closed-sourced one at a fraction of the cost. The strategy that they're suggesting to the White House is a combination of three different components, collaborative innovation, resource efficient models for smaller players, and transparent security. And basically what they're saying is they're saying that instead of following OpenAI's suggestion to the White House to be a light touch on regulation, they're claiming the open source world will provide a much stronger and safer foundation because everybody can see what is actually going on with these tools.
41:46Now, I must admit, I'm not sure which side of the fence I'm on. I don't think I know enough to make a fair judgment call. I definitely see the logic in arguments of both sides. One that's claiming that giving it to anyone to change and manipulate is scary. But on the other hand, I'm also thinking that not knowing what OpenAI and Cloud and so on are doing, and definitely similar companies from China and other places around the world, versus allowing everybody to see what's happening in the code and in the systems and in the weights and so on also makes sense to me. I don't think this is going to get resolved.
42:19I think we're going to have both universes running in parallel, just like we had with a lot of other tools like operating systems. Now, how will that impact the Trump administration team that pushes innovation and very little regulation? It will be interesting to follow that as well. Now, a few quick updates from the robotics world. So a new company from California called Dexterity just introduced a new robot called METCH. And it's an interesting combination of old school robotics and humanoid robotics. So it's a moving platform versus a humanoid robot, but it has arms that are mimicking human arms just with more capabilities.
42:53And these arms can lift 65 pounds each. So it's making it the robot that now can lift the most amount of thing and move around a factory. Their first use case is to manage packages on different supply chain lines and put it on trucks or other conveyor beds and so on. And it's very efficient at doing that in the demo that they were showing. It can lift high weights while still doing it very accurately. Now, the interesting thing is they're claiming that one worker can oversee up to 10 Mechs robots at the same time, which means from a labor perspective, it's not just replacing the human workers who are moving the packages around, but also the amount of managers required to do that.
43:34Their initial use case, as I mentioned, is truckloading, but they're definitely looking to scale for solution in the general logistics world. Staying on the US side of robotics, on March 25th, Figure, which is one of the leading humanoid robotics companies in the world, have released a video showing that their robot Figure 2 can walk like humans. So it doesn't have that robotics walk anymore. And the way they achieve that is through a physics simulator that allows AI to get reinforcement learning of years within just a few hours just by running its brain, if you want, in the sim. And now it is working a lot more like a human with heel strides, toe-offs, and arm swinging, which makes it look very human.
44:18The interesting thing here is not the fact that it works more naturally or what seems more naturally to us, but the fact that you can train these robots in a simulated environment to do tasks that previously took months and years to do, and now it can do it in a very short amount of time. And that means that they will be able to mimic literally any activity that we want and learn it very quickly. Moving from the US to China, so we mentioned many times before, China has some of the most advanced robots in the world. So Agibot, which is one of the Chinese successful startups, are planning to produce 3 ,000 to 5 ,000 humanoid robots in 2025.
44:51So this year, this is a jump from only 1 ,000 units in 2024, so a 3 to 5x increase. They are releasing three different models, Adjibot Link C2, Genie Operator 01, and Yangzang A2. The Yangzang A2 is very interesting because it has very delicate capabilities and precision. And one of the things they demoed is the ability to thread a needle by the robot, which requires obviously very gentle dexterity and understanding what's happening in its surrounding. So different robots for different scales, and they're scaling up their production. And the quote that they released is, we aim to deploy new products in industrial scenarios this year, replacing humans in specific tasks to deliver tangible customer value.
45:34A, I agree. It will provide more value because the products that they handle will be able to be cheaper, at least from a supply chain and operation perspective. But the question again is what will happen to the people who held these positions before? And if they don't get paid and don't have money, who will buy the products that the robots will pack and ship and do everything with. I don't think anybody has answers to that, but we're going in that direction faster and faster. And another interesting robotics thing that happened in China this week, Chinese influencer Zhang Zhenyuan rented a Unitree G1, which is their smaller robot that we talked about a lot in the past few weeks because it learned how to do ninja moves and kicks and do different cool things.
46:12Well, you can now rent them in China for about$1 ,400 a day. And he rented it and documented the whole thing and showed the world how it was doing kicking, cleaning, and even attempting dancing. It was cooking eggs, sweeping floors, and even running alongside Zeng in the park. It couldn't do the dance move, but it still was very funny to watch. So you can go and watch that video online. But the whole concept that you can now rent these robots because people bought them because their price is now$14 ,000. You can buy a robot for$40 ,000 and it created a whole new market of robot rentals. So think about you want to throw a party or you need to clean your garage or you need to do any other big project.
46:49And And instead of buying the robot, you can rent it for the day to help you with that particular task. Now, going back to what I said before, until I know for sure that these robots won't destroy my house or more importantly, hurt my kids, there is no way any one of these is coming into my house. But as soon as these safety measures are in place, for a$14 ,000 amount, being able to do all the chores in the house and around the house, plus some other things that we can't even think of, it's a no-brainer, even at today's cost, which will dramatically drop as these things scale up their production.
47:20Switching gears to a different topic, Rev, I'm not sure how to pronounce it, I must admit, is a new company in the field of image generation. They just released their first image generator, codenamed Halfmoon, and it is running in a different business model than most of the other models. You're just paying for tokens and it's one cent per image. So you're paying$5 to generate 500 images, which is cheaper than most of the tools out there today.
47:49So think about a monthly plan in Me Journey starts at$8 a month. Here you can pay $5 and until you run out of credits, you can keep on using it. Most people probably are not going to generate 500 images per month, which is actually really good. In head-to-head tests by independent evaluators, it actually was as good and better than Me Journey Flux and Ideogram across multiple types and styles of image generation. So going back to what we started with ChatGPT and Gemini generating amazing images as part of their current payment plan, and now with other quote-unquote professional tools that do it significantly cheaper, this will also drive the cost down of image generation across the board between open source and closed source models.
48:31Now the cool thing is you can try it for free. So you get 100 initial free credits and then 20 daily freebies. That should be enough for most users unless you're a professional image generator. I said that about a year and a half ago. I do not see a future in which image databases survive. There's absolutely no reason for them to exist anymore because you can generate any image you want in any style that fits exactly what you need in seconds. Why would you ever spend time searching an image database for an output? So the whole mess that came out last week with Google Gemini that can eliminate watermarks from existing images on platforms like Getty Images is completely irrelevant from my perspective because I haven't downloaded or searched an image on the internet for the last year and a half.
49:16And that was before the tools were as good as they are right now. That's another industry that is going to be probably wiped off the face of the earth as more and more people understand how to use AI image generators. And the fact that they're now built into the actual chatbots is going to make that a lot more widely available. Staying on the creator's universe on March 18th, the US Federal Appeal Court declared AI-generated music uncopyrightable, reinforcing its initial stating of that situation and aligning it with the Copyright Act that covers music from 1976. And it aligns with everything else that the Copyright Office has said.
49:50If humans are not heavily involved in creating something, it's not copyrightable. I think that's actually a very good thing as a musician myself and somebody who loves to play music, these tools can generate hundreds of thousands of new songs every single day. Even if very few of them become successful, it will wipe out any human generation of music very, very quickly as we will drown with music generated by AI and then musicians will not be able to make any living. Now you're asking, what does one have to do with the other? People can still create music and people will love it. Yes, but there's no way to monetize it.
50:22So if there's no way to monetize it, there's no reason for a lot of people to play it. And then there's no reason for this industry to exist, which means the human side of things will still be able to make money, maybe challenged a little bit, but not wiped out by AI-generated music. Staying on the topic of creation of things with AI, two very interesting releases when it comes to voice models. The first one is Maya, which is a model that was released by a company called Sesame. Maya is a very punchy, human-like, edgy, if you want, voice assistant, and it has blown up in the internet with really cool, crazy examples of how human-like and cheerful it sounds compared to the relatively dull other voice models that sound too robotic.
51:04They have released their model as open source for anybody to use. So you can go to Hugging Face or GitHub and get the code for Maya and run it as your own and get a very human-like voice used in any application that you want. In addition, on March 26, Grok with a Q, so Grok, the AI infrastructure inference company, has released that they are teaming up with Play AI to release what they call Dialog, which is a text-to-speech model that runs 10 times faster than real-time speech, meaning its ability to synthesize and create speech is significantly faster than the speech itself. They're claiming that it's very, very flexible, and it uses full conversational history to nail rhythm, tone, emotion, and make voices sound very natural.
51:46So it's currently available on Grok's high-speed inference platform, making very little mistakes, sounding very real, And it has English and Arabic as the languages that run on it. So far, the Middle Eastern languages, Arabic and Hebrew, were usually not available on any of these voice platforms. And Arabic is the fourth most spoken language on the planet. So it's obviously a big deal. So if you are an Arabic user or you have Arabic clients and you want to use an AI voice agent, you can now do this on this new platform. What does that mean? It means that we will see more and more voice agents working across multiple aspects of industries.
52:20I assume the customer service world is going to be the first one taken by a storm, but later on, it might be salespeople and a lot of other customer or internal facing capabilities. You just have conversations on any topic. Think about everything that we talked about before with agents becoming a part of the workforce. You'll be able to talk in natural language with an AI agent that is not human and not know the difference between that and a human employee on a conference call. Does that make sense from a productivity perspective? Yes. Does that connect very well to what we started with as the research from Harvard and Ethan Mollick?
52:52Yes. Is this going to be a very weird future? Absolutely. Now, we're talking a lot about moving forward to AGI and where we are in that process with all the different advanced capabilities that we're seeing almost on a daily basis right now. We talked about the ARK Foundation previously, where ARK AGI was the benchmark that was blasted by the leading models right now. Well, they just came out with Arc AGI 2, and they came out with it on March 24th. And the goal is to push the limits of what AI can do from a reasoning perspective. So similar to Arc AGI 1, it is a visual pattern puzzles that are geared to challenge AI in solving them.
53:31And it's very interesting to see the results with over 400 humans averaging 60 % success rate on this test, while DeepSix R1 varies between 1 % to 1.3 % accuracy. And GPT 4.5 and CLOD 3.7 Sonnet are around 1 % as well. So significantly far behind humans, showing that throwing more compute and even reasoning capabilities has its limits, at least right now, compared to the human brain's agility and understanding of specific problem solving. The ARC is also offering a prize called the ARC Prize 2025, and the prize is going to go to the first company or organization that's going to hit 85 % accuracy on this test, but only for 42 cents per task.
54:13So what they're claiming is that it's not just about being able to solve it, but also being able to solve it efficiently, which will actually make sense from a financial investment perspective. The last piece of news that we are going to talk about is actually an interesting trivia. So on March 20th, the Computer History Museum, together with Google, has dropped the original 2012 AlexNet source code on GitHub. This was the first neural network that actually worked because it took the error rates from 25 % to about 15%. And it was basically what started the current AI era that we know. So this is very, very cool.
54:53If you want to run the very first neural network that worked properly, that is actually really small and can run on a single computer with the right GPU, such as GTX 580, and you can run it in your house. If you're really geeky and you really want to play with it, it is now available. That's it for today. We'll be back on Tuesday showing you exactly how to create amazing videos, including all the different steps between research of what you need in the video to scripting the video, to creating the video, to creating the audio of the people in the video, to actually rendering the video and editing the video, all with AI, all within less than an hour.
55:29So if you need to create videos, and most businesses do, this is going to be a fascinating episode. We're also opening a spring cohort of our AI Business Transformation course. So if you have not had AI efficient and effective business-related training and you're looking for one, check out the course in the link in the show notes, share this podcast with other people who can benefit from it, and have an amazing rest of the day.
From the publisher
Is your business ready for an AI teammate who works faster, costs less, and never takes a sick day?
This week, the AI landscape didn’t just evolve—it exploded. From Harvard-backed research showing AI-powered individuals outperform entire teams, to OpenAI’s shockingly good new image generator, the future of business just took another quantum leap.
AI isn’t just an assistant anymore—it’s becoming the MVP. And the pace? AI capabilities are now doubling every 3–7 months. If you're a business leader and you’re not keeping up, you’re already behind.
In this tightly-packed, high-signal episode as I break down the most critical AI developments business leaders need to know—from real research, new capabilities, to the societal shifts no one’s prepared for.
In this AI news, you'll discover:
- How a solo employee using AI can outperform a 2-person team
- Why 37% productivity gains might be just the beginning
- AI's doubling speed — and what that means for your job, team, or company
- GPT-4o’s jaw-dropping new image generation capabilities (and what it kills off)
- Why Canva, Photoshop, and even designers should be paying attention
- Bill Gates’ prediction: AI doctors & tutors as common as smartphones
- The hidden emotional risks of using AI and how it’s impacting users’ mental health
- The rise of humanoid robots that lift 65 lbs, cook eggs, and might be your next warehouse crew
- Latest copyright battles (and contradictory rulings) shaking up AI law
- Deep dives into Claude 3, Gemini 2.5, DeepSeek, and what benchmarks actually matter
- Why you’ll soon have voice agents handling customer service, sales—and sounding eerily human
Registration is now open for the Spring 2025 AI Business Transformation Course, designed to help leaders like you implement AI company-wide with confidence.
🎓 Use code LEVERAGINGAI100 for $100 off https://multiplai.ai/ai-course/
About Leveraging AI
- The Ultimate AI Course for Business People: https://multiplai.ai/ai-course/
- YouTube Full Episodes: https://www.youtube.com/@Multiplai_AI/
- Connect with Isar Meitis: https://www.linkedin.com/in/isarmeitis/
- Join our Live Sessions, AI Hangouts and newsletter: https://services.multiplai.ai/events
If you’ve enjoyed or benefited from some of the insights of this episode, leave us a five-star review on your favorite podcast platform, and let us know what you learned, found helpful, or liked most about this show!



