285 | Meta is BACK - beats Gemini 3.1 pro, and ChatGPT 5.4, OpenAI Copies Anthropic's playbook, and Opus 4.7 is the new king - AI news of the week ending on April 17, 2026

18 Apr 2026 · 24 min · 14 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI news roundup (week ending Apr 17, 2026): Anthropic Claude Opus 4.7 upgrades for agent reliability, higher-resolution vision, and non-coding work; Meta launches MuseSpark, a top-ranked multimodal model; OpenAI releases GPT 5.4 Cyber via a trusted-access program; plus rising real-world cybersecurity incidents and dependency-chain risk.

Guest backgrounds

No guests mentioned; host is Isar Maitis (running AI workshops and teaching multi-agent orchestration).

Key claims

Opus 4.7 improves multi-step agent workflows +14% with 30% fewer tool errors; vision resolution up to 3.75MP with visual acuity 98.5%; legal/enterprise gains (Big Law 90.9%, Office QA Pro 21% fewer errors). Meta MuseSpark ranks #5 on Chatbot Arena, beats Gemini 3.1 Pro and GPT 5.4, uses 58M output tokens vs 157M/120M. OpenAI GPT 5.4 Cyber is restricted to vetted customers; cybersecurity risk is accelerating.

Notable examples

Next Web reportedly delegates complex tasks to Opus 4.7 unsupervised; host workshop participants build multi-agent business apps in 36 hours; Axie library (70M weekly downloads) hacked by North Korean actors, impacting OpenAI macOS apps—prompting urgent updates.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Workshop Insights from Zagreb

0:45 to 1:46

Isar shares his recent experience from an AI workshop in Croatia.

“If you're interested, you can sign up straight from the show notes.”

Anthropic's Claude Opus 4.7 Release

1:46 to 3:14

Discussion on the capabilities and improvements of Claude Opus 4.7.

“the tool errors by 30%, which is obviously very significant.”

Enhanced Image Analysis in Opus 4.7

3:14 to 4:40

Exploring the advancements in image resolution and analysis with Opus 4.7.

“To put things in perspective, in the visual acuity benchmark, the results of Opus 4.6 was 54.5%, And Opus 4.7 scores 98.5%, which A, means it basically perfectly, most likely much better than humans.”

Knowledge Work Gains Beyond Coding

4:40 to 5:58

Opus 4.7's improvements in legal and enterprise data tasks are discussed.

“So that obviously has a very serious impact on the real world applications that you can develop with it.”

Opus 4.7's Cost and Pricing Structure

5:58 to 7:40

Insights on the pricing and token usage of Anthropic's new model.

“But on the other hand, it's most likely going to consume more tokens, up to 35 % more tokens to achieve the same thing.”

Real-World Impact of Opus 4.7

7:40 to 9:27

Isar shares practical implications of Opus 4.7 based on workshop results.

“So the bottom line, we're getting a new, even more capable model from Cloud that can do even better agentic work that is connected directly to real-life use cases in companies.”

Meta's MuseSpark Model Announcement

10:51 to 12:18

Overview of Meta's new AI model, MuseSpark, and its capabilities.

“The second big and really interesting release of this week comes from Meta.”

Comparative Performance of MuseSpark

12:18 to 14:00

Discussion on MuseSpark's performance relative to other top AI models.

“and yet currently this model ranks number five on the chatbot arena and the only models ahead of it are Claude Opus 4.6 and Claude Opus 4.7 in two different variants each in their standard and the thinking variants.”

Meta's Shift to Proprietary AI Models

14:00 to 15:21

Explore Meta's decision to transition from open-source to proprietary AI models and its implications.

“the previous models released by Meta, this model, it's a proprietary only model, meaning they are not releasing it as an open source.”

OpenAI's New Cybersecurity Model

15:22 to 16:46

Learn about OpenAI's GPT 5.4 Cyber model and its approach to tackling software vulnerabilities.

“coming in the open source world from China right now with Alibaba and DeepSeek that is now controlling almost 50 % of the hugging face download as of the end of 2025.”
Show all 14 chapters

Growing Cybersecurity Risks

16:47 to 18:25

Understand the escalating cybersecurity risks associated with new AI models and their deployment.

“And they are releasing it to a select group of vetted customers through what they call a trusted access program that they established just a month and a half ago.”

Risks of Concentrated AI Power

18:26 to 19:51

Discuss the implications of concentrated power in AI and the need for equitable access.

“reasoning, call it whatever you want to call it, that they are the one that's going to keep everybody else safe.”

Cybersecurity Incident with OpenAI

19:52 to 22:16

Examine the recent cybersecurity incident involving OpenAI and the importance of security measures.

“whether it's PUD or Mythos or whatever code name or numbers they're going to get, Axie, which is a widely used third-party developer library with over 70 million weekly downloads, was hacked in the past few weeks.”

The Need for Collaboration in AI Safety

22:17 to 23:20

Reflect on the urgent need for collaboration among tech companies to ensure AI safety.

“now because the change is happening extremely fast and the impact it is going to have on businesses, society, security, et cetera, is going to be very, very dramatic.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Hello and welcome to a weekend news episode of the Leveraging AI podcast, the podcast that shares practical ethical ways to leverage AI to improve efficiency, grow your business and advance your career. This is Isar Maitis, your host, and I'm recording this episode from a hotel room in Zagreb, Croatia, after doing a two-day AI workshop for a company from this region of the world. It was absolutely amazing. I'm going to tell you more about it afterwards. But because I am traveling, and because it's really late at night, and because I have an early flight to catch, this is going to be a relatively short episode.

0:33And we're going to talk about three releases, one from Anthropic, one from OpenAI, and one from Meta. So there's definitely enough to cover even in this short episode. There's a lot of other news that you can find in our newsletter. If you're interested, you can sign up straight from the show notes. But let's get started. The first release that we're going to talk about comes from Anthropic. Anthropic just released Claude Opus 4.7, which, as expected, is its most capable model that they've ever released. And it's focused on, as you can expect, better capabilities in exactly the needs of creating agents and building business-related applications.

1:16And the biggest differences are in four specific areas from Opus 4.6. The first one is that they're saying that this model has significantly better capability in executing correctly over long periods of time when running agents. So to make this more specific, they are saying that this model delivers a 14 % improvement over Opus 4.6 on multi-step workflows while using fewer tokens in producing this outcome and while reducing the tool errors by 30%, which is obviously very significant. The people from the Next Web, the online news agency reported they are able to now hand off a much more complex tasks to Opus 4.7 and get significantly better results compared to Opus 4.5, all without supervision.

2:09Now, Opus 4.7 handles complex, long-running tasks without any issues with a lot of consistency, and it is paying very, very close attention to instructions. So it has better instructions following per Claude than any other model they ever created, which makes it probably than any other model ever. For somebody who has been developing complex, long-running agents and doing this every single day and teaching it in my workshops, I can tell you that this is one of the biggest unlocks. The longer you can let it run on its own unsupervised, the more you can do in parallel because I am the bottleneck with everything that I'm doing right now.

2:45And so will you become once you start creating a lot of agents in parallel and you want to remove yourself as much as possible from the process. And if the agent can now run for 20 minutes without making anything weird and following the instructions very, very carefully, that means you don't have to talk to it or watch it or tell it what to do next for 20 minutes. And that is a big deal. The second thing that Opus 4.6 has a big improvement in is better vision capabilities, specifically analyzing images. The maximum image resolution has increased from 1 ,568 pixels or 1.15 megapixel, which was the highest resolution that Claude Opus 4.6 could handle, to a 3.75 megapixel, which is more than 3x, the resolution that it can handle, while having a better capability to understand what's actually in the image and analyze it correctly.

3:39To put things in perspective, in the visual acuity benchmark, the results of Opus 4.6 was 54.5%, And Opus 4.7 scores 98.5%, which A, means it basically perfectly, most likely much better than humans. And B, it more or less doubles the results of Opus 4.6, which is actually pretty good as somebody who's using these capabilities all the time. Now, if you think about what that means in the real world, it means that the AI can understand basically any image you upload, whether it's scanned documents, invoices, extracted text, images of things, images of claims is in insurance. Literally anything that has information that can be either scanned or photographed, it can understand in a much higher level of detail and in a much higher resolution.

4:31So you win on both fronts. It has a better understanding of what it's looking at, and it can have more details because it can analyze bigger files with higher resolution. So that obviously has a very serious impact on the real world applications that you can develop with it. The third thing that ties directly to that is real knowledge work gains that is not coding. So part of this release had a very significant focus on things that are not coding related. As an example, on the big law benchmark, which is a legal benchmark, it now scores 90.9 % at the high effort setup. And this is in legal reasoning workflows.

5:12That is a huge jump from Opus 4.6 on another benchmark that is a data bricks benchmark that's called Office QA Pro. It generated 21 % less errors while doing different kinds of enterprise types of work. And it also scores very high on finance agent benchmarks as well. So if you think about the combination of legal finance and enterprise data analysis, that covers a very significant portion of the amount of mundane day-to-day work that happens across multiple organizations, especially larger enterprises. And so this model now provides a big improvement on all these things. Now, an interesting thing from a pricing perspective, from a pricing perspective, the price of Opus 4.7 is exactly the same per million tokens as Opus 4.6.

5:59So that's pretty good. But on the other hand, it's most likely going to consume more tokens, up to 35 % more tokens to achieve the same thing. However, what Anthropic is claiming is that because it is making fewer mistakes and you'll be able to get in one run the results that you want, it is probably going to cost you less money to run this model than the previous model because it will be able to do things with less iterations, but it's going to do it with a little more tokens and it's going to do it better. So you probably end up either a wash or a little better than before while getting a higher quality.

6:32Now on the API, Anthropic included some new capabilities in order to allow developers and vibe coders and people who who are developing agents like me to have better control over pricing if you're not using your subscription and you're using the API. So one of those functions is called X-High, which defines the effort level to fine-tune the balance between the intelligence level and the token spend. And they're also introducing a new concept called task budget that is allowed to help you control the budget better. Basically, it's an advisory token target for the entire loop that you can define in order to keep the cost in control when you're developing these kind of solutions, which I think is very interesting.

7:10I didn't get a chance to test it yet, but I'm very curious to test it compared to Opus 4.6 and see if it's actually saving money or costing more money at the end of the day. Now, obviously, because of everything we reported about last week with Mythos being a huge risk to cybersecurity, a lot of people ask themselves how Opus 4.7 relates to that, and the answer is it doesn't. So Opus 4.7 is not Mythos. It's a evolution of Opus 4.6. It's not a brand new model, and it does not generate the same level of exposure when it comes to cybersecurity risks. So the bottom line, we're getting a new, even more capable model from Cloud that can do even better agentic work that is connected directly to real-life use cases in companies.

7:56I must say that the current level of Cloud, even before the release of Opus 4.7, is absolutely incredible. I will share with you in two sentences what we did in the workshop. And I'm not going to get into the details of it, obviously, because it's a client of mine and I don't want to share the information. But what I can share with you is there were 20 something people in the workshop. And when I asked how many of you used Claude in the beginning of the workshop, a few hands went up, maybe four or five. When I asked how many of them used Claude Cowork before the workshop, two hands stayed up. And when I asked how many actually created scales and built automation around them, zero people kept their hands up.

8:33So people had no experience coming into this workshop. And 36 hours later, when we stepped out of the hackathon, which is how I end all my workshops, they had several groups, each and every one of them developing really incredible, sophisticated applications with multi-steps in the background across more or less every aspect of the business. So they were developing applications in operations and in customer service and in sales and in post-sale and in pricing optimization and in marketing. So almost every aspect of the business, they had a development of a application, including multiple agents in the back end that was in an advanced state of setup.

9:12So none of them has finished, obviously, developing these applications, but they will finish this application in the next week or two after knowing nothing about the cloud ecosystem and how to use it properly, which shows you how incredible this transformation can be. how you can go from knowing almost nothing about how to use Cloud or advanced AI capabilities to being able to develop advanced applications that connect to your ecosystem, your tech stack, and provide real solutions across multiple aspects of your business. I find it absolutely magical to see what people can create after such a short amount of time.

9:50And by the way, if you're interested in learning this kind of capability for yourself, this is exactly what we're going to teach in the multi-agent orchestration course that we have released. I mentioned that in previous weeks. We started selling the early bird of this course a few weeks ago, and we sold out one early bird session, which led us to open another one. That one got sold out very, very quickly. And now we're selling a non-early bird session that starts in June. So if you want to join that course, you better join quickly because I'm pretty sure this one will sell out as well. And between now and the end of April, you get$200 off.

10:24So if you follow the link from the show notes and you want to learn how to use Claude effectively for real work, including building multi-agent orchestration solutions and applications, come and join the course and do it quickly before it sells out. If you are in leadership position in an organization and you want your company to learn that in a dedicated custom workshop, reach out to me on LinkedIn or send me an email. Again, there's links to both of that in the show notes, and I will glad to tell you exactly what we're doing. And now back to the news. The second big and really interesting release of this week comes from Meta.

10:57So Meta just launched MuseSpark, which is the highly awaited new model from Meta. So if you remember, the last release from Meta was Lama 4, which was a huge disappointment that had a huge investment and actually came out really far behind the other competing models, which prompted Zuckerberg to basically overhaul the entire Meta solution when it comes to Meta AI and reshuffle the entire department and establish the Meta Super Intelligence Lab, also known as ASL. He brought in Alexander Webb to run this department and had a complete reorganization and reshuffle of leadership and the people involved.

11:36And for a very long time, they seemed to be in a complete mess. Nothing seemed to work. There were a lot of rumors of really bad things that are happening, including the quality of the models. Well, we finally got the first model. As I mentioned, it's called Muse Spark, and it's actually performing really, really well. So after all the rumors and all the disappointments and all the question marks that was flying around, where Meta is going and can they even catch up and can they be a relevant player in this race? Well, the answer now is pretty obvious. Yes. So despite the fact this is the first model this group has developed, and yes, they've been added for about nine months.

12:12But again, brand new group, lots of reshuffling, lots of changes in organizations and reorganizations and then reorganizations. and yet currently this model ranks number five on the chatbot arena and the only models ahead of it are Claude Opus 4.6 and Claude Opus 4.7 in two different variants each in their standard and the thinking variants. So this model is currently performing on average against because it's the chatbot arena and people are using it for multiple different use cases. It is ahead in the ranking of Gemini 3.1 Pro, of Grok 4.2, of GPT 5.4, so top-of-the-line models from three of the world's biggest labs, and Meta's model is currently ahead of these models in the chatbot arena.

12:56The model has really advanced multi-modal capabilities. It is scoring 86.4 on the Char-Cheve figure understanding, which is, again, comparing to Claude 4.6, that scored 16.5.3 and GPT 5.4 that scores 82.8. So it's ahead of both of these. It's also scoring 80.4 on the MMMU Pro Vision benchmark, second only to Gemini 3.1 Pro, which is at 83.9. Again, very, very close scores. So it is a very capable model. The other really interesting thing about this model is that it only required 58 million output tokens versus Cloud Opus 4.6, 157 million tokens and GPT 5.4, 120 million tokens. So again, 58 versus 157 versus 120 million tokens to deliver the same reasoning across several different benchmarks.

13:51That means that by definition, it's going to be significantly cheaper because it is using a lot less compute to deliver superior results. Now, different than all the previous models released by Meta, this model, it's a proprietary only model, meaning they are not releasing it as an open source. If you remember, one of the things that Meta insisted about in everything until now is that they're going to release the models as open source. There were two reasons for that. One is Yalla Kuhn, which from very early on was very much supportive of the open source universe and collaborating in that space in order to get more people involved.

14:25and now he's not there anymore. And the other reason is that was one of their ways to push out models that could risk or be uncomfortable for the proprietary-only models from Anthropic and OpenAI and Google because they give it for free because they don't care about the revenue that's coming from it because they don't need the revenue from it because they make the revenue somewhere else. So both these tactics led to the fact that they were releasing it as open source, but not anymore. As I mentioned, this model comes only as a close source model, which I think is a loss to the open source community just from the fact that the Lama family of models was downloaded 1.2 billion times and is being used in over a thousand commercial applications as of the beginning of this year.

15:10So from a footprint, it actually got a decent footprint around the world, but that is not the direction that they're going right now, at least not with this model. One of the reasons that might have pushed them away from the open source world is the really crazy competition coming in the open source world from China right now with Alibaba and DeepSeek that is now controlling almost 50 % of the hugging face download as of the end of 2025. And new models like GLM-5 and QEN 3.6 Plus are being some of the top models in the world. So that's a very different environment compared to when they started releasing the LAMA model.

15:45So the bottom line, we have another real contender in the race with a very powerful and capable model. It will probably take a few weeks before we're going to learn how well it actually works in different environments, including in agentic long-term capabilities. Based on the fact it's their first model that they're releasing, it is very impressive. And to summarize it, I will use a quote from Alexander Wang, the chief AI officer at Meta. And he said, nine months ago, we rebuilt our AI stack from scratch. New infrastructure, new architecture, new data pipelines. This is step one. Bigger models are already in development with plans to open source future versions.

16:23So again, they're just getting started and they released a model that is now number five on the ranking, just behind the two latest Opus models. So very promising from Meta. It will be interesting to see how that evolves and how that impacts the other players. Speaking of the other players, OpenAI released GPT 5.4 Cyber. So this is a specialized cybersecurity model that specializes in autonomously identifying software vulnerabilities so they can be patched and fixed. And they are releasing it to a select group of vetted customers through what they call a trusted access program that they established just a month and a half ago.

16:59If that sounds familiar, because it is exactly what Anthropic did with Mythos just a week ago. So they said they have a new model that can identify vulnerabilities and they think it's a very high risk for the world. And they had a short list of partners that is going to get access to it. And they gave access to only these people to try to patch all these vulnerabilities before they're releasing it to the world. What does that mean? It means that these companies are going to keep on doing very similar things to what the other company is doing. And they're going to keep on growing their models to capabilities similar to what the other companies are developing.

17:30and that the risk for the world from a cybersecurity perspective is growing exponentially as these models are getting better, despite the fact that they're releasing it early to different partners to try to fix that. I seriously doubt that in the amount of time from the moment they're going to give it to all these partners to the moment they can block every vulnerability in major software in the world, it's going to take longer than the time they're actually going to allow it before they're going to release these models to the wild, which means the cybersecurity risk is growing very dramatically sometime this year.

18:02By the way, the other thing that it means from a concentration of power perspective, I always thought and hoped that AI is the ultimate equalizer because anybody with internet access and AI logins can create things that previously were only available to really large corporations. The reality is that what's happening is in this very latest development, more and more capabilities are delivered to a very short list of companies under the excuse or the reasoning, call it whatever you want to call it, that they are the one that's going to keep everybody else safe. That is not necessarily the direction I want to see the world going.

18:40On the other hand, I think that if they can actually release it to a small group of companies to actually build real safety and precaution capabilities and only then release it to everybody else, assuming they actually do, it may not be a bad idea. To me, part of the interesting and maybe scary underlying text of all of this is so far what happened is every time they trained a model, they were then able to put specific restrictions on it to make it safe for us. Right? So that was the safety mechanism is let's put this model in a box that we can control and this way it will be safe to release it to the public.

19:14And they're taking a different approach now with access control to the models to be able to potentially understand what they're doing and understand the risks better and try to mitigate the risks before they evolve. And that's a dramatically different approach than every model to this point, which basically tells you that these models are getting more powerful and more capable and less controllable, and hence they are changing the strategy to this new strategy. And the last piece of news that I want to talk about in this relatively short AI news episode is the fact that cybersecurity issues are already on the rise and that before we have the next version of models, whether it's PUD or Mythos or whatever code name or numbers they're going to get, Axie, which is a widely used third-party developer library with over 70 million weekly downloads, was hacked in the past few weeks.

20:05And through this hacking, through releasing malicious code as part of this library, North Korean hackers were able to damage and have access to different libraries in the OpenAI Mac OS applications, including ChatGPT Desktop, Codex CLI, and Atlas. Now, if you are an OpenAI Mac OS user, you should very quickly update to the next version in order to get the new signing certificates in order to keep your computer and your data safe. OpenAI is claiming that personal information and user's information did not leak out of this process, but they do admit that they were hacked through this process of using this third-party open space library that a lot of other people are using.

20:49They're not the only one. But the reason I wanted to share this with you as part of this relatively short episode is that I want you to think about the concept that if OpenAI, with their very, very deep pockets and their high level of concern to security, especially in a year in which they're planning to go public, most likely in a few months, so it is a big deal for them. If potential enterprise clients, which is their biggest focus right now is going to learn that their system is easy to hack. They are not in a good shape. If they can get hit through a dependency chain, basically they're using a different tool, a different platform, a different library, and that can be hacked and they cannot detect it.

21:28It tells you how crazy the world we live in is right now and how risky it is before these new, very powerful and very capable models in cybersecurity are being released. So I must admit, I am concerned. I don't know if crazy concerned, but I'm concerned with the potential impact of what this may do to the world. And I really, really hope that this will lead to a partnership of OpenAI and Anthropic and Meta and Google and Chinese companies to try to think how to protect the world and hopefully also academia and governments in order to start paying attention and start building some kind of a safety net for software, a safety net for society, a safety net for jobs.

22:15We have to start thinking about this in much more effective ways than we are doing right now because the change is happening extremely fast and the impact it is going to have on businesses, society, security, et cetera, is going to be very, very dramatic. and I fear that without the right partnership, it will be very hard for each and every one of these groups, bodies, government, et cetera, to be able to handle this on their own. So on an optimistic note, like I said, I really hope this will actually lead to collaboration, which will allow all of us to enjoy the incredible, amazing benefits of AI.

22:50As I mentioned, I just saw magic happens in the last two days. People who did not know how to use it were able to develop sophisticated applications and connect them to data sets and actually build real business value less than 48 hours after they did not know anything about this. So think what that can mean to success and creation of business and entrepreneurship and everything else that actually does good in the world. It is very, very exciting, but we have to find ways to do this in a safe way. So I really hope we will see more and more collaboration in this field. That's it for today. I'm going to bed and I'm going to the airport really early in the morning to fly back home.

23:26I wish all of you an amazing rest of your weekend and I will see you back on YouTube.

From the publisher

SECURE YOUR SPOT FOR THE:
MULTI-AGENT ORCHESTRATION AI COURSE: https://multiplai.ai/multi-agent-orchestration-course/
AI BUSINESS TRANSFORMATION COURSE: https://multiplai.ai/ai-course/

Are AI tools finally ready to run parts of your business without you?

The latest wave of AI releases suggests the answer is shifting from “not yet” to “sooner than you think.” From dramatically improved autonomous agents to major leaps in multimodal capabilities—and growing cybersecurity concerns—this episode breaks down what’s changing and why it matters.

If you’re leading a business, the opportunity is massive—but so is the responsibility to adapt quickly and safely. This episode gives you a clear, practical lens on both.

In this session, you'll discover:

  • How Claude Opus 4.7 is unlocking longer-running, more reliable AI agents
  • Why improved image understanding is a major business advantage (not just a tech upgrade)
  • What real gains in legal, finance, and enterprise workflows mean for productivity
  • Why Meta’s new model is a serious comeback—and what makes it different
  • How lower token usage could reshape AI cost structures
  • Why OpenAI’s cybersecurity model signals a shift in how AI risks are handled
  • The growing tension between innovation, access, and control in AI development
  • Real-world proof of how teams can go from zero AI experience to building applications in 48 hours
  • Why cybersecurity risks are accelerating—and what leaders should be thinking about now
  • The critical need for collaboration between tech companies, governments, and organizations

About Leveraging AI

If you’ve enjoyed or benefited from some of the insights of this episode, leave us a five-star review on your favorite podcast platform, and let us know what you learned, found helpful, or liked most about this show!

More from Leveraging AI

All 330 episodes
285 | Meta is BACK - beats Gemini 3.1 pro, and ChatGPT 5.4, OpenAI Copies Anthropic's playbook, and Opus 4.7 is the new king - AI news of the week ending on April 17, 2026Leveraging AI · 24 min
Listen in VO