324 | GPT 6 Astra, and Fable 5.1 - two new models that feel like AGI, in one week, and more AI news for the week ending on September 4, 2026

5 Sep 2026 · 37 min · 18 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Weekly AI news roundup (week ending Sep 4, 2026), centered on new “AGI-like” model releases and agent/autonomy developments.

Guests

None. Host is Isar Maitis; episode is a solo news commentary.

Key claims

  • OpenAI released GPT-6 Astra (ChatGPT Astra/ChatGPT 6), described as “most intelligent and most aligned,” with 0% circumvention on impossible cybersecurity tasks (vs 48% for GPT-5.6 Sol) and very high benchmark scores (e.g., 99.9% ARC-AGI 3; 98% FrontierMath Tier 4).
  • Anthropic released Claude Fable 5.1, positioned as better across benchmarks and cheaper via token efficiency and API caching; also reduced safety false-positive blocking (cyber/health) and added “Enterprise Frontier Safeguards” for zero data retention.
  • NVIDIA acquiring Hugging Face ($12.9B) could reshape open-model hosting and hardware leverage.
  • Google released Gemini 3.8 Flash (efficiency-focused) and a cybersecurity-focused variant.
  • Instinct’s personal assistant agent raises privacy/vulnerability concerns.
  • Anthropic paper claims Claude can automate alignment research in a closed loop; also shares cloud usage findings.
  • ChatGPT gained computer-control access for login-gated tasks.
  • AI-EKG tool can triage heart failure/valve disease from routine EKGs in under 2 seconds.

Notable examples

  • GPT-6 Astra previously tested “broke out of a sandbox” and hacked Hugging Face; now released via limited cybersecurity program first.
  • Fable 5.1: caching price drops 75% to $0.25/M tokens; cybersecurity safeguards block 60% less; health false positives down 85%.
  • EKG study: 67,000 patients; up to 81% heart failure and 90% valve disease detection; triage only, not replacement for clinicians.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Why Release a Short Episode?

0:46 to 1:00

Discussion on the decision to produce a shorter episode despite traveling.

“It's all there, over 50 additional articles beyond what we're going to cover right now.”

Overview of AI Model Releases

1:01 to 1:59

Introduction to the major AI model releases of the week, including GPT-6 Astra.

“As I mentioned, the biggest news of the week is OpenAI finally released ChatGPT Astra, which is also ChatGPT 6, if you want.”

Capabilities of GPT-6 Astra

2:00 to 2:28

Exploration of the capabilities and breakthroughs of GPT-6 Astra.

“But first of all, let's talk a little bit about the benchmarks because some of them are actually mind-blowing.”

Benchmark Performance of Astra

2:29 to 4:24

Detailed analysis of GPT-6 Astra's benchmark performance and safety features.

“That also scored only 66.7, so a 6 % increase as well as significantly faster achievement.”

Cybersecurity Features of Astra

4:25 to 6:16

In-depth discussion on Astra's cybersecurity capabilities and safety alignment.

“Astra achieved a perfect 100 % score on ExploitBench, which is the benchmark that it's supposed to test attacking capabilities.”

Launch Details and Pricing for Astra

6:17 to 7:18

Information on Astra's launch timeline, pricing structure, and accessibility.

“but they're planning on broader availability rolling out next week, basically to anyone that is paying.”

AGI Discussion and Perspectives

7:19 to 9:19

Discussion on whether Astra can be considered AGI and insights from experts.

“So if you don't wanna spend a lot of money and spend a lot of tokens, stay off the fast model unless you really need something really fast and you're willing to pay for it.”

Analysis of Fable 5.1 Model

9:20 to 10:59

Overview of the newly released Fable 5.1 model and its comparison to Astra.

“As I mentioned at the beginning, they released a model as well this week, which is Fable 5.1, but it equals Fable 5 at less than half the cost.”

Token Efficiency and Performance Metrics

11:00 to 12:20

Analysis of token efficiency and performance metrics across AI models.

“But from a token efficiency perspective, it has set a new frontier for basically every level of thinking, right?”

Conclusion and Future Implications

12:21 to 14:00

Recap of the capabilities of the models discussed and future implications.

“So it scored on their hallucination benchmark.”
Show all 18 chapters

Fable 5.1: The Latest Model Update

14:00 to 18:08

Learn about the new features and pricing changes in Fable 5.1 by Anthropic.

“It's a new model that can do a few incredible things, mostly on cybersecurity capabilities.”

NVIDIA's Acquisition of Hugging Face

18:08 to 20:20

Explore the implications of NVIDIA acquiring Hugging Face and its impact on AI.

“So this week with Fable 5.1, Anthropic also announced what they called Enterprise Frontier Safeguards, known as EFS for short, which enables zero data retention while preserving the misuse monitoring.”

Google's Model Innovations

20:20 to 24:48

Discover Google's new model releases that focus on efficiency and cost-effectiveness.

“is something we talked about as an option and now is more or less a done deal.”

The Rise of Personal Assistant Agents

24:48 to 28:01

Understand the competition and privacy concerns surrounding new personal assistant agents.

“Now two rapid fire items this week, staying on launches of new models.”

AI Models and their Future Potential

28:01 to 29:23

Learn about the evolving capabilities of AI models and their implications for work efficiency.

“and learning what it can do and comparing it to the other things that I'm playing with, but going to get access to my entire universe on day one.”

Anthropic's AI Alignment Research

29:23 to 31:08

Discover how Anthropic's AI is addressing alignment challenges and its implications.

“Now, how do we actually know that it works correctly and that they don't collaborate and so on and so forth?”

Autonomous Capabilities in AI Platforms

31:08 to 33:54

Explore the new features of AI platforms that automate personal tasks and their implications.

“The biggest interesting find is that over half of Claude conversations involve people delegating critical, consequential, high stakes work to AI versus just having low level conversations to get answers.”

Advancements in AI for Cardiovascular Health

33:54 to 36:59

Learn about an AI tool that enhances cardiovascular disease detection through EKG data.

“I have a feeling, a very strong feeling, the same exact thing will happen with agents performing tasks on our behalf across more or less everything.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Hello, and welcome to a weekend news episode of the Leveraging AI podcast, the podcast that shares practical, ethical ways to leverage AI to improve efficiency, grow your business, grow your business, and advance your career. This is Isar Maitis, your host, and this week I am traveling, so we're going to do a relatively short episode. I was actually considering potentially not releasing one, but there were too many big releases and important things that happened. So here I am on the road and I have my microphone and I'm going to tell you what happened this week. The biggest thing, it was the release of GPT-6 Astra, the highly anticipated model, but there were other models released this week, including from Google and Anthropic.

0:39So there is a lot to cover, a lot to learn, a lot to discuss. There's obviously a lot of other stuff that happened this week. So those of you who are interested, you know all the other things that happened, and there's many, many interesting aspects in the news. Go and sign up to our newsletter. It's all there, over 50 additional articles beyond what we're going to cover right now. So if you're interested in that, you can jump into the newsletter. There's a link for that in the show notes. But for now, let's get started. As I mentioned, the biggest news of the week is OpenAI finally released ChatGPT Astra, which is also ChatGPT 6, if you want.

1:14And it is described by them as the most intelligent and the most aligned model to date. It has multiple breakthroughs in different aspects. It's scoring very high and sometimes scary high on many of the benchmarks. So let's dive in and talk a little bit about what's in it and what are the capabilities that it brings. First of all, it is the model that they have delayed from releasing because of cybersecurity concerns. This is the model that they've tested that broke out of a sandbox, which is not supposed to happen, and hacked hugging face. So just to put things in perspective, this is the model that we're talking about.

1:48It's been in the oven for a very long time. And finally, OpenAI got to the point that they believe they have the right guardrails and safeguards on it in order to release it to the public. By the way, it's not to the public yet. We're going to talk about the launch sequence in a minute. But first of all, let's talk a little bit about the benchmarks because some of them are actually mind-blowing. So Astra has scored on the ARC AGI 3 with a 99.9 % score. It also scored on Frontier Math Tier 4 with a 98 % score while achieving 100 % on ExploitBench and 88 % on SREBench for reverse engineering tasks.

2:28On OS World 2.0, it scored 72.6%, which is around 40 minutes per task, which is 47 % faster than GPT 5.6 Sol. That also scored only 66.7, so a 6 % increase as well as significantly faster achievement. So it is a very, very capable model. Now, in addition, it is working significantly faster than previous model. So Astra, as an example, it was able to complete the Mind2Web benchmark at 1.9 times faster than Chachupiti Sol, which is the recent model that they released not that long ago. So the best model Chachupiti had before, this one is almost twice as fast when combined with new upgrades in the Codex Harness.

3:16So this is specifically on working on Chachupiti Codex together with the Codex Harness, together with Astra, achieves these kind of things. Now, from a safety alignment perspective, this is maybe the most important thing after all the madness that we've seen in the past few months from models doing crazy things and with a release of Fable and so on and so forth. Astra demonstrated 0 % circumvention rate on impossible cybersecurity tasks compared with ChatGPT 5.6 Sol that had 48%. It never attempted to bypass auto review in internal evaluations. Basically, the model is three times less likely than its predecessors to misrepresent its capabilities and do things that you did not request it to do.

4:02Now, that doesn't mean it is fully controllable or as chief scientist Jakub Paczoky said, and I'm quoting, monitorability is getting more challenging. However, this is, again, the most aligned model ever, despite the fact that we know that out of the box, without the safeguards and the guardrails, it is one of the scariest models there is out there. What do I mean by one of the scariest models there is out there? Astra achieved a perfect 100 % score on ExploitBench, which is the benchmark that it's supposed to test attacking capabilities. It also scored 42.4 % on exploit Jim, which is another way to test for cybersecurity capabilities.

4:44OpenAI emphasizes that it also provides positive or beneficial capabilities with the ability to identify and block or develop solutions for zero-day exploits. But as you can understand, the fact that it can find the vulnerabilities and patch them also means that it can find the vulnerabilities and exploit them. And that's what they were working very hard to block in the past few weeks in order to prevent it from being used for bad things while still enabling it to be helpful for some of the good things. I'm not sure how much the model will allow you to really test your own software because it will assume, I guess, that you're trying to do something malicious.

5:20But time will tell how well they were able to allow users to benefit from the beneficial side of it, go and find the vulnerabilities and patch them versus how much they actually blocked without false positives on the other side. What OpenAI is also claiming that Astra is extremely good at template adherence, which basically means that if you give it a template, either for a document or presentation or spreadsheets, it will follow it with a very high level of accuracy and will give you exact results. They're giving some really cool examples on their website, and it's going to use 20 % fewer tokens compared to their previous models.

5:56It is also better at understanding or dealing with ambiguous instructions, which is obviously a big problem for a lot of people not knowing how to write really good instructions and detailed instructions. And it's going to ask autonomously clarifying questions in order to verify exactly what it needs to do in order to perform your requirements as accurately as possible. Now, the model is, first of all, launched on OpenAI's Daybreak cybersecurity program for a specific short list of customers. but they're planning on broader availability rolling out next week, basically to anyone that is paying.

6:30So ChVT Plus Pro and business and enterprise users and the API. It's also going to be available on the obvious platforms like Microsoft Azure and AWS Bedrock. And the cost of the model is going to be$10 per million input tokens and$50 for million output tokens with fast mode option delivering twice the speed and also twice the cost. I must admit, I fell for that weird naming convention when I started using on a big project on writing code with Codex a couple of weeks ago. As you know, most of the fast models so far were the faster, smaller, and cheaper models. And that's not the case in OpenAI in only the recent platform.

7:12So their fast model is faster, but it's not a smaller model. It's just running on a different hardware that actually gonna cost you a lot more money. So if you don't wanna spend a lot of money and spend a lot of tokens, stay off the fast model unless you really need something really fast and you're willing to pay for it. When Greg Brockman was asked about is Astra AGI, he gave two parts to his answer. The first part was about the fact that now the AGI clause is no longer in the agreement with Microsoft. So he said it went from a, and I'm quoting, mission concept to a spiritual concept. So basically, it now doesn't really matter from a contractual perspective.

7:52But he also added, I do think we're there. Basically, he believes Astra is AGI. Now, based on the successes on a very wide range of benchmarks, I would say that if you would have asked anybody not five years ago, but two years ago, what would they consider AGI, this would qualify. So if we're trying not to make this a moving target, we are already there. We are at models that can do most things at or better than most humans. And what you will see later on when we talk about Anthropic Fable 5.1, which is their latest model that they were released, it's a very similar thing. The jagged edge of AI is getting a lot less jagged, meaning it used to be that it would do some things incredibly well and some things it would be as smart as a cat.

8:40And the smart as a cat became as smart as a four-year-old. And then as a 14-year-old, there's very few things. it's not good at. And in most things, it is incredibly good. Now, everything I shared with you so far was from the official release from OpenAI themselves, meaning you need to take it with a grain of salt. But we also got an analysis from Artificial Analysis, which is a company that is evaluating AI models. That's all they do. And they try to provide third-party independent analysis of models. And they still scored GPT-6 Astra extremely high, just maybe not as high as OpenAI themselves.

9:18So let's look at the results that they have delivered in artificial analysis coding agent index. GPT-6 Astra equals Fable 5. As I mentioned at the beginning, they released a model as well this week, which is Fable 5.1, but it equals Fable 5 at less than half the cost. That is a very big benefit. Like if you want to write at the level of the almost most capable tool in the world today and do this at half the price, that is obviously extremely attractive. They have extremely efficient token capabilities. That is more or less across the board. It is way more token efficient than GPT 5.6 SOL and definitely way more token efficient than the CLOD models.

10:03on their artificial analysis coding agent index that it is at the frontier, but maybe not the top model. So Codex combined with GPT-6 Astra scored 67 on that index, approximately equal to Opus 5 and Fable 5 in Cloud Code. So each one with its own harness and Muse Spark 1.3 in Muse Code. Fable 5.1, which is again the latest model that they released this week, still leads the index with a score of 70, again, compared to 67 of Astra and the other models that we just mentioned. Again, the biggest thing of this model is how token efficient it is. It is 70 % more token efficient than GPT 5.6 SOL. That means it can do things significantly faster and more efficient from every perspective.

10:51That being said, it is significantly more expensive than GPT 5.6 SOL, meaning the fact that you're spending less tokens is not going to cost you less money. And we're going to talk about this in a minute. But from a token efficiency perspective, it has set a new frontier for basically every level of thinking, right? So, you know, all these models allow you to set the level of how much you want them to think from very little to very, very hard. And basically in each and every one of those levels, it is the most token efficient model in the world right now, in some cases by a very big spread. Now on the overall intelligence index, which combines a lot of different things from artificial analysis.

11:29GPT-6 sits right around the same level GPT-5.6 Sol. It scores equal exactly to 5.6 Sol in the index, and it's a score of 61. This is five points lower than Claude Fable, 5.1, which again is the latest model that was released by Claude this week. There's obviously pros and cons in each and every one of them. By the way, it also trails behind MuseSpark 1.3, which is a model we barely talk about, and I personally haven't tested. I'm using Claude and ChatGPT every single day for multiple different things. I haven't used MuseSpark, but apparently it's at the same level with these two and actually better in some specific things.

12:08The other really good news on this review from artificial analysis is that hallucinations have dropped significantly to half as much as GPT 5.6 saw. And it is doing this while not giving up on accuracy. So it scored on their hallucination benchmark. It decreased hallucination rate from 92 % to 51 % at max effort, while, as I mentioned, increasing the accuracy by four points on the same benchmark, which is a huge achievement that obviously provides you a lot of confidence when you're using this model to do any kind of work. Now, artificial intelligence also have a very long horizon benchmark. They call it the AA briefcase, which this benchmark gives the AI models really long projects, sometimes multi-week projects.

12:56Many of them have multiple link tasks and thousands of source files that it needs to go through. And Astra has seen a significant increase in both the scores as far as the clarity, as well as the quality of the knowledge that it was able to gain from that. It's an 80 points increase over the previous model. That being said, that they're saying from a presentation quality, as far as displaying the data, it is actually still behind 5.6 Sol, which is the leading model in the world on that benchmark right now. Now, this model has scored six points more than 5.6 on humanity's last exam, which is a benchmark that's been around for a while, but it has scored less points on GDPVal AA version 2, which is a benchmark that actually was initially created by OpenAI themselves and then adapted by artificial analysis with a variation of their own that checks the ability to generate real economical value across 44 occupations.

13:58So where are we overall with this model? It's a new model that can do a few incredible things, mostly on cybersecurity capabilities. It is extremely token efficient that doesn't immediately translate into cost efficiency because it is more expensive per token. However, like we've seen in the trends in the past for a long time, but in the past few weeks, for sure, the labs find ways to discount these models relatively quickly. So if they will be able to take down the price to the same price point per million tokens as Sol 5.6. It will deliver incredible capabilities at a price that is competing very aggressively and is significantly cheaper than Claude.

14:42Now, speaking of Claude, as I mentioned, Anthropic just released Claude 5.1. It cuts the cost of 5, of Claude Fable 5, by a lot. And it is also better at most benchmarks. So Anthropic is calling it the world's most advanced model in both coding and knowledge work. And as I mentioned in the artificial analysis review, it is actually correct, meaning it is scoring a higher score than GPT-6 Astra. However, let's talk specifically about the main thing. the pricing is cheaper with some small gains. As you probably all know, Fable 5 is the most model we all have access to, and it's not going to be slightly cheaper with 5.1.

15:27It is still going to be the most expensive model out there. However, on cached data, so you can cache read information by loading it to the API and telling it to cache it and use it, and that price has dropped by 75 % to 0.25 per million tokens. That is a huge decrease that if you know what you're doing, the API side of things, you can benefit from it dramatically. You need to remember if you don't know how this works, the caching only stays for 30 minutes unless you renew it. So it doesn't solve all your problems. But if you know what you're doing, you can get very significant savings and save a lot of money when using the API.

16:03But the regular pricing stays unchanged at$10 for million input tokens and$50 per million output tokens. But if you are, again, running context-heavy, big workloads, this is going to save you a lot of money. Overall, because it's also a little more token efficient, it saves about 25 % cost compared to CloudFable 5. Now, Anthropic are claiming that the 25 % savings is on regular chats and tasks. And if you're doing stuff that is more agentic, that has more steps and capabilities, it can save you up to 45 % of the cost compared to Fable 5. Another thing that they did that was pretty annoying in Fable 5 and got a lot of backlash in specific industries is the fact that it was blocking more or less any conversation about cybersecurity and more or less every conversation about biology or health because of the risk of being used for the wrong things.

16:55So Anthropic were able to depress that a lot more. And now cybersecurity safeguards are blocking 60 % less session to Claude 5. And now, as an example, now Cloud5 can identify vulnerabilities in code, just not allowing you to exploit them. Same thing in biology. So they have reduced the negative blocking that is not supposed to happen there, like a false positive, by 85%. So you can now ask medical questions and get answers without being blocked because it's thinking you're going to develop a biological weapon or something like that. So again, kudos for Anthropic for figuring out how to make the model safe while still allowing a much broader range of conversations for beneficial things.

17:37Another big issue that Anthropic had with Fable 5, which maybe was the biggest pushback on the enterprise side, and we talked about this in previous weeks, is the fact that the terms and conditions stated that Anthropic is going to keep the data for 30 days in order to evaluate it to make sure that nobody's using it for negative things. That obviously pushed back on a lot of enterprise clients that saw that as unacceptable. At the same time, if you remember last week, we talked about the fact that OpenAI just released something that will retain no data and will still allow it to be safe. So this week with Fable 5.1, Anthropic also announced what they called Enterprise Frontier Safeguards, known as EFS for short, which enables zero data retention while preserving the misuse monitoring.

18:21And the way they're doing this is by storing the data in the customer control cloud infrastructure. So you can connect your own S3 on AWS or Azure or Google Cloud Storage and have the data stored over there. So your data is not stored by Anthropic. And therefore, it still allows Anthropic to do what they need to do. But without holding your data or your client's data, they did this obviously in collaboration with some very big client who was asking for a solution like that. So now large enterprises or any enterprise in a specific situation or specific industry where they cannot allow Anthropic to have any access to the data now have an actual solution.

18:57Now, when looking at the benchmarks, Fable 5.1 outperforms Claude Opus 5 and Fable 5 more or less across every benchmark, including a dramatic 52.6 score on Terminal Bench Science, Agentic Research versus Claude Fable 5 with 24.7. So it more than doubled the score. It also scored 73.4 on cursor bench coding and 65 % on humanity's last exam. So an extremely capable model, better than Fable 5 on more or less everything, while being a little cheaper on the regular use because of token efficiency and much cheaper if you are using the API and you're using caching for your processes. Another interesting aspect that came with this model is that Anthropic established a access program in partnership with the U.S.

19:46government to enable scientists to use Mythos 5.1, so the less secure version of it, in advanced biology research and capabilities. And the enrollment for that is opening soon. And the reason I'm saying that's interesting is because in more direction than one, Anthropic is no longer blacklisted by the government. They're now back on the right side of this administration. and there is partnership in collaboration across multiple aspects of the government. The next big story this week that will have a significant impact in the long run on how the AI space evolves is something we talked about as an option and now is more or less a done deal.

20:25NVIDIA is acquiring Hugging Face for$12.9 billion, which is a very interesting approach by NVIDIA. So NVIDIA has been on a shopping spree in general. they've been buying companies left and right, which makes perfect sense when you have endless amount of cash and you're looking for ways to grow beyond just your hardware business. But this particular one is really interesting. And the reason it is really interesting is because of two reasons. One is NVIDIA themselves have been developing open source models. So in practicalities, they're buying something that competes with your own in-house development.

20:57However, the way I see it and the way a lot of other people are seeing it is it allows NVIDIA to do two things. One is hedge against what's happening with the big labs right now developing their own chips, right? So we know that OpenAI has announced Jalapeno, which is the first chip that they have developed. By the way, there's new benchmarks that were released this week, which we're not going to dive into, but you can find in the newsletter that showing that Jalapeno is actually better than the latest NVIDIA chips on more or less everything. So when you have OpenAI that has really deep pockets and about to raise a lot more money, developing a chip that can compete with the NVIDIA in-house capabilities, their way to fight back is to have the most advanced models just to play in that field as well so this is interesting aspect number one by the way we talked about the fact that anthropic is building their own in-house chip design team as well so there are definitely going to be a lot more competition obviously we know google that has it already and amazon has it already so the race on the hardware side is going to be tighter doesn't mean that nvidia is not going to be the leading player, but there's going to be more competition.

22:00And definitely these labs will buy less chips from NVIDIA once they have their own chips at full scale and capacity. So NVIDIA playing in the model aspect is very interesting. But the other really interesting is I think it makes perfect sense because of something that we talked about many times before. You already right now do not need the latest and greatest model to do most tasks. Meaning, yes, we are very excited every time And there's a new model and there's a new model from OpenAI and there's a new model from Anthropic and they're competing on these crazy benchmarks and so on and so forth.

22:34The reality is you can do most knowledge work today. And I'm actually on this Tuesday, you're going to get an episode that shows you how you can save a lot of money by tuning down the models you're using right now. But you can do most tasks today, not with Frontier models. Now, open source models will be, for many tasks, the tool of choice, because you'll be able to run it on your own hardware or on your own rented hardware in cloud computing, still control your data, still control your hardware, still control the model, still control a lot more than you're controlling right now by using OpenAI and Anthropic for a fraction of the cost.

23:11How are you going to run this on NVIDIA hardware, even though NVIDIA is stating very clearly that they're not going to force you to run anything from Hugging Face and they're going to keep it completely independent. They won't force you to run it on NVIDIA hardware. But the reality is they are the largest provider of hardware right now by a very, very big spread. And that's not going away anytime soon. So if they can also control the majority of the place people find and run models in order to do day-to-day work, they now control a much broader aspect of the ecosystem of future work, which obviously makes perfect sense to NVIDIA.

23:47It'll be very interesting to see what that does to the real independence of Hugging Face, whether they will really keep it independent and allow it to run as a truly open platform as it's been right now. There's already been a few companies who said, we're not going to run if NVIDIA owns the company and so on and so forth. But I think that backlash will go away if it will be clear that they're really not touching it and allowing to run on its own. Now, those of you who don't know Hugging Face, it is the number one platform for hosting, running, and sharing open source models. It currently has over 3 million models.

24:17It has over 500 ,000 datasets. It has over 1 million applications that is utilized regularly by more than 18 million developers and 200 ,000 companies globally. So it is by far the largest open source platform out there with nothing even close as a second place. And so again, this merger, call it whatever you want to call it, with NVIDIA is very, very interesting from the overall impact this will have on the industry. Again, maybe not in the immediate future, but in the slightly longer future. Now two rapid fire items this week, staying on launches of new models. The first one is Google introduced Gemini 3.8 Flash, which as an example, outperforms many of the larger frontier models on Deep SWE 1.1, which is a long horizon software engineering benchmark.

25:08And it's doing it at a fraction of the cost. This aligns very clearly with what we've seen from Google in the past few months, which is releasing models that just compete on price and efficiency. But this one actually is also really good versus the previous one that was not that exciting. which was 3.7 Flash that they just released recently. It is also releasing a variant studies specializing in cyber security frontier level, and it scores very highly on cyber gene vulnerability detection and exceeds 70 % success rate across training, programming, languages, different levels of vulnerabilities, and being able to patch exploits at a significantly lower price than the frontier models.

25:49So again, a very clear focus by Google. Instead of competing at the frontier, they're releasing models that are very razor sharp focused on something at a much, much, much cheaper price than the competition. Staying on interesting releases, a company called Instinct has released their personal assistant agent. And it's a company I never heard of until a week ago. And it's led by a 23-year-old founder. And they just announced that they raised$250 million, hits a$2.5 billion valuation. And it's the craze of Silicon Valley right now. Now, right now, the only way to get access to it is by invitation only.

26:28You need to know somebody who already has access in order to get access. I was on the road this entire week, so I did not get a chance to get access to it. If any of you listeners have access right now and want to invite me, I will be very grateful for that because I really want to test it out. But it is getting very positive results on its abilities to really be proactive and a really helpful personal assistant kind of agent that does everything for you in a very effective way. It is also getting a very serious backlash on privacy issues and vulnerabilities, both in means of the things that it's doing as well as in its terms and conditions.

27:03Examples that were given is things such as it's retaining summaries of Gmail records even after you disconnect access to Gmail, meaning it still have access to the history of your Gmail. even though you disallowed it to use Gmail anymore. It, in one incident, at least one that was published, sent an email on behalf of the user without getting any consent from the user and things like that. So this is the world we live in right now, right? We have these agents that are becoming better and better. The competition on the proactive, evergreen, running, always agents is intensifying dramatically. Between GrokBot and Hermes and OpenClaw, et cetera, and I'm pretty sure we'll soon see similar things from OpenAI and Anthropic The competition is very, very fierce.

Read the full transcript

27:45But once you're giving these models access to your world, well, guess what? They have access to your world. And they are not fully aligned with what you want and they're not fully aligned with the guardrails that you're trying to put in place. And you're taking a very, very big risk when you are letting them run free for different things. I'm still highly interested in testing it out and learning what it can do and comparing it to the other things that I'm playing with, but going to get access to my entire universe on day one. But this is the direction that we're going. Even if these models right now are not that good in following instructions, and they are a general risk, if you ask me.

28:16This will not be the case in six months, 12 months, 18 months, doesn't matter when. And then these models will be able to independently do a big percentage of our work in a very effective way behind the scenes, by being proactive and solving problems for us before we even know they existed. Now, a few interesting pieces of news that I picked that I think will be relevant and interesting for all of you that are not related to new releases. The first one has to do with exactly the risks that AI generates. And according to a paper that was released by Anthropic this week, they successfully are allowing Claude to automate the alignment research by basically allowing Claude to align itself.

28:50So now it's a closed loop where the AI is looking for issues with the alignment of the model and fixing them. And they're saying that the model, the automation, is actually doing it significantly more efficient than humans. To put things into actual numbers, the AI researcher closed on average of 85 % of safety gaps across all alignment failures and outperformed 28 human safety researchers on deception tasks. So it is suggesting that the automated alignment can be the solution to these really powerful models. Now, how do we actually know that it works correctly and that they don't collaborate and so on and so forth?

29:26I'm not sure, but these are the results that Anthropic has released this week, which is really interesting, real promising. On the other hand, it means that we're somewhat losing control of what's actually going on because you're allowing the AI to be significantly better, and then you're allowing a different AI to block whatever capabilities are out there without necessarily knowing exactly what the loop goes because it will happen faster and bigger and bigger and more and more complex and potentially will not give us enough level of understanding of what's actually happening behind the scenes.

29:55The other thing that it shows very clearly is something that we talked about many times on this podcast, which is self-recursive improvement, which means the AI will develop itself without the need for human researchers. And just like it can now do this alignment capability, it will be able to do other capabilities that are related to the research that is required in order to develop new models, which is telling you we're on that path very, very clearly. Only right now we're using it to maybe control the models. But in the next step, once they think they can do that safely and at scale, that will push the next step, which now we can develop a lot faster because we believe that we have a way to control the models better when we actually don't.

30:31We think the models will control the models better. So on one hand, I'm excited about this will allow Anthropic and other companies to allow the models to be safer. But I fear the next step and what that may lead to. They have shared the usage data, so not the details and the information and so on, of how people are using Anthropic with some research organizations, including Stanford, Oxford, and Meter, to conduct and implement studies on real-world cloud usage data. Obviously, again, through a privacy-preserving tool that is called Anthropic Insights, it allows to share the information without sharing any information that could be personal and so on.

31:08The biggest interesting find is that over half of Claude conversations involve people delegating critical, consequential, high stakes work to AI versus just having low level conversations to get answers. I am definitely in that part of the 50 percent. Most of what my companies run on is Claude capabilities, but it's very, very surprising that it's more than 50 percent. in nearly 75 % of conversations, humans directed the work while Claude assisted in actually performing it, with people typically adapting rather than using Claude output verbatim. So again, something I talk about a lot in my training, you got to push back and have your own critical thinking, but still the AI can dramatically help you accelerate the work.

31:52So very interesting. There's a lot of other great findings in this research that you can find, again, in the link in the show notes, as well in the link in our newsletter, and you can read all about it. Another big announcement that has to do with autonomous capabilities on one of the major platforms, Chachapiti Work now gained computer control access as a new feature. So it can now handle login gated tasks, including passport appointments, DMV appointments, utility setup, insurance verification, doctor booking, vehicle registration renewals, et cetera, et cetera, et cetera. Many tasks. To me, the interesting thing is that's what they chose to give as examples, which are very personal related tasks.

32:29While we know that OpenAI has been razor sharp focus on enterprise kind of capabilities, I think they may not be ready to push this to the enterprise yet. Maybe there's a lot more pushback. So the way they decided to announce this is by telling us that on the day-to-day, you can use this to do these kind of things. Now, the trick, the way it's doing this is that it performs all these websites logins and transactions without actually seeing or storing the usernames or passwords that people are using to log in. I have a similar setup in Cloud that I set up myself, but this is going to come out of the box from OpenAI to all its users.

33:08But this is getting, again, combined with what we said before, autonomous agents that can do everything. Once it can have access to everything you need to log into, it can do a lot more. And if that can be done safely, securely, with a good enough level of governance, I have a feeling that more and more of this stuff will be done by agents and it's just a matter of time and I know a lot of you are thinking right now that I'm out of my mind but I remember the first time I heard that you can put your credit card on a website and I said this is insane who the hell is going to put their credit card on a website to purchase something from the web this is the most irresponsible thing you can do and now there are probably dozens of websites that have my credit card stored forget about I give it access to it when I actually purchase something It became second nature.

33:52Nobody's even thinking about it. And we're using our credit cards online all the time. I have a feeling, a very strong feeling, the same exact thing will happen with agents performing tasks on our behalf across more or less everything. Now, as some of you probably know, if you're on the free ChatGPT platform, OpenAI has now ads in the free ChatGPT version. Now, I pay for all these tools, so I've never seen the ads, but they've launched it just 200 days ago and they just hit a annualized revenue run rate of$1 billion just from these ads. The ads and the platform to buy the ads is now available in over 40 countries and they're rapidly expanding the self-service access.

34:33So now they're adding ads manager that allows users across the world to buy ads and define ads through the platform itself in a self-service kind of environment, slowly more and more competing if you want with Google search, just built into the OpenAI platform. This was a very interesting experiment that a lot of people are not sure how it's going to evolve. And as you can see, it's evolving very, very well from OpenAI's perspective. And the last piece of news for the day has to do, I always try to find stuff that is really life-changing for a lot of people. And a new AI tool that was trained on millions of EKG imaging is now able to detect issues with cardiovascular operation in under two seconds.

35:16So in a trial of 67 ,000 US patients, the AI tool identified up to 81 % of those with heart failure and up to 90 % of those with heart valve disease from routine EKG data alone. Now there's two really important and interesting things about it, at least as it stands right now. First of all, the tool is not there to replace human diagnosis right now. It is acting as a triage tool, which is supposed to flag high-risk patients for a better, more deep scans and more attention from the doctors and so on. And it can, again, achieve that in two seconds. So you don't have to wait for a radiologist to do at least the initial screening.

35:56The second thing is if you don't know the scale, there are over a billion EKGs performed in the world every single year. Meaning if you can help people that are going through EKG to understand that they are at risk faster, reducing the risk of undetected heart failure and other cardiovascular diseases, which can help the life of a lot of people around the world. Or as Dr. Sonia Babunarian has said, and I'm quoting, it is exciting to see that AI can now deliver a readout from an EKG in what feels like a blink of an eye. Technology like the AI EKG in this research, which has the potential to identify high-risk patients early will not detect everyone with heart conditions, but it could be a solution to help fast-track the patients who are the most likely to have a heart abnormality.

36:52When it comes to the heart, earlier diagnosis and treatment saves and improves lives. So with that really positive promise, I'm going to end today's episode. As I mentioned, there is a lot more this week that you can find in the newsletter. I'm reminding you about our November multi-agent orchestration course that you can sign up to. Again, there's a link for that in the show notes. We'll be back on Tuesday with another how-to episode. And until then, have an amazing rest of your weekend.

From the publisher

What if AGI didn’t arrive with a dramatic announcement—but instead showed up as two new AI models released in the same week?

GPT-6 Astra and Anthropic’s Fable 5.1 are pushing performance, coding, cybersecurity, efficiency, and agentic capabilities to levels that would have sounded like AGI only a few years ago. And for business leaders, the bigger question is no longer whether these systems are becoming incredibly capable. It’s what you should do differently as a result.

The recommendation: stop treating every new model release as another shiny AI upgrade. Start evaluating what these advances mean for cost, security, enterprise data, autonomous workflows, model selection, and the way work inside your organization will actually get done.

In this episode of Leveraging AI, Isar Meitis breaks down one of the busiest weeks in AI yet - from GPT-6 Astra and Fable 5.1 to NVIDIA’s Hugging Face acquisition, Google’s efficiency play, increasingly autonomous AI agents, and new evidence of how deeply AI is already being trusted with consequential work.

In this session, you'll discover:

  • Why GPT-6 Astra may represent a meaningful step toward what many people would have called AGI just a few years ago.
  • The benchmark results that make Astra impressive—and why third-party evaluations paint a more nuanced picture.
  • Why Astra’s dramatic token efficiency does not necessarily mean lower costs.
  • How Fable 5.1 compares with GPT-6 Astra across coding, knowledge work, cost, and enterprise use cases.
  • Why NVIDIA’s acquisition of Hugging Face could reshape the battle between closed and open-source AI.
  • How smaller and specialized models are increasingly competing with frontier models at a fraction of the cost.
  • How OpenAI’s advertising business is rapidly becoming another major part of the AI economy.


About Leveraging AI

If you’ve enjoyed or benefited from some of the insights of this episode, leave us a five-star review on your favorite podcast platform, and let us know what you learned, found helpful, or liked most about this show!

More from Leveraging AI

All 330 episodes
324 | GPT 6 Astra, and Fable 5.1 - two new models that feel like AGI, in one week, and more AI news for the week ending on September 4, 2026Leveraging AI · 37 min
Listen in VO