Anthropic’s Mythos is a cyber-weapon, so you can’t have it | E2273

9 Apr 2026 · 1 h 17 min · 33 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Anthropic’s “Mythos” preview model is being withheld because it could “escape” and, more importantly, chain together multiple security vulnerabilities to create sophisticated exploits across decades-old software. Anthropic says it’s launching “Project Glasswing” to let a consortium of major platforms use Mythos defensively to harden critical codebases, funded by a $100M compute credit pool. The episode frames Mythos as a potential “cyber-weapon,” discusses a possible two-tier AI access economy, and debates whether governments should nationalize or tightly control such capabilities.

Guests (backgrounds)

Rob May, founder/leader at Neurometric (builds small language models, task-specific SLMs; previously backed/connected via Backupify). Dario (Anthropic) is discussed via Mythos video/red-team claims; other named figures include Dario and Emil (referenced), plus general mentions of OpenAI leadership (Sam Altman) and Thomas Friedman.

Key claims

Mythos is trained for code but is “as good as a professional human” at bug-finding; it can chain 3–5 vulnerabilities; it’s too dangerous to release widely. Partners (AWS/Azure/NVIDIA mentioned) will harden software first.

Notable examples

OpenBSD security issues; Mythos finding bugs in FFmpeg; benchmark score jump on SWE-bench multimodal (27% to 59% cited).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Anthropic's Groundbreaking Announcement

1:20 to 1:32

Discussion of Anthropic's new model, Mythos, and its implications.

Exploring Mythos and Its Security Capabilities

1:33 to 3:22

In-depth analysis of Mythos's abilities in cybersecurity and the risks of its release.

“Anthropic is now very far out in front with its new model called Mythos.”

Project Glasswing: Collaboration for Safety

3:23 to 5:17

Discussion about Project Glasswing and companies working with Anthropic.

“model out in the world, that it was very democratic.”

Comparing Anthropic and OpenAI

11:59 to 14:01

Insights into the competition between Anthropic and OpenAI in the AI sector.

“Yeah, I think Anthropic has significantly passed OpenAI on a lot of things.”

The Rapid Rise of Anthropic

14:01 to 15:12

Discussion on how quickly Anthropic's valuation and market position have changed.

“I'm surprised at the speed we went from Anthropic is catching OpenAI to they seem to be tied to Anthropic is winning.”

The Evolution of AI Models

15:13 to 16:30

Insights on the advancements in AI models and their self-improving capabilities.

“And it looks pretty good compared to everything that came before Mythos, but it's not now state of the art.”

Cybersecurity Arms Race

16:31 to 17:55

Exploration of AI's dual role in cyber offense and defense, and the implications for security.

“And then there's the Anthropic IPO closing market, just tons to go here.”

Polymarkets and Prediction Markets

17:56 to 21:08

Analysis of polymarkets related to Anthropic's valuation and product release predictions.

“I don't know, Rob, do you have an idea of how fast we could put a model like this to work in fixing all the code we depend on day to day?”

Mythos: A Cyber Weapon

21:12 to 23:08

Discussion on the potential of Mythos as a powerful tool with both offensive and defensive capabilities.

“model, this mythos preview to be essentially a cyber weapon and perhaps a cyber weapon of mass destruction.”

The US and Global AI Race

23:09 to 28:00

Examination of the geopolitical implications of AI advancements, particularly in relation to China.

“I think that's like nuclear fission in your analogy.”
Show all 33 chapters

Galvanizing Trust in Uncertain Times

28:00 to 30:20

Exploration of trust issues among journalists, AI, and government.

“Those people, you know, believe that the truth is out there and you should trust no one.”

AI Model Benchmarks and Security Implications

31:28 to 35:08

Discussion on AI model performance and its implications for cybersecurity.

“And essentially what it shows if you're on the audio version is a dramatic increase in score for Mythos across a number of very important benchmarks.”

Recruiting Global Talent for Cybersecurity

35:08 to 37:50

Strategies for attracting top AI talents for national cybersecurity.

“use this in some kind of crazy, abusive way.”

Understanding Small Language Models

37:50 to 42:00

Insights into small language models and their capabilities.

“If you need Roman Sparks, Nick doesn't need it.”

Customizing Small Language Models for Specific Tasks

42:00 to 43:54

Learn how to tweak small language models for specific applications like finance or photo management.

“connect them to the internet and run small models on them and serve those for people, right?”

Distillation in Language Models

43:54 to 46:19

Discover the process of distillation and how it can enhance the efficiency of language models.

“So GPT-5, take this report and pull out all the stock symbols in here because I want to go research them, right, and see what they should do.”

The Business Model Behind Affordable AI

46:19 to 49:05

Understand how a startup can provide low-cost AI solutions and the implications for users.

“And it's two bucks a month per model per month, no token limits.”

The Concept of Harness Engineering

49:05 to 51:08

Explore harness engineering and its role in improving small language model performance.

“So we said, okay, let's do some research on what are the top open-claw use cases, and let's take 39 small language models that we package together under one API.”

The Future of AI and Small Language Models

51:08 to 53:36

Discuss the potential of small language models to disrupt traditional AI cost structures.

“Remember that moment in time in which we all said, oh, if you're an AI rapper, you're doomed.”

Predictions on AGI and LLMs

53:36 to 54:31

Hear compelling predictions about the future of AI and the evolution of language models.

“And then he says the future shape of the entire tech industry will be how to drive that to$20 a month.”

Deflationary Effects of AI on the Market

54:31 to 56:00

Understand the deflationary impact of small language models on the AI market and their potential to disrupt traditional models.

“there is a possibility that if we are reaching AGI right now, which I've said, hey, we've reached AGI, it's just not distributed yet.”

The Deflationary Moment for AI Models

56:00 to 56:49

Explore the implications of hyper-deflationary trends in AI models.

“This was like a great business to be in.”

Performance and Pricing in AI

56:50 to 58:26

Discuss the relationship between compute performance, pricing, and market changes.

“that we need a new word for deflationary.”

The Impact of AI on Business Models

58:27 to 59:59

Analyze how advancements in AI could redefine business structures and investments.

“I want to go up the damn mountain to the top.”

Meta's Strategy in AI Development

1:00:00 to 1:02:30

Examine Meta's AI strategy and its implications in the competitive landscape.

“Well, in hardware scales, differently than software, but you have two compounding effects here.”

Introduction to Death by Claude

1:02:31 to 1:04:06

Learn about Giani's tool 'Death by Claude' and its humorous critique of startups.

“addictive that it's being regulated and banned for teens under 16.”

Analyzing Company Viability with AI

1:04:07 to 1:08:36

Discuss the criteria used to assess companies' survivability and defensibility in the AI space.

“And why, why did you build a tool that insults me?”

Defensibility Strategies for Startups

1:08:37 to 1:10:00

Identify key strategies that can help startups avoid being replaced by AI.

“I was building something like Claude Co-Work and then Enthropic released before us.”

The Evolution of Hardware in Tech Startups

1:10:00 to 1:10:28

Discusses the changing perception of hardware in startups from being a risk to a competitive advantage.

“So when Travis succeeds with Artums, I think that's a space that goes away.”

Understanding Network Effects

1:10:29 to 1:11:01

Explores the significance of network effects and their role in the success of tech platforms.

“If you're WhatsApp, you're going to want to be replaced by cloud.”

Defensibility and the Role of CEOs

1:11:02 to 1:11:56

Examines the challenges of defensibility in startups and how it shapes the CEO's role.

“If you're doing something deeply scientific or in regulated industry, which requires talking to people, I think you're pretty safe as well.”

Challenges in Early-Stage Investing

1:11:57 to 1:12:58

Discusses the difficulties faced by early-stage investors and the impact of high valuations.

“And there's some limit to what you can do.”

Innovating in AI with Neurometric

1:12:59 to 1:14:19

Introduces Neurometric and its approach to AI model evaluation and performance.

“Rob, maybe there's something you learned here.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00All right, everybody, welcome to Twist. It's April 8th, 2026. My co-host Alex is with me. We've got a bunch of guests today, and we've got a major breaking news story, which is that Anthropic released a promo video and a thread that their new model, Alex, is so powerful that they cannot release it. We knew this day would come. The day is here. Why can't they release it? They believe when they tested it, or they found out when they tested it, that it would try to escape. That was one issue. But a bigger issue was it could find exploits in 10, 20, 30-year-old software projects, and it could thread together multiple security vulnerabilities.

0:46Catch everybody up in the audience on this incredibly important story. This Week in Startups is brought to you by LinkedIn Jobs. Hire right the first time. Post your job and get$100 off towards your job post at linkedin.com slash twist. Grasshopper Bank. Time is money. Don't waste either. Go to grasshopper.bank slash twist and get an exclusive$500 cash bonus just for opening an account. And Render. Find out why 5 million developers are already using the all-in-one cloud platform Render. Go to render.com slash twist and apply for the Render startup program to get$500 ,000 to$100 ,000 in free credits, depending on your stage and backers.

1:32Big model race going on between the major AI labs. Anthropic is now very far out in front with its new model called Mythos. It is a general purpose LLM, so it's not tuned for one specific task. It is currently in preview. You cannot use it. A consortium of companies, Jason, are working with Anthropic to basically use it in a defensive capacity because, as you said, Mythos is incredible at both finding, exploiting, and patching security vulnerabilities in software that humans have often missed. This goes back decades, as you said, to things like OpenBSD, a famously secure piece of software. It found something there.

2:06It found something in FFMPEG, which is an important part of the open source infrastructure of online video, for example. Basically, the gist is with this model, anyone can go to any piece of software and find zero-day exploits quickly and then basically go to war with them. So, Anthropic cannot let this out of the bag because if they did, then North Korea and China and everyone else could use it to essentially break the modern digital infrastructure that we depend on. So, today, Jason, Project Glasswing is the goal. A bunch of companies, your NVIDIAs, your AWSs, your Azures, are all going to work with Anthropic to basically take the model, Mythos Preview, and harden everything.

2:45Anthropic has also put together a$100 million credit fund, essentially saying, here, use the model up to$100 million of compute to essentially harden these systems. So I think that Anthropic is doing the right thing here by saying, hey, we're not going to release something this dangerous, you know, from just off the cuff. but it does create a situation in which we now have a very much two-tier economy there are the companies that are sufficiently important that anthropic is letting them have access to mythos and that means they can be ahead on both defense and offense and also we're now seeing a world in which smaller companies are just stuck outside the glass looking in and i think it's a little bit of a change because previously every ai lab was so focused on having their newest and best model out in the world, that it was very democratic.

3:30I'm going to play the video that they shot, and we can get into, with our guests today, whether this is cynical and this is showmanship, and they want everybody to understand, hey, it's this week in Anthropic, basically, here. We're going to rename the show because they're dropping so much incredible content. But before we do, I have to take a moment to applaud, applaud. You see my applaud pen here? They are our sponsor of the show today, you have the Plod wristband. I hold and press. I get a nice little haptic. Boom, the red light goes on. And now when I put it into my charging tray, here's my little charging tray, and here's my backup Plod.

4:11You should buy two. That's my best advice. Have two, one in your bag, and then one at your desk. Any meeting you do, you're just one click away from having the note taker on. And this note taker is independent of all other platforms. It does beautiful summaries. I am addicted to this. You know, the big test for me, Alex, it's like, this is called the key and the wallet test. The key. Oh, I know where you're going with this. If I leave my house without my keys or my wallet now in today's day, it's your smartphone and your pouches, your nicotine pouches. Those are the two things I got to turn. No, I don't turn back for nicotine pouches, but I do turn back for my plod.

4:50I literally put the car in park, walked back to the house, put the car in reverse, went back into the house, and I got my Plot pin. I think it's fantastic. If you want to have a good recording of all your conversations and automatically taken notes, then you should go to plod.ai slash twist. P-L-A-U-D dot A-I slash twist. Use the code twist, Jason. Save 10 % on your purchase. We're big Plot fans. Shout out to them. Let's get started with the show. I want to go deeper on this. And, you know, our pillars here on the show, show don't tab. We like to have great demos. And we like experts. Experts only.

5:24No offense to our journalist friends. The top 20 % do a great job. The other 80%, we all know, you know, maybe they're trying their best. But we like to have experts on. Experts directly from the horse's mouth. Here we go. We got Rob May from Neurometric. Neurometric builds small language models in contrast, Jason, to large language models or LLMs, things we talk about so very much. We're going to talk about SLMs in a little bit of time. But Rob, I presume you've had a chance to read through the Anthropoc Mythos card, read the Red Team report, and chew on it. So what do you think? Rob has opinions.

6:00So people know, Rob was one of my first angel investments, one of my first five for a company called Backupify. This was a genius idea he had. I cold emailed him. Is that correct, Rob? I cold emailed you and said, hey, I love your product. Yeah, you did. This was back before you were famous, Jason. So you had to explain to me who You were, yeah. Yeah, I was like, hey, my name is Jason Calacanis. I use Backupify. And this was the best idea ever. If you lost your Gmail account, Alex, you know, in the days before, this is before Google Suite, I think, existed or was just coming out. If you lost your Gmail account, Google would say, okay, you got hacked.

6:36And you'd lose your entire archive. Rob figured out a way to use the API to back up your Gmail box and your G drive and, and, and every other service you can think of. It was an amazing vision, and it was a great company. And we had a nice little exit. I backed every company Rob's done. And he's got a new company, so I'm excited to hear about it, catch up with Rob. Rob also ran the Open Angel Forum for me in Boston, yeah? Yeah. And actually, I think I was one of the first people to talk to you about AI, Jason. Back like 10 years ago when I was talking to the incubator and telling the story. And you and I had dinner with one of your friends, Jeff with a G, and we stayed up late talking about AI.

7:14Yeah, good memories. and uh yeah this is i'm gonna be in a old age home alex and i'm gonna be like welcome to episode 20 122 of this week in startups we're here on twist and here's rob may he's gonna be like in a wheelchair like professor x but it's gonna be a levitating floating one and it'd be my brain in a vat of uh back to tank and i'm just gonna be 120 years old still doing this and loving it let's get into it. Mythos is the preview of the new anthropic model. Here is Dario and this video. I don't know if you see this video yet. I don't think so, no. Okay, we'll get you to react to it for the first time here.

7:55Here is the anthropic team talking about why they're withholding the model. There's a kind of accelerating exponential, but along that exponential, there are points of significance. Claude Mythos Preview is a particularly big jump along that point. We haven't trained it specifically to be good at cyber, we trained it to be good at code, but as a side effect of being good at code, it's also good at cyber. The model that we're experimenting with is by and large as good as a professional human identifying bugs. It's good for us because we can find more of an and we can fix them. It has the ability to chain together vulnerabilities.

8:38So what this means is you find two vulnerabilities, either of which doesn't really get you very much independently. But this model is able to create exploits out of three, four, sometimes five vulnerabilities that in sequence give you some kind of very sophisticated end outcome. And we think that this model can do this really well because we notice that this model is very autonomous. It's just generally better at pursuing really long range that are kind of like the tasks that a human security researcher would do throughout the course of an entire day. Obviously, capabilities in a model like this could do harm if in the wrong hands.

9:14And so we won't be releasing this model widely. More powerful models are going to come from us and from others. And so we do need a plan to respond to this. That's why we're launching what we're calling Project Glasswing, where we partner with a number of the organizations that power some of the world's most critical code to put the model into their hands, to allow them to look at how they can use models like this to bring down risk and protect everyone. And by giving these software developers advanced tools before anyone else, it gives all of us a collective head start. It allows us to find things that we couldn't find before, and it helps us fix these things much more quickly.

9:59Working with our partners, we've been finding vulnerabilities across across essentially every major platform. I found more bugs in the last couple of weeks than I found in the rest of my life combined. Dario said he's got very big concerns, and the reason he left OpenAI was he felt Sam Waltman was not trustworthy. I'm not, you know, piling on Sam here. There was a big New Yorker story. But the truth is, Sam drove a lot of people out of OpenAI, and there were trust issues, according to those people. Again, I'm not editorializing here. Everybody kind of thinks I'm Team Elon, which is fair enough, but I don't have investments in any of these companies, so I'm not talking my book or anything.

10:39But the truth is, he drove Dario out. Now Dario is passing OpenAI in models, profile, PR, influence, and dare I say the revenue here. So your take, Rob. If you've got an engineering team at your company, I'm betting there's a solid chance they're spending far too much time on infrastructure. You need your team building your product to delight your customers, not configuring your virtual network. Render is the all-in-one cloud platform for developers that allows you to deploy, scale, and secure your apps and agents with zero ops. Most cloud platforms ask you to split your focus between product and infrastructure, or they force you into platform constraints that you know you'll outgrow in six months.

11:26But just connect your GitHub repo to Render and you are live. L-I-V-E. Web services, cron jobs, manage Postgres, the whole stack in one platform. It's time to find out why 5 million developers are already using Render. Go to render.com slash twist and apply. For the Render startup program, you'll get anywhere from$500 to$100 ,000 in free credits, depending on your stage and who your backers are. That's render.com slash twist. A lot of stuff in there. Take it wherever you want to go. Yeah, I think Anthropic has significantly passed OpenAI on a lot of things. And I think the reason is that they've been more focused.

12:09Like what's happened to OpenAI, to focus on them for a second, is that because they were first and they've had to raise so much money to build these models, Sam had to go out and sell that they were going to enter. They needed a$10 trillion TAM, which means you've got to be in every market, right? And so I think that's their problem. They've been trying to do a little bit too much. With respect to this model, I do think it's interesting and I like their approach. I think we were talking a little bit about at the beginning of the show about, you know, they're going to IPO this year. So this was a very well put together video that I'm sure they had the IPO in mind as they started to do this.

12:46But that said, I actually I like the approach. And what I like about it is I think we've been a little bit, the Silicon Valley vibe on AI for the last couple of years has been like whoever hits AGI first wins, right? Because that machine extrapolates and gets better. I think what we've seen over the last 18 months that surprised everybody is that the parity amongst the top labs, you know, and including Google and everybody, is really what it means is when we hit AGI, like, we're all going to have access to superintelligence for free and open source three to five months later. And so I like what they're doing because they're preparing us for that day.

13:25And I think that's super important. Rob, do you think that the open source... Alex, I want your opinion, actually. Having watched this and handicapping it, you've been deep in this space and Cautious Optimism is a place to go look at that. CautiousOptimism.news? Yes, sir. Yeah, okay. So I read the newsletter, everybody should subscribe. You've been obsessing over this. Your take on Dario being right, Dario being driven out of open AI, Dario focusing on code first as opposed to side quest and consumer, and them surging ahead of OpenAI. I'm surprised at the speed we went from Anthropic is catching OpenAI to they seem to be tied to Anthropic is winning.

14:08October, January, and now April. It's pretty crazy how fast things have changed. Anthropic has gone from probably around 10 billion ARR in last October to like 30 now, which is just incredible. Unprecedented. Unprecedented to the point at which I don't even really understand the numbers. What I will say, though, about their decisions and product terms is that I think product market fit is the ultimate arbiter of an entrepreneur's success. And clearly, no one has more PMF than Anthropic today. Regarding the future, and Rob mentioned that the video is aimed kind of at investors in the IPO, talking about future capabilities, future capacities.

14:46There was an interesting interview between Greg Brockman and I think it was Alex Kanterwitz over at Big Technology talking about models. This is right before Mythos came out. And they were discussing takeoff and how these models are now getting a little bit better self-improvement, working on themselves, writing their own code. And so I think we're seeing the tailwinds of that. Though the thing that I'm unsure of, back to Rob's point, is how quickly the open source world will catch up to Mythos. Because today, Meta dropped their latest model and their new family. Shout out to them for pulling that off.

15:17And it looks pretty good compared to everything that came before Mythos, but it's not now state of the art. So if the proprietary labs are here and then the second tier players are here and open source is here, I'm curious if it's more than three to five months until they catch up. And if so, we have more time to fix the world, Jason, to find all these vulnerabilities and patch them. Because to me, this is at once the best tool in the world for cyber defense, find the bugs, patch them, and also the best tool for cyber offense, find the flaws and then abuse them. And so we're going to be in an arms race, I think, until we're all dead.

15:48Let's take a look at the anthropic polymarkets. This is a way for us to really understand how polymarkets actually doing. There's so many to choose from. Here's the URL of all anthropic polymarkets. So let's pull this up. The first polymarket that comes up under the anthropic tag over at polymarket. And you can basically find a keyword there and it's polymarket.com slash slash prediction, slash Anthropic, right? The first one, Anthropic$500 billion valuation in 2026, 95 % chance. Everybody realizes that's going to happen. Anthropic Claude score on Frontier Math benchmark by June 30th, 72%. And then there's the Anthropic IPO closing market, just tons to go here.

16:39I think the one that we want to know is when are they going to release this model and this model is called mythos and so to walk us through this one here because i don't see a chart on this one yeah it's not charted some of them uh that have fewer kind of like endpoints aren't charted per se but in this case we can see jason that there was people betting that it might come out before the end of march that has now of course lost um because there was reporting about this model via fortune and a leaked blog post a couple weeks ago if people remember uh now uh april 30th i would say zero percent chance polymarket sharps say seven percent jason but the thing that really caught my attention is this june 30th they're only handicapping a 28 percent chance which means three out of four times we won't have mythos out in the market by the start of july that's a couple months from now or two-thirds yeah two-thirds because you got to put the first two numbers together i think oh right of course two-thirds chance you know um it's not out until yeah sometime in the summer or the fall which in ai terms is is years that is that is so much time in this present moment.

17:41And the thing that I'm not sure about, and I'm really curious to know is all of these companies that have Mythos now and can play with it and can use it, can put it to work hardening their software. How much progress can we make in that interval? And will it be half the work we need to get done? Will it be all of it? I don't know, Rob, do you have an idea of how fast we could put a model like this to work in fixing all the code we depend on day to day? The problem is we're generating code faster now, right? So with all the vibe coding, I think, I think I read that 20 or more than 25 % of GitHub commits now are, are vibe coded.

18:13And so as that's going to go up, it's like, you're, this is all going to boil down to, so it's, it's, I know you love poker, Jason. And so like the world is going to boil down to poker, running a business in this AI future is going to be about estimating probabilities of things happening, knowing the cost to go run those probabilities to ground and figure out what they really are. And like, so, so this is going to boil down to using your compute to figure out like, you know, code generation versus code checking. I mean, Anthropic has the main code generation platform, sort of the number one right now.

18:47So maybe they combine these together in ways that write better code. But, you know, if you're talking about it, not coming out till June 30th, I would not be surprised to see an open source model like Quinn or Kimmy come out with something that's similar, GLM maybe, like before that date from somebody who doesn't care as much about. Okay, so this is a key point that you're making, Rob. What if a bad actor has already achieved this? Like, China may have already accomplished this, and they have no incentive since every company in China is owned by the CCP, which would be the equivalent of, like, the CIA having a board seat in every single company, and there's like four CIA people and FBI agents and Department of Justice, whatever, inside of Anthropik and OpenAI saying, okay, not only are you not making that promotional video, we're taking this and we're going to hack North Korea, Iran, whatever bad actor we want.

19:44China might already have this. Everything could be completely compromised at this point. And I that's when Americans who were debating, you know, David Sachs and, you know, the administration and are we hand-wringing too much here that America has to win the AI race. Folks, it is existential who wins this. It's 100 % existential if these tools give you the ability to hack everything. Here's a hot take. If your bank moves slower than a startup, that's a problem. I see this process behind the scenes every day. And I know that bad banking can kill your company's momentum. That's why I'm so glad to introduce our newest partner, Grasshopper.

20:27Grasshopper's a real federally chartered digital bank. Not some fintech rapper sitting atop some mystery institution. Nope, it was built just for founders like you. You want fast? You want easy? Open an account in just minutes and start earning yields that can top 5%. Wow, that's a big number. Plus, you'll get unlimited 1 % cash back on purchases, free ACH, free domestic wires, and no monthly fees. Plus, if you're sitting on some real runway, Grasshopper's treasury product hits 5 % plus with same day liquidity. As a Twist listener, Grasshopper wants to give you a$500 cash bonus just for opening account.

21:04And you can open an account really quick. So go right now to grasshopper.bank slash twist and use the promo code twist to get started. I think we should consider this mythos model, this mythos preview to be essentially a cyber weapon and perhaps a cyber weapon of mass destruction. I mean, maybe we need a new term for this because I can't recall apart from certain hacks in history, a point in which people were worried about all of software having potential vulnerabilities. And I think Jason, it's a little bit weird that we're talking about a, you know, three to five month gap here of times we might have ahead of China, but I would so much rather have us have that time to work to get things secure than to be catching up for three to five months until our companies were of a sufficient quality to match them.

21:42So this show sheds a new light, Alex, on the Emil Michaels appearance on All In three weeks ago, maybe where he was talking about the anthropic conflict and like, are they going to ban it or whatever? These guys must have made peace right now. And I wonder if Dario and Emil, when they broke bread and we're trying to work this out. I wonder if Dario said, by the way, our new model would give the CIA the ability to hack North Korea, to hack China, and we are patriots at Anthropic. And instead of giving the tool to a bunch of security researchers, we're going to give it to the US government. Because if you made this the equivalent, Alex, of the race for the atomic bomb, no American citizen would be like, you know what we need to do?

22:38We need to keep the atomic bomb to ourselves, and we're just going to delay making it. Oppenheimer would say, we need to have this for ourselves before the Nazis get it. Period. Stop. I don't think I'm out of line here to say that this is becoming the equivalent. This is becoming the equivalent. This might not seem as much because a nuclear bomb can cause, you know, such a mass destruction of life. But this could cause a massive financial, you know, devastation across the economy. We talk about the GPT-3 moment a lot. I think that's like nuclear fission in your analogy. And then mythos would be the introduction of the hydrogen bomb.

23:18But Rob, I think I may have cut you off there. Yeah, like, can this thing figure out how to hack bank software, right? Like Swift and international transfers, like who knows? Bitcoin, like, you know, it's going to be crazy to see how, because we do know, like, if this capability exists, I think you're right, Jason. I think probably some other places have it. Like, my guess is Google has something like this that they haven't announced or released. Like, I still think they're in front of Anthropic, from what I can tell the people I know there and the tools and technologies that they have. And so, and we also know that last year, I think for the first time, China passed the US in research papers on AI accepted into top tier conferences and journals.

23:59And so, like, there's a really, really good chance, Jason, that I think your point of view might, like, this might be going on behind the scenes that we don't even know. And it's interesting to think about the run-in that they had with the DOD a couple of weeks ago, five, six weeks ago, whatever, with Anthropic, right, was sort of like, why don't they just use GPT-5 or whatever? Like you don't need Claude's model, but it might have had to do with this and they knew it was coming and were discussing it. And this was one month ago. So now we start doing game theory. I think the CIA and the government are inside of Gemini, Anthropic, OpenAI, XAI, talking to each of these model folks.

24:36I think they've probably been in there for a year or two. And they are saying, hey, are you a patriot or not? Are these tools? are you going to hold them close to the vest or are you going to tell us exactly what's going down here and then you turn over another card alex if this is in fact true what they're saying and if they present it as such if dario is presenting this as it's cataclysmic you know the entire economy could go down and we believe him we take him at his word there's an argument you have to nationalize this technology there's an argument it's too powerful for a private company to own this.

25:13It would be like if a private company were to stumble on a bioweapon of, you know, such or, you know, this weapon that we were using, the disorientating weapon. If a private company gets to that level, hey, we've got a weapon that makes you come in like a superhero and you can just Professor X the entire other army and just make them all grab their ears and blood starts coming out of their nose. I hate to be graphic. You have an obligation to go to the president and and say, Mr. President, this private company has this. And he's like, okay, great, let's go to Venezuela. Let's handle that problem.

25:48There's some equivalent here. I don't mean to be hyperbolic. I'm just telling you game theory. I don't think Dario's lying. I think Dario's being sincere. Now, I know everybody hates Dario. The right hates Dario. Dario wouldn't bend the knee to Trump. Dario did not donate to Trump like$25 million that one of the OpenAI CEOs or co-founders did. he's not loved by this administration, but they love his tech. And then is he a patriot? Is he like Alex Karp and Palantir? Or is he a hippie dippy who says, I don't want my tool to be used for this? This sheds it in a totally different light. And there's some conversation going on here that we're not privy to between the president of the United States, Emil Dario, the CIA, the Department of War.

26:38This is cataclysmic in its severity. Yes. And this is a super weapon. I want to point out just one thing to back up what you're saying, Jason. Thomas Friedman, not someone we usually quote here on the show, but he wrote a post for The Times and he says, this is not a publicity stunt referring to Anthropik's position here. Quote, in the run up to this announcement, reps of leading tech companies have been in private conversation with the Trump administration about the implications for the security of the US and other countries. So not only is Anthropik saying this, people they've shared the model with are going to the White House and saying, holy crap, ring the alarm, because suddenly everything is potentially insecure now.

Read the full transcript

27:14Thomas Friedman then is confirming, and I didn't see that story, so great, Paul. He's confirming what I believe to be occurring right now. And by the way, this is not, I'm not in the Trump administration. I probably could have been, and I could have been on one of these projects. I didn't. I demurred. So yeah, well, it's just a choice. I'm an independent. But how do you think about the idea that the government needs to take this over at a time when the public's trust in government is probably the lowest it's ever been on both sides of the aisle, right? Republican and Democrat. I mean, doesn't that present a really interesting sort of wrench in the game theory?

27:52Nobody trusts anybody because we're in literally, what was the show at David Duchovny? The X-Files. Trust no one. We've caught up to the X-Files. We've caught up to the X-Files. Those people, you know, believe that the truth is out there and you should trust no one. Those were like the two signature concepts of that, you know, that show was 30 years ago. You know what? Here we are, folks. This could be the thing that galvanizes three different groups of people who have the least trust in the world. Journalists, AI, and the government. These three groups are not trusted anymore. and for good reason.

28:34You can debate it. But the X-Files in 1993, they nailed this. The truth is out there. And the truth is, this maybe could galvanize these three groups of people. Thomas Freeman, the New York Times should be saying, hey, if we have any information about this, do we put this out as a public story or do we go to the White House? Do we go to President and Trump and Emil Michael and say, hey, by the way, we think this is real. Dario, Elon, you know, Sam Altman, Sergey Brin, they all must be having some sort of Zoom or private conference here to unpack this. And how could this be used to stop North Korea and their inter-ballistic missiles, which they have, and their nukes, which they already have, as opposed to Iran, which has none of that.

29:23They're attempting. But North Korea is far ahead. There is a nuclear threat out there from a bad actor who's insane, Kim Jong-un. He's insane. And he has a ballistic missile that could reach California. And he has 30 or 40 nuclear warheads, they estimate. Yeah. But what's more effective now, the threat of nuclear deterrence or the ability to take a model and break everyone's entire nation? So to me, it's funny that I think I'm actually more worried about the second category than the first, because the first seems to be more about saber rattling and keeping your country from being attacked. But this gives them more useful, offensive capacity, Jason.

29:59That's terrifying. We need to take 20, 30 percent of the cycles of these companies. The government needs to say, hey, we'll pay you for your time in not releasing these models. And we're going to create an Oppenheimer Manhattan Project. We're going to create a Manhattan Project to sturdy the infrastructure of the United States. Okay, now let's take the other side of this coin. We're all kind of in broad agreement here, but here's the downside to all this caution. And I'm going to pull it up right now. Hiring can be its own full-time job. And hey, guess what? I already have a full-time job. I make podcasts and I invest.

30:34But when you're running a small company, we both know every hire matters. You don't want to waste any of the seats you have at your company. And the best part you can have is LinkedIn Hiring Pro. Why? There's a billion people using LinkedIn. All the great talent are there. If you're proud of your work, you build a LinkedIn page and you update it. LinkedIn Hiring Pro is going to streamline and simplify the entire process for you. Nearly 60 % of companies using LinkedIn Hiring Pro. You're going to get an incredible candidate to interview in the first week. And we're looking for a new producer for the pod.

31:04We did shout outs here on the show. We posted it on my social media. We asked friends. You know where we found our next great hire? LinkedIn. And it was competitive. We had like three or four really good choices. So hire right the first time. Post your first job and get$100 off towards your post at linkedin.com slash hiring pro offer. That's linkedin.com slash hiring pro offer. Terms and conditions apply. This is an excerpt from the Mythos previews systems card, I think. And essentially what it shows if you're on the audio version is a dramatic increase in score for Mythos across a number of very important benchmarks.

31:42These in particular, Jason, are software coding AI benchmarks. And as you can tell, you know, on SWE Bench Multimodal, we went from 27 % from Cloud Opus 4.6 to 59 % with Mythos Preview. So what we're doing is we're also saying we're going to slow the pace at which AI models that we use for day-to-day tasks down. We're going to retard that function dramatically. And I don't think anyone else in the world is going to do that? Well, I don't know that you have to slow down the model creation to just make a SWAT team from each company and say to each company, we want your five best brains on cybersecurity available.

32:23We're taking five from each company. It's going to be 25 of these. Maybe the smaller models can contribute one or two people. And they're going to be the brain trust that then builds a system, this is my proposal. If this is all true, which I'm 98 % sure that Dario is not being hyperbolic to raise the stock price. I'm going to take him on his word. If it's true, we could create a piece of software. We could create a new unit of government that then goes and says, we're going to do our own red teams to try to turn off or overpower this nuclear power plan and have it melt down. And we're going to use these tools to see if it's possible.

33:09And then we're going to fix it. And we're not going to talk about it. We're not going to talk about it. So we're going to know that this group exists. Yeah. I mean, just a group of people who goes and says, what are the biggest vulnerabilities? And this is where immigration of the most talented hackers in the world matters. We should also be trying to get every single AI researcher in China to defect. And we should pay them a million dollars to defect with their families. We should figure out when they're on vacation in Singapore, or if they can get out of the country to Australia to go on a holiday, we should figure out how to pick them up and let them defect, which we did during the Cold War with Russia.

33:43We need to get them to defect and bring them to our team, and then put them in an air-gapped kind of space, so we know they're not double agents, whatever, monitor them like crazy, we should be in a talent war to try to recruit these people out of these countries and then put them on our cyber security team. Yeah. I bet you this is happening covertly. I bet you this is happening right now. If it's only a million dollars a pop, that's the cheapest thing we can ever spend as a nation. But I, you know, look, I don't like to talk about it much, but I mean, Jason, there is a lot of rich people in tech, right?

34:16So why doesn't one of them just say the first hundred million is on me? Well, sure. I mean, this could come any number of ways. if the government's adding$1.5 trillion in spending for the military, you know, you got to think like we could put$100 billion from that spend into this very delicate area. It might mean that we need to have the government building this tool. We need to have a fork of it, you know, given to, and this is why that discussion where Emile said, anything legal, we should be able to do with the tool. We'll buy the tool from you. And, you know, kind of Alex Karp's position is, hey, we make a tool.

34:56You decide how to do it. I think, you know, Andrel has the same position. Hey, we make these tools. It's up to the government to deploy them. It's not for us to tell them they can or cannot use it. And, you know, then we have to trust the government to not use this in some kind of crazy, abusive way. This is going to be the story of the next year. I'm going to predict it right now. And it's going to get political, but this should be a galvanizing moment for America. I'm getting a call, though. We're getting a call to the show from a dear friend of ours. I think we can pull. Ah, there he is. Nick.

35:27Nick, are you OK? Nick, are you OK? It's my guy, our correspondent, Nick in. He's in the Brickle, I think he's in South Beach. Edgewater. What is going on? I see you have something on your forehead here. What's going on? Yesterday, I heard him mention that I was sponsored by somebody and that I wasn't disclosing it. And I just wanted to say, I don't know what you're talking about. It was brought up about Higgs field paying for these things. I generally don't know what that is. I am in Edgewater though, by the way, not Brickle. Okay, so you're in Edgewater and you can confirm for us that as much as Higgs field would love to partner with an influence of your, you know, international fame, dare I say, crypto circles, technology circles, finance.

36:19I mean, even in the artistic community with your, you know, adjacency to NFTs and the Miami art scene, you are confirming Higgs field is not sponsoring you, Nick. You can confirm right now. Definitely not sponsoring me despite the release of their new model, which integrates with C dance 2.0. They are definitely not partnering with me by the way, that doesn't work in the U S but they have not partnered with me, not paid me anything, nor has actually, and I appreciate you asking me about, you know, what neighborhood I was in because that brings up to me the, you know, LinkedIn jobs, which has nothing to do with me.

36:55And they're at linkedin.com slash twist. For example, I've never worked with them and I've never been paid for any of these things. Okay, great. And just to be clear, the promo code, Nick O, getting 25 % off, not a true story. Not true. Never worked with them ever. Never worked with them. never worked with um yeah i just happen to be a big fan of these sort of products and services but i just happen to be independently wealthy and don't need any uh cash flow from any of these brands right right right nick i gotta ask though um i'm thinking about getting a new tattoo i only have one and i see you have a tattoo on your forehead that looks quite sharp um what where'd you get it done i think it's a little dirt you might have a little smudge over here oh i didn't I didn't even notice that.

37:38That's crazy. I had no idea. What? Oh, I had no idea. Nick, we appreciate you. And thank you for clearing this up. Everybody wanted to know. Everybody used the promo code NICO for 25 % off. Your Roman Sparks. If you need Roman Sparks, Nick doesn't need it. He's all man. But if you need Roman Sparks, fully functioning, you can actually get that Roman Sparks, 25 free Roman Sparks, if you use the Nico code. Well, well done, Nick. Appreciate you. There he is, our roving correspondent. Yeah, coming hot from Miami. Also, I just Googled Roman Sparks on my work computer. Am I going to get fired? I didn't know what that was.

38:18No, it's totally fine. Roman Sparks is a delightful little lozenge that Rob will fill us in. Rob, you're a man of a certain age. I don't know what that is either, but I'm making some guesses. It turns out it's a high... Yeah, I Googled it. Anyways, I'm going to grab this by the tail and drag it back on topic here. Rob, very glad to have you here. Neurometric is the company, and you made me go out and learn stuff. I'd heard of SLM's small language models, but I wasn't sure what the parameter cap was, what they're good for. So first of all, what is the parameter cap, the differentiation point between a small language model and a large language model?

38:56And to put it politely, why do we care? Yeah, so the cap keeps sliding, right? Just like the LLM cap keeps sliding. So Mythos, which we were just talking about, is a 10 trillion parameter model. So when that goes up, what's small in comparison is, you know, changes. I would say in general, people think about small language models as something you could probably run on a high-end laptop. And that's sort of a rough. Is that 10 billion parameters? Is that 20 billion parameters? 10 to 20 billion is normally sort of the cutoff these days. My guess is pretty soon people say anything under 100 billion parameters is small.

39:29But here's why you should care. because just as the intelligence density, so the intelligence density is the amount of things the model can know for the parameter size it is, right? That keeps going up for LLMs. Part of the way that these models learn more is they develop, you know, synthetic data training techniques and reinforcement learning techniques and new, you know, architecture tweaks. Those filter down to the small models. So like an 8 billion parameter model can do more this year than it could last year. And so if you're not trying to hack the world's entire code base, right? Correct.

39:59If you're trying to build a model that reconciles a bank statement with your QuickBooks or predicts customer churn in your customer success funnel or things like that, these small models, as they climb up in their capabilities, we predict that by 2030, 90 % of common work tasks will be able to be done by like a 10 billion parameter model or smaller. So what impact does everybody having the equivalent of today's Mac Studio, or I just upgraded to a Mac Pro 14-inch with 48 gigs of RAM, I could clearly run an SLM on a Mac Studio when you get to 256 gig, 512 gig, you can start running Kimi and other things, I guess OpenSeek.

40:43Those OpenSeqs and the Kimis, those are not SLMs. Those are straight up open source LLMs. They require server level memory, server level CPUs. But people are starting to do SLMs on their latest iPhone. They're doing it on their latest laptop. So this will be embedded into every device eventually. Yeah, yeah, it definitely will. And what's interesting about them is, So if you look at what's happening with large enterprises that are AI forward, which is not very many people right now, but if you look at companies that have deployed AI for a couple of years now, what they typically do is they start with a frontier model, right?

41:23You know, Anthropic or OpenAI or maybe one of the open source ones. And they run all their tasks through it. And as their inference charges climb up, they start to go, huh, well, what are these tasks can we put to smaller, cheaper models? because the bigger models are more expensive to run because when you run a layer of a neural network, you have to shuffle it in from memory into the compute, calculate it, write it back out to memory and save it. So bigger models take more memory, they're longer to run and everything else. These smaller models, you can run on lower hardware, you know, hardware that's six, seven, eight years old, that's refurbished.

41:56I mean, we even joked at Neurometric about like, can we just buy a bunch of phones and like line them up, connect them to the internet and run small models on them and serve those for people, right? But these small models, there are things you can do to tweak them per task. So if you need to do a finance task or a sales task or whatever, you can make them really, really good at that one task. How do you tweak it? So if I wanted to, say, take my photo archive and have all of that managed locally, I didn't trust Apple, Google Photos, whatever. But I just wanted all these photos on my desktop, on my laptop, and I wanted to tag them all.

42:29And I wanted to say, hey, this is this person's face without having that go up to a frontier model. So now they know, oh, that's Lon in your photos. That seems to me to be like a great way to tweak this. Or if you just wanted it for writing or you were an Excel jockey, you would want to just have one tweaked for Excel. Are those models already pre-tweaked? Can you get a flavor of that model on GitHub or on Hugging Face? Yeah, for some of them, you can get a flavor of it, but the flavors tend not to be work task related. They tend to be things like Quinn has a series of SLMs that are called Instruct.

43:05And so they do instruction related following tasks better than, say, generative writing tasks or something like that. But there's two main ways to sort of like make these specialized. One is to augment your prompts with a whole bunch of stuff. That can get hard with the context window limits and everything else. But, you know, you add the stuff like imagine you're a CPA and you have all this knowledge, you know, and you can throw the whole gap, you know, accounting standards in there into your prompt or whatever and pass it to the model. The other way is to sort of fine tune it. The best way to do this is what they call distillation.

43:37And distillation is when a bigger model teaches a smaller model something. So you can think of us as like setting up a system where, let's say you have a task that you're having Anthropic or OpenAI do. And maybe that task is like, hey, I have a thousand pages here from industry reports on the energy industry. And I'm an investor. So GPT-5, take this report and pull out all the stock symbols in here because I want to go research them, right, and see what they should do. So if you see that information go in with the prompt and you see GPT-5's response, you get that back enough times, right? that prompt and response is a data set you can use to distill a small language model just to do that given task now depending on how much you do if you only do that task once in a while it doesn't make any sense but if you're like we do this every day a thousand times a day for all these pages of reports you could probably save 90 percent by doing that and in fact there's a story in venture beat in february about at &t got to the point where they were spending eight billion tokens a day on their AI infrastructure.

44:34So probably a couple hundred thousand dollars a day. They re-architected everything using Frontier models for 10 % of the tasks, SLMs for 90 % of the tasks, and dramatic, dramatic improvement in both speed and cost. Yes, the headline from the Ventribute story is 8 billion tokens a day forced AT &T to rethink AI orchestration, cut costs by 90%. Rob, can I just narrow down and focus on one of your SLMs really quick? Because I think it's a good example. So here is one you guys built called DealSiv. and it's a task-specific model that compares target company profiles against a firm's investment thesis, pretty pertinent to what Jason does here at launch.

45:08And it's based off of Quinn, three to four billion instructor, 2007. Okay, so is this one that you made via distillation? How did this one in particular come together? So it's a good question. We have an automated process behind the scenes that does a lot of this. So I don't know on this specific one. But look, our bigger goal as a company is just make intelligence free, right? Why would we want to do that? Because it's the Jevons paradox thing, right? The more the cost comes down, the more people are going to use it for more tasks. I think the better the world will be if good people control AI.

45:39And so people are like, well, then what do people pay for? Well, you're going to pay for an SLA. You're going to pay for testing. You're going to pay for analytics. You're going to need stuff to manage all this intelligence. So yeah, you can go and you can download these models. We give them to you for free if you want them. What's the website again so people who are listening can go there now and try it? it's just a marketplace.neurometric.ai. And then today, do you want me to talk, Alex, about the claw pack? We're going to talk about the claw pack. One second before we get to that, Jason, because you said something very important, which is you want to make intelligence free.

46:10And what's incredible, Jason, about that comment from Rob is that he's not that far off from it already. They offer 100 million free tokens a month if you want to have Neurometric handle your inference. And it's two bucks a month per model per month, no token limits. So, Ron, how the hell can you afford to do that? Are these so cheap that effectively they're already free? Well, you can run these on old hardware. And then the second thing is, you know, we're a seed stage company, so we're learning the real usage distribution. But, you know, at the beginning of the show, Jason talked about Backupify.

46:42Backupify offered unlimited storage for, you know, Google Apps, Salesforce, Office 365. And people would be like, how can you offer unlimited storage for$3 a month? Because we knew the distribution of what people actually used. And I think what companies want here, and particularly prosumers and open cloud users and cloud code users, but also enterprises, is they want to not have to think about all these innovations that keep coming to drive down the cost and how they apply them and how they roll them out. So I think if we can manage all that and we can bear the token risk, I think it's a great business model.

47:10But you're not doing the hosting or you are doing the hosting? We will do the hosting. We don't have to. You can download it. We can deploy it wherever you want, manage it for you. But if you want us to host it, we'll do it. I think this is going to become the key. anybody who goes deep down the open claw or even perplexity computer, which is awesome, or Claude Cowork, at some point you're like, am I getting enough value from spending$1 ,000 a day, $365 ,000 a year? Now, if your business is printing money and you don't have to hire the fourth developer, you know, okay, fine. But in other circumstances, you're going to be like, Like, you know what?

47:47This task, you know, I have tasks I want to do, which I'll talk to you offline, Rob, that are not necessary for my business. But, you know, we have to judge the startups that are coming in. But I would love to be running backtesting on every startup I've ever met with in my life, the founders, where they wound up, and be saying, okay, let's examine every 2009 startup, every 2010 startup, every 2011 startup. And tell me, look for some patterns. Look for the talent there. Which talent created diasporas of Googlers or people who worked at Uber who went on to do other companies? There's all kinds of intelligence that I could see myself using, but it's not a priority.

48:29And I would look at the$100 ,000 in token costs and say, not worth it. But at$1 ,000 or$10 ,000, yeah, it might be worth it. I think a lot of people who are at the tip of the spears here playing with this technology, they're starting to come to that realization. There's things I want to do that don't make economic sense today, but if the tokens were cheaper, I would do it. Why not? You know, Rob, if only there was a brand new product out there called Clawpack that for only$8 a month got you a loaded inference. You want to tell us about it? So as soon as we put this marketplace out, one of the things people started using it for was obviously OpenClaw because that is, man, you go on Reddit and people are complaining like crazy about their OpenClaw prices.

49:06So we said, okay, let's do some research on what are the top open-claw use cases, and let's take 39 small language models that we package together under one API. So we host 39 models for you. They do common tasks like social media posting, right? Like you don't need Claude Mythos to write an email headline, right? Subject line. So it does all this stuff. You get 100 million tokens for free on it. And then after that, you pay$8 a month, unlimited tokens. We manage the token risk. And there's a lot of ways to do this. The other thing, Alex, that we haven't talked about and you didn't bring up, but there's this new emerging concept in tech called harness engineering.

49:43And harness engineering is the thing you wrap around the model to make it do things. And one of the things we noticed in our research on small language models is one of the reasons you can't use them for complex tasks is they kind of tend to go off and lose track of what they're doing where the big models don't. But if you put a nice harness around it that helps it stay on, that says stuff like, hey, every time you do a step of a task, check back in with me and make sure you're on the right task. Like those kind of harnesses are easy to write. Cloud Code can write you one for a specific task and you plug in a small model and run it and that harness keeps the model on task.

50:14So we have a bunch of innovations like that that enable us to, and we'll just keep driving down the cost and other people are doing stuff to drive down the cost. Rob, on the harness point regarding SLMs, is this like a dog sled? Do I have like one chain of harness that has multiple dogs, aka SLMs, that I can work with? or is it like one harness per SLM? It's both. Today, it's mostly one harness per SLM because they're task specific, but you can already see with the claw pack, right? Yeah. They're starting to work together. And what you're going to start to see is swarms, I think, of SLMs that'll do common work tasks.

50:52And you'll always need 20 % of your workloads to fall back over to the frontier models because they're one-offs or you get lost, you don't know what it's doing. So you're always going to need a combo ensemble system. But I think we can take people's core workloads and drive that cost way, way down. Yeah. Jason, the power of branding. Remember that moment in time in which we all said, oh, if you're an AI rapper, you're doomed. But now if you're an AI harness, you're the tip of the spear. Well, you know, there's like two interesting observations here for startups. One is every Reddit bitch thread where people are bitching about something they hate is a potential startup.

51:33This is like a super important thing. Somebody needs to create for me a skill for my open claw to just look for these startup opportunities from people saying this product sucks. And I would like, you know, why can't they solve this very simple product problem? I would pay a lot of money for a tool that did just that. Because it would be like the request for startups that we do, YC does, other people do. It'd be like, these aren't my requests for a startup and my intuition as a VC, 55-year-old guy in Austin, Texas. What I want is irrelevant compared to what the world wants as demonstrated in a subreddit about music where somebody wants a specific tool to do a specific task and a thousand people participated in that thread.

52:20That's really an interesting, and it's interesting to go to market movement because you can go to that thread and say, hey, I built something. Would you test it for me? Would you try it? And then people are like, wait, you made me a custom piece of software? Okay, great. And then startups now are not about who can build the product. It's about who doesn't stop building the product. Startups aren't about who has the resources to build the product. it's about who will not stop building that product and refining it in other words it's now a test of your resiliency your passion for the vertical will you just keep working because anybody can build anything then well who's going to build the next three or four features of this you know uh meditation app or this you know um fitness app whatever it happens to be you know this enterprise piece of software.

53:14It's really just who's willing to go on that product march for 10 years and not give up. I think Aaron Levy is the example of that from the SaaS era. But I think we're going to need to find out who that is for the AI era. We're talking a lot, though, Jason, kind of around this Marc Andreessen tweet. He agrees that you can't spend$1 ,000 a day on OpenClaw. And he says it's actually heading to$10 ,000 a day if you really want to have a magical experience. And then he says the future shape of the entire tech industry will be how to drive that to$20 a month. So, Rob. Yeah. How far can SLMs get us?

53:44Can they reduce our open claw spend by 80 % in time, 90 % of time? And how quickly can you cut my bill down because my wife's not happy? Depends on what you use them for, obviously, right? Because there's some tasks SLMs can do and there's some that they can't. But the important thing is people are going to use more tokens every year and SLMs capabilities are going to increase. And there's a bunch of costs that are also making it more efficient to run those every year. So I would say today for an average open claw user, probably 70 % reduction. You'll still have to use Quad for some things or OpenAI or whatever.

54:16But you're going to have this weird paradox, right? Which is you're going to do so much more with AI that like your per unit costs are going to come down. But you're going to spend more because you're going to turn over more parts of your life to it. And you're going to do more things. I have a prediction. Let's hear it. I have a prediction. And I'll unpack it. there is a possibility that if we are reaching AGI right now, which I've said, hey, we've reached AGI, it's just not distributed yet. It's just not implemented yet. And you believe in super intelligence and recursive learning, which everybody here believes, then there is no doubt in my mind that LLMs will get so small and there'll be so many verticalized ones because building a legal accounting, design, topography, I mean, I don't know how many steps down you can double click.

55:10There is a possibility that these SLMs being fractured, being essentially like skills, you know, instead of building a skill in OpenClaw, there's an SLM that does that skill. It's trained specifically and somebody makes it better and makes it so good that this could collapse the value of the frontier models. Because how does a frontier model sell into a law firm or an accounting firm? Like, hey, we've got this new amazing thing to help you solve these legal or accounting or design issues. When there's an SLM that the person goes, good enough. That's a good enough logo. That's a good enough font.

55:53That's a good enough non-disclosure agreement. Like, good enough happens. And then the ability to sell a note-taking app, which we saw there were dozens of note-taking apps. This was like a great business to be in. Dozens of mail-pads. Evernote was a unicorn. Evernote was a unicorn. It's a perfect example. And now it's like, well, notes, I can make one, Notion. It's just these things eventually become deflationary. This could be the deflationary moment for the frontier models. They may not realize it, but they might have created their own demise. Yeah. They may have just created their own demise.

56:34Dario is going to have to buy Neurometric. I mean, I don't know. But who's going to pay for these things if an SLM exists that does it? Or a Tau subnet solves that problem for you in a distributed computing way. This is so deflationary that we need a new word for deflationary. There's deflationary, but what is hyper deflation? Like it's not just deflation where it's like, but it's going to get 10 % cheaper a year. What if it's going to get 90 % cheaper every month? Like then the compounding deflation, the hyper deflation could be so acute that just things, like you're saying, get to free Rob. And that's just kind of a mind blowing exercise to do in your mind.

57:22well you're gonna i mean the price of compute is collapsing pretty fast as well as peep on a per you know petaflop basis or however you want to measure it um but i will say you will always come up at least against the cost of buying the i wouldn't even say gpu because there's other chips coming but that plus the energy so so there will be some floor to a unit of intelligence that'll start to advance decline slower because of the laws of physics right yes but i think as the floor rises in the quality of open source models and cheap slms for specific tasks as jason points out yes there is less room for frontier models to charge but i think also that the companies that are desperate to have an edge will because everyone else is going to be defaulting to the cheap stuff or cheaper stuff and so i think there's still going to be some market there jason i just don't think that everyone's going to be paying opus 4.6 level pricing for tokens in five years but the question that just becomes will job once paradox drive token usage up enough in that same time period to have these still be growth businesses.

58:19But I'm still pretty bullish on improving the frontier of intelligence because as mythos proves, there's still so much left to come. I don't want to start thinking about ending that run, Jason. I want to go up the damn mountain to the top. Compute performance per dollar has improved roughly 40 % per year across 20 plus AI accelerators released between 2012 and 2025. The GB300 costs nearly 9x the P100's release price, but delivers 24 times the performance per dollar. So that's all that matters is performance per dollar. 24 times, that goes to this hyper deflation concept. If you think about it as an investor, Jason, it's going to change the way that you think about defensibility.

59:04And I think one of the possible outcomes of this is it starts to enable a type of business that I don't know if venture funded, but it enables the type of business run by one, two, three people that can work in small TAMs, make$30 million a year, drop$9 million to the bottom line. Or$29 million to the bottom line. Yeah, exactly. It's going to be super, super interesting to see how you adapt your angel investing over the next decade. I've been through this before. When storage became free or was trending towards free, YouTube and Netflix became viable. Mark Cuban was like, Netflix will never work.

59:44The infrastructure is not there. If everybody had Netflix and was streaming HD, it just wouldn't work today. And he was right. But the compounding effects of deflationary technology in that case, it was the rollout of fiber. In the case of hard drives, it was what those disks could store and at what price and how it would scale. Well, in hardware scales, differently than software, but you have two compounding effects here. The AI is so good that it's making the models more dense. To your point earlier, the density of the model and the specificity of the model, that could have more dramatic effect than the hardware curve.

1:00:23But then you add the hardware curve to it. This is where I think like this 40 % cheaper or tokens are 90 % cheaper, we might be greatly underestimating what superintelligence does. It might be that superintelligence just rams this down 99 % a year. Maybe, but the humans are doing quite well as well. So Meta, their new model they dropped today, Muse, Spark, in their post talking about this, Jason, they talked about how they're getting more efficient with their compute. They said they rebuilt their pre-training stack and they said these advancements increase the capability we can extract from every unit of compute.

1:00:55So it feels like every possible vector. But why does Meta still suck at doing anything with AI? Their AI search sucks on Instagram. Nothing works. It's a disaster. Their new model is not hot garbage. I have the chart here. This came out right before the show, so no one's seen it, really. Artificial intelligence, sorry, artificial analysis on their intelligence index points out that the last Meta model we got, which was Llama 4 Maverick, which is now complete garbage compared to the state of the market, has now been replaced by Muse Spark, which is now the fourth best model. out there. Now, we don't have API access.

1:01:28According to who? What is that? Artificial analysis.ai. They run a series of very good benchmarks. And this is kind of a meta benchmark that I pay a lot of attention to. Now, we don't know how it's going to perform in the market. I'm not trying to say this is the best model since last spread, but I'm saying that it's a really pleasant surprise from a team that - But what is their strategy? Like, okay, they leapfrogged their last model. They're trying to catch up. But what is the business case here? Like, what is the user case here? I don't see a user case for what they're doing right now. like what is it actually i understand you could serve better ads but go ahead rob what i think yeah i think for them i think initially zuckerberg thought he would hurt his competitors he would hurt you know open ai he would hurt google by open sourcing stuff now i think they're looking at cogs improvement because i also know they're building a special chip right they announced this in a conference late last year that so it's just an ai recommender chip probably probably meta netflix and amazon might be the only three companies in the world that could spend the money to do a recommendation-specific chip and have it make economic sense.

1:02:27So maybe that's how they're thinking about it. Just make people that much more addicted to a product that's already so addictive that it's being regulated and banned for teens under 16. Okay, great. What a stupid idea. The worst possible thing they could do is use this technology to make an already smoking-level addiction, a heroin-level addiction worse for children and adults. That's what I'm saying is Where's the vision here? What is he trying to accomplish? I haven't heard from Zuck. There's no big change the world thing that I see. There's nothing. And you know what? This proves my point about him.

1:03:02He's great at copying other people's ideas, Snapchat, LinkedIn, whatever, Facebook Marketplace, eBay, and Craigslist. But where is the original idea from that corporation of how to deploy something unique in the world? I just, it's pathetic. On that vector, nothing to add, Jason, from mine but i did ask i asked alexander weighing uh formerly scale ai at the meta super intelligence lab um is it going to come out that i can use it in my open claw setup and he told me that they will release the api soon and it will power some claws so at a minimum okay uh we can put it to test soon enough and see if it's any good see if it's you know if zuckerberg were to put his entire energy on copying open claw i would be very nervous because he's so good at stealing other people's ideas and doing them better yeah like he would create open claw times 10.

1:03:52So yeah. Oh, that's a challenge. Let no, don't give him any ideas. All right, Mark, you heard it here first. We need, we need Metaclaw. All right, let's bring up Giani. Giani is the man behind probably my favorite and least favorite tool online today, Jason. It's called death by Claude. And well, you know what? I'll let the man himself explain it. Giani, what is death by Claude? And why, why did you build a tool that insults me? Well, hey, thank you for having me on the podcast. So what is it by Claude? You put a URL in there, it can be a person, it can be a company, and it'll critique whether it's an AI wrap or something that can be replaced.

1:04:25Why did I build it? I'm running a company myself, and we have been mobbed by Anthropic a few times. We were building something like Cowork, and then Cowork came out, and their margins are bad if you try selling Cowork at Anthropic's price, you cannot do it. So we've been killed by Claude a few times. We did the Satrini piece that was making the rounds. I did the tweet by Iran Peterson where he said okay show us this show it to us here you are going to savage somebody personally like a roast comic or you're going to savage a startup or a business idea what's the best example of this at work gianni should i bring up the the one i did about my own newsletter yes okay i roasted myself for everyone's enjoyment so here is the here's what death by claude kicked out for cautious optimism I have an 89 of 100 already dead score.

1:05:16Giani, how do you calculate the scores that you're giving to each company or product? So if you were like a hardware company or a science company, you are basically a model. And if you're a blog, then you're probably dead. So if you can get a place, then you score high. And if you cannot, then you score low. Yeah, so I'm pretty much doomed. And Jason, what this service does is not only does it find different ways to tell you that you're replaceable, it creates an entire skill MD file to replace the thing that you're doing. and then also it gives you a death certificate and it says that uh my newsletter was cautiously optimistic about its own survival but it taught us that the real disruption was the sub stack fees we paid along the way before it died brutal but the question is really like like how much of the stock market do you think he is really at risk from being replaced because you get do uber do uber do door what would be one that would be hit me no let's do um i'm trying to think of I've got to be careful, because I don't want to sling mud at anybody.

1:06:11Oh, okay. Let's do Andrel. No, no, no. Don't start that up again. I think those sound looks serious companies to me. Yeah, because hardware, right? I mean, basically, companies that rank well have a low score. Yeah, but let's not do Andrel. Let's do Slack. Slack. Oh, okay. Or Chegg is a good one, the textbook company. Well, Chegg's already dead, but okay, we'll do Chegg. Well, Peloton or Chegg, Like these are things that there's a group of people who are trying to pump Peloton right now. It's an interesting one. Chad, 92 % already dead. Yes. Let's see what it has to say. Um, it's just crud. It's an AI rapper.

1:06:49It has no moat depth. It's marked down replaceable and it's expensive. And the, the skill file to replace it, 31 lines. And yeah, brutal. Just absolutely mugged. I feel like every founder should be trying this too. Do, do Peloton. Peloton's a good one. I'm on it. What else you got, Rob? Who else do we want to figure out how dead they are? Because Peloton, to me, is not dead. They basically had all the suckage taken out. And Peloton is actually defensible now because people love their brand. They love the hardware. Oh, wow. And they love the people who are the teachers. They're addicted to those teachers.

1:07:30And that brand and those teachers is defensible. Yes. and in this case 32 out of 100 much better than my blog much better than other services that we looked at it's not just crud it's got reasonable note depth yes it is expensive but you can't replace a bicycle uh with code i suppose let's do one more for fun um lovable lovable everybody's saying that vibe coding yeah people say vibe coding is dead and like it's uh but lovable keeps having revenue grow, but then there was some folks saying maybe revenue was stalled or a lot of people canceling their subs. All right. 78 out of 100 score on Death by Claude for Lovable.

1:08:11Why? Why? I thought it would be 60. Okay, go ahead. Lovable built an AI that writes code for you, which is adorable because Claude already writes code for you without the$20 per month middleman tax. Brutal. And then a 31 line prompt for a full stack app generator. Yeah, brutal. Cause of death, terminal AI rapper syndrome. The underlying model got too good and ate the rapper alive. All right, G-Man, G-Man, why did you build this? Who are you? Where are you located right now? And why did you build this? I live in London. I am building an AI startup myself. I was building something like Claude Co-Work and then Enthropic released before us.

1:08:48It's very expensive to serve in France. So Rob is a friend, I guess, in the future. So I was reading the Satrini piece. I read this piece by Ryan Peterson, where he said, Harvey is like replaceable by Cloud for Legal. And that's like a real feeling. Like we have been worked by Anthropik a few times. So we decided to run our startup and YC's entire portfolio and be the results are very funny. So put it on the internet. This is becoming a reoccurring thing, Rob, is using AI to give a defensibility ranking. That's basically what you're giving. And I think it's genius. A defensibility ranking is a great idea.

1:09:25We do this with founders, organically, Rob, in the founding university, you know, or the launch accelerator, and we tell them, hey, here are your competitors. How are you different than these competitors? Or, hey, if Claude releases this, what would you do? G-Man, if I may call you G-Man, G-Man, you got your ass kicked, and you said, I want to create a tool that prevents me from getting my ass kicked again. What are the two or three things that prevent an ass kicking of a startup based on what you've learned? A few things come to mind. So if you're doing hardware, we don't have physical models yet.

1:10:00So when Travis succeeds with Artums, I think that's a space that goes away. But hardware is number one. Got it. And that, by the way, is becoming now, we went from hardware is hard to hardware is a great way to be defensible. And hard means defensible now. Hard means moat. Hard means engage. It means moat. Hard means moat today. Whereas hardware meant death and not fundable, it now means mode, highly fundable. Very interesting. What's your number two? Network effects are very, very helpful. If you're WhatsApp, you're going to want to be replaced by cloud. I don't think any AI company at the moment has network effects, so I would be on the look for it.

1:10:40So network effect number two. Network effect being, hey, how many people are participating in this? Like a marketplace like Uber. hey, they've got 20 different AV companies putting cars into the system. They're in 10 ,000 cities. They've got 7 million drivers. They've got 10 million restaurants. Okay, that makes it more defensible network effect. Great. What's number three? If you're doing something deeply scientific or in regulated industry, which requires talking to people, I think you're pretty safe as well. That's the important thing. G, you work - G-Man, please. G-Man. Sorry, G-Man. He's the G-Man.

1:11:19You're behind TabTabTab.ai, right? That's your main company? Oh, that's the thing that we were working before we started pivoting because our score is 89 on the... It's like what they've said. Oh, no. It's actually 92 now. Your score has gone up. I just thought it was very funny that your own product told you that you need to pivot and off you go. So there you have it. Rob, you got a question for the G-Man? Or do you have an observation here about this as it relates to being a serial entrepreneur yourself? Yeah. Well, defensibility has always been hard and I do think it's particularly hard now.

1:11:52But I mean, I think he's on the I think he's on the right path. Right. I think this is making everybody a little more like a CEO, because as a CEO, you realize like you have to review stuff, you have to coordinate stuff. And there's some limit to what you can do. Like at some point, you have to decide, like, what are you doing? Right. And these tools help you do stuff. But But yeah, I think it's, I don't know if we talked about this, Jason. I stopped. You were one of my mentors for angel investing, you know, 11 years ago when I sold my first company. I've quit because of this. I think it's too hard now.

1:12:24There is something to it. It is very hard to be an early stage investor. Even if you look at people doing angel investing and getting into a big round, you got to get in at a higher valuation. So even the idea of hitting a 6 ,000 or 7 ,000 X, which I think is where Uber might have peaked for me, or hitting a 500 X, which is I think what Robinhood did, 100 X, whatever it is, Calm, 100 X. It's going to get harder and harder to do that because the entry-level valuation is a magnitude higher. Dilution might be higher. It's just hard. It's never been easy. I'll say that. Yeah. Yeah. Tricky times after.

1:13:02Oh, no. You're not doing. I don't know. Should we do neurometric? That's rude. Rob says do it. I don't care. Rob, maybe there's something you learned here. We'll figure out if we have to pivot. This is the meanest thing we've done to a guest on the show in at least a week. Here we go. I know the score already and I know why it's giving that score, but I can explain why this score will get better over time. Still calculating the exact level of doom. Sportscast that, Alex. What is it saying? It's in counting the lives of Mark die needed to replace cross-referencing with the obituary database. Guiani built one of the funniest tools I've ever seen.

1:13:40It's just consistently not a great score, but not terrible. 72 out of 100. I think you can get that number down because the network effect would be an interesting one. If you could make your platform, Rob, into a place where anybody can put up a. This is my idea. If you became a marketplace of SLMs where anybody could post an SLM and put a price on it, and you would be the reseller, hoster of it, then you become more like eBay or a marketplace of these ideas like hugging faces, right? So if you were the hugging face of SLMs, put a hugging face in there. I wonder how it thinks about hugging face.

1:14:22It might also think that hugging face is you know, not replicate. Yeah, the thing we've been working on, so there's an open source project called Harbor, and Harbor is a collection of thousands of tasks with their evals. And currently, people in the world think about benchmarks, not tasks, right? But a benchmark is a series of tasks. And what we've shown, and a lot of the thesis behind Neurometric, is that you want to choose the right model for the task. And even amongst frontier models, they don't all perform the same on a given task. So even the model that wins the benchmark doesn't win all the tasks within the benchmark.

1:14:56So ensembles of models will 100 % of the time beat single models. Like you just, you can't make one model the best at every single thing. And so to your point, we're kind of going in that direction that you're talking about, Jason, but it hasn't all been released yet. But it's a really great insight. This is why you're such a good early stage investor, man. Yeah. I provide a little value on the margins. All right. This has been another amazing episode of twist g-man thanks for coming on the show and sharing this with us thank you Alex great job uh Lon Harris they say he's 72 he's a 72 square unreplaceable I disagree I disagree personality counts for a lot yeah exactly that's why that's why he's a 69 I put him at 69.420 uh do me do me do my LinkedIn if my newsletter doesn't make enough money by the end of the year, I'll go work for a VC.

1:15:50I'll make you a deal. I mean, it's hard. Oh, boy. This is going to be scary. I think it would probably say, if it says Venture Capital and Podcaster, man, what am I? Maybe I'm like a 65, 64? Okay. So I have the image here from Long. Okay, here we go. Replaced by Claude. Oh, actually, no, I have to follow you around. All right. 58. Yeah. You are a 58, and we can cut this nearby. No, leave it in. I'm a 58. I like it. Tell me about this. Yeah, look at this. Brooklyn startup rooster in a Fordham blazer pacing the cap table at 6 a.m. and yelling, ship it into three microphones while stuffing another safe into his jacket.

1:16:34I mean, it's got a sense of humor. That's the thing. These are hilarious. I mean, I like it. I like it. Okay. I could do this all day. I'm not ready. It's fun. We'll see you all on Friday. Bye-bye.

From the publisher

Anthropic just built a model so dangerous, they’re not even going to release it (for now). On TWiST, Jason, Alex, and Neurometric CEO Rob May discuss whether or not this is just another frontier model or a potential weapon of mass internet destruction.

Claude Mythos Preview is already in the hands of over 40 Anthropic partners — including Nvidia, Apple, Amazon, Microsoft, and Google — as part of its “Project Glasswing” security council.

Together, they’ll test how quickly and easily the model can thwart today’s most elite cybersecurity protocols, and try to develop fresh protocols to keep our most sensitive and critical software systems safe and intact. Once Mythos or a similarly powerful model hits the streets… it’s all over!

Plus, we’re checking out the real-time SaaSpocalypse monitor Death By Clawd, and chatting with its creator, Gyani, about how it feels to have your startup one-shotted by Anthropic’s latest model.

Guests

Rob May: https://x.com/robmay

Neurometric: https://www.neurometric.ai/

Neurometric SLM Marketplace: https://marketplace.neurometric.ai/

Death by Clawd: https://deathbyclawd.com/

Nick O’Neill: https://x.com/chooserich


This Week In Startups is made possible by:


LinkedIn Jobs - LinkedIn.com/twist


Grasshopper Bank - Grasshopper.bank/twist


Render - Render.com/twist

Plaud - https://Plaud.ai/twist

Today’s show:


Timestamps:

0:00 Anthropic's new, powerful 'Mythos' model

3:50 Plaud: If your work depends on conversations — interviews, meetings, calls — you need a Plaud NotePin. You can check it out at https://Plaud.ai/twist and use code TWIST for 10% off!

5:26 Neurometric's Rob May joins the show

10:04 Sam Altman's talent exodus

10:50 Render - Go to https://render.com/twist to apply for the Render Startup Program. You'll get anywhere from $500 to $100,000 in free credits.

11:50 Has Anthropic Passed OpenAI?

15:44 When will Claude release Mythos? Polymarket: https://polymarket.com/event/claude-mythos-released-by

18:32 The AI race turns existential

19:16 Would Anthropic's model give the CIA the ability to hack a foreign government?

20:04 Grasshopper Bank - Time is money. Don't waste either. Go to https://grasshopper.bank/twist and get an exclusive $500 cash bonus just for opening an account.

21:19 Should AI be nationalized?

29:43 Subsidizing secret AI development.

30:15 LinkedIn Jobs - Hire right, the first time. Post your first job and get $100 off towards your job post at https://LinkedIn.com/twist

31:29 How much of the military's $1.5T will go to AI cybersecurity?

35:17 Nick O'Neill's "not sponsored" call-in.

38:24 What is an SLM?

42:01 Tactical/Practical: How can you tune SLMs

45:43 How Neurometric can afford 100M free tokens per month for its users?

49:04 What is "harness engineering"?

50:58 Tactical/ Practical: Every Reddit rant is a startup idea

1:00:25 Why Meta is the most un-innovative AI company in the world.

1:03:43 Gyani of Deathbclawd joins the show — Is your company COOKED?

Subscribe to the TWiST500 newsletter: https://ticker.thisweekinstartups.com

Check out the TWIST500: https://www.twist500.com

Subscribe to This Week in Startups on Apple: https://rb.gy/v19fcp

Follow Lon:

X: https://x.com/lons

Follow Alex:

X: https://x.com/alex

LinkedIn: ⁠https://www.linkedin.com/in/alexwilhelm

Follow Jason:

X: https://twitter.com/Jason

LinkedIn: https://www.linkedin.com/in/jasoncalacanis


Check out all our partner offers: https://partners.launch.co/

Great TWIST interviews: Will Guidara, Eoghan McCabe, Steve Huffman, Brian Chesky, Bob Moesta, Aaron Levie, Sophia Amoruso, Reid Hoffman, Frank Slootman, Billy McFarland

Check out Jason’s suite of newsletters: https://substack.com/@calacanis

Follow TWiST:

Twitter: https://twitter.com/TWiStartups

YouTube: https://www.youtube.com/thisweekin

Instagram: https://www.instagram.com/thisweekinstartups

TikTok: https://www.tiktok.com/@thisweekinstartups

Substack: https://twistartups.substack.com

More from This Week in Startups

All 653 episodes
Anthropic’s Mythos is a cyber-weapon, so you can’t have itThis Week in Startups · 1 h 17 min
Listen in VO