In short
Cal Newport critiques mainstream AI coverage for “directionally true” but unverified claims, and for a nihilistic “head-shaking doomerism” style that stresses people without offering actionable guidance. He also argues that “agents” and the “agentic web” are widely overstated: LLMs generate plausible text, but reliable autonomous planning and execution require extra symbolic/if-then tooling, and most agent visions (e.g., agents browsing the web or using GUIs) aren’t real or aren’t practical yet. He further challenges the hype around Anthropic’s “Mythos,” saying similar capabilities existed before and that marketing exaggerates novelty and impact.
Guests
Cal Newport (computer science professor and commentator; discussed as “commsci professor” and media commentator). Host Ed Zitron (commentary/podcast host).
Key claims
AI reporting often uses stress-inducing framing; CEOs/media may be marketing or morally hazardous; agentic systems don’t exist as described; LLMs aren’t trustworthy planners; Mythos’ “breakthrough” is overstated.
Notable examples
job-automation reports using vague “at risk” buckets; Axios-style headline misquoting; COVID coverage comparisons; “agents booking plane tickets” claims; Mythos “thousands of zero-days” vs prior Opus-4.6 claims; FreeBSD exploit examples; need for APIs/rewiring apps for agents.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOAI Reporting and Its Implications
3:16 to 6:06
Discussion on the nature of AI reporting and its consequences.
“So I kind of wanted to start with, I asked you for a quote a few, like a week ago, maybe two weeks ago.”
Moral Hazards in Tech Reporting
6:06 to 9:07
Exploration of the responsibilities of tech CEOs and media.
“Like, I've read, I think, every AI jobs report now.”
Comparing AI Coverage to COVID Coverage
9:07 to 14:02
Analyzing parallels between AI and COVID media coverage styles.
“It keeps us seeming inevitable and important, in which case that's a huge moral hazard because you are making many, many people, normal people, stress the hell out.”
Conflicting Perspectives on AI and Media
14:02 to 14:24
Explore the confusing narratives surrounding AI's impact on jobs and media.
The Dilemma of AI Usage
14:24 to 15:05
Discuss the paradox of using AI tools despite fears of job replacement.
The Confusion of Anthropomorphism in AI
15:05 to 16:06
Analyze the implications of humanizing AI interactions and its drawbacks.
“Like what's the, surely chat GPT would be seen as like a rock versus a shotgun at that point.”
Natural Language Processing vs. Human Interaction
16:06 to 17:34
Debate the effectiveness of natural language processing in AI without conversation.
“And it's just naturally illogical stuff.”
Redefining AI Responses
17:34 to 18:22
Propose how AI should respond like traditional computing interfaces.
The Challenge of AI's Verbosity
18:22 to 19:56
Examine the challenges of AI's verbose nature and its adaptability.
“Or indeed, if I'm being told that, I need to be told if it's a bad idea.”
AI's Limitations in Action
19:56 to 21:04
Discuss real-world examples of AI's limitations when providing information.
“And tools that are built upon LLM as the digital brain are still way more scarce than you would imagine outside of maybe computer programming and coding harnesses.”
Show all 26 chapters
Narration Insights from Cal Penn
23:58 to 24:54
Cal Penn shares emotional experiences while narrating a sci-fi audiobook.
“I'm the host of Earsay, the Audible and iHeart Audiobook Club.”
Critique of the AI Agent Discourse
25:31 to 28:00
Critically analyze the conversation around AI agents and their capabilities.
“This whole agent conversation, I've never seen anything like it in my life.”
The Limitations of LLMs in Planning
28:00 to 29:20
Discover the constraints of language models in generating effective plans.
“So all it wants to do is produce a story that sounds reasonable.”
The Reality of AI Agents in the Workforce
29:20 to 31:00
Explore the current and future role of AI agents in knowledge work.
“Just asking an LLM, tell me, give me a plan for doing X.”
Challenges of Implementing AI in Organizations
31:00 to 34:20
Understand the complexities involved in integrating AI agents into companies.
“And they never say, because the answer is when we come up with something else, because I don't even think neurosymbolic makes sense for this.”
Misconceptions About AI and Productivity
34:20 to 37:30
Learn why AI is not the silver bullet for productivity bottlenecks.
“it'll write a Python program that'll call hooks into Excel.”
The Future of AI in Data Analysis and Research
37:30 to 42:00
Examine the impact of AI on data analysis and research bottlenecks.
“I think there's a lot of that going on right now with AI and productivity as we look at what the AI can do and then try to make that thing into somehow being the key to getting things done.”
The Power and Vulnerabilities of Mythos
42:00 to 43:30
Discussion on the perceived power of Mythos and its vulnerabilities.
“They had a psychiatrist or a psychologist, I can't remember, talk to it and be like, yeah, we found these emotional features.”
Mythos vs. Opus: A Comparison
43:30 to 45:50
Analyzing the differences in marketing and capabilities between Mythos and Opus.
“because it was absolutely brilliant what they did there.”
Skepticism Towards AI Claims
45:50 to 47:26
Critique of the hype surrounding AI claims and the importance of skepticism.
“They didn't count on a lot of security researchers said, well, wait a second.”
Critique of AI Safety Institute's Findings
50:37 to 56:00
Discussion on the credibility and implications of the AI Safety Institute's report.
“The headline was massive increase in AI scheming is detected.”
Skepticism in AI Claims
56:00 to 56:40
The discussion highlights the need for greater skepticism towards AI advancements and their implications.
Understanding AI Harnesses
56:40 to 58:50
An exploration of the concept of AI harnesses and their role in enhancing AI capabilities.
“the direction the true media coverage it's like well this is scary right I mean that system card it's like 180 pages long.”
Incremental Improvements vs. Fundamental Leaps
58:50 to 1:01:25
The conversation critiques the incremental nature of recent AI advancements and the implications for the future of AI.
“It's all in let's build better, just hand coding, no machine learning, no intelligence, no Skynet here, but just hand coding these programs that we'll call LLMs.”
Future of Distributed AGI
1:01:25 to 1:03:40
Discussion on the potential future of AI as distributed systems rather than monolithic models.
“But the real fear then is like, well, wait a second.”
The Economic Landscape of AI
1:03:40 to 1:06:20
Insights into the economic implications and competitive landscape of AI development.
“And if that's not the moat, if it's just, oh, if I want to build a poker playing AI that's really good, I just need people who are good at poker and to spend a couple of years and figure out a cool custom system.”
Transcript
Automatic transcript. May contain errors.0:00This is an iHeart Podcast. Guaranteed human. Hey everyone, it's Cal Penn. I'm inviting you to join the best sounding book club you've ever heard with my podcast, Earsay, the Audible and iHeart Audiobook Club. Every episode, I nerd out with amazing guests and dive into the best new audiobooks available on Audible. It's the book club for your ears. Listen to Earsay, the Audible and iHeart Audiobook Club. on the iHeartRadio app or wherever you get your podcasts.
0:35Shake it up with Vital Protein's Collagen and Protein Shake. It's a high-quality, ready-to-drink shake with 30 grams of protein and 10 grams of collagen to support healthy hair, skin, nails, bones, and joints. With zero grams of added sugar, no artificial sweeteners, and absolutely no carrageenan, it's a clean, delicious way to fuel your day. so you don't just age gracefully, you age powerfully. Vital Proteins, stay vital. Learn more at vitalproteins.com. Support for the show comes from Public, the investing platform for those who take it seriously. On Public, you can build a multi-asset portfolio of stocks, bonds, options, crypto, and now generated assets, which allow you to turn any idea into an investable index with AI.
1:21It all starts with your prompt, from renewable energy companies with high free cashflow to semiconductor suppliers, growing revenue over 20 % year over year. You can literally type any prompt and put the AI to work. It screens thousands of stocks, builds a one-of-a-kind index, and lets you backtest it against the S &P 500. Then you can invest in a few clicks. Generated assets are like ETFs with infinite possibilities, completely customizable and based on your thesis, not someone else's. Go to public.com slash podcast and earn an uncapped 1 % bonus when you transfer your portfolio. That's public.com slash podcast.
1:53Paid for by public investing. Brokered services by Open to the Public Investing, Inc., member FINRA, and SIPC. Advisory services by Public Advisors, LLC, SEC Registered Advisor. Generated Assets is an interactive analysis tool. Output is for informational purposes only and is not an investment recommendation or advice. Complete disclosures available at public.com slash disclosures. Apple Vacations, where your story starts. Need a vacation? The Apple Vacations foundational sale is here. So you can save up to$150 on your next escape. Book by April 30th for instant savings on trips to Mexico, the Caribbean, Hawaii, Central America and top U.S.
2:24destinations. If you're ready for a break without breaking the bank, save up to$150 at AppleVacations.com or contact your travel advisor. Apple Vacations, where your story starts.
2:40CallZone Media. Hello and welcome to Better Offline. I'm, of course, your host, Ed Zitron.
2:58As ever, support your neighborhood Zitron by subscribing to the premium newsletter, discount link in the episode notes, of course, buy a t-shirt, download a blog, whatever it is you want to do, okay? It's not up to me what you do, but today I'm joined by the incredible commsci professor and commentator Cal Newport. Cal, thank you for joining me. Always a pleasure, Ed. So I kind of wanted to start with, I asked you for a quote a few, like a week ago, maybe two weeks ago. I can't remember how time works anymore. But it was around the way the reporters cover AI and how it seems that a lot of the reporting is kind of directionally true rather than actually true.
3:35Yes. And I want to add something to it since. So I've been thinking about that quote. Yeah, I've been thinking about it. So what I said, if I remember that quote properly, what I was saying is I was picking up a lot in the reporting on AI that you would lean into a story without having necessarily verified that the details are true and that this is what's actually going on, say, with the new AI model, you would lean into it anyways because it was what I call directionally correct. It makes the general point that you see it as your job as a reporter to make, which is, hey, you need to be worried about this or this is a big deal.
4:09Right? And so I think that is a problem. There's another issue I'm seeing, though. I've sort of been refining my thinking on this. I'm also wondering if some of what I'm seeing in some of the the reporting on this is it's just a embrace of the form of, I'm going to give you a stress wave with no relief. Just like we're all going to take turns. Just, I will choose an area you haven't thought about. How about math, math, mathematics are going to go away. Mathematicians are going to be, okay, I'll take that one. Yeah. Let's go. Negative clickbait. Yeah. And they, but there, there's this weird sort of passivity to it where it's like, I'm just going to sort of, it's, it's a, I call it like head-shaking doomerism.
4:49You're just like, it's us. This field's just going away. What can we do? Like just this sort of like passive head-shaking. It's a very specific style. You don't see a lot of other reporting historically, I think, that takes on this resignation of, I'm just going to make the case that like you're screwed and then kind of give you a shoulder shrug. And then we're going to drop the mic and walk off. And I'm kind of getting tired of this. I think there is a cost to stressing the hell out of people. I mean, I'm getting letters all the time now from people. They'll say things like, I feel like I'm trapped in a cage just being hit with wave after wave of stress, and there's no outlet.
5:25There's no door or possibility of making things better. And I think the CEOs are doing it, and I think increasingly we're seeing commentators doing it as well. This is not good in many different ways. So I don't know. I'm adding that to my list. Some of it's directional true reporting. They really are worried that people aren't worried enough. And I think it's just sport now. Can you find an area to come in and just write a head-shaking article that's only trying to undermine the existence of this important human activity or this job or our lives or whatever? It's a very unusual style that quickly became a standard.
5:57And I see it a lot with anything to do with AI and job studies. like i've been sent this tufts report where it's like oh yeah ai affected or i they find these weird weasel words where it's like jobs that could be at risk from ai at some point and we put them in one bucket and then jobs that might one day be we'll put that in another bucket and there you go don't know what we're like you said don't know what we're meant to do with this don't know what anyone's meant to do with this information but it's just like well there you have it there you yeah, but we're all fucked. It's the end. Even though the data does not say that.
6:36Like, I've read, I think, every AI jobs report now. Every single one. And they're all the same. They are all, right now, AI can do this. And then you look at what it says, it's like, it can do law. Well, it can't really do law. It can do one sigma within law, kind of. And even then, it isn't really obvious. And the people saying it can do that are partners at law firms that don't write motions, but don't do like the grunt work. So it's almost, it feels like the reporters have either given up or are just looking for clicks. And it's hard to tell sometimes. This is what I'm trying to figure out because I'm realizing if it's entirely just, I think this is directionally true and that's good enough, then they should be way more upset and in the streets and sparking a revolution, right?
7:24Like if you actually really believed 50 % of the economy was going to be automated, that we're going to have to have government checks just so we can afford to buy the cat food to eat after all the jobs are gone. If you really thought that our entire infrastructure was about to collapse, that superintelligence was going to emerge suddenly and be a threat to human existence, you wouldn't just write a sort of too-cool-for-school head-shaking resignation article. You would be like, we got to, where are the John Connors, right? Like we need to get on the cool trench coats and get out there and go against the Skynet revolution, you would be on your feet.
8:01Nothing would be more important to you. So this is my case about the tech CEOs. I think there's a moral hazard that I don't think that we're putting our finger on properly here, right? So you have the tech CEOs in the AI space that'll just come out and just drop these bombs. Yeah. White-collar blood. You know, he never actually said that. That's Axios putting words in people's mouths. Oh, that was Axios. I thought that that was it. Dario Amadei, Wario, He did say 50%, but I thought he said the blood bath. That's my bad. Well, I trust, I, I, this, the New York fact checkers figured that out for me, but Axios does a lot of this where they put like these really quotable quotes in the headlines about articles on interviews or speeches given by AI people.
8:43And it turns out the thing in the headline wasn't what they said. It was directly what they said. But anyways, so they're out there making these big statements. The jobs are going away. The internet, as we know it, is about to all fall apart because of mythos is going to have this new capability. The super intelligence is coming. I don't even know what's going to happen. There's two possible things going on here, and both of them are morally bad. One is, which is the one I think is true, which is this is largely marketing. This works. It gets reported. It keeps us seeming inevitable and important, in which case that's a huge moral hazard because you are making many, many people, normal people, stress the hell out.
9:21actively scaring them, actively scaring them. The other option is you actually believe it's true. Well, this is an even larger moral trap that you've just fallen into because you are now perpetuating something that's going to cause exponentially more harm. You should be the very first person shutting down your company and trying to get the other ones to do it as well. So it's this weird moral trap they've set up where whatever is actually going on here. If they're coming out here saying these things, it is bad. This, this can't possibly normatively speaking, be the right ethical behavior to be out there saying these scary things all the time.
9:53Because either you need to be building the barricade, or you're just scaring people for the marketing. Neither of these, I think, is something that's defensible. I have a third and worse option, which is I choose Axios. I think Axios, there are some good reporters there. I think the leadership over there is disgusting. I think that they are aligning themselves with the companies. I think that what, like, if you watch, there was a Jim, what's his name, interviewing Sam Altman. These, I think that there is a level of and i i would put this across people like kevin rose and casey newton these are my words not cows um that they're aligning them that they're saying we think this is going to happen and we're here to tell you great news this is good news for me the writer because i will be safe somehow i will be fine you will not you should be scared but it's also a good thing because economy marketing market good and it is it's a very incoherent message because it's like to your point yeah if this was a virus like a pandemic you wouldn't be writing hey millions of people are gonna die what pretty good right hey it'll be good we'll have less people that'd be good right it would be seen as peculiar someone did write that someone did write that by the way someone did write that they did say i remember early pandemic someone did write hey you know what this is good for the planet did it go like hey we're driving less this is great and we're overpopulated like oh i mean i mean that's a different conversation that maybe i but in all seriousness you didn't have mainstream media being like well covid's gonna kill everyone the end i guess i guess you know maybe we'll just be inside forever you didn't have this kind of straight in fact you had the direct opposite it was we need to get outside again who cares about this thing well it's just yeah go on yeah i think that's an interesting now i want to just pull on that thread a little bit because i think covid covid gives an interesting i think it gives two different interesting observations that go in both directions, right?
11:46So I think you're definitely right what you're saying is when the pandemic was coming or it was getting bad, really a lot of the coverage was about what should we be doing or who are the people doing the wrong thing. But it was very much coming from this angle of like, okay, we need to do whatever it is. Like we need to be better about this. It's got to be vaccines. It's got to be masks. It's got to be pickier mitigation, whether you like it or not, it was very focused on what should we be doing or who is it that's getting in the way of a plan that maybe would get us out of this, which is where I think you're very right, is that you did not see a lot of COVID pieces that were just, well, I'm just going to kind of walk through like all the different ways.
12:26You might die and the morgues are going to fill up and that's COVID. That's just how life goes. But I also think the other thing we saw in a lot of COVID coverage is something that we are seeing in the AI coverage. That's where I saw a lot of the directionally true, not factually, but directionally true. There was definitely a period early on in COVID because I was following that coverage quite carefully where the papers were thinking, okay, this is the right behavior. And they were probably right about a lot of these things, but I just would notice this. There'd be a lot of like, okay, we need people to buy into, for example, the lockdowns or whatever.
13:02And there'd be a lot of directionally true reporting where maybe they would like put on a photo of a mass grave that was sort of unrelated to COVID, or you would see a lot of, there'd be pushback from like conservatives about schools. And then they put a lot of articles in the paper about teachers dying of COVID, even though it was, they weren't in school, they got COVID elsewhere. And if you really pushed on it, it was because it's directionally true. Like the general or more general truth here is like, we need to be worried about this or these mitigations work. It doesn't matter if this photo is actually right or if this teacher who died in Orlando, the fact that they hadn't yet been back in a school building yet, it's serving the directional truth.
13:45So it's like it highlights something. COVID highlights something we're seeing now, that the reporters that are doing directional reporting, like we should be scared about it. I dare you not to be scared now. I dare you not to be scared now. Just trying to ratchet it up. But then you also get the contrast, which is this new style of just like head shaking resignation. and actually i don't think the reporters think they're going to be safe they're also like writing is going to go away the media is going to go away so it's a it's an almost like a nihilistic type of approach to this like yeah i'm screwed we're all screwed what are we going to do and that is definitely different than we saw during that last crisis which was obviously much more actually severe than what's happening now so it's really confusing me to be honest well the directionally reporting during covert yeah probably shouldn't have but at the same time it was in it was actually in pursuit of something good like it was an attempt to make people take this seriously because that's ultimately what it was is take this seriously don't go outside stay like don't don't meet with people don't be indoors with people blah blah blah blah great in this case it's like yep you should be scared of this and what should you do fuck knows use chat gpt i guess yeah and what's what's really confusing to me as well as you say oh these people don't think they'll be safe for the most part i just don't i actually take back what i said i think a lot of them just don't acknowledge it they don't acknowledge the core ridiculousness of being like well everyone's jobs are going to get replaced don't know like the garfield meme with him looking at the the garfield with the the cross out on the tv yeah flawlessly described there um it's just it's frustrating as well because it is terrifying people without like i'm not saying literally axios or however but stories like this are what made that made a mentally unstable person throw a molotov cocktail at sam altman's house like it's obvious that these people were scared of the ai doom partly because to your point what the fuck are we meant to do about it because using these tools is not i don't really see how that works because if going along that line of logic if the answer is you need to use this stuff now but the eventual end point is that it's intelligent enough to do everything for you how does using it now matter at all?
15:54Like what's the, surely chat GPT would be seen as like a rock versus a shotgun at that point. Like it's just technologically irrelevant if they get to AGI, which they probably won't. And it's just naturally illogical stuff. Yeah. And I'm with you. I've been making that same argument. This idea that you need to learn how to prompt some generation of a chat bot that exists right now is going to be the key to your long-term, I mean, even if, as you say, AI ends up playing a major sustained role in the economy, it's not going to be everyone typing on a web interface to a chat bot that's sycophantic and has a personality.
16:31Like, I think I've heard you say this recently, and I agree with it as well. Like, I don't think we should be chatting with technology. We should not be chatting in a sort of anthropomorphized, humanized way. It doesn't mean you can't do natural language processing. I mean, Google is natural language processing. You're writing your Google searches in natural language, but no one's having a conversation with Google. It's you, you, you list the keywords as quickly as possible. And Google's pretty good at figuring out, you know, population, Spain, 1982, and you press enter and you get that information.
17:02You're not like, Hey, so I'm wondering what the population is of Spain in 1982. Can you help me find that question mark? There's, there's something odd about that, uh, anthropomorphized conversational interface. I guess we saw a lot of Star Trek growing up and that's, you know, what we, what we think the future is supposed to be like but it has all sorts of problems remembering star trek when he would go computer do this the computer didn't go that's a great idea jean-luc what a great idea thank you for the computer just did the thing that's like i don't have any trouble with natural language queries because i think the whole reason that say chat gpt has grown comes from search i think it is the core of it because chat gpt and claude and all them are better at understanding what you asked for not saying the data output is necessarily great but just they understand the the the inference they make from what you say is better than google or at least better than google has been i feel like it was better before and i think that had google not kind of boofed it on this one we wouldn't be in this spot but even then using google now it forces you it forces the ai summaries and you could do minus ai and all that but sometimes i don't remember to and it's just it's just turn search into this nightmare but nevertheless back to what you were saying i agree i think the anthropomorphization needs to go i think that these things need to respond like terminal windows or what have you they need to respond like computers and go okay here you go just don't need all that cludge i I don't need to be told, oh, what a great idea.
18:35I know, I had it. Or indeed, if I'm being told that, I need to be told if it's a bad idea. But I don't even necessarily need an answer. I just need stuff to look at so that I can come to my own conclusions. I think it's hard, actually. I think it's actually hard to get a language model to do that, right? Because if you think about, when you go back to the base layer of what's happening in the pre-training, is that you're building a language model that's trying to win at the token guessing game. So I'm trying to guess what word or part of word actually comes next to what I assume to be a real piece of text.
19:05And then if you do that auto-regressively, so you call it again and again and again, adding the answer to the input so it grows out an answer, what you're going to get is text expansion. You've given me a text that I'm trying to expand as if there was a real text that exists and I'm trying to match it. You get that like kind of indirectly. So really its idiom is the type of text it's trained on, which for the most part is more sort of prose style text. So you can tune it away from it. Like you can tune its mood, you can tune its sycophancy, but it might be hard to actually tune an LLM because it deals with human written prose as its main training data.
19:42It might be harder than we think to tune that away from being verbose and to just give a table. Now, I guess you could take its output and then maybe run that through another thing that then strips away the other piece. It's like, it's possible. But I think the anthropomorphized verbosity we see in language models is also, that's kind of the native tongue of this particular, which is why we still have a lot of chatbots being emphasized. And tools that are built upon LLM as the digital brain are still way more scarce than you would imagine outside of maybe computer programming and coding harnesses.
20:16We just don't have a lot of other examples where we just use the LLM as a general person's digital brain. Because I think this verbosity is okay. Humans can interpret that, but it's not great if the LLM is just a digital brain that's interfacing between you and another computer. It doesn't need to hear that their idea is great or wants to try to parse the different types of text. There's some interesting things going on there about the fundamental nature of these things. but even then with google ai mode it still seems kind it still like actually seems like it can give fairly short answers yeah but if you mess if you argue with it as i have it will just provide you with it even google's will provide you with just hot dog shit yeah like it will just claim something is true my why one i just did a private equity thing on private credit even and my favorite thing is being like what fund is this part of and i go it's part of this fund that fund was funded was founded after this happened and he goes okay well maybe it's this one different fund three years old doesn't not involved do you have proof of that well this is what you don't know this is what you don't see in star trek is you know captain kirk or whoever i'm going to mix up the episodes here you know say like hey computer we are approaching deep space nine prepare docking procedures and computer is like photon torpedo fired yeah destroyed and you're like well no i said we're supposed to dock oh you're right kirk i shouldn't have fired the thank you for holding me accountable captain kirk that was i i did the opposite thing you know yeah that didn't happen in star trek it's spring which means i'm back talking to you about buying your clothes from quince who just launched a line of new wrinkle resistant european linen dress shirts i'm looking forward to buying for the few times a year I have to dress like an adult along with some fleece joggers that I'll wear the rest of the time.
22:06Their clothes are great. They make high quality everyday essentials, bags, coats and even perfumes. I swear by their hooded down parker as New York is a nasty habit of randomly getting cold and when it's a little warmer out I love their various leather jackets I've bought over the last year. The best part is their prices are 50-60 % less than similar brands. All because they work directly with ethical factories and cut out the middlemen. so you're paying for quality, not brand markup. And everything is designed to last and make getting dressed easy. I really love their stuff, and I think you will too.
22:34Refresh your wardrobe with Quince. Go to quince.com slash better for free shipping and 365-day returns. Now available in Canada too. Go to quince.com slash better for free shipping and 365-day returns. quince.com slash better. Support for the show comes from Public, the investing platform for those who take it seriously. On Public, you can build a multi-asset portfolio of stocks, bonds, options, crypto, and now generated assets, which allow you to turn any idea into an investable index with AI. It all starts with your prompt. From renewable energy companies with high free cash flow to semiconductor suppliers growing revenue over 20 % year over year, you can literally type any prompt and put the AI to work.
23:21It screens thousands of stocks, builds a one-of-a-kind index, and lets you backtest it against the S &P 500. Then you can invest in a few clicks. Generated assets are like ETFs with infinite possibilities, completely customizable and based on your thesis, not someone else's. Go to public.com slash podcast and earn an uncapped 1 % bonus when you transfer your portfolio. That's public.com slash podcast. Paid for by Public Investing. Brokered services by Open to the Public Investing, Inc., member FINRA, and SIPC. Advisory services by Public Advisors, LLC, SEC Registered Advisor. Generated assets is an interactive analysis tool.
23:51Output is for informational purposes only and is not an investment recommendation or advice. Complete disclosures available at public.com slash disclosures. Hey, everyone. It's Cal Penn. I'm the host of Earsay, the Audible and iHeart Audiobook Club. This week on the podcast, I am sitting down with Ray Porter, the narrator of Andy Weir's audiobook Project Hail Mary, massive sci-fi adventure about survival and science and what happens when you wake up alone very far from us. earth. I really had to make a decision because I caught myself getting that frog in my throat and starting to get teary as I'm narrating some of these sections.
24:29And it's like, okay, yo, yo, yo, is this indulgent? And I really thought about it. I was like, no, at this point, it would kind of be betraying the trust the author and the listener have in telling this story if I don't go through it. But there's places in this book that deeply emotionally affected me. And I left it on the mic. That's great. Because it served the story. People will say like, oh my God, I cried at the end. It's like, yeah, dude, me too. Listen to Earsay, the Audible and iHeart Audiobook Club on the iHeartRadio app or wherever you get your podcasts. Shake it up with Vital Proteins Collagen and Protein Shake.
25:06It's a high quality ready to drink shake with 30 grams of protein and 10 grams of collagen to support healthy hair, skin, nails, bones, and joints. With zero grams of added sugar, no artificial sweeteners, and absolutely no carrageenan, it's a clean, delicious way to fuel your day. So you don't just age gracefully, you age powerfully. Vital Proteins. Stay vital. Learn more at vitalproteins.com
25:37So one thing that's really been driving me insane, by which I mean going on Twitter, is looking at people like Aaron Levy of Box and Brian Armstrong of Coinbase talking about like agents spending money and the agentic web and how we need to prepare the web for agents doing stuff and the agents will do this fantastical doesn't exist agents don't do that just like not they don't have the ability to like oh they'll use computers computer use is basically non-functional in AI and it takes insane amounts of compute it feels like a conversation keeps happening in theory, in the media, on social media, about something that's possibly completely impossible, but the certainty they discuss it with is insane to me.
26:21This whole agent conversation, I've never seen anything like it in my life. I mean, it does feel a little bit like crypto to me. I think that is kind of a fair comparison where if you had a blockchain-driven software, like in theory, that software would kind of work, but it just gave you a worse version of what you could already do for pennies using the actual, you know, Amazon server somewhere. And all you were really gaining was some sort of cyber libertarian philosophical feel of goods about like, yes, but this was purely decentralized. I got worse versions of software to be decentralized. But now no one can control it.
26:59This is what like early agents, I mean, okay. So here's what I've been writing about agents. I've been thinking a lot about it. I mean, the issue is I don't think people understand what they are. I think people think that it's a new type of digital brain that is now able to go on and do more autonomous activity. I always see this get mixed up. It's just like people talking about Mythos breaking out of its sandbox to do XYZ. Mythos is a language model. You can give it an input and it can give you a token. You're talking about a program that is calling Mythos and then taking actions based on what it called.
27:33And this is really what we're talking about with agents is the digital brains are LLMs. And then you write a program that will say to the LLM, give me a plan for doing X. And then the LLM spits out what seems like a reasonable text that seems like a reasonable plan. And then you execute that plan. The program executes that plan on behalf of the LLM. And I wrote about this. Yeah. And I wrote about this earlier this year. LLMs are bad, you know, as a digital brain are bad planners. It's not really, you're not going to get consistently usable plans because what an LLM is actually trying to do is finish the story you gave it.
28:06So all it wants to do is produce a story that sounds reasonable. So it's giving you reasonable sounding plans. Like, yeah, that's what a plan for doing this would more or less sound like. But what it's not doing is actually doing step-by-step evaluations. It doesn't have a clearly isolated goal that it's trying to measure how close you're getting to it. It doesn't have a world model to evaluate what's going to happen with the steps that are going to unfold next. And so in almost every context, it turns out, oh, a digital brain by itself being an LLM doesn't lead to good agents. In programming, it seems to work a little bit better.
Read the full transcript
28:39But I do think Gary Marcus, I don't know if it was a scoop, but Gary Marcus captured in a recent newsletter something really important. When Anthropic leaked the code for their cloud code coding harness that sits on top of their LLMs to do coding, It turns out they've added a huge amount of old-fashioned, hand-coded, symbolic AI-style rules and pattern recognizers and special if-thens. So they've just been sitting there tuning this program for specifically doing computer programming. And the LLM is being a little bit more isolated to just the code production. So they've kind of just gone back to old-fashioned.
29:16That's just like an old-fashioned system that is plussing up an LLM. But I'm with you. Yeah, it's very hard. Just asking an LLM, tell me, give me a plan for doing X. For almost any scenario of X, you really can't trust a plan from a model whose goal is primarily to finish text, to finish the story you gave it in a reasonable style way. That's not how we plan. That's not how we think about planning. And it doesn't give you consistently usable plans. So, yeah, but you're right. It's magic. Like the agents are coming. They've been saying this. I mean, I wrote the article I wrote, you know, in January.
29:52What happened to the year of the agent? 2025 was the year of the agent. All we had was coding agents. That's the only thing that we worked on that whole year. It was supposed to, I mean, I have the receipts. Early 2025, all of these executives saying, your work as a knowledge worker, not as a computer programmer, but just as a knowledge worker, is going to be largely done with agents. You're going to have agents are going to be a major part of your workforce in just a normal office setting. And none of that happened because it turns out just asking an LLM, give me a plan for doing x doesn't often actually produce a workable plan and as a result the only way to make agents work which they do not is to build a bunch of symbolic or if this then that just like scripts like i mean if you use manus for example it's just writing a shit ton of python and it's writing it to do stuff that it it's like oh yeah let me just do this and it just writes a python tool to fill out a spreadsheet it's insane it's really insane but what's more insane to me is that the conversation around agents is as if they're already here i'm about to read you something from box ceo aaron levy the ceo of a of a public company one corollary to the fact that ai agents take real work to set up in a company at scale is that the role of the forward deployed engineer or whatever it gets called in the future isn't going away anytime soon when a vendor sells any kind of agents into an organization you're no longer just selling a software tool that gets implemented and you're done you're fundamentally selling some sort of actual workflow being done by your technology what are you fucking talking about what are you talking you are a cloud storage and collaboration what do you sell and the answer is nothing they don't sell any agents agents don't oh agents are going to do this what you are describing is a different kind of technology just yeah that's it like it's something else that doesn't exist but this is everywhere you go you look at any consultancy right now any conference right now there will be a speech about agents even meredith whitaker who i deeply deeply respect went on stage last year was like yeah ai agents using money they're booking plane tickets no they're not they're not that's not happening and i say i say this again deeply respect meredith i said this online people flip their shit at me it's like oh she's directionally correct yeah she's directionally correct it's like let's be scared of the things that exist because i think it's perhaps scarier for a different reason that we have large swaths of the tech industry talking about something that doesn't exist like just like agents don't like they don't they don't exist they don't like people are talking about the agentic internet i keep reading about even on the verge i read about it I read it all over the shop where it's like, oh yeah, well, the internet needs to be rebuilt for agents to use.
32:39It's like, what do you mean? And they never say, because the answer is when we come up with something else, because I don't even think neurosymbolic makes sense for this. I mean, neurosymbolic being the one where it's, they have a deterministic system that they access from what I understand. Like the other thing as well, now that I think about it out loud is how would they actually browse the internet? Where are they being housed? Are we using GPUs to make them browse the internet? That's insanely, insanely, that's very, very convoluted and probably quite expensive to do. And to what end? That's the real question, right?
33:15I mean, I've seen these proposals. I mean, basically where a lot of these proposals go, I mean, it's the agents were supposed to, we thought that we could just make AI do anything. So we'll just, we'll have it use the mouse and just use our computers for us. Oh, that's hard. We don't know how to do that. All right. So what we'll do is we'll rewire all applications that anyone uses in the internet so that we don't actually have to use the mouse. It can have a text interface so that an LLM, like they do, the coding agents do, can give a description of how to do something in Excel in text without having to actually move a mouse or click things around.
33:48And then these evolve to say, okay, well, what's the one type of instruction that we're good at producing? because when LLMs produce plans, they're directionally correct plans. They don't actually get the thing done. But they said, oh, what LLMs are good at is producing code that compiles and we can actually check that it works. And so this is where this whole vision has changed is that all applications and internet websites should have a code accessible API that you can expose and then an LLM can write a program that will then access that API. So we don't need to teach the LLM how to use Excel.
34:21it'll write a Python program that'll call hooks into Excel. The problem with this is no one wants to open up their application to just agents in general. If I'm Microsoft, I was like, I want to write a custom tool for my program. Why would I expose my program for anyone else to use it? But your original question, it's a big one, to what end? I've been writing about this recently, especially with work and AI. You got to find the real bottlenecks, right? Yeah. It's the drunk looking for the keys under the streetlight. There's a lot of this going on where this is what we can do with AI right now.
34:58Then this now becomes like the key to productivity. But the real bottlenecks in people's work is often not the things that we're trying to aim AI at. Like I don't know people are super frustrated at booking a plane ticket online. Yeah, it's really easy. How often do you book plane tickets? You kind of want to know. Like let me see. Maybe this time will be better. What seat's available? It takes five minutes. So it was a huge jump to go from a travel agent to a web interface. But this is not a bottleneck in people's life now where I want to give complicated. I book flights all the time. Yeah. And they're easy.
35:28They're so simple. I can do it while sitting on the toilet. I don't want an agent to choose. And people are like, oh, your calendar will tell it. My calendar doesn't lay out my entire day. I don't have every single thing I do on there. It's just strange. Well, I had the same argument with social science researchers who are like, if you're geeky enough to learn coding agents, they're like, this is revolutionizing science research because now, for example, you could have it write a program to process a data file and then format it into a plot. And that might've taken you four hours to do and you work with it for a half hour and you get that result.
36:09This is revolutionizing research. And I'm saying, well, it's not. The bottleneck for social science researchers is not analyzing data and producing plots. You're not sitting there doing that eight hours a day, every day, and if I could do this twice as fast, I'll produce twice as many papers. I might write one paper in a three-month period. Yeah, in there, there's like four hours I spent making a plot, and sure, it'd be nice if that four hours became 30 minutes, but that's four hours out of like a multi-month process of sort of thinking about this paper. What is a plot, by the way? Like a graph.
36:39Oh, right. Yeah, the computer science term, but yeah. It's like, that's nice, that got a little bit faster, but that's not the bottleneck. That's not what's going to unlock a lot more research. It's like, man, I would write more papers if it wasn't for how long it took me to draw a graph. And if you could, I have five graphs. Isn't the problem data? Getting the data. Like actually collecting data? That's what it is. I wrote about this talking to a well-known business school professor years ago for my book, Deep Work. And he talked about, he just realized, oh, being a business professor, publishing papers is about data access.
37:08I have to spend most of my year talking to people, building relationships, trying to set up an agreement with a company where I can get good data that I can get three papers out of. In all of that work, there's one day in there where you're crunching the numbers and making a plot. And it's nicer if you could do a little bit faster, but it's not a productivity bottleneck. It's a marginal efficiency. I think there's a lot of that going on right now with AI and productivity as we look at what the AI can do and then try to make that thing into somehow being the key to getting things done. I just, my productivity problem is that the UI and UX and everything sucks.
37:43Everything's disjointed. Setting up Riverside is always fun. They move the menus around. Projects are in a different place. That takes up time. Moving files places also takes up a lot of time. This morning when I put out my private credit piece, I had to do these threads. I had to click around a website and put in the alt text, but I had to tweak it slightly. It's like, I don't know how AI would possibly help me here. And they're not working on that. Well, they tried. They tried, though. Did they? I thought that was going to be, this is what I was excited about earlier in the Gen AI revolution.
38:15I was like, okay, here is the real value prop. Is natural language interface into advanced features on software? Where I can just say, all right, I want you to go take this column in the spreadsheet and get rid of all the rows that have values before this. And then I want to make a pie chart. Because I don't want to learn how to do all that in Excel. I don't know how to do that. And they tried it. I mean, this was Microsoft Copilot. But it turns out we underestimated the degree to which when we as humans are interacting with a chatbot that we're incredibly gracious, we're able to adjust and kind of get the gist of what it means and filter out the part of the chatbot response that's not really relevant or ask the follow-up question.
38:52And when they tried to just use LLM responses to automate actions within programs, it's just not accurate enough. So they wanted that to be the case, that you could just be talking to a Riverside bot and you never would have to press a button ever again in Riverside. It's just not accurate enough. LLMs, it's fine for human conversation. It's just not accurate enough in this general case. Also, that thing you're describing with how they want the agentic web to just be a series of APIs so that every agent writes Python or what have you to use them, that's a massive computational increase for no reason.
39:28Because you're basically saying instead of someone clicking a mouse and hitting a keyboard, we will write code for everything. yeah what an insane what a truly insane idea i mean it's it's just very like salesforce today i don't know if you saw they announced that they're doing salesforce headless 360 mark benioff needs to fire everyone in marketing but they've made it so that you can do everything with salesforce via an api which is i mean the first question i always ask is what does salesforce do because i've talked to so many people they can't tell me there's like 21 different features no one knows what they do but it's like it's just a very bizarre thing it's very much a cart horse thing but also what agent like that's what this is the thing that really drives me insane they're talking about we built this api for the agentic web for agents to use it which one what agent what are you talking about well it will be in the future what do you you change something materially with your publicly traded company worth$300 billion because it might happen while we're getting ahead of it.
40:36What the f - and it's, you talk to members of the media about this and they just go, yeah, you know, yeah, yeah, yeah, you know, it'll happen. It's obviously going to happen. They wouldn't put this much money behind it if it wasn't going to. It's like, I don't know, especially with Salesforce. And I'm like, you don't think Salesforce would spend a bunch of money for no reason well buddy you've you've not been following salesforce at all then i mean yeah go on yeah i was gonna say how much did meta spend on the metaverse over 70 billion dollars where did that money go where did it go where did it go customizing floating dinosaur avatars building legs but let's change that's the second 50 billion right did it if they had gotten the second half the investment they would have got to the legs they're just not there yet another 100 billion will have toes um so changing subject a little mythos has been one of my favorite media hysterias recently i genuinely wonder like if they ran war of the worlds again today i think axios would have a headline two minutes and it'd be like there are aliens they're attacking i heard it on i heard it on a podcast i've looked through the system card i don't know if you have for mythos it's wacky it's wacky it's wacky i can't believe we're we're letting people get away with having a psychologist talking to the chatbot in your system.
41:59It's nuts. It's all gone so marketing. They had a psychiatrist or a psychologist, I can't remember, talk to it and be like, yeah, we found these emotional features. We need regulators to stop this stuff because I've heard, and people's response to this is, well, banks are having meetings about it and the government's having meetings about it. Governments have meetings about NFTs. there was a way gavin newsom signed an executive order about uh web 3 these people will meet and talk about anything oh it's scary and they're not talking about it which means it's powerful well how is it powerful what does it do because i think you probably saw this as well it didn't list how many false positives there were it also didn't mention that the free bsd bug that they talk about that they found the wasn't actually exploitable i think it was something about like about the level it was at.
42:53I forget. I don't do programming other than very simple Python, a dog's Python. Yeah. I mean, FreeBSD kernel is full of bugs. All these things are full of bugs. Because they're open source. I had to have this conversation with someone recently where they were like, Mythos, can you believe, of all the places it found a bug, in the kernel of Linux, like in Linux they found, like are you kidding me? all day long is just bug fixes having to be pushed into that repository. Yeah, the Mythos story, I think, I mean, A, someone needs to get a Nobel Prize in marketing because it was absolutely brilliant what they did there.
43:34I've spent a lot of time on it. It's complicated because, again, you can't really trust the system. The system cards are just gonzo that Anthropic puts out, and it's not publicly available. But there were, I think, a few very telling things. So there's two features they say Mythos has. One is finding vulnerabilities in source code, and two is writing programs to exploit them. It's first really important that people understand this has been something that people have been doing with LLMs since the beginning of publicly available LLMs, right? Not only is there nothing new about that, but I found – they put this on my podcast – almost word for word from the Anthropic system card, they said in the Opus 4.6, rather, systems card, right?
44:15a publicly available model that's already been out for many months, almost word for word for what they said about Mythos, except for no coverage of it and no fear. They said, we have found 500 zero-day vulnerabilities, including some that had been existing for decades without having been discovered. That is what they said about what Opus 4.6 could do. For Mythos, they said the same thing. They just replaced the word 500 with thousands. But when Opus 4.6 came out, there was no, oh my God, They have found many hundreds of zero-day exploits, many of which have been around for decades because they didn't push that marketing button.
44:51No one particularly cared about it. I went back to my podcast and showed multiple papers. This has been a huge concern. And it's a real concern, by the way, right? Yeah. Is that partially what slows down slightly cracking, right, the breaking into systems, is the fact that it's annoying and hard. And LLMs have made it easier. GPT-4 was good at finding exploits, right? And this was a big deal. They were like, GPT-3.5 wasn't great at it, GPT-4 is. And then as we got the more recent models, they've been much better at writing code to exploit them because we had better agents for it and they're better able to produce multi-step software goals.
45:27And so they can better build software to exploit them. This is a real issue, but it's not new with Mythos, right? Yeah. But Mythos was presented as if some Rubicon had been passed. But there was a couple things I noticed right off the bat. One, they made the mistake of listing a bunch of the exploits that they, vulnerabilities they had found to try to brag. Look at this thing in FreeBSD. Look at this thing in FFP &G or whatever. Like they showed all these exploits they found. They didn't count on a lot of security researchers said, well, wait a second. Why don't I get like a much smaller, cheaper model aiming at that same source code and say, can you find any vulnerabilities?
46:00They could find the same ones. So the evidence that it's finding vulnerabilities better, we don't have any way of knowing that's true. And if anything, we actually are getting a lot of reports that they were paying big bounties for security researchers. I'm going to give you access to mythos. I'm going to pay you for any bugs you can report that you found with it. So they had security researchers just who knows how many false positives were coming out of that. And then on the exploitation side, we only really have one study. It comes from AISI, who I do not trust, but it's the only independent study.
46:32The fact that they gave them access itself should make us maybe a little bit suspect. But it basically just showed like normal progression. No massive leap. Model by model gets a little bit better on some of these tests and benchmarks. And Mythos has no out-of-scale leap. It's just like on some it's about the same. On some it's a little bit better. And yet it got covered as if we had just turned on, you know, Whopper from the movie War Games. like we had just some new entity that was like on its own undermining security and I do not think that I think that was highly credulous coverage of what almost certainly is just like a standard slight jagged move forward on these various capabilities that we've been seeing for the last three years
47:25The show comes from Public, the investing platform for those who take it seriously. On Public, you can build a multi-asset portfolio of stocks, bonds, options, crypto, and now generated assets, which allow you to turn any idea into an investable index with AI. It all starts with your prompt, from renewable energy companies with high free cash flow to semiconductor suppliers growing revenue over 20 % year over year. You can literally type any prompt and put the AI to work. It screens thousands of stocks, builds a one-of-a-kind index, and lets you backtest it against the S &P 500. Then you can invest in a few clicks.
47:57Generated assets are like ETFs with infinite possibilities, completely customizable and based on your thesis, not someone else's. Go to public.com slash podcast and earn an uncapped 1 % bonus when you transfer your portfolio. That's public.com slash podcast. Paid for by Public Investing. Brokered services by Open to the Public Investing, Inc., member FINRA and SIPC. Advisory services by Public Advisors, LLC, SEC Registered Advisor. Generated assets is an interactive analysis tool. Output is for informational purposes only and is not an investment recommendation or advice. Complete disclosures available at public.com slash disclosures.
48:26Hey, everyone. It's Cal Penn. I'm the host of Earsay, the Audible and iHeart Audiobook Club. This week on the podcast, I am sitting down with Ray Porter, the narrator of Andy Weir's audiobook project Hail Mary, massive sci-fi adventure about survival and science and what happens when you wake up alone very far from Earth. I really had to make a decision because I caught myself getting that frog in my throat and starting to get teary as I'm narrating some of these sections. And it's like, OK, yo, yo, yo, is this indulgent? And I really thought about it. I was like, no, at this point, it would kind of be betraying the trust the author and the listener have in telling this story if I don't go through it.
49:10But there's places in this book that deeply emotionally affected me. And I left it on the mic. That's great. Because it served the story. People will say like, oh my God, I cried at the end. It's like, yeah, dude, me too. Listen to Earsay, the Audible and iHeart Audiobook Club on the iHeartRadio app or wherever you get your podcasts. Shake it up with Vital Proteins Collagen and Protein Shake. It's a high-quality, ready-to-drink shake with 30 grams of protein and 10 grams of collagen to support healthy hair, skin, nails, bones, and joints. With zero grams of added sugar, no artificial sweeteners, and absolutely no carrageenan.
49:50It's a clean, delicious way to fuel your day. So you don't just age gracefully, you age powerfully. Vital Proteins, stay vital. Learn more at vitalproteins.com. Apple Vacations, where your story starts. Need a vacation? The Apple Vacations foundational sale is here, so you can save up to$150 on your next escape. Book by April 30th for instant savings on trips to Mexico, the Caribbean, Hawaii, Central America, and top U.S. destinations. If you're ready for a break without breaking the bank, save up to$150 at applevacations.com or contact your travel advisor.
50:36Also, when you said that the difference between Opus 4.6 and Mythos 500-2000s makes me ask the very simple question of did they look as hard to your point about the security researcher like did did they did they spend as much time probably not so they probably could have found them also by the way i immediately was looking at our ai safety institute is of course heavily linked to effective altruism can i say why i'm upset at ais i talked about them two weeks ago i did a or through i don't know this is coming out but i did a podcast in whenever march where i looked at this report and mainly i looked at the guardians coverage of this report done by ais i but it was just the most inane thing.
51:15The headline was massive increase in AI scheming is detected. And they had a chart. Jesus fucking Christ. And they had a chart and bad line went up. And it went up in like January and it goes up. And if you read this article about this study, they're like, something's going on. Scheming has been increasing rapidly recently. And they like gave some examples of it or whatever. And so I look at this, like, I want to look at it. What is going on here? So I look at this chart. What are they charting? oh they're charting tweets per day that they've detect tweets about ai doing things that you didn't want it to do and i said huh so when does this line start going up the week that open claw was released to the public and everyone just started building their own bad agents and then tweeting about how bad they were and you know what word was not mentioned in that article open claw and even though the examples they were giving so they just said scheming just started rising i I guess AI is becoming sentient.
52:13And all they were measuring was people - Multiple tweets paraphrasing the same viral story to use their own fucking language. And then I looked at the biggest spike. I was like, well, this day in February on this chart had the biggest spike. It was like, oh, there was this one tweet about OpenClaw erasing someone's emails and then it got retweeted. It went super viral. I was like, okay, great. You just, the real headline of this article, letting people write their own agents leads to terrible agents. That's it. But the whole thing, so that's AISI. I'm looking at the tweets as well. One of them is from a 47 follower account with AIR called underscore underscore just underscore underscore Lisa.
52:55And it's, this is really bad. Opus is editing files and making up reasons it's deleting adult content. So hallucinations. And also Opus is not doing that. the stupid open claw program you wrote that's prompting Opus and then taking action on your computer based on what it says is deleting your files. The program you wrote that you gave access to your files and just said, whatever we get from this prompt, execute it, is erasing your files. Opus can't do anything. It can produce tokens. But here's the other point I want to make about mythos that I don't think is being made. And it reminds me of the Sherlock Holmes story of the dog that didn't bark, right?
53:30Where the actual piece of evidence that mattered is not what you heard, but what you didn't hear. This is what I think the real story here is, is you did not hear Dario Amadei in the lead up to the Mythos release in the last year, let's say, or the last two years. You did not hear him talking about what we're working on and why AI is important is because we're going to be able to find vulnerabilities in software that have been long hidden. We're going to build the ultimate cybersecurity machine. This was not discussed. That's old-fashioned stuff. That's boring stuff. That's stuff that we were worried at.
54:01Even GPT-2, people were worried about that. What we've been hearing about steadily was jobs are going to be automated. We're going to have whole creative industries wiped out. We might have sentience coming, and at the very least, AGI and these massive disruptions. This is what they've been focusing on again and again. And then their biggest, best model, their newest, greatest, bestest model that they train forever and use all the electricity, what did they say about it? None of those things. They didn't talk about any of the things they said the key AI was, the things they were afraid of, the things they were excited about.
54:34Instead, they went back and talked about a boring parochial old feature that has been an issue that nerdy security researchers have been talking about for a half decade now. That to me is if I was an investor, I would say, take off your like Greek helmet cosplay mythos is coming to destroy. Hold on a second. Is this better at automating jobs? Is this better at producing code? Is this AGO? Why are we talking about finding bugs? We're worried about that with GPT-4. That's a problem. But it's like that's not something new. Uh-oh, something must be going on. You just put a lot of money into a new model, and the best thing you could find to emphasize was it's good at finding bugs.
55:17I think that is a problem. It's what they didn't say about this model. They would have much, much, much rather be able to brag, this model is now much better at any of those things that they've been saying is the key to the AI future. And you didn't hear them talk much at all about any of those. Yeah. And that's the thing. If it was so powerful, like here's the thing. I don't know what would make me convinced that LLMs were the future, but a step toward it would be, we typed create a Slack competitor, which they claim they did once and then didn't show it and refused to. And they said, oh, it worked autonomously for 30 hours, but then wouldn't talk about it.
55:50If they were like, we created the Slack clone, here it is. and it was bug-free. Like, it actually just worked and we're like, we now, we have done this. Because theoretically, if this SaaSpocalypse story was true, which it's not, that AI is going to replace all software, if they actually did that, if they, because someone from Anthropic just left the board of Figma and they created a Figma clone and the stock went down because the market's run by toddlers, if they were like, we've released a clone of Microsoft Word, it's like, we've done Anthropic Word and we now sell that as part of our subscription, that would actually be quite something but the thing is they're not it's kind of it gets back to the old talking point of if they made AGI why would they sell it wouldn't it be a massive competitive advantage to keep this and I think you're right I think maybe Mythos is not as powerful as they say and they've just had to dress it up but it gets back to the thing of the direction the true media coverage it's like well this is scary right I mean that system card it's like 180 pages long.
56:50I don't got all day. I have to write three 100-word blogs a week. I couldn't possibly spend time reading this. And it's just... We need so much more skepticism. We need so much more skepticism, right? I mean, this is why, again, like the most skeptical... We're not skeptics, but like the... I call it the East Coast computer scientists. So those of you, we're technically minded and we're not near Silicon Valley. So we're not in that world. It's very hard to be a professor in a world where there's hundreds of millions of dollars being handed around than they try to like ignore. But the East Coast computer scientists are all baffled by, you talk to any East Coast computer scientists, they're all baffled by, like oftentimes there's claims that are just not true or widely exaggerated.
57:31Why are we so credulous? I mean, it'd be one thing if it was like a government agency we didn't realize was like trying to, you know, protect the fact that there was UFOs and they're just straight up lying. We've never encountered that before. Like I didn't realize that, you know, no, it's a business, right? And the credulity with which we're taking these claims. Like Mythos is, I think, the most important story there is, yeah, this is another example of what I wrote last summer about AI has hit a bit of a wall in the sense that all of the improvements that have come really since over the last two years have almost all been either on post-training or, more importantly, on the harnesses that you built.
58:06So it comes in the software you're building to take advantage. What is a harness? I've seen this word used a lot. I think it's good for me and the listeners to hear the exact definition. Think of it as like a computer program that can do stuff. You can talk to it, can do stuff, and it'll prompt or talk to an LLM as like its digital brain. So the harness might actually be able to touch your file system, write the files, compile code, move things around. But to figure out what actions to take, it will also then prompt an LLM and say, okay, what should I do next? And you can put it on different...
58:38Is that just a wrapper? Yeah, it's a wrapper. So this... But that's where all the progress has come. All of the progress in coding agents since about a year has come, especially starting this fall, has come from better wrappers, better harnesses. It's all in let's build better, just hand coding, no machine learning, no intelligence, no Skynet here, but just hand coding these programs that we'll call LLMs. Let's just keep tuning and tweaking those to be better and better. And of course, the programmers building those particular programs, they're building them to do their type of work. So it's a field they understand really well.
59:11So they can really just sit here and twist and tune. And also, programmers are very adaptable. They like tools, and they'll adapt around the weaknesses or not. So it's kind of like a best-case scenario. But this is another indication of we're not getting these fundamental giant leaps in the capabilities of the digital brains. It's either some bench-maxing, like we tuned it to do better on a particular benchmark, or we built better programs around it. So when you put the money that they put in the mythos, and if really the best thing you had to emphasize when it was done is we have a cybersecurity benchmark where Opus 4.6 was at 66.7 and this is 83.1, that doesn't necessarily going to justify what's going on.
59:50Or that AISI has this – there's only one thing in there where they see a leap from Mythos at a particular contrived security scenario they came up with. And this big leap that got them all worried was Opus 4.6 could on average complete 16 out of 32 steps in this challenge. And Mythos on average could do 22 steps out of 32. That's hundreds and hundreds of millions of dollars of training, electricity or whatever. I think that's an issue. I just – I think that – and maybe this is a simplistic point. i don't think they know what they're doing at this point like i don't get the sense that anthropic or even open ai has a strategy because today as we're speaking so this will be out next wednesday but they released anthropic design the thing i mentioned the figma clone it's like why are you fucking cloning figma what are you doing you're trying i thought you're going to automate the economy yeah you're going to replace a so you've made a figma clone what like we heard the rumors last year that they were going to do a product um and open ai was going to do a productivity suite it's like why it's like they're doing everything they can to ignore the core problem which is the core technology is not going anywhere like because mythos appears to be they called it a step change but that's a nice way of saying incremental improvement it's 100 correct yeah and let me tell you why i would be worried if i was them here's the worrisome thing about mythos right is again they talk about these vulnerabilities hidden for decades that you know mythos found or what have you and they replicated multiple different independent security teams were able to find most of those vulnerabilities using three to five billion parameter open weight models so let me put that in perspective right a a model like mythos is going to have hundreds of billions if not a trillion parameters and they use a three to five billion parameter are off the shelf you could run this model on a chip inside your sorry 10 trillion 10 trillion oh okay that's crazy love the number bro is that true yeah that's what it's oh my god 10 trillion parameters is insane like you better be uh that better be either gaming the stock market and creating billions of dollars a days and like fancy option returns or changing lead in the gold because to run something that has 10 trillion parameters to do almost anything else, it's like we're going to launch ourselves in this space to do something and then land every time.
1:02:22That's so incredibly expensive. But the real fear then is like, well, wait a second. If they could do most of this stuff with a free, cheap model that I could just run on a machine at home, that's what keeps, I think, Dario Amadei up at night. That's what keeps Sam Altman up at night. It's the future. Look, I've been pitching this, right? I think the useful and the only ethical and sustainable future for AI is what I call distributed AGI. And I think it's just what the future is going to be, which is you have specialized applications for different things. Where, oh, we want to do this thing over here.
1:02:57We built something that has some AI in it. And maybe it has an LLM or it's a modular architecture and it has a billion parameter model in there and a world model. And it's really good at doing this thing. And it's small and it mainly runs on chip. and now this program can do this thing that I used to have to do. And you multiply that across 10 ,000 different use cases and you're like, oh, we kind of have AGI, right? There's all these different things that have AI tools that like do pretty well. That's like a completely, probably the most probable future. It's a future I really like for a lot of reasons.
1:03:27There'll be a lot of things that we can't make progress on. A lot of things we will, but it's a much more heterogeneous future. There's no giant how 9 ,000 brains. It's economically more interesting and diverse. It doesn't have all the sustainability issues. That has to be the future. But the problem about that future, if you're Sam Altman or Daryl Amadei, is that their entire moat is, unless you need 10 trillion parameters, they want that to be the key to the AI future because that moat is something that no one can cross. And if that's not the moat, if it's just, oh, if I want to build a poker playing AI that's really good, I just need people who are good at poker and to spend a couple of years and figure out a cool custom system.
1:04:03And that thing now does well. well, if that's the future, you don't need open AI and you don't need Anthropic. And I think that probably might be the future. And I think that's terrifying. They're trying to race to an IPO and they're marketing out of their butts. What can we do to keep things going so at least we can get our stock on the market? That's what would keep me up at night if I was them. Actually, there might be a lot of AI in the future and it's not going to be nearly as sexy as they're hoping. What if there's also, by the way, that 10 trillion number, I can't source it to Anthropic.
1:04:32I've seen it reported multiple places. This is a problem. They never talk about it. We have an issue with news right now. We're just like mythology spreads, ironic considering the name. But the other thing is as well, it's like hundreds of billions of trillion parameter. You're just using a nuke to kill a single gopher. Yeah. You're just like, we're going to throw everything we have at it. To the point that I don't know if you've been seeing the amount of trouble anthropic has had keeping its service online and how they're making the models dumber yeah it just feels like we're in this weird hysterical moment where no one knows why they're doing this but everyone's ready to accept whatever anyone said like it's just like oh we're all doing this insane thing so we're just going to repeat what kind of informs the bias and makes us look less dumb i think the more excited we are i think the frontier models are like f1 cars and the equivalent of points on the F1 circuit are your positioning on the benchmark leaderboards.
1:05:34So you do this, you build these giant models and you spend all this money in electricity and they're so big, they're not even economically viable to have people use, which might really be what's going on with mythos. It's like, we have to make this seem super premium because otherwise people are going to get charged $5 ,000 a month. And just like if you're Red Bull or Ferrari, your F1 car doing well on this leaderboard just lets people know, this company builds good cars, and then you can sell your normal cars. I think that's a lot of what's going on here, is that they want to be high on that leaderboard means we know how to do AI.
1:06:07We AI smart. Even though the future of actual consumer deployed products is going to be much more like a Honda Odyssey minivan than it's going to be like a top Formula One car. Well, Cal, it's been an absolute pleasure having you as ever. Where can people find you? You can find me at calnewport.com. my podcast is deep questions on Thursdays. The Thursday episodes are all AI reality checks, right? Take a fun story. Actually, Ed's coming up or he, he's, he may have already been on it by the time this comes out, or maybe it's the day after this comes out. So now you have to check it out. Now AI reality episodes get a double dose.
1:06:43You bring this out of me. by the way, you bring out my sort of ornery side. I'm normally like the very, very kind of stayed, uh, professor, New Yorker writer, just like, well, on the one hand, on the other, you bring this out of me. I love it. But the thing is, you're critical only of things that need to be, you're still willing to humor these things, as long as there's something to humor. And that's why I like having you on because people claim I'm just a hater. So we've got to have people for a little balance. But thank you for joining me. Thank you everyone for listening. You have a monologue coming up as well on Friday.
1:07:14Thank you all.
1:07:23Thank you for listening to Better Offline. The editor and composer of the Better Offline theme song is Matt Ossowski. You can check out more of his music and audio projects at mattosowski.com. M-A-T-T-O-S-O-W-S-K-I dot com. You can email me at ez at betteroffline.com or visit betteroffline.com to find more podcast links and, of course, my newsletter. I also really recommend you go to chat.wheresyoured.at to visit the Discord and go to r slash betteroffline to check out our Reddit. Thank you so much for listening. Better Offline is a production of Cool Zone Media. For more from Cool Zone Media, visit our website, coolzonemedia.com, or check us out on the iHeartRadio app, Apple Podcasts, or wherever you get your podcasts.
1:08:30Apple Vacations, where your story starts. Need a vacation? The Apple Vacations foundational sale is here, so you can save up to$150 on your next escape. Book by April 30th for instant savings on trips to Mexico, the Caribbean, Hawaii, Central America, and top U.S. destinations. If you're ready for a break without breaking the bank, save up to$150 at AppleVacations.com or contact your travel advisor. Apple Vacations, where your story starts. And now, the No Panic Party Save, brought to you by Grand Appliance. Your party is safe. Wow, Amy, great party. And the food, incredible. Thanks. And did I tell you my stove died two days ago?
1:09:12What? Did you panic? Nope. I called Grand Appliance and got a great deal with next day install on the Frigidaire gallery I wanted. Next day? Wow. GrandAppliance.com, right? That's it. My family's shopped there for decades. Shop Grand Appliance. Appliance experts since 1930. Hey everyone, it's Kel Penn. I'm inviting you to join the best sounding book club you've ever heard with my podcast, Earsay, the Audible and iHeart Audiobook Club. Every episode, I nerd out with amazing guests and dive into the best new audiobooks available on Audible. It's the book club for your ears. Listen to Earsay, the Audible and iHeart Audiobook Club on the iHeartRadio app or wherever you get your podcasts.
1:09:59Tyler Reddick here from 2311 Racing. Another checkered flag for the books. Time to celebrate with Chumba. Jump in at Chumbacasino.com. Let's Chumba. No purchase necessary. BTW Group. Voidware prohibited by law. CT &C 21 Plus. Sponsored by Chumba Casino. This is an iHeart Podcast. Guaranteed human.
From the publisher
In this week’s Better Offline, Ed is joined by computer science professor and writer Cal Newport to talk about the Claude Mythos marketing scam, the lies around AI job loss, and why LLMs shouldn’t talk like they’re people.
Please support me by subscribing to my premium newsletter - here’s $10 off your first year of annual https://edzitronswheresyouredatghostio.outpost.pub/public/promo-subscription/84rt762qen
Podcast & Videos: https://www.youtube.com/@CalNewportMedia/
Newsletter: https://calnewport.com
New Yorker archive: https://www.newyorker.com/contributors/cal-newport
YOU CAN NOW BUY BETTER OFFLINE MERCH! Go to https://cottonbureau.com/people/better-offline and use code FREE99 for free shipping on orders of $99 or more. Buy our new “FUCK DATA CENTERS” shirts today!
---
LINKS: https://www.tinyurl.com/betterofflinelinks
Newsletter: https://www.wheresyoured.at/
Reddit: https://www.reddit.com/r/BetterOffline/
Discord: chat.wheresyoured.at
Ed's Socials:
https://www.instagram.com/edzitron
https://bsky.app/profile/edzitron.com
https://www.threads.net/@edzitron
Email Me: ez@betteroffline.com
See omnystudio.com/listener for privacy information.
