In short
The episode covers the “model mayhem” moment in AI: OpenAI’s GPT-5.6 (general-purpose, expanded coding/agent abilities) and Meta’s Muse Spark 1.1 (agentic coding with tool use). Key claims include: GPT-5.6 Soul scores 7.78% on ARC-AGI v3, suggesting improved generalization/spatial reasoning; the “spiky frontier” means different models excel for different tasks; coding models enable fast creation of interactive “vibe-coded” browser minigames (e.g., a sailing mini-game on the OpenAI blog).
Notable examples
GPT-5.6 Soul autonomously post-trained GPT-5.6 Luna; comparisons between “Fable” and 5.6; a “bulletproof” magic trick and joke that GPT-5.6 reportedly solved/laughed at.
Guests
none are clearly identified in the transcript.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOAI Model Launches Overview
0:03 to 0:21
Discussion of various AI models being launched, including XAI's Grok 4.5.
“Today on TBPN, we're talking about model mayhem.”
Meta's Muse Spark Announcement
0:21 to 0:41
Announcement of Meta's Muse Spark, a new coding model.
Mark Zuckerberg's Presence
0:41 to 1:40
Discussion on Mark Zuckerberg's return to X and his activity online.
“he has not been an active user but the ai vortex sucked him in and he's got a post oh i think he's an active user john you think so he's just not an active post he's just not an active he's not He's a lurker.”
OpenAI's GPT-5.6 Features
1:40 to 1:56
Introduction to OpenAI's new GPT-5.6 model and its capabilities.
“A new general purpose model with expanded coding and agent capabilities alongside GPT Live, which we talked about yesterday.”
Benchmarking AI Models
1:56 to 2:18
Exploring the performance benchmarks for AI models and comparisons.
Analogies in AI Models
2:18 to 2:52
Using hip-hop analogies to describe differences between AI models.
“people are drawing analogies between Fable 5 being some recluse genius and 5.6 being a collaborative co-worker that you love chatting with or something like that.”
Insights from Arc AGI V3
2:52 to 3:08
Discussion on insights gained from Arc AGI's benchmarks.
“I mean, the funny thing is that will be very explicit for like 100 people in the whole world.”
AI's Progress and Puzzle-Solving
3:08 to 4:14
Exploration of AI's progress in puzzle-solving and reasoning.
“And 5.6 soul scored a massive 7.78%, which is tiny, considering that the whole point of Arc AGI is that a human should be able to get 100 % on it, and basically any human.”
Engaging with GPT-5.6 Games
4:14 to 4:44
Engagement with games created for the GPT-5.6 launch.
“The blog post is also very, very fun because it includes games.”
Interactive AI Developments
4:44 to 6:04
Discussion on the new interactive capabilities of AI and mini-games.
“I want production team to see what they can get.”
Show all 18 chapters
Magic Tricks and Humor in AI
6:04 to 8:34
AI's ability to understand humor and perform magic tricks discussed.
“Despite all the research achievements, we are still very, very early in exploring the tech tree for model training.”
Comparing AI Model Capabilities
8:34 to 10:34
Comparison of capabilities between different AI models including GPT-5.6.
“Only way to know it's funny is a first principle sense of humor.”
Market Dynamics in AI
10:34 to 12:06
Exploring market dynamics and growth within the AI industry.
“Although AI 2040 launched today, the sequel to AI 2027, That's something that's more of a thought-provoking piece that you can debate and interrogate and talk through.”
Meta's Keystroke Logging Discussion
12:06 to 14:01
Discussion on Meta's workplace keystroke logging initiative.
“And we're sort of like duking it out between those.”
Monitoring Work Processes at Meta
14:01 to 16:56
Learn how Meta is experimenting with monitoring work processes and the implications for privacy.
“It doesn't seem that crazy to go to keystrokes because everything is already so monitored.”
Meta's Muse Spark 1.1 and Pricing Strategy
16:56 to 18:34
Discover the features of Meta's Muse Spark 1.1 and its competitive pricing strategy in the AI market.
“In a crowded market for AI tools, Mark Zuckerberg wants to win on price.”
Internal Workloads vs. External Models
18:34 to 21:13
Explore how Meta is balancing internal workloads with external AI models and the economic implications.
“basically validate whether or not they should be using this model themselves, right?”
Ben Thompson on AI's Importance to Meta
21:13 to 24:08
Understand Ben Thompson's insights on why AI is crucial for Meta's future and its business model.
“And these are the same tradeoffs and decisions that every lab is having to make is how much compute do we allocate towards research, towards internal use, towards the API, to subscriptions, to free plans, et cetera.”
Transcript
Automatic transcript. May contain errors.0:00John Coogan:I need some soundboard. Here we go. Yes. Today on TBPN, we're talking about model mayhem. Everyone's launching new models. Slow summer, but not for the AI race. You got XAI unveiling Grok 4.5, the first model built specifically for coding and AI agents, developing collaboration with Cursor. Talked about it a little bit yesterday, but we have some more benchmarks, some more discussion on the timeline about where this model fits in on the Pareto frontier. here also why it might be outperforming so well on cursor bench uh lots of debates there meta announced muse spark a new agentic coding model with mark zuckerberg returning to x for the first time in basically a decade three years ago he posted one joke post about launching threads but he has not been an active user but the ai vortex sucked him in and he's got a post oh i think he's an active user john you think so he's just not an active post he's just not an active he's not He's a lurker.
0:57He's not an active contributor. You're calling him a lurker. I'm calling him a lurker. You're calling him a lurker. I'm calling him a lurker. I think he's absolutely glued. You think so? I think so. You really think so? I think so.
1:07John Coogan:I feel like, I don't know, so busy, so much other stuff going on. I feel like most people are on that level. The busiest people I know are not active on X, but they are on X a lot. Sometimes. But there's a different class of person. You can just quiz them. Screenshots come to them via Slack or via text message because they have a team that's monitoring the timeline and then is delivered. This is the important stuff. They're calling him Mark Lerkerberg. But the other big news, OpenAI just released GPT 5.6. Let's go. A new general purpose model with expanded coding and agent capabilities alongside GPT Live, which we talked about yesterday.
1:53John Coogan:A new real-time interactive voice experience. Reactions are great to 5.6. Bunch of interesting details here. People have been identifying that while there is a frontier and there are just a few companies that are actually on the frontier, the frontier is spiky and they have different flavors to them and different reasons to pull different tools off the shelf. people are drawing analogies between Fable 5 being some recluse genius and 5.6 being a collaborative co-worker that you love chatting with or something like that. I said, I don't know how else to describe it, but Fable 5 is like Kendrick on Good Kid, Mad City, and 5.6 Soul is like Chief Keef on Finally Red.
2:42Now it makes sense to me. Thank you for breaking it down. wanted to put it into 2010 hip-hop terminology.
2:49John Coogan:Really, really clear there. Thanks for clearing that up. I mean, the funny thing is that will be very explicit for like 100 people in the whole world. This one's for you. The most interesting benchmark to me has always been ARK AGI V3. We've interviewed the team over there many times and had a lot of fun understanding what goes into that benchmark. And 5.6 soul scored a massive 7.78%, which is tiny, considering that the whole point of Arc AGI is that a human should be able to get 100 % on it, and basically any human. So it is a true test of AGI in the sense of, you know, can you give this test to just actually anyone, not, you know, the crazy math projects, the crazy hard programming projects, the hacking, all of that stuff is very economically valuable, of course.
3:40John Coogan:But there's a more interesting question where when there's less of a spiky frontier and there's just this question of what is something that anybody can do that AI can't? Because we've been searching for those and the Arc AGI team has done a fantastic job building out these puzzles that AI has historically struggled with. Arc AGI, one, the model sort of climbed. Two, became a little bit more complicated. And now three, we're starting to see glimpses of progress, although 7.76 % isn't 99%. We're nowhere near saturation, but it's still a huge jump. Opus 4.8 had 1.5%, so GPT 5.6 Soul is showing more generalization, more spatial reasoning, more puzzle-solving abilities.
4:24John Coogan:So fun, fun stuff. The blog post is also very, very fun because it includes games. I'm a big fan of the GPT 5.6 launch games. I got immediately sucked into the to the sailing mini game, which is very high fidelity, but also delightful to actually play. Should we play it? Yes, we should definitely play it. Yeah, Saltwind. You guys play it. I want production team to see what they can get. I think my time was 25 seconds. And is this hosted on a site? I think this is, I mean, this is hosted on the OpenAI blog, but I think the idea is that you could vibe code this in the latest GPT 5.6 in the app, in chat GPT and then deploy it and have someone.
5:07John Coogan:Are you trimming the sails appropriately? Because it looks like you're losing speed. You're losing wind. It's not working. I'm going to smoke you. I got 25 seconds. Wow. Amateur hour over here. Look at this. Can you boost? Yeah, yeah. Well, the whole game, which you probably missed, is that there is a little bar there where you have to trim the sails to be in the sweet spot of the wind while you're turning. So as you see the bar, there's a recommendation for where you put the sails. You've got to keep that line in. See? It's moving over. You've got to press the down. I see. I see. Yeah, exactly.
5:42John Coogan:Keep trimming those sails while you steer the ship. This stuff is very, very fun. One interesting data point from the live stream, which was just an hour ago, they said, already Sol has been transforming our research program. As one example, GPT 5.6 Sol autonomously post-trained 5.6 Luna. Yeah, that's fair. A lot of people are having fun with that. Dylan Field says, a lot of people want to compare Fable versus 5.6. So this is a mistake. They're apples and oranges. Despite all the research achievements, we are still very, very early in exploring the tech tree for model training. Cool. Sorry. I'm just getting set up again.
6:20John Coogan:Oh, yes. I do think that – didn't Dylan Abbrascotto write something about this? What was the essay he wrote about interactive memes and this idea of generative AI enabling these vibe-coded minigames? We've been seeing a bunch of them with the Copybara simulator, the Coconut simulator, where it's something that's just a joke that's funny for a few people. And normally, you would instantiate that in a tweet. Or maybe if you were getting really crazy, you'd do a Photoshop edit of a meme. But now you can go and create a full minigame, something that runs in the browser, and soon something that runs in Unreal Engine and can actually be distributed on the Steam store.
7:03John Coogan:We're already seeing that with, like, the data center simulators and all these funny simulator games that are going on Steam. All the advances in the coding model certainly speeds up the ability to actually deliver polished software. I'm particularly excited for, like, next. Yeah, Dylan's title was The Future of Entertainment is Interactive. Yes, yes. But yeah, that's part of what I honestly love about AIs. There's a lot of things you can make now that never would have made sense to make because they would have taken you four days and it was good for like a small laugh. Now you can do it in four minutes and it's just fun.
7:38John Coogan:Yeah, I think there's going to be, there's, if you have some sort of like small custom, some sort of custom functionality in your business, it feels like there's a huge. Is this the David Senra simulator? later why is this david center late nights in a miami abandoned apartment complex rooms in 2015 just recording podcasts and reading this is very creepy like uh horror backrooms liminal space game stanley tang uh co-founder and cpo over at doordash says i have an insane magic trick that so far none of the models can figure out including mythos it's a bulletproof trick that i've shown to a hundred plus people including magicians that couldn't figure it out it's not anywhere on the internet only way to know it is through first principles reasoning told everyone i'll believe in agi when it can crack this trick well gpt 5.6 just did how i want him to i want him to actually open like well now okay like give us now that now that a model cracked because i feel like a lot of magic tricks are like sleight of hand so is he uploading a video or something like well yeah so john palmer says i have a hilarious joke that so far none of the models think is funny it's a Bulletproof joke that I've told to 100 plus people, including comedians, and no one laughed.
8:52It's not anywhere on the internet. Only way to know it's funny is a first principle sense of humor. Told everyone I'll believe in AGI when it tells me a joke. The joke is funny. Well, 5.6 just did. Huge, huge news. Huge news.
9:05John Coogan:GBP 5.6 is a Porsche. Fable's like warp drive. I had a different experience. Fable is an F1 car. 5.6 sole at Ultra as a Tesla Model X Plaid. Does it find things that Fable misses during plannings and coding? Yes. Yes, most of the time. But for the hardest problems, does Fable routinely find things that 5.6 doesn't? Also, yes, some of the time. Is 5.6 way faster and affordable? Yes. With an unlimited token budget, what am I currently using 95 plus percent of the time? GPT 5.6 from Siki Chen. So interesting take that the Parade of Frontier is alive and well, and everyone's duking it out for their slice of the AI opportunity.
9:44John Coogan:Very interesting seeing how the market share is shifting during a time of acceleration. You have multiple companies that are growing revenues, even accelerating revenues, while market share is declining because the overall market is growing so fast that if you're only growing at 300 % and someone else is growing at 400%, you're losing market share, but you have one of the greatest businesses by modern metrics. Very, very interesting dynamics in AI. It's also funny because yesterday with Ben Thompson, you were like, slow summer. And then in the span of 24 hours, you get Rock 4.5, Muse 1.1. Yeah, I mean, this isn't as dramatic as the AI talent wars.
10:24John Coogan:It's not as dramatic as... Or rippling deals. Yeah, yeah. This is new technology. And there's only so much of a take to be given around these things. Although AI 2040 launched today, the sequel to AI 2027, That's something that's more of a thought-provoking piece that you can debate and interrogate and talk through. I'm sure we'll go through some of it because they pose a couple interesting ideas of where AI might go and where they want it to go and how they want the industry to develop. Sort of advocating for a slowdown generally, but it's an interesting way they puzzle piece all the different geopolitical chips on the table.
11:05John Coogan:Of course, people are joking about the lead is widening because the Anthropic and OpenAI version numbers over time, GPT-6 is predicted. And it is, the model numbering, we were talking about this this morning, that the numbers, they sort of don't mean anything anymore. Do the model numbers mean anything in particular? It used to be the model number was the pre-train, and then the version number was the post-train, but then that sort of got flipped around. And now it's just like, do you feel like you're competing at a four-class or a five-class? So I wouldn't be surprised if we saw like Muse Spark, not release Muse Spark 2, but Muse Spark 6 or 5 and jump straight.
11:51John Coogan:I mean, Samsung wound up doing this where they jumped to the year, like sort of like the car manufacturers, where, you know, there's a 5 Series BMW, but then there's also just the 2027, because that's the actual model year that's relevant. The 2027 5 Series. Yeah, which is sort of odd. And we're sort of like duking it out between those. Yeah, I mean, I think post-reasoning models, you just have like a different way to scale the models besides just pre-training. So it's hard to bake that all into one number that is evocative of both those two ways. Yeah. So the number is becoming closer to the year in the second decade of the 21st century, basically.
12:30John Coogan:It's just like, is this on the frontier in 2026? You'll probably see a six by the end of the year in front of the models that are leading in the year 2026, something like that. I'm very interested with Google strategy because the rumor is that 3.5 Pro will be coming out this next week, I believe. But it's very odd going into the Gemini app right now and seeing that there's 3.5 Flash, but then you have to go back to 3.1 Pro. I think 3.1 Pro is the most advanced model, but they default you to 3.1 Flash Lite. and I would expect them to jump just forward to four, but I think that they're going to do 3.5 Pro, but it's been a little bit of a slower cycle there.
13:17John Coogan:As silly, I mean, obviously all these numbers don't really mean anything. They're marketing terms, but I still think they do actually stick in people's mind. And so there should be some strategy around them. Mark Zuckerberg is on a press tour. He's talking to the legacy media for the first time in a long time. Andrew Bosworth, the CTO of Meta, I also did an interview with the head of the Atlantic, dug into some of the launches around the glasses, and then also had a whole discussion in that podcast around the goals of the keystroke logging thing. It was interesting. I mean, it was framed as like a tough interview around surveillance in the workplace.
14:00John Coogan:And certainly the headlines were very scary. I don't know where I sit on it because I kind of always assume that everything you do at work is logged in the sense that, like, if you're on a work computer and every web page you visit is going through the network and monitored for traffic and security purposes and all the code you write and all the emails you write and all the documents are stored in the shared document. It doesn't seem that crazy to go to keystrokes because everything is already so monitored. But he was framing it as more of an experiment, something that they weren't sure was going to pan out, something that they allowed everyone in everyone at Meta.
14:41John Coogan:So there were there were certain sections of the workforce that were by default opted out. So anyone who was working on confidential or sensitive information was opted out of that program by default. He said he himself, Andrew Bosworth, was opted out of that program because he has a bunch of legal holds because they're getting sued all the time. So they can't be recording everything, I guess, that he's doing because then that would be admissible in court. And so all of a sudden the lawyer who's suing him would say, okay, great. In the email, you said, you know, we don't want to do this. But before you type that.
15:15Let's see your writing process. Exactly.
15:17John Coogan:Yeah. Let's see what sentence you typed and then deleted. Like, what word did you use before minimal impact? Did you say medium impact or whatever? So he was opted out. And apparently, I think all of the meta employees who were part of that program were able to turn it off indefinitely. Like, you could toggle it on and off. And the idea was that they wanted to collect information on how work plays out over a 12 to 18-month period. And they couldn't get that from any sort of data labeler because they needed to have very high skilled workers actually chopping wood on projects for a long, long time to see how projects go from start to finish.
16:01John Coogan:So basically, how do you compact the longest possible rollout, not just a single chain of code, but an actual series of meetings and decisions and trade-offs and everything that goes into making a decision in a white -collar workplace? Like, how do you actually reason through all of that? It's hard to distill that from just, oh, well, the code got written this way, so that's the right way to write the code. The code might have gotten written that way because a lawyer said, hey, oh, we have to do this. And then the marketer said, oh, well, you know, we have an activation with this person, so we need to integrate it this way.
16:36John Coogan:And then the business people came in and said, oh, well, like the margins will be better if we write it this way. And so it's not entirely first principles software engineering all the time when you're actually building real products. So interesting to see him sort of step into the, you know, a tough interview and sort of lay out his side of the story. But Mark Zuckerberg is in Bloomberg today pledging aggressive pricing with Meta's first pay-to-use AI, which is a funny framing for just an API for a model. But that's the way Bloomberg put it. In a crowded market for AI tools, Mark Zuckerberg wants to win on price.
17:14John Coogan:Meta Platforms unveiled a version of its most advanced artificial intelligence model, Muse Spark 1.1, that includes a new paid tier for developers, marking the first time. meta has charged businesses for access to its models and providing a new revenue stream it'll be among the most affordable options on the market zuckerberg said in an interview ahead of the release quote since this is not an open source model this is i think the first time that we're doing a real serious api and the pricing is going to be very aggressive and attractive makes sense i mean they own the data centers they're very efficient at building data centers they should be able to serve a model efficiently the new model standout improvement is is is in it's a agentic capabilities, the meta chief executive officer said.
17:56John Coogan:Agents are a big theme of AI this year with the label applied to systems that can complete multi-step tasks on behalf of the user. Zuckerberg described MuseSpark 1.1 as having quote, state-of-the-art or very close to it, agentic reasoning and tool use. The model is also greatly improved when it comes to coding and meta employees are using it internally to build products and features for various apps. Yeah, my big question is how quickly do they move all of their internal workloads onto their own models. So they're buying, they're getting access to models through Google, Anthropic, and OpenAI. I think that a lot of companies will look to Meta's own actions as a way to basically validate whether or not they should be using this model themselves, right?
18:38Because it was just within the last month that Google had said like, hey, we don't have capacity. We don't have enough capacity for all of Meta's demand for our models. And so, yeah, they can't get enough AI elsewhere, at least from some providers. And so how much of their workloads will they be able to run themselves is a big question.
19:00John Coogan:Yeah, Meta was one of the first companies to sort of reportedly be token maxing and have a leaderboard and all of that. If you have your own model and your own data centers, the incentive to token max is much, much higher because you're just paying the electricity on the cards that you're already depreciating. So you should sort of lean a little bit back into that, not that you want to be fully token maxing, but you do want your employees using the tools that you've built as efficiently and as effectively as possible. And it's just way cheaper to explore when you're not paying margin on another closed source model and you're not paying anything else and you're actually improving the model.
19:40John Coogan:So it makes a lot of sense for them to roll this out broadly. The interesting take that Ben Thompson had, which we didn't get to yesterday because we wound up spending the whole interview talking about Xbox. But the interesting dynamic is that when you are willing to sell API access, you're willing to sell compute directly, and then you're also using your own tool internally. It creates this economic incentive internally that you have an incentive to always go with the most profitable, the most economically efficient outcome. That can be very good for business, very good for the investments. they made.
20:18John Coogan:The trick is that you can wind up in a little bit of a situation where your business team or your enterprise sales team goes and sells all your compute capacity or all your chips, and then internally your team is frustrated that they're not making enough progress. So there's a little bit of a dance there, but in general it's a forcing function on the internal use of their tools to say, hey, why is someone willing to pay five times as much with the value that we're creating here? We spent a billion dollars on energy consuming our own LLM, and someone showed up and said, wait, we'd pay you five billion for that same compute power to run a different model and do a different task.
21:03John Coogan:It's like, why is their model not economically valuable internally? That would be the question. The flip side is that they do have low cost, so they should be able to To say, oh, yeah, we actually did. Yeah, we inferenced MuseSpark 1.1 internally, and we improved the ad model, and boom, we made a bunch of money. And these are the same tradeoffs and decisions that every lab is having to make is how much compute do we allocate towards research, towards internal use, towards the API, to subscriptions, to free plans, et cetera. Yeah, there was that funny semi-analysis, deep dive into anthropics forecast.
21:40John Coogan:and in there, I mean, some staggering numbers, really, really optimistic. But the flip side was, who was Ed Zitron, was taking shots at the fact that they had EBTIT. EBTIT. Earnings before. Training. Training. Interest. No, training inference and everything. No, earnings before, training, interest, and taxes. And what was odd about it was that Ed Zitra was saying it's like the new community adjusted EBITDA, and it is always odd when a new non-gap metric pops up. In this case, I think it makes a lot of sense because training runs do fit a depreciation profile. It's a little bit different. I don't know why you wouldn't just put it in depreciation, though, like just figure out how to account for training runs through a depreciation schedule.
22:33John Coogan:And then maybe it's like a non-gap depreciation metric, but it's still in there instead of trying to get everyone up to speed on a different sounding phrase entirely. Yeah, I was looking back at Ben Thompson's earnings transcript or a script that he wrote for Mark Zuckerberg. He has a good segment on why AI matters. Ben writes, forgive the long preamble, but this is necessary context for me to properly explain why AI is so important to Meta and why I'm making the right choice to invest so heavily in both talent and infrastructure. And he goes on and on and on. But he says, what I've come to realize as I've embraced our status as an entertainment provider and ad purveyor is that our nature as a digital business, nonwithstanding, we are remarkably well placed to thrive in an AI era.
23:20Remember what we learned about humans. They are obsessed with other humans and they want to connect with them. That obsession and desire are only going to increase as we interact more and more with AI. AI is going to make our properties more essential, not less. Moreover, AI is a productivity tool, but productivity is not the end-all, be-all of the human experience. I've talked over the last year about building superintelligence that helps you get things done, but that's a business story. What we can do uniquely is give people the experiences they want from connection to entertainment to shopping when they are off the clock.
23:50The fact that we are investing in AI but not selling solutions to businesses is actually one of our business biggest advantages. So, of course, this is just a sort of fan fiction for an earnings transcript. Meta is, in fact, selling to businesses now. But who knows over time how big will the API business be relative to how much value they can unlock across their broader business with all of their infrastructure.
24:19John Coogan:Well, leave us five stars on Apple Podcasts and Spotify. Sign up for a newsletter, tbpn.com. And we will see you tomorrow at 11 a.m. � Arum Ashmade Eisen Aber
24:33Elizabeth Aber Elizabeth
From the publisher
Diet TBPN delivers the best of today’s TBPN episode in 30 minutes. TBPN is a live tech talk show hosted by John Coogan and Jordi Hays, streaming weekdays 11–2 PT on X and YouTube, with each episode posted to podcast platforms right after.
Described by The New York Times as “Silicon Valley’s newest obsession,” the show has recently featured Mark Zuckerberg, Sam Altman, Mark Cuban, and Satya Nadella.
TBPN is made possible by:
Ramp - https://ramp.com
Public - https://public.com
Cisco - https://www.cisco.com
Console - https://www.console.com
CrowdStrike - https://www.crowdstrike.com
Figma - https://www.figma.com
MongoDB - https://www.mongodb.com
NYSE - https://www.nyse.com
Railway - https://railway.com
Shopify - https://www.shopify.com/
Follow TBPN:
https://TBPN.com
https://x.com/tbpn
https://open.spotify.com/show/2L6WMqY3GUPCGBD0dX6p00?si=674252d53acf4231
https://podcasts.apple.com/us/podcast/technology-brothers/id1772360235
https://www.youtube.com/@TBPNLive



