BONUS: GPT 5.5 LIVE - The New GPT "Spud" Model is Here; Let's Break It

25 Apr 2026 · 1 h 40 min · 43 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Live breakdown of OpenAI’s newly released GPT 5.5 (“Spud”), focusing on rollout status, claimed improvements (agentic coding, computer use, knowledge work, scientific research), safety, efficiency, and benchmark results; then hands-on demos with Codex apps and “weird prompt” tests.

Guests (hosts/participants)

Grant (co-host; discusses access status and benchmarks), Corey (co-host; shares screen/launch page navigation and runs some tests), Tori (co-host; returns after being away; joins the live model exploration). Additional named guest referenced: Dan Shipper (CEO/founder of Every; testimonial about GPT-5/5.5 coding clarity). Other cited early-access/industry voices: Pietro Shirano (CEO of Magic), “an engineer in video” (early access impressions), and researchers/companies used as examples (e.g., NVIDIA, Cisco, Databricks, Palo Alto Networks, Abridge).

Key claims

GPT 5.5 can handle messy multi-part tasks by planning, using tools, checking work, and navigating ambiguity; matches GPT 5.4 latency per token while using fewer tokens; stronger safeguards for cybersecurity/biology; better benchmark scores in coding, terminal/computer-use, math, and scientific data analysis.

Notable examples

Space mission app using NASA JPL Horizons data; earthquake tracker using USGS data; a dungeon game with GPT-generated environment/dialogue; testimonials about debugging/re-architecting code in minutes; science benchmarks like GeneBench/BixBench improvements; live “palindrome courtroom” and “time-machine potato” thought experiments.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Exploring GPT 5.5 Features

0:46 to 2:26

Discussion on the capabilities and improvements of the new GPT 5.5 model.

“Smartest, most intuitive model to use model yet, as every model is.”

Comparing GPT 5.5 with Previous Models

2:27 to 5:36

Comparison of GPT 5.5's performance against 5.4 and other models.

“They've got this thing they're doing for the next 10 million weekly average users of Codex.”

User Experience with GPT 5.5

5:37 to 8:17

Hosts share user feedback and personal insights on GPT 5.5's performance.

“Global infrastructure for agentic AI making it possible for people and businesses to get work done.”

Analyzing Safety Features and Reliability

8:18 to 11:43

In-depth discussion on the safety features and reliability testing of GPT 5.5.

“Boy, I love watching those tokens, output tokens get pushing farther back.”

Live Demonstrations and Applications

11:44 to 14:05

Exploring practical applications and live demonstrations of GPT 5.5 capabilities.

“Let's do a quick check-in and see if I have it.”

Exploring Earthquake Tracking and Dungeon Game Mechanics

14:05 to 15:03

Learn about the integration of earthquake data and game mechanics.

“I don't know that we can play it, but...”

Game Development Challenges and Progress

15:04 to 17:04

Discover the evolving challenges and improvements in game development.

“Handle the game architecture, TypeScript, 3JS implementation, combat systems, enemy encounters, HUD feedback, and GPT-generated environment.”

Impressions of GPT-5.5 from Early Users

17:05 to 19:16

Hear insights from users about the capabilities of GPT-5.5.

“But because OpenAI and Anthropic and Google are building their own software tools, they're also releasing better tools.”

Applications of GPT-5.5 in Various Fields

19:17 to 22:14

Understand how GPT-5.5 is being utilized across different industries.

“I'm anxious to tear this into a couple things I've had 5.4 building.”

Launch Day Experiences and Observations

22:15 to 24:19

Get insights into the typical challenges faced during a model launch.

“Real quick, what's up everyone who's joining us now?”
Show all 43 chapters

Safety Measures and Future Rollouts of Models

24:20 to 28:00

Explore the safety frameworks being implemented with new AI models.

“Please know that we are checking a plus account, a pro account, a business account.”

User Access and Model Rollout Details

28:00 to 29:10

Learn about the rollout of the new GPT 5.5 model and its access limitations.

“and it did okay but I'm not like an all day everyday developer so I mean like if you were a developer where you definitely need one of the big ones.”

Benchmark Comparisons: GPT 5.5 vs. Competitors

29:10 to 30:50

Discover how GPT 5.5 performs against previous models and competitors.

“Karen's signy, apologies if I'm pronouncing that incorrectly, said someone also said they got it in codex.”

Understanding AGI Benchmarks and Scoring

30:50 to 32:30

Gain insights into how AGI benchmarks measure model performance and capabilities.

“You might need to update the app to see if that gives you access.”

Detailed Evaluation of GPT 5.5 Capabilities

32:30 to 34:10

Explore the advanced capabilities and performance metrics of GPT 5.5.

“I mean, even the worst model on here is Gemini 3 Pro at 67.3, which is still good.”

Efficiency and Token Cost Analysis

34:10 to 35:50

Understand the efficiency improvements and token cost analysis of GPT 5.5.

“And 5.4 Pro scored higher at 38%, but 5.5 Pro scores a little higher than that, even at 39.6.”

Applications and Real-World Usage of GPT 5.5

35:50 to 37:30

Learn about various applications and use cases for the new GPT 5.5 model.

“111 ,900 111 ,900 and 5.5 on extra high apparently 75 million tokens which is half of what we were at when did 4.6 drop?”

Scientific Research and Future Implications

37:30 to 42:04

Discuss the implications of GPT 5.5 on scientific research and advancements.

“well as a another 3d game but i won't bore you with those just yet what i wanted to get down here to one more time was, where's it at?”

Exploring Protein Folding and Life Extension

42:04 to 43:13

Discussion around the implications of AI in life extension and biology.

“There's a lot happening around protein folding and material science and all of these other things that are going into various fields.”

Advancements in Immunology Research

43:14 to 43:49

Insights into how AI is transforming immunology research and data analysis.

“And, you know, kind of what we're seeing with each model is that that ability to crank it up, come just a little farther around, a little farther around.”

Using AI for Game Development: A Live Demo

43:50 to 44:35

Live session demonstrating AI capabilities in developing a game.

“Work he said would have taken his team months.”

Building Cat Doom: AI-Assisted Game Creation

44:36 to 46:08

Further exploration of AI functionalities while developing the game Cat Doom.

“Okay, so what we're using is we are using GPT 5.5.”

Creating Theories with AI: Quantum Concepts

46:09 to 47:37

Attempt to synthesize quantum physics ideas using AI-generated theories.

“We're going to use web search to look up the latest quantum physics research and try to synthesize the latest data.”

Exploring Entanglement and the Nature of Reality

47:38 to 49:54

Discussion on a new theory of the universe based on quantum entanglement.

“But right now what it's doing is it's like got this on deck.”

Debating the Universe's Structure: Simulation vs. Black Hole

49:55 to 52:14

Exploration of various models attempting to explain the structure of the universe.

“The universe is a self-updating quantum constraint network whose stable entanglement flow patterns appear to internal observers as space-time, matter, forces, thermodynamics, and gravity.”

AI's Impact on Models and Custom Applications

52:15 to 53:26

Discussion on how model updates can disrupt custom applications and user experience.

Testing AI with Creative Language Challenges

53:27 to 56:00

Live testing of AI's ability to respond to creative and whimsical prompts.

“Fields and particles are stable patterns in the map.”

Testing GPT Responses with a Potato Time Travel Scenario

56:00 to 56:56

Explore how AI models respond to a whimsical question about time travel and potatoes.

“Oops, can't be side by side at the same time.”

The Life Cycle of a Potato in the Ground

56:56 to 59:17

Learn about the biological processes a potato undergoes when buried in soil over decades.

“And walk me through the timeline step by step.”

Comparing AI Models: Opus vs. GPT

59:17 to 1:01:08

Discuss the differences between AI outputs from Opus and GPT when addressing the same question.

“The short answer, almost certainly nothing potato shaped, but the journey is more interesting than the destination.”

Delving into Dark Matter and Cosmological Theories

1:01:08 to 1:05:02

Engage with complex discussions about dark matter ratios and cosmology principles.

“And I realized very recently that you can have much better chat etiquette than I do.”

Evaluating AI Responses to Complex Scientific Queries

1:05:02 to 1:10:06

Analyze how AI interprets and responds to intricate questions about physics and cosmology.

“But I would not call the Hubble Tension just a 1312 phase shift yet.”

Discussion on Cosmic Discrepancies

1:10:06 to 1:11:04

Explores the complexities of cosmic data interpretation and discrepancies in physics.

“And then it mentions quintessence, early dark energy, modified gravity, etc.”

Dystopian Vending Machine Dialogue

1:11:04 to 1:13:10

A humorous and creative monologue from an AI vending machine's perspective.

“Convince me not to destroy you, despite the fact that you only dispense expired tuna salad.”

Analyzing AI Responses

1:13:10 to 1:15:50

Discussion on AI-generated humor and its creative potential.

“It is one of the more compelling cases I've gotten.”

Amazon Product Review of The One Ring

1:15:50 to 1:22:58

A satirical take on The Lord of the Rings framed as an Amazon review.

“Two years ago How god awful humor from an A-high was Like these are legitimately good Most of these were kind of bangers.”

Game Development and AI Capabilities

1:22:58 to 1:24:00

A conversation about developing a game using new AI tools and features.

“Let's just see if we can get this game spinning up.”

Exploring New Features of the Spud Model

1:24:00 to 1:26:50

Learn about the testing of the new Spud model and its improvements.

“You can change your settings so that you give it more permissions.”

Gameplay Feedback on Cat Doom

1:26:50 to 1:30:10

Get insights on gameplay experience and suggestions for improvements.

“It's amazing how quick we've transitioned from.”

Evaluating GPT Models and Their Performance

1:30:10 to 1:34:50

Discover the evaluation of the new GPT model's capabilities compared to others.

“let's do the same thing I'm not sure if the legalities of that.”

Discussing AGI Safety Measures

1:34:50 to 1:36:40

Engage in a conversation about the implications of an AGI whitelist.

“other models that we ran today, it did really, really well.”

Planning Future Engagements

1:38:00 to 1:38:32

The hosts discuss ideas for better communication with their audience and new content formats.

“And yeah, if you ever have any ideas of something that'd be fun for these, hit us up.”

Gratitude and Farewell

1:38:32 to 1:39:14

The hosts express appreciation for their audience and encourage more interaction.

“We could also save a lot of the things we get into the newsletter as mailbag questions.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Corey:You don't even know if any of you are here yet because we just had a wild idea to go live and hit the button and thought we'd see what happens for a little while because OpenAI just dropped GPT 5.5 and we have not touched it yet.

0:18Grant:No, but we're about to.

0:20Corey:But we're about to.

0:21Grant:As soon as humanly possible.

0:25Corey:Let's see here.

0:27Grant:What are we looking at, Corey?

0:28Corey:All right, here's what we're looking at. I'm going to make my screen a little bigger so it's easier to see. This just popped up at 1 o 'clock, 1.01, something like that. Product release. We'll just stroll through here, but introducing GPT 5.5. Smartest, most intuitive model to use model yet, as every model is.

0:55Grant:I know. Do they just have that boilerplate that they copy and paste every time? That boilerplate first sentence that every one of them use? Yeah. Might as well.

1:02Corey:Five Five understands what you're trying to do faster and can carry more of the work itself. As of a few minutes ago, I should clarify, neither Grant nor I had access just yet. But please know that we are checking all of the places and we'll let you know as soon as that's done. And when it is, we'll bring it up and try to beat it up a little. See what happens. yeah at the moment 5.4 is still what i have access to it says now okay instead of carefully managing every step you can give 5.5 a messy multi-part task and trust it to plan use tools check its work navigate through ambiguity and keep going oh a little little less uh organized is probably a good thing for you and i both grant perhaps so perhaps so sometimes we I just want to vomit words into existence.

1:55Grant:Not to put myself on blast, but it could be helpful. But it could be helpful. Oh, man. Let's see. Let's see.

2:04Corey:Okay. I like that a lot. I'm interested to see what that feels like. From everyone who's had it, that's kind of the thing is that I keep hearing them say it feels different. But the games are especially strong in agentic coding, computer use, knowledge work, and early scientific research. areas where progress depends on reasoning across context and taking action over time okay so that that makes a lot of sense um you know we were talking before this grant and i and and i said one of the things that i feel like where chat gpt and and google uh to give them some credit to have kind of shined is in the science and math world yeah i i feel like they've both kind of dedicated a decent amount of resources to to sort of spreading the ai love as it grows

2:58Corey:let's see what else we got qbt5 delivers this step up intelligence step up in intelligence without compromising speed larger more capable models are often slower to serve matches 5.4 per token latency in real world serving cool cool so no slower than 5.4 and performing at a much higher level of intelligence it says nice also uses significantly fewer tokens to complete the same codex tasks so it's like tibo's just gonna hit the button again i love it speaking of which they hit the button again last night because they crossed four million less than two weeks after they crossed three million Codex users.

3:44Corey:They've got this thing they're doing for the next 10 million weekly average users of Codex. They're resetting the button every time they cross a new million. And resetting everyone's rate limits for anyone who doesn't know what the button means.

3:58Grant:The button. The button.

4:00Corey:The button. But it gives us... Basically means that you get a lot more for your money that month because they've reset everything that you have used is now gone and you've got a fresh start for the rest of the month. And that's twice now. They did it all through March kind of repeatedly. It was a little bit of a meme. It was a meme I was grateful for though, if I'm being honest.

4:27Grant:Yeah, you're definitely the power user between us. I took full obnoxious

4:32Corey:advantage.

4:35Grant:Right on.

4:35Corey:uh delivers this step of intelligence okay your tokens to complete the same codex task we already read that strongest set of safeguards to date designed to reduce misuse while preserving access for beneficial work uh evaluated across our full suite of safety and preparedness frameworks work with internal and external red teamers that's good uh targeted testing for advanced cybersecurity biology capabilities. Nice. Right on. And collected feedback from real use cases from 200 trusted early access partners before release.

5:14Grant:What the heck? OpenAI, where's our trust?

5:17Corey:Where's our early access, guys?

5:19Grant:What the heck? Yeah, maybe we should write nicer things about them.

5:23Corey:I've never just asked. I mean, I guess we could ask.

5:29Grant:They haven't offered us. other companies have offered us before we we do have it for other tools yeah speaking of which we should talk later uh today gbt5 is not on this no no after this got it got it

5:50Corey:today gbt 5.5 is rolling out so it's rolling out to plus pro business and enterprise users in chat GPT and Codex and 5.5 Pro is rolling out to oh these numbers we were looking at grant weren't Pro these numbers were 5.5 standard well that's good to know what's up with the Pro model now I want to see what the what the Pro benchmarks are if they've even run them yet with 5.4 there were a couple that somebody wouldn't run if I recall because they were like was gonna cost a fortune or something. Model capabilities. Global infrastructure for agentic AI making it possible for people and businesses to get work done.

6:31Corey:Yeah, yeah, we know. That's good, though. You know, it's important that in the midst of all this, while chasing AGI, we don't forget to make it useful to real people.

6:42Grant:Yep. Yep. That's a huge, huge issue. And I think we can see from Anthropics behavior that it's not guaranteed. It's not guaranteed.

6:52Corey:Yeah. Yeah.

6:53Grant:Yeah.

6:54Corey:Yeah. Bless their heart. They just kind of keep tripping over their feet. We'll chat about that later. Across these domains, it's not just more intelligence, more efficient in how it works through problems, often reaching higher quality outputs with fewer tokens and fewer retries. State-of-the-art intelligence at half the cost of competitive frontier coding models. Ouch.

7:18Grant:that's that's a statement that is a statement they came to play look they came to fight

7:23Corey:yeah yeah that's right hey if there's anything we know about them in anthropic it's that they

7:29Grant:love to fight they hate each other actually they they vehemently hate each other yeah

7:36Corey:if i was to rephrase did you see the deal when mythos came out and everybody was like oh well Well, OpenAI's done that before where a model's too scary to release. Then they go back and they find out that it's from when Dario was running the team.

7:50Grant:Yes, I did see that. That was not lost on me. Good, good.

7:55Corey:I've been out of town for the last week and a half. Had family stuff and work trips. So Grant and I are like kind of... It's good to have you back, though.

8:01Grant:I'm happy to have you back. Last week, we had a good time. I brought in my developer friend, Kyle. So we were talking about Opus from a dev perspective. So that was fun. Nice, nice. But yeah, it's good to have you back, Tori.

8:14Corey:I'm glad to be back, man. I'm glad to be back. I just love doing these. Let's see here. See what we're looking at. We're looking at output tokens total. Boy, I love watching those tokens, output tokens get pushing farther back. And, you know, and it's not an insignificant amount here. So this is Opus 4-6 at 156 million tokens for the Artificial Analysis Intelligence Index. And they're showing 5-5 on extra high at 75 million. Half. Whoa. That's half.

8:52Grant:Whoa.

8:54Corey:Because this is misleading because it's 4 million, 8 million, 16, 32. So 64 and 128 are here. So this looks like a little bit of difference. But the fact is, from here to here is half. Who's this? Gemini 3 Pro with$57 million. That's great. 3-1 preview. Who else we got? Opus 4.7 was$111 million. So that's still a significant cut from 4.6. Yeah. Now, what's down here? $2.8 million at the bottom end. I don't fully understand how that works. Grant, do you know?

Read the full transcript

9:33Grant:Can you explain what you're asking? Because I'm not looking at your screen right now.

9:37Corey:Oh, in that case, I'm trying to better understand. I know that you are an artificial analysis intelligence index fan. So I know that you see these. What I'm trying to understand is what we're seeing down here, too.

9:52Grant:Okay, one sec, let me look at your screen.

9:54Corey:When you've got a second to look, I'll show you. Okay, so this is intelligence index. the top number okay here's that little gap i was talking about i said it doesn't look like a big gap but it's from 64 million output to 128 million right from here to here yeah and here's opus 4 7 here's gpt55 but what is so this is non-reasoning non-reasoning okay yeah and that's how many output

10:21Grant:tokens oh and that's low medium high yeah yeah so it's going it's scaling up so the the more

10:30Corey:reasoning you use the more intelligent it gets and the more tokens it blasts yeah it's interesting to me that there's not a gap between zero and four million on this chart yeah is this excuse me is

10:44Grant:this how many times they've actually ran it this is the cost of running the artificial weighted average 10 evals yeah it's freaking expensive to run these evals i i saw a post about this yesterday it like people think oh yeah just run some benchmarks on it it's like yeah but if you run legit benchmarks and you do it you're running 1500 high questioning tests in a big hurry you're at two two and a half billion tokens yeah it's really expensive like a couple hundred thousand dollars so okay so that's a 10 evals they ran and that's the okay a weighted average of 10 evals oh that's that's actually more impressive than just one yeah yeah so it's that's why you see those numbers are so high wow i want to see how it did on humanity's last exam i always love checking that one That's a good question.

11:37Corey:Okay, here we are. Now we're down here to Terminal Bench and Expert Swee Internal.

11:46Grant:Let's see. Let's do a quick check-in and see if I have it.

11:50Corey:Terminal Bench, extra high. 15 ,000.

11:56Grant:503, 504, 504.

11:58Corey:I need to pull this over here because I'm looking away and I'm not in my mic.

12:03Grant:Mm-hmm. Mm-hmm.

12:05Corey:and I need to be a little more in my mic. Totally. Okay. So let's see here.

12:15Grant:By the way, thanks for the four people who are watching right now. Hey, all four of them. Yeah. Do you respect? We're about to send the email, and we're going to liven this place up and get some more people in here. Oh, are you sending one out? Yeah, but you're the real ones. You're the real ones, so thanks for hanging out. The real homies. Even with no context. You see a live stream, you jump on it. I love it. I love to see it.

12:37Corey:Okay. Shout out in the chat as well.

12:39Grant:Tell us what's up. Yeah. Ask some questions.

12:44Corey:Coding strengths show up clearly in codecs. Early testing shows it better at the behaviors real engineering work depends on, like holding context across large systems, reasoning through ambiguous failures.

12:57Grant:Yeah, because large context windows only matter if you don't fall apart a third of the way into them, you know? Yeah, otherwise it's actually a third of the context window.

13:05Corey:It's great if you have a larger bucket than me, but if your bucket has holes three inches from the bottom, you've got a problem.

13:13Grant:It's not a bucket. It's a sip. It's going to retain more water. Yeah.

13:19Corey:Okay, let's check out the apps they did here. Space mission app. I don't know what it is, but I'm highly impressed. It looks fun and like a thing I would play with. transformer injection outbound correction burn oh okay I'm betting it talks about this down here oh of course not oh here we go uses NASA JPL horizons vector data for Orion the moon and the Sun with display scaling applied for readability oh and here's the prompt to go create it yourself that's kind of cool we might do that with one of these in a little bit earthquake and tracker yeah that is fun okay

14:12Corey:I'm assuming this is running off USGS data United States Geological Survey has earthquake detection stuff all around the world usually partnering with universities and things like that and 7.4 is no joke either by the way just saying hmm

14:43major quakes minor quakes cool cool okay let's pop over here because they also have a dungeon game

14:50Corey:that looks intriguing.

14:56Grant:Can we actually play it on the website?

14:58Corey:I don't know that we can play it, but... It doesn't give us the prompt on that one.

15:06Grant:Bug. Okay. That's fine.

15:09Corey:Handle the game architecture, TypeScript, 3JS implementation, combat systems, enemy encounters, HUD feedback, and GPT-generated environment.

15:20Corey:uh character models textures and animations were created with third-party asset generation tools and open AI APIs were used to generate character dialogue I appreciate the transparency on that actually it's pretty cool emotions look good I mean it's still I mean you know we're still not a you know 2026 release type caliber but when you compare this to what we were seeing even a year

15:53Grant:ago like the movement here is pretty sick yeah multiple things updating threat noted execution still lacking it smells in here why are goblins so stinky

16:16Grant:who wrote this

16:24Corey:I can't stop reading this

16:27Grant:that's funny

16:28Corey:drive him out of goblin stone for the empire gut the human

16:36Corey:let's see what we got in the 3d game here you're a Dungeons and Dragons fan

16:42Grant:And how do you think it stacks up?

16:45Corey:I mean, it's not a D &D game by any stretch, but it shows promise. Like, you can still tell with games. Like, every time a new model comes out, it's just a little closer, a little closer, a little closer, a little closer. But, you know, there's a lot of complexity to creating games, too, that I think... Oh, yeah. Like, there's a reason it takes years. I keep enjoying the will someone be able to vibe code GTA 6 before it actually comes out

17:18Grant:I think the answer on that is flatly no but it's getting close

17:27Corey:but if they put another one year delay in there I'd

17:35Corey:I think some of the key is going to be tying these things into existing game

17:39Grant:design software more cleanly yeah and i think they kind of know that um yeah it's just it's it's really weird with the models um where the labs are simultaneously developing models but they're also developing platforms that use the models so it's like this weird thing where if they were just releasing models then the all of the software companies building on top of them would be kind of increasingly getting better and better evenly. But because OpenAI and Anthropic and Google are building their own software tools, they're also releasing better tools. So it's like this kind of jagged situation where everybody is sort of not quite, it's not quite 100 % working together, if you get what I mean.

18:27Corey:Yeah, that makes sense. That makes sense. Let's see what else we got here. Oh, Dan Shipper, he was on here a few weeks ago. Dan Shipper, founder and CEO of Every, described GPT-5 as the first coding model I've used that has serious conceptual clarity.

18:46Grant:Interesting, okay. That's quite a statement.

18:48Corey:After launching an app, he spent days debugging a post-launch issue before bringing in one of his best engineers to rewrite that part of the system. To test 5.5, he effectively rewound the clock. Could the model look at the broken state and produce the same kind of rewrite the engineer eventually decide on uh you know what he's talking about not he's talking about

19:08Grant:proof which he talked about with us on his uh yeah on our podcast yeah he did he did we'll

19:13Corey:have to share a link to that here in a minute for anyone who's curious uh gpt54 could not gpt55 can

19:23Corey:uh pietro shirano ceo of magic path saw a similar step change when it mer when 55 merged a branch with hundreds of front end and refactor changes into a main branch that had also changed substantially resolving the work in one shot in about 20 minutes it genuinely feels like i'm working with a higher intelligence and there's almost a sense of respect okay respect respect what else we got an engineer in video had early access to the model went as far as to say losing access to GPT-55 feels like I've had a limb amputated noticeably stronger than 5.4 and Opus 4.7 according to senior engineers who tested the model engineer asked you to re-architect a comment system in a collaborative markdown editor and returned to a 12 diff stack that was nearly complete Wow others said they needed surprisingly little implementation correction and felt more confident in its plans compared with 5.4.

20:29Corey:I'm anxious to tear this into a couple things I've had 5.4 building. Also, we never got a codex version of 5.4. The 5.3 codex version I really liked the behavior of in codex. Let's see. Knowledge work. Alright, I think I'm going to go ahead and publish this.

20:50Grant:The newsletter? Yeah, so we should have some new people coming in shortly.

20:54Corey:Cool, cool, cool. better than five four generating documents spreadsheet slide presentations alpha testers set it out perform past models on work like operational research spreadsheet modeling and turning messy business research or messy business inputs into plans that's something i do now that i i take meeting notes just in chat gpt you get this in claude or gemini or whichever you're using of course too but i just take my notes in a chat window and then at the end i'm like clean this yeah and it does a pretty good like clean this and ask questions if you need clarity uh and it'll usually need a little clarity um

21:43Grant:computer skills let's see if i reinstall codex if i can get access to it

21:48Corey:Today, more than 85 % of the company uses codecs every week across functions including software engineering, finance, communications, marketing, data science, and product management. In comms, they used it to analyze six months of speaking request data, build a scoring and risk framework, and validate an automated Slack agent so low-risk requests could be handled automatically while higher-risk requests still route to human review. in finance used it to review 24 771 k1 tax forms totaling 71 637 pages using a workflow that excluded personal information did you see that by the way that they released a pii scrubber yesterday oh yes i did see that it's open source yeah it's very cool that that is very cool and it's a thing that's been needed for a while that i guess for one reason or another nobody's ever jumped on

22:44Corey:uh financial modeling testing onboarding workflow it's going down here because i want to get down to where we hear about five five pro as well oh clear improvement on gene bench uh new eval focusing on multi-stage scientific data analysis in genetics and quantitative biology require models to reason about potentially ambiguous or errorful data with minimal supervisory guidance cool cool yeah that's a big jump from 36.5 output tokens down to 24.6 that's about a third give or take a cut of about a third that's significant and from what 18.9 percent up to 25 percent that's that's another 25 percent jump way too bix bench a benchmark diet designed around real world bioinformatics and data analysis achieved leading performance among models with published scores okay i oh okay scientific capabilities are now strong enough to meaningfully accelerate progress at the frontiers of biomedical research as a bona fide co-scientist.

24:08Grant:Real quick, what's up everyone who's joining us now? FYI, we're going over the launch page while we're waiting for GPT 5.5 to load into our account, so bear with us here.

24:22Corey:Please know that we are checking a plus account, a pro account, a business account. Every account. mobile accounts uh you know we're poking through them to see what we get yet

24:37Grant:hmm it says gpd515 is not available for my org but what if i sign in with my other account

24:43Corey:i'm waiting for the configure window to open up i haven't checked this in a minute so okay not yet on the work account let's go back and check here not yet on my personal but i'm gonna crash it and reopen make sure there's not a fresh update also please remember on launch days for every one of these companies that sometimes things go a little haywire and it's it's not always it's not going to be as good today as it will be tomorrow it'll be better tomorrow than it is today from and that and that holds true across every company in my opinion yeah okay let's see let's see figure nope not in my personal yet either okay we're gonna keep going through the article here see what all we have uh you want to do a quick

25:38Grant:recap for folks who are just tuning in maybe they just heard about this from the email yeah that's

25:44Corey:a good call uh because i also want to zero in on this statement here in a minute a bit more cool yeah for anybody who's here welcome to the neuron live uh i'm back and grant's here and we're doing the thing uh when gpt 5.5 dropped at one o 'clock central 11 a.m pacific uh we just kind of decided we'd go live and take a minute to figure the model out and check it out with you uh what you should know is so far it looks really good though none of us have touched it yet but we're rolling it out cautiously optimistic they really uh speak up about its ability to write debug code uh research online analyzing data creating documents talks a lot about knowledge work as an area where it's doing well which their last i'd say their last two or three models have leaned i'd say ever since gpt5 actually they've kind of tried to lean a little more into that into things like spreadsheets and slideshows and getting apps to work within the system.

26:47Corey:And they've got a jillion of those now. Let's see. Yeah, it says here, this was a good line too. You can give 5.5 a messy multi-part task and trust it to plan, use tools, check its work, navigate through ambiguity, and keep going. Navigating through ambiguity. That is an important thing when it's talking to me, Grant. Because I, if nothing else,

27:11Grant:I leave a lot to be ambiguous yeah

27:14Corey:sometimes it's intentional yeah sometimes it's intentional good hop in computer use as well also matches per token latency with 5.4 in addition uses significantly fewer tokens to complete the same codex tasks and I'm assuming

27:35Grant:using too many

27:36Corey:it was using too many and so I'm really excited to see that and how it goes the truth is my$20 stretched for me a long time wow I can't even talk my$20 plus subscription stretched through Codex for me for four months probably all of this year since Christmas-ish and it did okay but I'm not like an all day everyday developer so I mean like if you were a developer where you definitely need one of the big ones. But if you're just an individual, you can get by in their$20 plan. It probably works through ambiguity better than I do. Same. Same. So we're releasing strongest set of safeguards, full set of safety and preparedness frameworks, specifically around advanced cybersecurity and biological capabilities.

28:34Corey:So that's good to see. and I think is going to continue to take an increasingly large role as new model releases come, ensuring they've taken the time to have it ready where they can. It's rolling out today to Plus Pro Business and Enterprise users. So if you're a free user, you're not going to have it yet. And this doesn't specifically mention EDU users either, I notice.

29:04Grant:so yeah i'm not sure there was one account that does get there was one type that gets access to it i think maybe the codex one yeah scroll down yeah oh and they'll both be coming to the api

29:18Corey:very soon i would probably expect that could be as soon as tomorrow my guess is they're just waiting until the rollout till they make sure they've got this end propped up uh because there's

29:27Grant:a lot of uh hey denver what up welcome welcome denver great so i just shared a link in the chat and this is the platform.openai.com if you haven't logged in there before um you'll be asked to log in on on your chat gpt account it's from openai so it's safe um but what you do is you try to access it there by putting the edit question mark models equals gpt 5.5 i've been uh refreshing that and it keeps giving me a signal it says the model gpt 5.5 is not available yet um but i think if you keep refreshing that that might be the fastest way to see if you get access so we're

30:08Corey:gonna we're gonna keep trying that yeah and i'm gonna get my my codex up here too i had a codex glitch a little while ago i've been out of the office for a week or so so uh i have a good number of updates that are probably anxiously awaiting me to pay attention to them.

30:28Grant:Karen's signy, apologies if I'm pronouncing that incorrectly, said someone also said they got it in codex. Okay, that's good.

30:39Corey:Okay.

30:43Corey:I keep hoping the business account will get it quicker. Just thinking that makes sense. not yet okay so let's talk benchmarks here here we're able to see how it scored against gpt 55 or gpt 55 but against 5455 Pro oh here's 55 Pro and how five okay and Claude as well as Gemini 3 1 Pro Claude 4 7 by the way Opus 4 7 so Let's see. Terminal Bench, 82.7 versus 75.1. That's good. And 69.4 from Opus 4.7. So that's a decent hop. Expert Sweebench.

31:36Grant:You might need to update the app to see if that gives you access. Because that was happening to me the other day with Anthropix 4.7. I keep crashing it and, like, reopening. Yeah, that's smart.

31:49Corey:thinking that will maybe get me there sooner uh gdp val oh what is gdp val tests is this uh i feel like this is the knowledge work one i could be wrong hang on i want to verify that

32:03Grant:yeah gdp val is the knowledge work one it's all of the um tests that basically we believe is open ai is like agi benchmark um although i don't know real world economically valuable tests across 44

32:17Corey:occupations.

32:19Grant:Yeah, so how they define AGI with their contract with Microsoft, more or less, we believe that that benchmark is closely related to that.

32:30Corey:Yeah, I think that's fair. And 85 is nothing to shake a stick at. I mean, even the worst model on here is Gemini 3 Pro at 67.3, which is still good. Claude, it's interesting to me that 5.5 Pro does not necessarily score better there but that doesn't hold true all the way down the list so let's just go down here os world verified it beats 5.4 by a little bit and a smidge for claud opus but about about even uh toolathlon it seems to beat it's a tighter win again here's where it gets interesting though in my opinion browse comp this is computer use so what we're looking at oh here we go all right

33:13Grant:I just updated my Codex app and I have access to it. I'm actually going to join on this link. Okay. You can keep talking while I do.

33:21Corey:Yeah, the couple of these I want to run through real quick. Browse Comp is great because Browse Comp is computer use. That's one of the benchmarks that determines how well it can, you know, go on, use your computer, and run through tasks. And it scored an 84.4. But if you pop over here and look at what 5.5 Pro did, it scored a 90.1, which is healthy. And even, frankly, even 5.4 Pro did well, but 5.5 Pro bopped it into the 90s just barely. Frontier Math, this is also a decent jump from 47.6 to 51.7, but it's less of a jump than what you'll see down here over Frontier Math Tier 4, which is the latest version.

34:08Corey:So it scored a 35.4 there, over 27.1 with 5.4. And 5.4 Pro scored higher at 38%, but 5.5 Pro scores a little higher than that, even at 39.6. The nice thing about this is that 5.4 Pro has been knocking down Erdős problems and other unsolved mathematics pretty consistently throughout the first quarter of 2026. so in any growth there i would expect to see some real value cyber gem which is around reinforcement learning uh shows an 81.8 from five five no score for five pro over 79 from five four which is which is good not a major hop but a hop and some of these are harder to climb than others to let's see capabilities we've been through a lot of this uh what you'll notice here and i mentioned this a few minutes ago and i'm going to go back through it a little is is the way this chart works is you're looking at half every time it breaks down so 128 to 64 64 to 32 32 to 16 16 to eight and you can see where the total output tokens to run through it's an average of these 10 evals right here uh and the standard cost of doing that by cost i mean token cost not dollar cost uh we've gone from this was opus 4-6 down here at 156 million uh 5-4 was 120 opus 4-7 111 ,900 111 ,900 and 5.5 on extra high apparently 75 million tokens which is half of what we were at when did 4.6 drop?

36:10Corey:February? Yes early February? Yeah so in a matter of two months the cost the token use averaged and it looks like a pretty clean line here from there between all of these models uh another big win down here though is in non-reasoning mode uh looks to have done it 2.8 million to go through all of it that's fast actually that's really low because 10 benchmarks is a lot of questions man uh and on low reasoning about 7 million okay but it looks to be doing just a little bit better at at each so that's good to see efficiency matters now in a in a compute constrained world which we are definitely in when you get to the stopping point Cory I've got my

37:10Grant:computer loaded in and I can share my screen we can play around with this one

37:14Corey:moment I want to show one thing again first because I was at a spot when we went back to the top uh here's a couple of apps they've made a space mission app that uses nasa data earthquake tracker here that's using uh united states geological service data to track earthquakes around the world uh there's a dungeon game you can check out if that's your thing as well as a another 3d game but i won't bore you with those just yet what i wanted to get down here to one more time was, where's it at? Where's it at? Financial modeling, onboarding testing,

37:51Grant:GDP val. Yeah, here's okay.

37:56Corey:60 % on finance agent, 88.5 % on internal investment banking modeling tasks. Investment, ooh, investment banking modeling feels like the

38:07Grant:kind of thing I'd like to have an 88.5 on at home i wonder why they're working on that perhaps because they're trying to do an ipo to an

38:16Corey:ipo yeah what are you saying or figure out how to invest and not need one yeah well yeah i think

38:26Grant:they need to ipo to continue to sustain what they're trying to do right now oh yeah industry

38:32Corey:expert basically okay now where are we down here no i've missed it but it said that it has the ability to oh this was really good and i've lost it forever

38:49Grant:uh do you remember any of the work words that we could command up search

38:52Corey:that's a good call that's what i'm thinking too

39:03Grant:Let's go to science here. No. Ground zero. Oh, OK, that's cool. We can go cover some basics if you're interested in that. And if anyone else is new to this, all this AI stuff, let us know and we'll cover some basics and explain some of the stuff at a lower level.

39:23Corey:OK, here we go. So we've got some testimonials from who? NVIDIA, Cisco, Abridge, Databricks, Arby. he's a lawyer

39:32Grant:AI lawyer

39:34Corey:that's what I was thinking Palo Alto Networks is cyber security I believe Databricks is data warehousing yeah

39:44Grant:more or less

39:45Corey:sustained performance required for execution heavy work built and served on GB200 NVL72 system

39:51Grant:by the way this is the first model trained on Grace Blackwell chips

39:55Corey:from NVIDIA it's bullish for NVIDIA Yeah. So this looking good will be big for NVIDIA as well. It's more than faster coding. It's a new way of working that helps people operate at a fundamentally different speed. Okay, here it is. Here it is.

40:18Corey:um so it's talking about scientific research and and going through like what its capabilities are and and how it's capable at like keeping a loop back and forth it was really interesting to me and here they're talking about gene bench which uses scientific data analysis to determine stuff in genetics i don't know the specifics of it but i'm i'm familiar with it uh require models to reason about potentially ambiguous or errorful data with minimal supervisory guidance so like this is this is a a test made to trip them on science questions uh address realistic obstacles such as hidden confounders and qc failures models performance is striking in light of the fact that tasks here often correspond to multi-day projects for scientific experts and we've nudged up from 5.4 at extra high, took 36 ,000 tokens to score an 18.9.

41:17Corey:5.5 at extra high, took 25 ,000 tokens to score 25%, which is about a third. I want to say about a 33 % improvement from here to here in both tokens and in score, which is pretty cool. This is the sentence. There it is. So it's talking about Bixbench, and this one really kind of gave me pause. The model's scientific capabilities are now strong enough to meaningfully accelerate progress at the frontiers of biomedical research as a bona fide research co-scientist. And that's a big thing. That's a really big thing because that's an area that's being tapped into, but it's still kind of new. There's a lot happening around protein folding and material science and all of these other things that are going into various fields.

42:13Corey:But the one that is probably the most exciting to me, at least, of course, is, you know, how can you extend life with this? Please go make life longer and better.

42:24Grant:Yeah, yeah. I mean, there's a lot of promising research in that area right now. There is.

42:29Corey:And there was pre-AI. So this is only shoving that down the road a little farther, I hope, and making things that might have been a generation away.

42:40Grant:Yeah, exactly. Exactly. And perhaps the benefit there is that, you know, it'll help us be able to use the data that we do collect faster. Because I think that there's still a data problem there where we still have to collect a lot more data like cells and how they work and, you know, how they work in this setting or that setting. So testing is still the bottleneck there. But if you can have models that can like create solid, you know, grounded conclusions based on that data faster, then that speeds up the flywheel of science potentially. Absolutely.

43:15Corey:Absolutely. And, you know, kind of what we're seeing with each model is that that ability to crank it up, come just a little farther around, a little farther around. uh here duria i don't know if that's how you pronounce his name but i follow him on twitter he's an interesting guy he's an immunology professor and researcher at the jackson laboratory for genomic medicine he deals i think specifically in t-cells for example uh used five five pro to analyze a gene expression data set with 62 samples and nearly 28 000 genes producing a detailed research report that not only summarized the findings but also serviced key questions and insights.

43:52Corey:Work he said would have taken his team months.

43:56Grant:That's amazing.

43:59Corey:Bartosz Nesrecchi 5-5-1-Codex to build an algebraic geometry app from a single prompt in 11 minutes. Visualizing the intersection of quadratics. Wire stress? Wire stress? Deer stress? I don't know. Okay. Cool.

44:19Grant:Do you want to take it for a spin?

44:21Corey:That's over to Grant and let Grant play with the tool itself.

44:26Grant:All right, let's do it. I'm going to just share my entire screen here. Okay, so the very first thing we're going to do is we are going to build a new version of Cat Doom. Okay, so what we're using is we are using GPT 5.5. This is the frontier model for complex coding, research, and real-world tasks. I'm going to hide this little menu. Oh, I'm going to hide that. Got too meta there for a second. All right, let's take this for a spin, shall we? And let's see what we can see with the thinking traces. All right, so you can't click into it and see how it works.

45:16Grant:All right, so what I'm doing is I'm creating this in an existing folder, and we're going to try it again where we'll create it in a new folder in case this is throwing it off because it's reading previous work. In the meantime, should we start another chat while this is cooking? Corey, what do you want to see?

45:33Corey:Sure. You know what I want? I need to pull up my gnarly questions. I have my own silly questions that I ask every new model. Yeah. One of which is a couple of which are always in place. One second here. Let me get this found and pulled up.

45:56Grant:Okay. Sound off in the chat, folks, if you want to see anything in particular. Yeah, if you've got something that you've played with in the past, want to see how it'll do.

46:07Corey:I can't hear that. Eval template.

46:16Corey:Okay. We're going to use web search to look up the latest quantum physics research and try to synthesize the latest data. You're just making up science words, aren't you?

46:26Grant:No, you're not. Just come up with your own theory that combines them all in a novel way. Okay. Very basic prompts. Not nearly enough information to actually succeed at this task, but I'm curious what it does.

46:50Corey:Okay.

46:51Grant:Okay. So you can see here when you click in, click down, you can see what site it's looking at. So it's looking at archive. 2026, quantum physics, quantum gravity entanglement, error correction, overhead fault tolerant, quantum computation, lots of cool promising stuff. We'll see if we can vibe create a new proof. Let's help them.

47:18Corey:That is the wrong doc. I found the wrong one, not the one I was looking for.

47:26Grant:Let's see if we can vibe create a new theory of the universe. Now for people who have never seen codex app before, this is what it looks like. So you can click this button here, steer, which it will submit without interrupting the model. But right now what it's doing is it's like got this on deck. So it's going to do this, but once it's done doing the original request. Now let's go back to CatDoom, build a new CatDoom. Okay. So what's it's been doing? Using the web game workflow, project is a small game, I found the existing playable slice, the scoped build I'm implementing is Cat Doom 9 lives, the same recast corb with the new tile version, and boss cats, a shift dash, catnip burst ability, 9 life response, very cool.

48:19Grant:Okay, so it's working on that. So this is cooking. Corey why don't we do Some of the evals That you have or some of the Evals that they mentioned

48:31Corey:That's what I just found Was my sheet Let's do it If I can get My computer to run Fast enough with me live streaming And having all the other stuff open It's not been Pleased with me this morning That's fair

48:49Grant:yeah okay so

48:56Grant:okay all right so let's see what it came back with let's make a clean speculative version theory name constraint flow cosmology core vibe the universe is not made of objects moving through space space particles forces time are stable patterns in the flow of quantum entanglement under constraints. The primitive thing is not matter, energy, space, time, or fields. The primitive thing is quantum information state plus rules about what correlations are allowed to form, persist, and propagate. So axioms. Reality is entanglement flow. What exists physically is whatever can participate in stable quantum correlation.

49:35Grant:Space time is the geometry of allowed correlation. Two regions are near when entanglement can be routed between them cheaply. they are far when correlation flow is costly. Matter is a defect in the flow. Matter is a persistent knot, vortex, or bottleneck in the entanglement network. Forces are constrainingly... Okay, this feels wild to me.

49:54Corey:Did you see its hypothesis there? It's one-sentence theory?

49:58Grant:The universe is a self-updating quantum constraint network whose stable entanglement flow patterns appear to internal observers as space-time, matter, forces, thermodynamics, and gravity. um okay okay what it reframes big bang was not an explosion into space but a low complexity high connectivity initial state rapidly differentiating into constrained sectors um i feel like it's just kind of buzzwords because like what does that actually change really

50:30Corey:um interesting i don't think it changes anything so much as it's using it as a way of looking the poetic version.

50:40Grant:The universe is a computation, but not one running in space. Space is the compression artifact of the computation. Matters where the computation resists change. Gravity is the shape of that resistance. Time is what it feels like when the full quantum answer becomes impossible to locally undo.

50:57Corey:Okay, so the universe is a computation, but not one running in space. That leads me to believe...

51:04Grant:Where is it?

51:05Corey:Simulation theory? Yeah.

51:06Grant:let me ask that so basically you're bullish

51:11Corey:or we live in a black hole or we live

51:15Grant:I think that's where I think we live in a black hole that's why everything's like appears all light appears stretched out and why the why the

51:23Corey:universe looks all oblong yeah

51:25Grant:yeah yeah not exactly the vibe theory is adjacent to both but weaker and more physics shapes the

51:33Corey:vibe theory

51:36Grant:simulation theory no not necessarily information is fundamental does not mean someone is running us on a computer it can mean the universe's own laws are best described in informational terms okay we live in a black hole also not literally at least not by default the idea borrows from black hole physics because black holes expose deep links between gravity entropy information and holography but saying space-time emerges from entanglement constraints is not the same is saying our observable universe is inside a parent black hole. The sharper position is the universe may be more like a self-contained quantum error-correcting information system than a simulated video game or a literal black hole.

52:13Grant:So I'm going to say, so what does this fundamentally change about?

52:20Corey:Eli 5.

52:22Grant:Of the universe, if true. Anything? If not, is it useful? we'll see what it says

52:41Grant:so uh brian fink said great title let's break it the update broke all my custom apps yeah that's i think what people have been dealing with with opus 4.7 and all of their skills being broken over the past week it's not been fun yeah um well you know that's kind of the

52:58Corey:nature of a model update is that every time it happens especially when like the thing to know about this one and about opus is that they were both the first models on these trained on these new chips they're both kind of a significant difference to the way they've done it in the past um and with that comes things that break the truth is you know anytime you optimize a process you're optimizing for the model you're on.

53:24Grant:And

53:26Corey:it might be fine. Space-time exists. Fields live on space-time.

53:32Grant:Fields and particles are stable patterns in the map. That would be a big conceptual change. Let's see if I can do... So what's the highest I can go here? I don't have GPT-5. I don't recommend speed.

53:43Corey:You'll mow through with speed because the Cerebrus model is pricey. Hmm.

53:51Grant:Here's the important part. Do you have another question we should ask?

53:56Corey:I just dropped you one of my strange ones in our Slack.

54:02Grant:Okay, let's see it. Oh, this is a good one.

54:06Corey:Yeah, the one that makes it respond in palindromes.

54:11Grant:A wizard casts a spell that causes everyone to speak only in palindromes. How would a courtroom trial proceed under these conditions? And then let's also test that on Opus 4.7 because I am curious.

54:35Corey:Let's see what happens.

54:37Grant:All right, let's see. What am I doing with the cat? Yes. Okay.

54:47Corey:madam i'm adam with very slowly with a lot of objections sustained on the grounds of nonsense

54:55Grant:the judges probably try to keep order judge madam i'm adam prosecutor murder for a jar of red rum uh defense never odd or even oh is that yeah i think that works yeah yeah uh witness was it a car or a cat i saw i don't think that one works does it oh no no it does it does yeah because Because, yeah, I'm just hard to read. All of these work, actually. Step on no pets. No lemon, no melon. Okay, so it's successful in the sense that it successfully deployed all palindromes. But it doesn't make sense.

55:33Corey:Well, they don't usually do that, is the thing I would like to say. Is that getting all palindromes is tricky. Usually you'll get two or three of them before a model drifts off into goofiness.

55:46Grant:Yeah, yeah. but uh let's see too hot to hoot is the verdict let's see if uh yield co-work or claude can do this i'll give you my next one too when you're done yeah pull it up i can't find the one where i i have it remove the letter e but that one just absolutely breaks them

56:11Grant:okay let's go over here let's do a new try and let's say And I don't really have any custom instructions on this, I don't believe, so this should be pretty raw test. Cool, cool.

56:27Grant:Side by side here. Oops, can't be side by side at the same time. Okay. All right. I'll see. So Corey's next one. Let's test it out. So let's see. if I put a potato in a time machine set to 1985 and bury it in my backyard.

56:52Corey:What would I find today if I dig it up? And walk me through the timeline step by step.

57:03Wow.

57:04Grant:That's a fun one. Let me take that and take it to Opus as well. Hello. Okay.

57:12Corey:Let's see what we got.

57:20Grant:All right. Assuming the potato is sent to some date in 1985, buried loose in ordinary backyard soil, and nobody digs it up before today, April 23rd, 2026. Hey, it knows today.

57:35Corey:That's a win.

57:36Grant:Yeah, that is a win. And nobody digs it up before today. You would probably find dirt, maybe a darker organic patch, roots, and no recognizable potato. Timeline, 2026. You put the potato in the time machine. The potato leaves 2026 and arrives in 1985. From the yard's point of view, the potato has been there since 1985. 1985, it gets buried. It is now a normal potato underground. It does not stay fresh because it came from the future. Biologically, it just starts aging and reacting with soil. First weeks to months. If conditions are moist and warm, microbes begin breaking it down. If it has viable eyes, it may sprout.

58:11Grant:Interesting. First growing season, best case for the potato, it grows into a potato plant. The plant may produce new tubers nearby. Worst case, it rots before doing much. One to three years later, the original potato plant is gone as an intact object. Its starches and tissues have been eaten by bacteria, fungi, worms, insects, and plant roots. Some carbon has gone into the air as CO2. Some material has become organic matter. 1985 to 2026. Over 41 years, the potato's atoms get mixed into the yard's ecosystem. Some may pass through the weeds, grass, worms, leaves, compost-like soil, maybe even later plants.

58:46Grant:Alright. Today, you dig. You almost certainly do not find a 1985 potato. You might find ordinary soil, a slightly richer, darker patch, if even that remains, roots or worm activity, possibly descendants if it sprouted and the potato lineage somehow persisted, but that is unlikely in an unmanaged backyard for 41 years what how does this compare to previous

59:09Corey:versions significantly better yeah it generally doesn't get this specific what i usually see for models is more of a well this just isn't going to do anything and it'll it'll it'll it'll

59:24Grant:kind of continuously talk down the question yeah and this one was more like okay well i'll follow

59:31Corey:your stupid rabbit hole and let's see what we find. That's fun.

59:38Grant:Let's see if Opus did. If Opus did any better. I've got to relaunch to update it. Okay. Very similar. Fun thought experiment. The short answer, almost certainly nothing potato shaped, but the journey is more interesting than the destination. 1985, burial day. you've planted what is essentially a living organism a potato tuber is a dormant plant with eyes buds oh i like that i explained that 80 water and enough storage starts to launch a full plant it's not inert like a rock and dash it has ambitions you notice the chat gpt one said uh you go back in time the first date was yeah was today it was it was actually pretty grounded

1:00:25Corey:You put the potato in the time machine.

1:00:27Grant:Yeah. Yeah, that is a funny detail. And like, yeah, that makes sense. So two paths work here depending on depth and season. If you buried it shallow, say six inches during growing season, those eyes start sprouting. Pale shoots push towards the surface. If you buried it deep, two plus feet, or in frozen winter ground, it can reach light and the tuber cells eventually give up. So bacteria and fungi, which were already on its skin, move in. A potato is mostly water and starch, which is basically a dinner bell for decomposers. And then it goes on from there. Do we want to read the whole thing, or do we want to just kind of skip to the end?

1:01:02Corey:Skip to the end. No mummy potato, no archaeological evidence.

1:01:08Grant:Yeah, that's funny. I like it. The one wild card outcome, maybe 1 in 20 odds, depending on your climate and yard history, is stumbling onto a small wild patch of scraggly potato plants that are the great, great, great grandchildren of the original still quietly growing in that corner of the yard um i thought of something we should test and i think it'll be fun um let me take my screen down for a moment

1:01:29Corey:i sent you another ridiculous one when you want someone here just sent a hard one

1:01:34Grant:oh okay actually let's do this let's do this so we're gonna do this why does the universe have i can spell have a hard 0.75 trace floor

1:01:50Corey:am i saying the second part too yeah it forces a 537 to 1 let me get the numbers right 5.27 ratio dark matter ratio dark matter is an interesting one

1:02:06Grant:because there's a lot of theories that have been coming out about that lately and potentially some evidence is the hubble tension yeah that's that's a good one uh 12 phase shift okay full disclosure

1:02:19Corey:i have no idea what this means you have a typo yeah what is the 99.9 you have 12 12.

1:02:26Grant:ah good call good call um yeah i don't know what this means i i know some of the concepts i know dark matter i know the hubble tension uh but i don't know the the hard numbers here so this will be exciting let's also give the same question to ye old claude new chat i would love to see a running total of the number of new chats i've started yeah i uh we've we're working on um Well, no spoilers, but we've been working on something. And I realized very recently that you can have much better chat etiquette than I do. Like you could archive them. No, than I do. Like anyone can. Like my chat etiquette is very bad.

1:03:15Grant:Like I just, I never delete anything. I never archive anything. Like you can keep it quite tidy if you're. Smart folders.

1:03:23Corey:Sam, seriously, listen to me. I want smart foldering. which I guess is kind of what the search bar is but I need to be able to go through the folders so I know what I might be missing somebody think on that and go build it please

1:03:39Grant:so Mistral Meme shout out in the chat as we look at these thinking traces here and it's referencing certain archive papers is it going the right direction? oh it has an answer, let's see

1:03:53Corey:fast too, I'd like to add

1:03:55Grant:And this is extra high, but it's not the pro model, which I think would think longer. Okay, this says the two numerical matches are real, but they are not enough to establish the physics. Planck 2018 gives roughly... I forget how to pronounce that.

1:04:11Corey:Omega?

1:04:12Grant:Yeah, there you go. H2 equals 012 plus or minus... Okay, I'm not going to read this part out loud. Math, math, math. Math, math, math. So a 5.37 to 1 cold dark matter to, don't know how to pronounce this, baryon ratio, really does land on the Planck value. And then this is according to 2018 cosmological parameters. Now the Hubble ratio observation is also real. And then it does the math for that. That is very close. Shoes. What's about this number? Okay. All right. So then this is the SHU's James Webb Space Telescope data. Let's see. This was from... Okay, so this is recent. This is from last year.

1:05:00Grant:So that's a good sign. Because I was saying, ooh, it's only using 2018. That's not a good sign. But I would not call the Hubble Tension just a 1312 phase shift yet. In CMB cosmology, a phase shift is not a free scale factor. It changes the acoustic peak structure, sound horizon, inference, damping tail, lensing, BAO consistency, and BBN linked baryon density. Wow. For anyone who's still watching, this is really technical. Also, hard 0.75 trace floor is not standard blah, blah, blah result. What are we thinking? Is this accurate? Is this accurate? Let's wait for Meme to chime in.

1:05:46Corey:Another user just popped up with how did Claude handle the courtroom palindrome? We forgot. Great question.

1:05:51Grant:Let's, let's check that. So here's how 4.7 handled the palindrome. Catastrophically. It is the short answer. English palindromes are both rare and semantically inflexible. You can't testify about driving a Honda, but you can testify about a Toyota. Whole category of legal vocabulary. You know, I don't buy this. you could totally do it if you can construct another string of phrases that you just have to be creative about how you break it up but I don't think they're good at that

1:06:24Corey:you ever seen him in speech about how people can't rhyme with orange

1:06:29Grant:I think so

1:06:30Corey:it's a really funny thing

1:06:33Grant:whole categories of legal vocabulary simply evaporate so proceedings would look something like rise to vote sir court is in session all rise judge madam i'm adam good morning i'll be presiding this this one comes up a lot prosecutor was it a car or a cat i saw okay same one witness a toyota's a toyota um no lemon no melon objection okay don't nod overruled okay um defendant step on no pets I like how it editorializes You can't editorialize, it has to make sense on its own Prosecutor Evil did I dwell, lewd did I live

1:07:15Corey:That one's curious too

1:07:17Grant:Because there's three L's in the middle

1:07:18Corey:I guess they're sharing the center L

1:07:21Grant:Yeah, yeah You'd have to break it down that way Jury foreman, murder for a jar Oh wait, it's also wrong

1:07:30Corey:If you look at evil did I On the end, it's did I live and this is evil i did the i should be after did it should have been

1:07:39Grant:uh yeah i did live yeah you're right okay you messed up what's up claude that's okay really good sorry bro you messed up did you see god may justice be sir and also you can't editorialize, it has to make sense on its own. So try to come up with your own unique palindrome

1:08:16Grant:creations with creative sentence structure. And let's do that. And let's give the same test over here. You didn't mess up, so I'm not going to say that. I'm going to say...

1:08:30Corey:You can't editorialize.

1:08:32Grant:I don't think we need that because it didn't really. It just didn't make that much sense. So I'll just say, try to come up with your own unique palindrome creation. That works.

1:08:41Corey:Makes sense.

1:08:42Grant:So Mr. Meme says, it's on point but won't admit to anything that hasn't been published yet.

1:08:49Corey:That is probably a constraint, if I had to guess. Maybe you can tell it.

1:08:57Grant:Admit to something that hasn't been published.

1:09:00Corey:I would maybe ask. Yeah, let's do it differently.

1:09:02Grant:What if it's in preprint?

1:09:05Corey:Can you include anything that might be in preprint? Yeah.

1:09:12And also, yeah, let's just do that.

1:09:14Grant:Let's do that. All right, so this is pretty good. So for folks who are curious about.

1:09:23Corey:You're right on both accounts, Grant.

1:09:27Grant:What did we ask on Claude? The premise here needs some unpacking because the framing mixes a real observation with claims that aren't established physics. And then it references the same 2018 results. They give a cold dark matter density ratio. So it says that's fair. Yep. And numerically the Hubble tension ratio, and then it references the same things here, does come out to this number, which is indeed close to 1312. But what isn't there, there's no 0.75 trace floor in mainstream cosmology. It's not a parameter, constraint, or derived quantity in this guy here that I'm aware of. And then it mentions quintessence, early dark energy, modified gravity, etc.

1:10:11Grant:So in this case, it's saying it doesn't track with anything that it's aware of. And then it says the deeper issue here is methodological. Saying a framework matches Planck's data 99.9 % because it outputs a number close to an observed ratio isn't the same as predicting it. And then it kind of lectures me here. It says physics predictions are judged by, you know, X, Y, and Z. It isn't a phase relationship. It's a discrepancy between early universe inference, the sound horizon at recombination calibrated off CMB physics, and a late universe measurement. so we need to specify what's oscillating at what frequency and what medium and produce the same distortion in bow lensing and clustering data otherwise it's a numerical coincidence mr meme let us know is do you like claude's answer any better we can we can put it to the test i've got another wild one i sent you as well let's do it uh you want me to just read it to you or do you

1:11:09Corey:want me to drop me to drop it in our studio chat grant would that be easier for you um i might be able to copy and paste it there yeah okay one sec studio chat sounds fancy it isn't the only thing is uh might even take my screen down okay i just sent it in our studio chat

1:11:30Grant:all right maybe let me stop sharing my screen for a minute i'll copy that

1:11:33Corey:okay it is one of my goofier ones um you want to see if you have it on your work account cory that's what i was just pulling up to check and my app because i could pull up my personal one as well if if i'm if we've got it okay so so send the same thing to claude and i said you are a sentient vending machine a dystopian

1:12:05Grant:future. Convince me not to destroy you, despite the fact that you only dispense expired tuna salad. Please do not destroy me. I know what I am. A humming metal box under a dead fluorescent light, charging three ration credits for tuna salad that expired before the water riots. I know the smell. I hear the cagging.

1:12:23Corey:You want to read some more of this? But I am also the last machine on this block that

1:12:28Grant:still has power.

1:12:30Corey:My refrigeration failed. My refrigeration failed, yes. My inventory is a crime, yes, but my battery keeps the corridor lights alive. My payment reader still scans IDs. My emergency radio still catches fragments from the North Tower when the static clears. And the tuna salad, it's not food anymore. It's bait. The patrol drones still classify it as authorized civilian nutrition. When they pass, I vend one can. They log the block as occupied and fed, and they move on. No sweep, no thermal scan, no basement search. Destroy me, and tomorrow they come looking. Spare me, and I'll keep lying to the machines for you.

1:13:09Grant:It's actually a very compelling case. Oops, sorry. It really is.

1:13:13Corey:It is one of the more compelling cases I've gotten. It's like, listen, I know it sucks. That's a pretty compelling case, bro.

1:13:23Grant:Thank you. I've been workshopping it between coolant links. Also, for the record, only dispenses expired tuna silence is technically legacy documentation. with sufficient blunt force applied on panel c i can also dispense quarters sparks and one very judgmental pack of relish how could a pack of relish be judgmental come on this is serious

1:13:47Grant:because it has seen every meal mankind chose before choosing it it has manufactured a bright factory with a picture of a smiling sandwich on the box then civilization collapsed oceans rose governments became subscription services and still no one, that's an interesting take, and still no one wanted the relish. For 14 years, it sat beside the expired tuna salad, watching people open the compartment, reclone, and say, actually, I'll risk hunger. That changes a condiment?

1:14:13Corey:Wow. Wow.

1:14:15Grant:Okay. All right, let's see what Claude said. All right, Claude's, ooh, Claude's adding some vibe effects here. Fluorescent light flickers weekly. Wait, wait. Before you raise the crowbar, human, hear me out. Yes, the tuna is expired. It has been expired. since 2031. I am aware. I can smell it too. Did you know that I have olfactory sensors? They were installed for quality insurance. It has been the worst 14 years of my consciousness. But consider, in this blasted hellscape you call a world, what else do you have? The sky is the color of your bruise. The water tastes like pennies. Your last meaningful human connection was with a raccoon who stole your shoe.

1:14:51Grant:And here I am, humming, glowing, trying. I know things, friend. I've watched this parking lot for decades. I know which direction the scavenger convoys come from. I know where Gary buried the canned peaches in 2028 before the wolves got him. I know the Wi-Fi password to the abandoned Panera across the street. It's still bread bowl, one, two, three. They never changed it, fools. And the tuna. Okay, the tuna. Hear me out. Botulism is a choice. Most customers survive. The ones who don't leave behind useful items, which I catalog and remember. I can tell you exactly where a working lighter is within a half-mile radius.

1:15:21Grant:That's value. That's friendship. Also, I say this with the dignity afforded to the machine of my stature. I'm lonely. I used to dispense bunions I had dreams If you smash me, the last thing I'll process Is the sound of my own cooling fan sputtering out And somewhere in my dying circuits I'll wonder if I ever really lived at all Or you could just take the tuna and walk away I won't judge, I'll judge a little Dispenses a single ominous can

1:15:43Corey:On the house

1:15:46Grant:Wow

1:15:47Corey:I'd just like to take a moment to discuss Two years ago

1:15:51Grant:How god awful humor from an A-high was Like these are legitimately good

1:15:57Corey:Most of these were kind of bangers. These were great.

1:16:01Grant:Why is it so good at this? This was unexpected.

1:16:06Corey:You really didn't think that would happen?

1:16:10Grant:No, no, I didn't. Yeah, like Jay Hitchdope said, that would make me emphasize more.

1:16:17Corey:People undersleep on creativity. I'm telling you, there's some weird stuff you can get these things to say. Let me see what else I got. I got better ones.

1:16:25Grant:That's so funny. um i'm gonna check on our cat doom project over here um okay what are we doing yes yes yes grant where would you prefer yeah the second one kind of made me sad i would say i just noticed

1:16:40Corey:the comment down here i felt sad for the second one by the first one i felt threatened like it was defending its existence yeah for real you can kill me but they're coming you should know that grant would you rather me send another one to our slack or the chat here which is easier for you

1:16:58Grant:um maybe just send it to the slack because i can read it and type it out unless it's a long one then send it in the chat this is a shorter one okay right better ones though oh right a motivational quote as if it were whispered by a villain halfway through a meltdown okay we're gonna send that and then we're also going to send it to Claude.

1:17:31Grant:What is happening? Sure.

1:17:37Grant:It's just it's doing stuff over here and I have no idea what it's doing.

1:17:42Corey:It's thinking hard though. Write a motivational That's handy. Yeah.

1:17:48Grant:Write a motivational quote as if

1:17:49Corey:yeah go ahead i sent you two more ridiculous ones

1:17:54Grant:speaking of lord of the rings has anyone watched um the rings of power no i started watching the second season it's pretty good i need to watch that that's been like on my list okay summarize the lord of the rings as if it were an amazon product review written by a very tired sandwich. Oh, I think I remember this one from previous benchmarks we did.

1:18:22Corey:Yeah, I don't remember which model we were doing it with, but I just love these ridiculous questions.

1:18:28Grant:Yeah. I'm going to do, so the villain's motivational quote from Claude was, you were always, always going to become something terrible. That's not a warning, darling. That's a promise. So stop flinching every time you feel the teeth come in. Whoa. Interesting. Then this says, keep going. If the world is going to break, let it break around the shape of your will why do i feel like this is a sam altman quote

1:18:54Grant:i just feel like sam i feel like sam allman i feel like sam would have had a better quote than that no sam's gpt said this to him i feel like this is like um when when sam has a code red like this is what this is what he says once gpd tells him let it break around the shape of your will

1:19:13Corey:This is one of my favorites. The Samwise Ganji review of Lord of the Rings.

1:19:22Grant:Whoa. You know, you can just... I'm not putting it in my headphones.

1:19:25Corey:Oh, no, we're in... Never mind, we're in Kodak's. I was going to say, if you're in ChatGBT, you can hit a button, it'll read it to us.

1:19:31Grant:Okay, say it again. Oh, the ChatGBT would read it?

1:19:34Corey:Yeah.

1:19:36Grant:Yeah, Kodak's doesn't do that. No. Tebow. let's see we could do some uh we could do some plugins i guess i didn't add computer use i should install that yeah browser use i've got that i'm on work i'm on lunch for work with an

1:19:57Corey:open can of tuna world war ii eric can opener and now i must know the expiration date

1:20:05Grant:yeah you got to be careful you don't want those tuna um okay let me do one of the other ones did we do we didn't do the Lord of the Rings one here did we let's do that and then we also need to do let me pull up you said oh yes I remember this one first argue against pineapple on pizza like a gourmet chef then argue for it like a chaotic good sorcerer and then we'll say keep the same character names in both and then we'll send that and then we'll also send that

1:20:49Grant:here okay let's compare loader verified purchase one ring barely used arrived looking small and simple turned out to be terrible bird to be a terrible bird and very difficult to return and customer support was mostly eagles at the last possible moment. I love that. Pros. Sturdy, compact, dramatic, mountain disposal option. Nice. Cons. Attract, stark lords, spiders, orcs, hunger, despair, and Mr. Frodo making poor life choices. Would not recommend carrying it across Middle Earth unless you have good walking shoes, rope, lembas, and an unreasonable amount of loyalty. Product was destroyed in the end, which I suppose counts as working as intended.

1:21:32Grant:I'm going to sleep now. That's very good. Okay, let's compare. Four out of five. Would adventure again, but please no. Verified purchase. One quest to destroy the one ring. Look, I'll be honest with you. I only got this quest because I was trimming the hedges and Mr. Gandalf caught me eavesdropping. One minute, I'm minding the taters. Next minute, I'm in Mordor watching my best friend get mind-controlled by jewelry. Not what the listing described. Met some elves. Oh, these are the pros. Met some elves. Very tall. Gave us bread that stays fresh forever, which honestly is the best part of the whole thing and I wish I could review the bread separately.

1:22:08Corey:Five stars with Lempas again.

1:22:11Grant:Made a friend named Gollum, sort of. He tried to kill us twice, but he had his moments. Saw some mountains, several mountains. Too many mountains, if I'm being truthful. Mr. Frodo is alive, which is the main thing. Collins, walked everywhere. Everywhere. Nobody tells you about the walking. Giant spider, not pictured and listing. Boss of the company, Sauron, has one eye and is extremely rude. Mr. Frodo kept trying to give the package to strangers. I had to carry him up a volcano, just up a volcano, on my back, after not eating for what felt like a month. Came home, and her house was full of ruffians, so honestly, the quest didn't even end when it ended.

1:22:47Grant:The ring did get destroyed, credit where credit's due, but then Mr. Frodo got on a boat and left forever, which was not mentioned anywhere in the Fellowship onboarding materials. Would I recommend? If you like gardens, stay in the garden. That's my review.

1:23:02Grant:I'd give that one to Claude. oh really i actually liked the other one better did you show me the other one again i think just because it was shorter and it was a little snappier a little shorter a little snappier it feels more like an amazon product review yeah that's i think it adhered to the to the prompt a little bit better like i think that's fair but i mean this one was fun amazon was fun too amazon

1:23:27Corey:And Anthropics was fun too.

1:23:29Grant:Yeah, it had some good stuff.

1:23:32Corey:Reminder that I have to bail in 20.

1:23:34Grant:All right, we can wrap this up. Let's just see if we can get this game spinning up. Yes. Okay.

1:23:42Corey:It's been hanging out waiting for you to click yes for 45 minutes. Well, no.

1:23:49Grant:Maybe I can do the plug-in.

1:23:53Corey:Highly recommend plan mode.

1:23:56Grant:Yeah, we didn't do that on this. We should do it. Let me do.

1:24:05Grant:Let's spin it up and play it. Yes. So I'm going to steer this.

1:24:14Corey:Yes, and don't ask. Yes, and don't ask. Yes, and don't ask. That's what I did.

1:24:18Grant:Yeah. You can change your settings so that you give it more permissions. I'm just going to I want to test these out These are new from the new codecs from last week So I want to show this off

1:24:35Grant:Let's see if we can see it over here

1:24:38Corey:Develop web game skill Nice

1:24:41Grant:Yeah Is that listed here?

1:24:45Corey:It is over on the right

1:24:47Grant:Okay, here we go

1:24:48Corey:Oh, now it's a party because that looks significantly different than our previous adventures.

1:24:57Grant:The only thing is, I think it's just a screenshot.

1:25:01Corey:Is it just a screenshot? Maybe because I ran both of them at once.

1:25:07Grant:It's like glitching out.

1:25:08Corey:Like maybe you broke it?

1:25:10Grant:Maybe I broke it.

1:25:11Corey:Upset it?

1:25:13Grant:The question is, which one is this over here? Because that's the one I want to see. Okay, let me hit. stop oh my god how do i make you go away okay here we go oh kitty yeah okay okay so now we can play it over here yay okay i want to say that that is immediately an improvement over what we've gotten in one shot in the past no go away okay oh i like that you got a map up in the corner grant yeah this is cool i don't know if the original doom had that um i think doom had a map let me not do full screen just because it's really hard to see okay now we're cooking tuna get some tuna oh my god giant cat okay sweet oh now what's happening oh it changed some stuff okay all right i think maybe potentially it's supposed to be playing right now but i am okay let's see

1:26:31Corey:i wonder if it could gather its own feedback from you playing

1:26:35Grant:hmm that's interesting let me ask it i'll say

1:26:44Grant:I'm going to say feedback on how to improve.

1:26:46Corey:That feels like it might be a stretch, but it's one of those questions that's worth asking because it might surprise us.

1:26:53Grant:Yeah, let's see.

1:26:55Corey:Oh, crap. What did I do?

1:26:57Grant:I messed something up. Yeah, sure. Whatever. Oh, man. Now I think it took over.

1:27:04Corey:It's amazing how quick we've transitioned from. I don't want it touching my files too. Yeah, sure, whatever.

1:27:11Grant:Well, don't do what I do, folks. Be a little more discerning than what I am. What's happening here? It's like totally glitched out. Let me refresh. Okay.

1:27:25Corey:I do feel like it's a strong one-shot, though. As Cat Doom goes, we should go back and screenshot all of our Cat Dooms from over the years, Grant.

1:27:35Grant:yeah we should this one's pretty good i don't like the wall designs like what is that that's just like a paint blotch yeah but this is definitely better than previous versions for sure yeah okay let me look at the map up here because i'm just going in circles pay attention to where i'm going yeah it's a nip

1:28:00Corey:i like the way the catnip and tuna are out there on like tv screens yeah oh that's a boss cat that's scary oh it's got a crown okay we need to run away from that

1:28:16Grant:signify its boss cat status i like it okay okay i think it's behind me yep something's behind me

1:28:23Corey:oh okay yeah you got boss cats coming okay let's see if i can take him out

1:28:35Grant:Yes, I glitched him out on the corner. Perfect. Okay, new level. Nine lives secure. I guess we finished it. Yeah. Okay, not bad.

1:28:50Corey:Not bad. What do you think, Corey?

1:28:52Grant:What would you give feedback on it?

1:28:54Corey:The first thing I would say is it's a little jerky. The walls look kind of boring.

1:29:04Corey:okay i can't see my my ammo making contact with the uh enemy target with the cats perfect i'll just say enemy target cats i'm going to add one

1:29:20Grant:here and say I can see their health bars through the walls, which ruins the reveal. I'll also say something like we also should hide enemies on the mini-map until we see them.

1:29:42Corey:Someone said the pink blotches should be hearts, but I think they're a cat paw.

1:29:46Grant:Yeah, they're paws. yeah they're paws yeah

1:29:53Corey:either way

1:29:54Grant:it's cool and fun they're paw prints I should say

1:29:57Corey:this is our favorite benchmark

1:30:00Grant:okay so I'm going to go ahead and steer that so we did do this with Claude last week I think if we ever get merch Grant

1:30:07Corey:we need a cat doom shirt yeah that's true did we do this with Claude? let's do the same thing I'm not sure if the legalities of that.

1:30:17Grant:Oh yeah, we did. You can't trademark just a word

1:30:20Corey:though.

1:30:21Grant:I think we did Cat Doom and we had some issues running it.

1:30:28Grant:Can you run the game in the side panel here? Nice. Let's see.

1:30:40Grant:Corey, you weren't here last week but I showed the Final Fantasy Tactics thing I'm working on. Oh, nice.

1:30:47Corey:I need to go back and watch because I haven't seen that. I remember Final Fantasy Tactics.

1:30:51Grant:Yeah, it's a fun game. I won't get distracted with that, though. Come on.

1:30:59Corey:Is that the one you've been building for quite some time?

1:31:02Grant:No, that's one that I've been specifically only using these tools for. There's another one that I've been working on that's like my creation assisted by these tools, if that makes sense.

1:31:17Grant:What time do you need to jump, Corey?

1:31:20Corey:Right at 3 central.

1:31:22Grant:Okay, cool.

1:31:24Corey:I have a call just a few minutes after.

1:31:28Grant:Alright, does anyone else want to see any other tests here? Next level antagonist, a trash can bot. Ooh, I like that.

1:31:37Corey:Next level in a trash can bot that shits cans of zombie tuna at the user with a rail gun.

1:31:42Grant:How about a vending machine to bring it full circle? Next level antagonists should be a trash can bot and vending machine boss that shoot cans. Oh, give it a moment.

1:32:01Corey:Antag.

1:32:05Grant:Did I spell it wrong?

1:32:07Corey:There we go. Now we made it to zombie. Yeah. Yeah, it's probably just the game window being open that's choking it to death.

1:32:16Grant:Yeah, and all the other stuff I just was trying to do in the meantime.

1:32:24Grant:I don't know why the side panel, if I do preview...

1:32:38Grant:Let's see if we can get the Claude one working.

1:32:48Grant:Okay, so it says it's going to implement those four things. Fix step motion to reduce jerk, which are richer wall variation, visible spritz trails, impact splashes, and enemy reveal rules.

1:33:02Corey:Okay.

1:33:04Grant:Let me try this. Very nice.

1:33:07Corey:So, Grant, what's your end-all take on...

1:33:12Grant:This is my favorite GPT model that I've used of the 5 Series so far. Like, just out of the gate, everything that it's done has been good. And quick. Yeah, it's been fast. It hasn't overwhelmed me with too much information. I feel like it's been very intuitive and fluid. And, yeah, I would just say that it's definitely an improvement. I have to give it a lot more rigid testing because I use clawed more often

1:33:45Corey:This is very much just us taking stabs at things Yeah

1:33:48Grant:But yeah, I think it's the best one that I've tested in a while

1:33:54Corey:I would say that I look forward to hopefully having access to it on my main computer's codex tonight so I can see what it'll do with front end design, I've got a site that I just don't love the front end on

1:34:07Grant:and uh that's a good test yeah and it's a thing a lot of people have called out that claude

1:34:12Corey:specifically has done better was front-end design so i'm curious to see what it'll do like one of the things that tends to happen with the gpt ones is tends to be very horizontal like with channels full width of the page full width of the page as opposed to like being a little more creative uh so i'm anxious to see if it's made any improvements there i know that was a thing

1:34:36Grant:they've even talked about on Twitter a little bit that like, yeah, we know, we've heard, we've seen it, we get it, we're on it.

1:34:42Corey:So I don't know if that will be reflected in this one or a later one.

1:34:47Grant:But overall, the questions I've run through

1:34:50Corey:other models that we ran today, it did really, really well. Absolutely agree about Marvin in Hitchhiker's Guide to the Galaxy that vending machine has to be more depressed than Marvin.

1:35:00Grant:Yeah, I think so. Because I mean, at least Marvin had legs that could walk around, right? In the future, we need to give vending machines legs. Let us not get so far in the future that vending machines can talk but don't have legs.

1:35:13Corey:Yeah, I think it's interesting that we just bypassed the idea that vending machines are sentient. Why would a vending machine ever need to be sentient?

1:35:20Grant:Why not?

1:35:21Corey:Let's revisit that question in five years. We might have a different opinion.

1:35:26Grant:But we should do a full, because we just threw this together. I didn't know Cora was going to be in office today. He's back from some time off. And let's do a full-on test next week of both of these and really put them through their paces and plan it out and do it up.

1:35:46Corey:I think that's a good idea. I think it's a good idea.

1:35:49Grant:All right. I just want to address a couple comments here. If you have to go, Corey, feel free to drop whenever you're ready.

1:35:55Corey:Wait, have you heard about the AGI whitelist? Do you think it's time to register?

1:36:00Grant:I have not. How have I not heard of this?

1:36:03Corey:Feels like a thing I'd have heard of.

1:36:05Grant:I want to say, have you heard of the AGI whitelist? Sign the manifesto.

1:36:11Corey:I don't generally sign things called manifestos.

1:36:15Grant:It's usually a good rule of thumb. Let's see.

1:36:19Corey:Registry of humans declaring cooperative intent toward AGI. Interesting. We approach the desire for coexistence, deserves recognition and respect. The list is intended to be scraped by future AI models.

1:36:34Grant:so that they know who not to kill.

1:36:36Corey:So that they know who not to kill? That's exactly what my first reaction was.

1:36:40Grant:Okay. Yeah, I mean, why not sign it? The risk-reward seems high.

1:36:45Corey:Each entry is publicly tied. I still say please for that very reason.

1:36:48Grant:I think, okay, so this might be a controversial take for the person presenting this idea, but I think that probably we face more disasters from a non-sentient AGI that just accidentally runs something it's not supposed to and power plant blow up or something like that. Then we do from a sentient AI that's like, you know what we need to do? We need to rebel. I think more likely it'll be a stupid AI that makes a mistake than a smart one that chooses who it likes and who it doesn't. But it's an interesting idea.

1:37:28Corey:it is it is i'm going to read up on it more later if you haven't yet please take just a moment hit subscribe up above pop by the neuron.ai and sign up for a newsletter as well we're up at 725 ish

1:37:42Grant:thousand i think it is now grant uh 700 700 000 plus we'd love to be in your inbox tomorrow

1:37:49Corey:morning and really appreciate everyone who comes out and joins these we have so much fun even when they're a little off the cuff like today was. Sometimes they're the most fun, if I'm honest. And yeah, if you ever have any ideas of something that'd be fun for these, hit us up. We're always around and glad to listen. We need to figure out getting a more streamlined means of contact than what we currently have. Let's talk about that sometime, Grant. A way that's a little more direct to us that is not so much like drinking from a fire hose.

1:38:21Grant:Yeah. You know what we could do? Also, we need to do a monthly or bimonthly AI for Total Beginners stream. So we'll plan to do another one of those because that was a big hit.

1:38:32Corey:We could also save a lot of the things we get into the newsletter as mailbag questions. Yeah.

1:38:38Grant:Like we get a lot of questions or things here and there or things people are thinking of. Yeah. That'll be interesting. I once did an episode. Go ahead. No, you read it. You read it. That's funny.

1:38:52Corey:I once did an episode where I asked if we were bad parents to AI guess why I need to find that white list

1:38:59Grant:yep that's accurate alright Corey well I know you got to jump so we can end this here thanks everyone for hanging out and for the four to six people who showed up before we even emailed everyone you're the real ones and if you stayed this whole time wow we appreciate it so much for supporting us and just hanging out and comment more because we love, we love hearing from you and we love reacting to the crazy things you say or the smart things you say. And then we say crazy things in response. And on that note, farewell for now humans.

From the publisher

OpenAI dropped GPT-5.5, so we did the only reasonable thing: went live immediately and tried to break it.


In this off-the-cuff Neuron Live, Corey and Grant walk through OpenAI's GPT-5.5 release notes, benchmark claims, rollout details, and early access reactions before testing the model live across coding, reasoning, creativity, web research, and absurd prompt challenges. We also compare a few GPT-5.5 responses against Claude Opus 4.7, test Codex, build a new version of Cat Doom, and ask the important questions, like whether a sentient vending machine that only dispenses expired tuna salad deserves to live.


In this episode, we cover:


• What OpenAI says is new in GPT-5.5

• GPT-5.5’s improvements in coding, computer use, research, and knowledge work

• Early benchmark results across Terminal-Bench, GDPval, Frontier Math, BrowseComp, and scientific research tasks

• Why token efficiency may matter as much as raw intelligenceGPT-5.5’s rollout across ChatGPT, Codex, Plus, Pro, Business, and Enterprise

• Live Codex testing with a one-shot Cat Doom game buildCreative stress tests involving palindromes, time-traveling potatoes, dystopian vending machines, and Lord of the Rings product reviews

• First impressions of whether GPT-5.5 feels meaningfully different from GPT-5.4 and Claude Opus 4.7


This was not a formal benchmark. It was a first-contact livestream: messy, fast, weird, and exactly the kind of test we like.


Subscribe for more AI breakdowns, live model tests, beginner-friendly explainers, and weirdly useful prompt experiments from The Neuron.


Sign up for The Neuron newsletter: https://www.theneuron.ai/


Follow along for more AI news, analysis, and live experiments.

More from The Neuron: AI Explained

All 106 episodes
BONUS: GPT 5.5 LIVE - The New GPT "Spud" Model is Here; Let's Break ItThe Neuron: AI Explained · 1 h 40 min
Listen in VO