In short
The Startup Ideas Podcast - Episode Summary
Episode Title
Claude Opus 4.6 vs GPT-5.3 Codex: Live Build, Clear Winner
Episode Description In this episode, Greg Isenberg interviews Morgan Linton (Cofounder/CTO of Bold Metrics) to discuss the same-day release of Claude Opus 4.6 and GPT-5.3 Codex. They explore how to set up Opus 4.6 in Claude Code, the philosophical split between autonomous agent teams versus interactive pair-programming, and conduct a live build to see which model outperforms the other while creating a Polymarket competitor from scratch.
---
Key Takeaways
Introduction (00:00)
- The episode begins with anticipation surrounding the releases of Claude Opus 4.6 and GPT-5.3 Codex.
- Greg emphasizes that this episode is geared towards technical audiences looking for actionable insights.
Setting Up Opus 4.6 (03:26)
- To use Opus 4.6:
- Update Claude Code to version 2.1.32+.
- Set the model in `settings.json` and enable the experimental Agent Teams feature.
- The standout feature of Opus 4.6 is multi-agent orchestration, allowing parallel processing for various tasks.
Philosophical Divergence (08:32)
- Claude Opus focuses on autonomous agents, while GPT-5.3 Codex emphasizes interactive pair-programming.
- Different engineering philosophies reflect broader trends in how developers approach coding tasks.
Live Demo Setup (15:27)
- The hosts design prompts to create a Polymarket competitor using both models.
- They prepare for a head-to-head comparison while ensuring a controlled environment for both models.
Race Begins (18:26)
- Codex finishes its build in under 4 minutes, while Opus takes longer but aims for a more polished output.
- Observations on token consumption highlight significant differences in resource usage between the two models.
Model Performance Analysis (21:02 - 45:40)
- Codex:
- Achieves completion quickly.
- Excels in interactive collaboration, allowing for mid-task steering.
- Results in 10 tests for its build.
- Opus:
- Takes longer, consuming significantly more tokens due to multiple agent orchestration.
- Completes 96 tests, indicating thoroughness in testing.
- Produces richer results and a more refined UI.
Final Comparisons (44:22)
- When comparing outputs, Opus produced a more polished and feature-rich application despite taking longer.
- Conclusion is drawn that while Codex is faster, Opus's results are superior in terms of quality.
Final Takeaways (45:40)
- The choice between models depends on user preferences:
- Codex: Better for interactive, fast-paced development.
- Opus: More suited for delegating tasks and requiring less user involvement.
- The episode emphasizes the importance of understanding both models' capabilities to leverage them effectively in various projects.
---
Additional Resources
- [30+ Startup Ideas Database](https://gregisenberg.com/30startupideas)
- [Bold Metrics](https://boldmetrics.com)
- [The Vibe Marketer](https://www.thevibemarketer.com)
- Follow Greg Isenberg on:
- [Twitter](https://twitter.com/gregisenberg)
- [Instagram](https://instagram.com/gregisenberg/)
- [LinkedIn](https://www.linkedin.com/in/gisenberg/)
Social Links for Morgan Linton
- [Twitter](https://x.com/morganlinton)
- [Personal Website](https://linton.ai)
---
This summary encapsulates the major discussions, observations, and conclusions drawn from the podcast episode, providing a structured overview for readers keen on understanding the nuances of AI model performance in the context of coding and development.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroducing Morgan Linton and the Episode Goals
0:45 to 3:02
Morgan Linton discusses the new models and what to expect in the episode.
“So we put the models head to head, and there's a winner at the end.”
Understanding Opus 4.6 and Setup Tips
3:02 to 8:24
Detailed tips on how to set up and utilize Opus 4.6 effectively.
“All right, so I took some notes and essentially, you know, with 5.3 Codex, I'll be showing that in the desktop app on Mac because they're super excited about that.”
Philosophical Divergence of AI Models
8:24 to 9:30
Discussion on the philosophical differences between GPT-5.3 and Opus 4.6.
“I thought I would actually read this because this was posted on Hacker News four hours ago.”
Comparative Analysis of Opus 4.6 and GPT-5.3
9:30 to 14:01
Exploration of the strengths and weaknesses of each model.
“And also, hopefully, you know, everybody wants to pick a winner where it's like, oh, no, no, Opus 4.6 is better.”
Task-Driven Autonomy: Claude vs Codex
14:01 to 15:20
Explore the differences in task-driven autonomy between Claude 4.6 and Codex 5.3.
“And then you can restart, you can really fix things in line.”
Prompting for Team Building
15:21 to 16:58
Learn how to use prompts effectively with Opus and Codex for team building tasks.
“So this is now, I have like zero canned demos because I thought it'd be more fun just to try something together and see what happens.”
Live Coding Challenge: Building a Competitor
16:59 to 22:58
Join a live coding session to build a competitor to Polymarket using Opus and Codex.
“It's just about letting everybody have a fair shot to play the game.”
Performance Insights: Testing and Tokens
22:59 to 28:00
Discover the testing results and token usage during the coding challenge and implications for investors.
“So coherent pricing, slippage, bound loss behavior, domain trending engine.”
Token Usage and Monetization Strategy
28:00 to 29:20
Explore the implications of token usage and pricing in AI tools.
“I mean, this is literally, I mean, if you add that all up, you're talking about over 100 ,000 tokens.”
Building a Competitor: Token Calculations
29:20 to 30:40
Learn about the calculations involved in building AI competitors and token costs.
“do you remember how many tokens you get approximately?”
Show all 18 chapters
User Experience in AI Development
30:40 to 32:40
Discuss the importance of user experience and design in AI applications.
“But now, as I say, I'm watching the tokens creep up.”
Jack Dorsey's Design Influence
32:40 to 35:00
Investigate how design influences can shape AI-generated outputs.
“We don't have to take out second mortgages on our...”
AI Understanding and Interaction
35:00 to 38:00
Examine how AI interprets prompts and interacts with user input.
“It just seems like it's really just taking this part of your question and going, oh, okay, major refresh.”
Comparison of AI Systems: Codex vs Opus
38:00 to 40:20
Comparison metrics between Codex and Opus in terms of functionalities.
“I mean, this should be a work of art, whatever cycle this comes off.”
Initial Impressions of the Final Product
40:20 to 42:00
Get initial impressions of the output from AI systems after extensive usage.
“So first off, one really interesting thing here is, you know, Codex created 10 tests, right?”
Exploring the User Experience of Opus 4.6
42:00 to 44:00
The hosts dive deep into the user interface and design of Opus 4.6, discussing its effectiveness and features.
“Yeah, this doesn't even, it doesn't feel like an MVP.”
Comparative Analysis of Opus 4.6 and Codex
44:00 to 46:20
The hosts compare Opus 4.6 and Codex, evaluating their performance and user experience during a live build.
“In the end, it kind of, it's acting a little bit like data from Star Trek.”
Morgan's Insights on AI in Fashion
46:20 to 48:30
Morgan shares insights on how AI technology, particularly Opus, is revolutionizing the fashion industry.
“And then it has like a sample prompt, more of the details on the display mode.”
Transcript
Automatic transcript. May contain errors.0:00Morgan Linton:Today's a massive day because Anthropic just dropped Opus 4.6 and OpenAI answered with GPT 5.3 codex. But what is the better model and how do you get started and what are some tips and tricks to get the most out of them? Well, this episode is all about that. This is for the technical person who's trying to get the most out of these models, who don't just want hot takes, who want tactical sauce for getting the most out of these models. This episode of The Pod is with my dear friend, Morgan Linton. Morgan is one of the best engineers I know. He was an executive at Sonos. He's invested in a lot of AI companies, and he's building an AI company of his own.
0:41Morgan Linton:He's one of my first calls when I'm like, hey, which model's better? So we put the models head to head, and there's a winner at the end. We rebuild Polymarket, a multi-billion dollar app, but we use these models. So which is the better one? You'll find out by watching this episode. but you'll also learn to become a better AI developer because you'll have these tips and tricks in your back pocket.
1:12Morgan Linton:I'm with one of my favorite people, Morgan Linton. You might not know him, but he is just an incredible developer, founder, entrepreneur, investor. He does it all, but today, what I needed him to help me understand is Opus 4.6 just came out. GPT 5.3 Codex just came out. Morgan, help me understand. By the end of this episode, what are people going to get out of this?
1:38unknown host:Yeah, well, Greg, thanks for having me. Super exciting day. It's moving fast today. Opus 4.6 came out, and then Sam Altman put together a quick tweet, I want to say maybe 18 minutes later, announcing GPT 5.3 Codex. and me, I think everybody else has been jumping on it, playing around, figuring out the differences, you know, all the little neat new settings that there are in each of these. By the end of this, you're going to know first how to make sure that you are running Opus 4.6 and all of the little details you can change in the settings.json file to use some of the cool features in Opus 4.6, especially agent teams, which is probably the feature I'm the most excited about.
2:19unknown host:it'll also understand why you might use one versus the other because they both kind of tackle different engineering methodologies and then hopefully you'll see some cool stuff as we build some demos together that I've put together that I haven't tried myself so I'll be trying just live with you so we'll see how that goes
2:37Morgan Linton:I think one of them is we're going to try to recreate Polymarket and see which model performs best
2:45unknown host:They're going to do a head-to-head to try to each build their own version of Polymarket.
2:51Morgan Linton:So by the end of this episode, you will have a pretty good understanding of how to use the models, when to use the models, how to get started. Morgan, let's get into it.
3:01unknown host:Cool, right on. All right, so I took some notes and essentially, you know, with 5.3 Codex, I'll be showing that in the desktop app on Mac because they're super excited about that. I'm excited about it. I think if OpenAI was wanting a demo to be done the right way, they would want me to do it in their app. Whereas with Opus 4.6, I would say the Anthropic team would want me to do it in the CLI. And so there's a few different configuration settings that you do want to make sure that you get right when you're using Opus 4.6. We're trying to use Opus 4.6 today, tomorrow, whenever it is that you're jumping in to use it.
3:39unknown host:I've seen a lot of people online today on Twitter saying, it's weird, I'm having a problem. Like it's supposed to be agent teams, but I don't see them. Or how do I know what version I'm running? So I thought, let's start by just giving everybody a level playing field to know, okay, I want to be able to use quad code with Opus 4.6. How do I make sure I'm doing that and doing that correctly? So here's kind of the initial to-dos that everyone should have on their list. Just do an NPM update. See if that does the trick. If that doesn't and you're running an older version, then run clod update. But you should see, as of right now, it's 2.1.32.
4:19unknown host:If you see 1.something, you're running an old version. And then what you want to do is go into your settings.json, and I'll just show this here. So if you just do like cd tilde slash clod.
4:37Morgan Linton:so I bet that there's people who are running the old model they don't even realize it I have no idea
4:45unknown host:yeah yeah so I mean make sure you go in here cd-slash.clod here's your settings.json if you view this here's essentially what you should see now it's okay it can be model if you want to like really be specific about it you can put in clod-opish-4-6 that'll lock it in But because 4.6 is the newest model, you can also just put in model and just Opus and that'll work. The key thing that you want to do is, in my opinion, the coolest feature that they added with 4.6 is Agent Teams. I'm super excited to demo that with you. You have to make sure to turn that on because it is an experimental feature. And that's probably the biggest confusion I'm seeing people have today with Opus 4.6 is that they are running Opus 4.6.
5:35unknown host:They keep hearing about agent teams and they're giving it prompts like, build a team of agents, do this and this. And it's not quite doing it. And that's because you do have to enable this. So you do have to add in env this Claude Code Experimental Agent Teams and then set it equal to one, okay? Nothing too crazy. Once you do that, that will make all that possible. So with that in mind, you're pretty ready to go there. Then you can just run Claude in the terminal and you're good. For people that are using the API, the one thing I did want to point out is there's a pretty cool new addition, which is called Adaptive Thinking.
6:14unknown host:Also, just to be clear, because I'm seeing confusion on this too, this is in the API. This is not in Cloud Code itself. But Adaptive Thinking, just to show it here, you're able to essentially pick the level of effort that you would like the model to use. This is only going to work in 4.6. by the way, if you want to use like an effort level of max. And so here's kind of a different level. So with max, Claude always thinks with no constraints on thinking depth. It's Opus 4.6 only. So requests using max on other models are going to return an error. So if you're calling the API and you set the effort level to max and you get an error, then you're probably not using Opus 4.6.
6:56unknown host:But here's the example where you can see if I'm calling the API, I set the model to Claude Opus 4.6 and then here's where I can set the effort. And this is another thing. If you're using existing API code, you may have the model of Opus 4.5 and now you adjust the effort to max. It gives you an error. All you need to do is just bump the version and you're good. But this is kind of a neat thing they've added to the API with 4.6. It's worth mentioning. And then kind of the last thing I would say is just if you want to use split panes for agents. So if you want agents to show up in different panes and you're using something like warp, just make sure to install Tmux.
7:39unknown host:You can do this with install Tmux. And then if you do that, it's going to default to auto, which usually means in process, which means in that same terminal window you have, the agents are going to be working all together. If you want it to split pane, then you just need to update that setting in the settings.json to split panes. I'm not going to go into super details on that, but those are just like, I think, good housekeeping to start with for anyone using Opus. But don't worry about it. Really, all that anybody needs to do, especially if you don't even want to use Teams, agent Teams, is just make sure you're updated using the newest version and that the model is Opus and you'll be using Opus 4.6.
8:20Cool. So that's that.
8:23unknown host:Before I get into kind of the differences between Opus 4.6 and Codex, I thought I would actually read this because this was posted on Hacker News four hours ago. And I was reading it. I was thinking that's like the best way to explain it. So I'm just going to read this little section here because I think they do such a good job with it. This person is saying, what's interesting to me is that GPT-5 through Opus 4.6 are diverging philosophically. and really in the same way that actual engineers and orgs have diverged philosophically. I think this really nails it. With Codex 5.3, the framing is an interactive collaborator.
9:00unknown host:You steer it mid-execution, stay in the loop, course correct as it works. With Opus 4.6, the emphasis is the opposite, a more autonomous, agentic, thoughtful system that plans deeply, runs longer, and asks less of the human. That feels like a reflection of a real split and how people think LLM-based coding should work. Some want tight human-in-loop control. Others want to delegate whole chunks of work and review the results. And I honestly, I think that says it beautifully. And I think that nails the differences. And also, hopefully, you know, everybody wants to pick a winner where it's like, oh, no, no, Opus 4.6 is better.
9:40unknown host:Codex, it's different. It depends on what your methodology is. And I think what we're seeing now, not just with Vibe coding, but also with like overall like AI-powered engineering is how do you want to work with agentic coding? Do you want to have a totally autonomous experience where you're sending agents out to do work? Or do you want to work with an LLM like another teammate and pair a program with the LLM? And that's where you're now seeing a divergence where I think you're going to see a lot of teams using both because Codex really is your collaborator And what they've added with 5.3 is like really good, like mid-execution steering.
10:23unknown host:Whereas with Opus 4.6, it's probably the best of the best now being able to say, I want to spin up three or four agents. I want them to go do stuff. Hey, don't bug me. I want to trust they're going to do good stuff. And it's able to deliver.
10:38Morgan Linton:So are you saying that there, in some ways, it's just a preference? like depending on how you know there's no right or wrong basically you know not wrong to be an opus person or you know it just like might feel yeah it's just a preference
10:56unknown host:yeah well and you might be both right that's true it might turn out that you're both that's why like not to disappoint people here but we're not going to end this with me saying and so the winner is it's like well depends on what you want to do everyone has a different methodology for it so I'll dive in and try to try to make this part fast because I know the fun part is probably us going in and playing around with both of these and having them do a head-to-head and try to build a competitor to Polymarket in however much time we have. But I'll just start kind of going into these at a high level just so for anyone who wants to know, like, what are the core differences?
11:36unknown host:Why is this so interesting? Just going to what that is. So with Opus 4.6, much bigger context window, So you have a million token context window here. Very strong coherence over entire documents and repos designed for, you know, like load the whole universe and reason over it. Five, three, they talk about large context, but it's not a headline feature. And I actually went back and forth with it to get it to actually give me a number. And the number is around 200 ,000 tokens, which is not that impressive. That's smaller than I was thinking it would be. but that's okay. It's optimized, you know, for progressive execution rather than total recall.
12:18unknown host:So that's why that's not as important and, you know, optimized for deciding like what to keep in working memory. So high level, what that means is Claude is better when the task is understand everything first and then decide. GPT-5.3 codecs is probably better when the task is decide fast, act, iterate, more of that, you know, pair programming, you know, mid-task change behavior. For coding benchmarks, you know, Opus 4.6 is really good at code-based comprehension. Refactors with like architectural sensitivity, explaining why a system behaves a certain way. And then, you know, a little less tendency of this like YOLO write code, right?
13:03unknown host:That's good. Which is, I think, something everybody wants. Yeah, exactly. So, you know, that's good for everybody, but especially for Vibe coders that are getting started and they may not be able to identify hallucinations, Opus 4.6 is definitely going to perform better there. But then for teams, you know, building in large code bases like me and my team are doing, that's also really important. So kind of a win for everyone there. 5.3 Codex did win on SWD Bench Pro, Terminal Bench. Overall, it's like scored better on coding benchmarks. So probably better end-to-end app generation. and, you know, Claude's kind of like senior reviewer, staff engineer, GPT-5.3, probably like your founding engineer, right?
13:49unknown host:Agentic behavior, Opus 4.6, this is the key one, right? It's like the multi-agent orchestration. That's probably like the bleeding edge feature in 4.6. And then with 5.3 Codex, really like task-driven autonomy, build, test, modify without being asked, but then this task steering, you can watch it, you can go in, it's like your buddy's coding and you can say, oh, wait, wait, man, wait, why are you doing this? And you can stop it and it'll go, okay. And then you can restart, you can really fix things in line. Much harder to do that with Opus. With Opus, you'll kind of be stopping it and then starting somewhat fresh, but it has a pretty big context window so it knows what it did.
14:33unknown host:But, you know, Claude's really asking like, should we do this? GPT-53 is like, how fast can I ship this?
14:40Morgan Linton:It's really, I mean, it's so cool because it almost feels like they're different people. You know what I mean? Like they have different styles.
14:48unknown host:Yes, totally. Yeah, it's a good way to look at it. It's like a different personality type, right? And then, yeah, failure modes, you know, Claude 4.6, it might overanalyze. It's got a much bigger context window. It can hesitate when requirements are ambiguous. and then it can stop short of full execution. 5.3 codecs could be overconfident, can lock in a flawed assumption early, but you can steer it back in the right direction if that happens. So that's kind of a high level overview on the two.
15:25Morgan Linton:Cool, that's helpful.
15:27unknown host:Yeah. So should we just dive in? I haven't tested any of this. So this is now, I have like zero canned demos because I thought it'd be more fun just to try something together and see what happens. So should we try it?
15:41Morgan Linton:Yeah.
15:42unknown host:Okay, so let's see. I'm going to start with Opus, and I've got these prompts preloaded. So I'm giving different prompts, just like I think you said it really well. It's like you're talking to different people. And so, you know, when I'm talking to Opus, I can tell Opus, build me a team, and here's what I want each member of the team to do. When I'm talking to Codex, I can't really tell it to build me a team, but I can tell it to think about stuff. So the prompt that I'm going to give to Opus is build a competitor to Polymarket. Create an agent team to explore this from different angles. One teammate on technical architecture, one on understanding Polymarket and the ins and outs of prediction markets, one on UX, and one that just works on building really good tests to make sure everything works.
16:29unknown host:For Codex, I'm going to give it a little different prompt, but very similar. So still build a competitive polymarket, but now think deeply about technical architecture. Understanding polymarket and the ins and outs of prediction markets. Good, clean UX. Make sure it builds really good tests to make sure everything works. And to be fair, I'm going to try to pace these in around the same time. You're a fair guy, Morgan. I'm trying to keep it fair here, right? It's the only way to do it. Like I said, no winners or losers. It's just about letting everybody have a fair shot to play the game.
17:05Morgan Linton:Yeah.
17:06unknown host:All right. So let's see. I'm going to make different directory streams. I'll do, let's just call this opus45.com market competitor. All right. So let's fire up Claude in here. By the way, if you want to check when you're running, just to really make sure that you're in a good place with the model, If you type slash model, I can see here, right, cloud opus 46, right? So I'm good there. I'm going to take this prompt, copy it. Make sure this is all copied incorrectly. Okay, got that. I'm not going to hit enter yet. I'm making this totally fair. I don't want anyone at Anthropic or OpenAI to get upset with me.
17:52unknown host:So I want to be in good terms with both of them. Totally.
17:56Morgan Linton:Smart guy.
Read the full transcript
17:57unknown host:let's see oh wait actually you know what I do want to create a new folder for this
18:01Morgan Linton:but we are keeping it real we're being objective neither myself or Morgan are affiliated with either well actually I don't know about you I'm not affiliated with graphic or open AI nope I love them both equally
18:16unknown host:how about that okay and I'm going to try to start them as close to the same time as I can enter go So, all right, they're going. Off to the races.
18:30Morgan Linton:So, what do you think is going to happen?
18:33unknown host:That's a great question. Well, I know right now, because I told Opus 4.5 to build using different teammates, it's going to do that. So, you can see here it says, I'll build a Polymark competitor by launching parallel research agents first, then synthesizing their finding to a comprehensive implementation planning code base. This is brand, brand new, right? Like if I did this with Opus yesterday, wouldn't be possible. That's kind of the difference here is that the way that Codex is working is the way things have kind of always worked, right? So if you see, this is like the individual person, right?
19:15unknown host:It's not saying, okay, I'm going to launch all these different agents and compare what they say. It's like, okay, I'm going to inspect the workspace. This is your, you know, really detail oriented, really senior, like founding engineer, like that example gave, right? Whereas over here, you can see it's already launched these agents and now it wants to do web searches and I'm going to let it do that. So multiple agents were asking to do web searches. So now, now launching all four research agents in parallel. So this is off and I've got, you know, my technical architecture agent. I've got this other agent that these are both doing web searches right now.
19:54unknown host:So one is looking at like prediction market order book matching engine architecture. So this one's learning about engine architecture for prediction markets. This one's looking at polymarket, how it works, barnyard predictive market mechanics. And then I've got the UX design is doing some design research. And then we've got some test research. Okay, now it's going to go to polymarket. And let's really hope the polymarket doesn't block it because that'll make things harder for it. Meanwhile, over here, this has discovered codex has figured out the repo is empty. So it's going to scaffold it from scratch.
20:28unknown host:And it is starting to, I'm now wiring the core market math and trading engine. So it's interesting, right? So you've got codex is out here building and is like building the engine, with Opus 4.6, it still has agents out there doing research work.
20:48Morgan Linton:You really start to see just how different they really are as they make progress.
20:55unknown host:Yeah. Like I said, I haven't tested it before, so we don't know how long it'll take each of these.
21:00Morgan Linton:Totally. Yeah, and I think one question I have is It's like, is one model better for being more of a beginner, non-technical live coder? Or, you know, it doesn't really matter.
21:16unknown host:Yeah, it's a good question. I mean, I think the fair answer would be probably Codex. Because Codex edged out Opus 4.6 a little bit on some of those coding benchmarks and is kind of known for writing better production code, probably codex in that way. At the same time, one of the downsides, and like I said, I could only do this in a totally balanced way because they're so different. At the same time for a vibe coder, knowing when to interject and stop codex and say, oh, wait, you're doing this this way. Can you instead look at doing it this way they're probably not going to know how to do that.
22:01unknown host:And so that's where maybe Opus 4.6 is better where you could say, okay, spin up four or five agents and let them work with each other. Yeah. Okay, Codex is done. All right. So Codex built a competitor to Polymarket in three minutes and 47 seconds.
22:22Morgan Linton:And to be clear, Polymarket's a multi-billion dollar company.
22:25unknown host:Yeah, I don't think this will work quite as well. Um, but, uh, but we'll see, let's see. So, um, let's just check out if it worked first. I'll let this keep running here. Um, so, you know, it'll tell you at the end here, it actually did the testing. So you can see it built a test suite. So it has an LMSR math unit test suite, an engine behavior unit test suite and an API integration, uh, test suite. And it passed with 10 out of 10 tests. As far as what it built, it has this core LMSR market maker engine. So coherent pricing, slippage, bound loss behavior, domain trending engine. It built a REST API router, which is kind of interesting because I didn't tell it that it would have to build obviously any of this in any way.
23:14unknown host:It figured out the architecture on its own. Clean responsive front end. All right, well, let's see. Let's see if it is actually. So let's go here. I'll let this keep running this has got these four agents just running away here and I'm in here I'm going to do npm test all right test 10 past 10 that looks good to me npm start all right it's running let's see okay here we go so So this looks like it has the ability. So let's, Greg, let's make you, we'll make you the first trader. All right. Say add. Okay. You got a thousand bucks.
24:05Morgan Linton:Okay.
24:06unknown host:There we go. Not bad. All right. What market do you want to create?
24:12Morgan Linton:Well, Bitcoin, I think as we speak, has crashed to what, 63 ,000 or something?
24:17unknown host:Something like that. Yeah.
24:18Morgan Linton:So I do like the, I mean, will BDC be above 110K by...
24:25unknown host:Yeah. Okay.
24:28Morgan Linton:By DESC 31, 2026. That's pretty good.
24:30unknown host:Yeah, okay.
24:30Morgan Linton:That's pretty good.
24:31unknown host:So let's... I mean, that's almost double. Yeah. That'd be pretty good.
24:37Morgan Linton:It depends. Depends when you bought it, you know? If you bought it at 125K, then...
24:43unknown host:Yeah.
24:44Morgan Linton:You're not so happy, but... Let's see.
24:47unknown host:So then I don't know. I don't even know what resolution criterion source would be. I mean, I think I know what it's getting at, but I guess you could say like, why don't we say use CoinMarket cap as the source and resolve by looking at the price on the last day of December, just before midnight, I guess.
25:16Morgan Linton:I guess like
25:19unknown host:the price of BTC yeah alright okay it looks like it's okay so we've got it now so we'll use CoinMarketCap okay so then you can do a yes 50 % so what do you think yes or no
25:38Morgan Linton:I mean this isn't financial advice this is just purely purposes but I think so So I think that.
25:47unknown host:All right. That's a yes for Greg. Buy. Let's see. How many shares you want to buy? You've got a thousand bucks.
25:56Morgan Linton:I want to put it all. I'll put it all.
25:59unknown host:I don't know how much it is per share. Let's see if it's a thousand, if that's right. Okay. Yeah. Okay. Trade executed. Okay. So, I mean, it seems like it built something, you know, as a prototype, relatively functional here. I guess that it actually has decremented. So, okay, 1 ,000 shares was not. That ended up being, you know, about$24 that you spent. So, you've got more money if you wanted to create another market. But it worked. It's not returning an error. It shows the volume here. Interesting. All right, so let's go back. Let's see. So far, so good with that, I'd say. Let's see what's going on here.
26:40unknown host:So, we've got, okay. So, first off, look at how many tokens. People have been talking about how token hungry Opus is. And it's very token hungry. Each one of these agents has used over 25 ,000 tokens. So let's see though. So they finished, right? The technical research around architecture is done. Prediction market research is done. The UX design research is done. The testing strategy is done. Now it's going to go and build. so it's writing the package JSON
27:17Morgan Linton:did you see the ad that Anthropic launched about ads
27:23unknown host:yes I watched them all they're hilarious although actually I guess Sam was not very happy about them today I saw a tweet from Sam that was less than happy so I found them hilarious but I also understand his side as well
27:38Morgan Linton:So basically, it seems like Anthropic is sort of anti-ads for now. And obviously, ChatGPT is going to be introducing ads.
27:47unknown host:Yes.
27:48Morgan Linton:And, you know, when I'm watching this and I'm seeing you're going through 25 ,000 tokens, 25 ,000 tokens, 25 ,000 tokens. Yeah. I'm like, yeah, of course, Anthropic doesn't really.
27:58unknown host:Yeah.
27:59Morgan Linton:Exactly. You know? Yeah.
28:02unknown host:Yeah. I mean, this is literally, I mean, if you add that all up, you're talking about over 100 ,000 tokens. used in doing this. So I think that's one of the very good things for investors in Anthropic is with agents and agents now being I think probably the new killer feature in Opus you're going to take whatever token usage and multiply it by the number of agents.
28:27Morgan Linton:Exactly. It's actually really smart and I wonder if that was the thinking they're like how can we get people to use more tokens oh, we'll just spin up agents and we'll design it like that? Or did they think, okay, how can we design a system that is best for the use case? And then they're like, okay, then we'll monetize it like this. I don't know.
28:48unknown host:Yeah, probably a combination of the two. I can tell you I've never used so many tokens in one day as today. So it's working.
28:57Morgan Linton:100 ,000 tokens is roughly how much in US dollars?
29:02unknown host:I don't know because I have a Claude Max plan. $200. Yeah, so I'm not paying. We're not seeing it hit any limits right now, right? So I'm not paying more than$200. I can tell you that.
29:16Morgan Linton:Yeah. My guess is it's, you know, we're talking like in the$200 Max plan, do you remember how many tokens you get approximately?
29:29unknown host:That's a good question. Let me fire up Claude and ask it. Let's see here. How many tokens do I get? Estimate. Let's see. Okay, so here you go. Estimate. So 45 million tokens per month of Sonnet. But let's see. What is your estimate for Opus 4.6? Let's see.
30:00Morgan Linton:It's like they don't really want you to know.
30:02unknown host:No, they're trying to make it a little harder. Okay. They're not even going to tell me, actually. They're just going to say, there's no public data. It's very new. Opus is roughly 5x more expensive. So then if it's 45 million, that's 5x. So 10 million is probably the answer about, right?
30:21Morgan Linton:So then if we're doing quick math, let's just say we spent 100 ,000 tokens, you know 100 ,000 divided by 5 million is you know we're going to spend
30:35unknown host:more than that because look at this we're now over 17 ,000 tokens on top of that in this next build okay so but still let's say you know even if we use a million tokens building a competitor to Polymarker right now we're still only using a tenth of what it but it can do that's not terrible
30:55Morgan Linton:No, I mean, it's$20, which is like the price of a cocktail in Miami.
31:00unknown host:Yeah, yeah, exactly. Yeah, yeah. So let's see. But now, as I say, I'm watching the tokens creep up. All right, so it's building the API routes now.
31:14Morgan Linton:I have a feeling this is going to be a better end result.
31:18unknown host:I was actually just going to say that. This feels, and then maybe it's just because there were four agents that were doing all the work beforehand and now it's doing the work. It feels like we're going to see something very different when we load what it builds.
31:35Morgan Linton:Yeah, I don't think we gave it any design, like visual design. Do you recommend for folks to just get the MVP out, play around with it on localhost, click some buttons and then update with the visual design from there?
31:55unknown host:That's a good question. I do like 50-50. Sometimes if I have something in mind, especially if I want something like on brand with something like suppose I'm building something that is going to be in the like open claw multi-book ecosystem, I would probably say, hey, I want to design a site that, you know, looks somewhat similar to or is inspired by, you know, open claw. dot AI and moltbook.com. Take a look at those sites and get inspiration. These models are great at doing stuff like that. I'm really excited to see what this is doing though. I think we're now like well over 200 ,000 tokens. They found what I could tell.
32:41unknown host:But we're not at 10 million.
32:43Morgan Linton:We're not hitting any limits. We don't have to take out second mortgages on our...
32:50unknown host:Yet. It's still going. I guess, you know, we can tell, like, here's an interesting thing in a comparison. Like, this is still going. Why don't we say, like, the design, because the design looked kind of bland to me, right?
33:06Morgan Linton:Yes, it did.
33:11unknown host:Can you spruce it up and make it look nicer? Because, like, we may as well have codecs working away, too, right?
33:19Morgan Linton:Yeah, so you didn't really give it any specific, it should look like square.com. No, no. We'll basically see if Codex, if 5.3 has a little bit of taste.
33:35unknown host:Yeah, yeah. So that's what it's saying now. It's saying, okay, I'll upgrade the visual system without changing functionality, stronger typography, richer color direction, better card hierarchy, and purposeful motion. I don't know what that means, but we'll find out. All right. Okay. So now it's editing. Index.html. It looks like it's going to add motion hover polish. Okay. Totally that. This current task for over 30 ,000 tokens building the front end UI. Okay. It's done. Yeah. Codex is fast, by the way. Right. I mean, that's pretty darn fast.
34:11Morgan Linton:Yeah.
34:12unknown host:So we should be able to just go here. It should have already automatically reloaded. Okay. All right.
34:19Morgan Linton:I mean, it's not that different. It's not that different. I think, can I try something?
34:28unknown host:Yeah, go for it.
34:29Morgan Linton:I'm going to say, I would say, okay, thank you, but this was a minor design refresh. I'm looking for a major one.
34:41unknown host:There you go, yeah. And then I'm going to say, pretend you are Jack Dorsey. And how would he design this website to be clean, elegant, and full of interesting interactions? Yeah, great. Yeah.
35:07Morgan Linton:Jack Dorsey for people who don't know co-founder of formerly known as Twitter and Square Block now he's a design guy I don't know he's the first guy who came to mind or first person who came to mind
35:22unknown host:yeah that's a good one that's a good prompt let's see so I'll do a full visual re-architect not an incremental tweak new layout language stronger typography monochrome first palette interesting interaction driven cars that's a okay you know what's interesting is it didn't i would have kind of hoped and maybe we you know we're not quite at agi yet i would kind of hope that it would say let me go find some art like if you told me that greg like hey morgan can you read it i would be like yeah let me go look at some articles about jack dorsey's design aesthetic exactly um i'm surprised it's not doing that Instead, it's going like, I am assuming it knows who Jack Dorsey is, although I don't know if it actually does.
36:09unknown host:It just seems like it's really just taking this part of your question and going, oh, okay, major refresh. I'll do that.
36:18Morgan Linton:Well, can't you ask it? Can't you say, do you know who Jack Dorsey is?
36:22unknown host:Let's see. I can actually, I'm supposed to be able to, in the middle, cut it off. So let me see. Yeah. Do you know who Jack Dorsey is? Let's see. Okay. So here we go. This is the midstream test. It's thinking about it. 43 ,000 tokens over here. Okay. Yes. Okay, here we go. Yes. Jack Dorsey is the co-founder of Twitter, form of the Doctrine Square. Okay. With a design style that's typically minimal, restraint, and interaction focused. Okay. All right. So, touche. It showed us.
37:01Morgan Linton:Yeah.
37:03unknown host:now here's the weird thing it looks like it's
37:05Morgan Linton:is it complete or do we have to say
37:07unknown host:complete like right when I was saying that are you done or did you stop because I asked a question
37:13Morgan Linton:yeah this is really interesting I will say like oh I pause when you ask the question the major redesign mostly
37:25unknown host:so that's weird so you ask a question it just stops but like so like yes of course continue. Okay.
37:35Morgan Linton:So that's actually some weird UX. Like it obviously should just continue after. Right?
37:40unknown host:Yeah. I would assume that it's such a weird thing because it said, yeah, if you want, I'll resume now.
37:47Morgan Linton:I will say, I do like that you can, in midstream, like kind of edit things. Yeah. Like that's how my brain works. Totally.
37:56unknown host:Yeah. Yeah. Truth is using a ton of tokens. It is amazing to see the detail. I mean, this should be a work of art, whatever cycle this comes off. All right, this is done. So now let's see, we can go back to this. And okay. I mean, I'm not blown away, but it's okay.
38:22Morgan Linton:Opinions become price in milliseconds. Trade conviction, not noise. Signal market is dying for fast. thesis iteration with transparent pricing. I mean, I would push it more, I think.
38:38unknown host:Yeah, I guess. I would say that's not the Jack Dorsey I know. I don't know Jack Dorsey.
38:53Morgan Linton:I was looking for a Caps Lock major upgrade. Uh, that, that might mean, um, way more copy, way more images, way more storytelling. Yeah, exactly. Et cetera, et cetera. Yeah.
39:15unknown host:I'll just say seriously, take your time. Go nuts.
39:19Morgan Linton:Yeah.
39:20unknown host:What are credits?
39:23Morgan Linton:Famous last words, Morgan.
39:24unknown host:I know, right? It's like, oh, perfect. Okay. that's like a signal within the opening headquarters like we finally got someone
39:30Morgan Linton:totally it's a whale alright so Opus has finished I have no idea
39:36unknown host:how many credits I've used but probably actually let me ask it how many tokens in total did you use to put all of this together including the four agents and then we can Opus is using tokens to answer this okay Okay. Doesn't know. Let's see. Okay, here we go. Okay, it's estimating. It actually doesn't know, which is weird, because it should know. Although, I wonder if I can actually do slash cost. Oh, here we go. Yeah, okay. Oh, it doesn't. Okay, no. No need to monitor cost. Okay, so they really don't want you to know. Okay, it's guessing 150 ,000 to 250 ,000 tokens total.
40:22Morgan Linton:Yeah, that's probably right.
40:24unknown host:Okay, sure. Yeah. Okay, so here's what it's done. So first off, one really interesting thing here is, you know, Codex created 10 tests, right? Opus created 96 tests. So definitely a lot more detail on the testing side. And it's called it Forecast, whereas Codex called it Signal Market. So different names. A Polymark competitor is built and verified. Here's what each team member delivered. So the architecture technical lead decided modular monolith, Next.js 14 app router, central limit order book, database schema, Russell API. Okay, the prediction market domain expert, binary yes no market where yes no is always a dollar.
41:14unknown host:Okay. Seated markets across crypto politics. Okay, the UX design lead, dark mode trading platform. pages it is a green for yes red for no okay testing qa lead did order book tests okay so here's the tests are breaking down order book tests matching engine okay all right npm run to start the app so let's go in here and oh interesting okay i don't want to say anything
41:47Morgan Linton:actually i've already given to him what's your initial take i mean my hello jack dorsey you know what i'm saying like this is this is what i expected it to look like when we pushed codex yeah me too this looks really clean what happens when you hover over uh oh yeah look at that yeah it's got hover states hover state yeah um it's obviously got it organized like when i go sports
42:16unknown host:It's, you know, will the next two will have over 120 million viewers? Will AI pass the Turing test by 2027? Will a movie preserve a three? It's got stuff in there, yeah. Yeah, this doesn't even,
42:28Morgan Linton:it doesn't feel like an MVP.
42:30unknown host:Yeah, this is pretty wild actually. And it created some stuff, you know, that we never talked to it about, right? Like a leaderboard, which it's already populated with some initial stuff. Portfolio section, yeah. Interesting. So let's see now. So, I mean, I'm more impressed with the, it was maybe worth the 150 ,000, 250 ,000 tokens. Feeling better about it. Will SpaceX land humans on Mars before 2030? Only 8 % of things, huh? Oh yeah, look at this actually. Whoa.
43:05Morgan Linton:This is insane, bro.
43:06unknown host:That's clean.
43:08Morgan Linton:That's clean.
43:09unknown host:I wasn't expecting to click in and actually get a well-designed page like this. Huh. so if I were to do that I have to sign in a tray I don't know if I'm going to be able to sign in because I haven't set anything up let me just check
43:21Morgan Linton:well you can sign up it says don't have an account
43:25unknown host:oh yeah sign up I don't know if it gets all connected though let's see though alright I'm snagging you know what actually I'm going to take the username Greg I'll steal your username alright let's see okay so yeah it's probably because I was going to say the database isn't wired up yet so I'm not I'm not surprised that I would actually have to do. I wasn't expecting to do that. But I get it. I mean, it's clean. This is pretty neat. Yep. Yeah. All right, let's see. So then, can this, all right, we've given, I don't know about you, but this is the last chance I'm going to give Codex on the design side.
44:02unknown host:It's out of opportunities here. So, oh, here, it's funny, though. In the end, it kind of, it's acting a little bit like data from Star Trek. Yeah. When you question, what are credits? In this context, credits usually means... That's quite it. That's good, though. Okay, so let's see. Let's take a look and see. All right, here we go. The new version of Signal Market.
44:26Morgan Linton:Boom. Oh.
44:28unknown host:Okay. This is getting a little bit interesting. Let's see here. Read the manifesto.
44:36Morgan Linton:Yeah, I mean, I don't hate it.
44:39unknown host:It's definitely better. Yeah. I mean, it's got a lot going on. Yeah.
44:47Morgan Linton:But it's different. Terminal. It's different than any sort of prediction market app I've seen from a UX perspective. I just feel like this is just so clean, though.
45:03unknown host:This is so good. It's so fast. Yeah. I mean, I would say, like I said, I'm not going to say which one is, It's not that Opus is better than Codex or vice versa, but I would say in this test, Opus won.
45:16Morgan Linton:Yeah, in this test, Opus won.
45:18unknown host:Yeah.
45:19Morgan Linton:That's just the truth. Yeah, yeah.
45:22unknown host:But we could give it another, I mean, you know, you never know. I think that what's interesting about this is, I mean, Codex built it like, I don't know how much faster, we can look at the timing on this video, but like 20 times faster or something, right? Yeah, yeah. Yeah.
45:40Morgan Linton:Well, anything else you wanted to cover? I don't think we'll have time to do another example, but anything else you wanted to cover that you want to leave people with?
45:49unknown host:Yeah. Let me see if there is anything else in here. I covered the adaptive thinking. Oh, I guess just on the orchestration, I would say this is a feature I'm probably the most excited about with Opus. And clearly we saw in this example it working really well. Just make sure to look at the documentation. It's all in the docs now. And it gives some examples as well because it has this idea of like compare with subagents of like context and communication and coordination and kind of breaks this down. And then it has like a sample prompt, more of the details on the display mode. There's a lot of other stuff that I didn't go into there.
46:31unknown host:That's probably what I would leave people with because I think a lot of people are going to want to dive in and use agents with Opus 4.6 and they've got pretty good details on all the little tweaks that you can make with it.
46:45Morgan Linton:Amazing. Well, Morgan, I can't thank you enough for coming on. I hope people love this episode. I love talking to you.
46:54unknown host:Thank you for having me, Greg. It's a total honor.
46:58Morgan Linton:Yeah, it's just, I love how clearly you communicate and to technical people, but also non-technical people. And you're criminally under-followed, so I'm going to include links where you can find Morgan and follow him on X. He talks a lot about vibe coding over there. And Morgan, anything else you want to, places that you want to leave people to go and check you out?
47:25unknown host:Yeah, I mean, I'm the co-founder and CTO of Bold Metrics. I'll just give a little plug for us. We have AI technology that's used by apparel brands and retailers. So if you're shopping online and want to find the right size, you see a find my size button. We a lot of times power that and have really powerful machine learning models that update and adapt over time to help people find the right size and give lots of really interesting data to lots of amazing brands and retailers that you probably all know and love. And then me and my team, you know, we're using all of this tooling. Like I had a meeting with my team this morning about Opus 4.6 and Codex.
48:06unknown host:And I've given everybody access to both of these. And I actually have multiple teams of mine that are trying current things we're working on and are actually testing with each to see which performs better. So, you know, the one thing I encourage all engineering teams to do is like, and engineering leaders to do is like, let your teams loose with this stuff. Let them try it. some of this stuff is really cutting edge and really performing and gives us the opportunity to do better, more creative work.
48:34Morgan Linton:Yeah, stop listening to us right now. Yeah, exactly. Go and get us. X out of this YouTube or Spotify link, but actually give us a like, a comment, and subscribe. Let us know if you like this episode. Morgan, thanks again for coming on the show. This was a lot of fun. See you next time.
48:52unknown host:Greg, thank you so much. Total honor.
From the publisher
I sit down with Morgan Linton, Cofounder/CTO of Bold Metrics, to break down the same-day release of Claude Opus 4.6 and GPT-5.3 Codex. We walk through exactly how to set up Opus 4.6 in Claude Code, explore the philosophical split between autonomous agent teams and interactive pair-programming, and then put both models to the test by having each one build a Polymarket competitor from scratch, live and unscripted. By the end, you'll know how to configure each model, when to reach for one over the other, and what happened when we let them race head-to-head.
Timestamps
00:00 – Intro
03:26 – Setting Up Opus 4.6 in Claude Code
05:16 – Enabling Agent Teams
08:32 – The Philosophical Divergence between Codex and Opus
11:11 – Core Feature Comparison (Context Window, Benchmarks, Agentic Behavior)
15:27 – Live Demo Setup: Polymarket Build Prompt Design
18:26 – Race Begins
21:02 – Best Model for Vibe Coders
22:12 – Codex Finishes in Under 4 Minutes
26:38 – Opus Agents Still Running, Token Usage Climbing
31:41 – Testing and Reviewing the Codex Build
40:25 – Opus Build Completes, First Look at Results
42:47 – Opus Final Build Reveal
44:22 – Side-by-Side Comparison: Opus Takes This Round
45:40 – Final Takeaways and Recommendations
Key Points
Opus 4.6 and GPT-5.3 Codex dropped within 18 minutes of each other and represent two fundamentally different engineering philosophies — autonomous agents vs. interactive collaboration.
To use Opus 4.6 properly, you must update Claude Code to version 2.1.32+, set the model in settings.json, and explicitly enable the experimental Agent Teams feature.
Opus 4.6's standout feature is multi-agent orchestration: you can spin up parallel agents for research, architecture, UX, and testing — all working simultaneously.
GPT-5.3 Codex's standout feature is mid-task steering: you can interrupt, redirect, and course-correct the model while it's actively building.
In the live head-to-head, Codex finished a Polymarket competitor in under 4 minutes; Opus took significantly longer but produced a more polished UI, richer feature set, and 96 tests vs. Codex's 10.
Agent teams multiply token usage substantially — a single Opus build can consume 150,000–250,000 tokens across all agents.
The #1 tool to find startup ideas/trends - https://www.ideabrowser.com
LCA helps Fortune 500s and fast-growing startups build their future - from Warner Music to Fortnite to Dropbox. We turn 'what if' into reality with AI, apps, and next-gen products https://latecheckout.agency/
The Vibe Marketer - Resources for people into vibe marketing/marketing with AI: https://www.thevibemarketer.com/
FIND ME ON SOCIAL
X/Twitter: https://twitter.com/gregisenberg
Instagram: https://instagram.com/gregisenberg/
LinkedIn: https://www.linkedin.com/in/gisenberg/
Morgan Linton
X/Twitter: https://x.com/morganlinton
Bold Metrics: https://boldmetrics.com
Personal Website: https://linton.ai
