In short
Mike Krieger (Anthropic Labs lead) explains how Anthropic Labs builds “AI-native” frontier products, why Fable’s temporary access was pulled amid rapid Trump-administration backlash, and what Labs is working on next (more agent agency, interoperability, and token/effort efficiency). He also discusses how Anthropic balances safety, transparency, and product-vs-platform strategy, and why consumer “breakout” AI apps remain hard.
Guests
Mike Krieger, co-founder of Instagram; now runs Anthropic Labs at Anthropic. Lauren Good, journalist at Wired.
Key claims
Labs exists to prevent product teams from lagging model capability; it closes the gap between what models can do and how people use them. Fable’s removal reduces delegation/scope and makes work less reliable, though other models (e.g., Opus 4.8) still work. Token usage doesn’t correlate strongly with productivity; Anthropic targets token efficiency.
Notable examples
Cloud Code and Cowork; “CloudCode Artifacts” (text plus drawing). A Labs task: converting a large Python codebase to TypeScript via dynamic workflows, planned/executed/verified in about an hour. Model safety risk framing via “uplift” comparisons in model cards (bio domain).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOMike's Role at Anthropic Labs
0:51 to 1:19
Insights into Mike's transition to his current role and the focus of Anthropic Labs.
“In the face of ongoing disruption and opportunity, TMT leaders need to deliver tangible results, not just ideas.”
Mike's Role at Anthropic Labs
1:26 to 2:49
Insights into Mike's transition to his current role and the focus of Anthropic Labs.
“First of all, I mean, we want to talk about what you're working on at labs and explain your role to folks.”
Fable: Model Release Reactions
2:49 to 4:38
Discussion on the public reaction to Fable's release and its implications.
“I've learned to not really trust day of or even week of model reactions.”
AI Safety and Marketing Concerns
4:38 to 8:00
Exploring the balance between AI safety and public perception.
“We're dealing with new situations and they can develop really quickly as well.”
Practical Use of AI Models in Development
8:00 to 12:35
How AI models are changing the development process and productivity.
“So is this ban essentially now limiting your ability to do that within labs?”
The Evolution and Purpose of Anthropic Labs
12:35 to 14:00
Understanding the need and evolution of Anthropic Labs over time.
“So that was, so you're basically saying it was faster.”
The Evolution of Anthropic's Labs
14:00 to 17:08
Learn how Anthropic's Labs have developed to keep pace with model advancements.
“Would you say that's an accurate description of what you're doing at Labs?”
Navigating Product Development in AI
17:08 to 19:25
Discover how Anthropic balances between being a platform and creating products.
“just a text box and a big text response is not going to cut it anymore.”
Ethics and Culture at Anthropic
19:25 to 22:42
Explore Anthropic's ethical positioning and its cultural impact on Silicon Valley.
“that has social messaging, video, if all.”
Future Innovations at Anthropic
26:19 to 26:45
Hear about Anthropic's ambitious plans for AI and product developments.
“Lately, though, the shop's been quiet, so Hank decides to bring back the$1 slice.”
Show all 19 chapters
Future Innovations at Anthropic
26:50 to 28:00
Hear about Anthropic's ambitious plans for AI and product developments.
“And with labs, what you try to do is get ahead of that so you can show people what AI might be able to do now and six months from now.”
Interoperability and AI Agency
28:00 to 29:42
Exploring the interoperability of AI and its ability to understand its environment.
“Until yesterday, I would have said the same thing about Cloud design and Cloud code, where if you're in Cloud code, you're like, cool.”
Bridging the Gap in Work Processes
29:42 to 31:09
Discussing the importance of aligning understanding and execution in work tasks.
“It was eight different steps of copying and pasting, of manually moving, pretty annoying to have to do, probably error-prone.”
The Future of Tokens in AI Economics
31:09 to 33:06
Debating the role of tokens in measuring usage and productivity in AI.
“I think I get very, very excited about those things.”
Outcome-Based Pricing Models
33:06 to 35:38
Exploring the concept of outcome-based pricing in AI services.
“Yeah, I think both of those are really interesting questions.”
Challenges in Consumer AI Adoption
35:38 to 38:18
Analyzing why breakout consumer AI applications are challenging to achieve.
“or we have an outcome-based mode where you can say, here's what good looks like.”
Ethical Considerations in AI Development
38:18 to 42:00
Addressing the potential harms and ethical implications of AI technologies.
“It's just like it's harder in many ways than ever to break through, even if you can code more quickly.”
Introduction and Thanks
42:00 to 42:16
The hosts express gratitude for Mike's insights and welcome the discussion.
“Mike, it's always great to be can we thank you again for bringing your insight today Let's hear from Mike and Lauren Thank you great job Thank you.”
Introduction and Thanks
42:19 to 42:36
The hosts express gratitude for Mike's insights and welcome the discussion.
“It can help you with practically anything on the web, like restoring a vintage motorcycle from a 50-page restoration block, or finally break down that long article you've had open for weeks.”
Transcript
Automatic transcript. May contain errors.0:00Anthropic is not just a massive model builder, it's a massive product builder as well with products like Cloud Code and Cowork that have taken off like crazy over the past few months. And Cloud Code came out of a place that few of us know much about called Anthropic Labs. And Anthropic Labs is a organization within Anthropic that is working on building the next level of frontier products with AI at the center. And so today we are lucky to hear from the person running that lab, Mike Krieger, who is the co-founder of Instagram and now the lead of Anthropic Labs at Anthropic. We're going to welcome him on stage along with Lauren Good of Wired, who will join me as a co-interview.
0:44Mike and Lauren, let's hear it for both of you guys.
0:49Thank you. You're here. In the face of ongoing disruption and opportunity, TMT leaders need to deliver tangible results, not just ideas. When pace and performance matter most, PwC combines market insights and deep sector experience with AI, cloud and emerging tech to accelerate your transformation and drive measurable ROI from strategy to execution. PwC can help you anticipate what's next, outpace disruption and compete. For more information, visit pwc.com. All right. All right, Mike. So chill times in an anthropic land. Nothing going on. Slow week. Slow week. You want to start, Lauren? Yeah. First of all, I mean, we want to talk about what you're working on at labs and explain your role to folks.
1:41But I want to ask you first, how close are you right now to the situation with the White House? Less than in my CPO role. So I transitioned about like five months ago into this labs role. I think in the CPR role, I think would have been deep in it. Now I'm more, you know, obviously, you know, we want to restore it. And as a like product person, I want to make sure that that gets access, but less close to it than, you know, in that sort of C-level role that I had before. Okay. Alex, do you have a follow-up? Oh, yeah. I have like eight follow-ups. Definitely. Well, there was a guy named Ben on X who said, I will mail Anthropic an original copy of my long-form birth certificate if they will enable Fable for me again.
2:18I sound like those lunatics who are obsessed with 4.0 now. Will you take Ben's long form? I don't know that we'll take his long form, but it has been interesting. I mean, Fable was only available for a few days, but I've definitely, every time I've tweeted since then, they have not read whatever I was tweeting, and they've mostly been like, bring back Fable, which, like, on Instagram, we got rid of Gotham. Do you remember Gotham, the filter? This is like, and then for the rest of, like, the next eight years, all I heard was bring back Gotham, so it struck a nerve. Fable will come back before Gotham did.
2:48But yeah, it's clearly the folks that have gotten to use it and started incorporating it. It's actually really interesting. I've learned to not really trust day of or even week of model reactions. You don't really know until you've put it through its paces. And so I almost just completely block out the noise in the first couple of days of any new model release. Because, I don't know, everybody has maybe their toy example thing that they like to do with a new model. but it's hard to actually put it through its paces until you've actually had real work done with it. And I think people were just starting to do that, and then we had to sort of pull back Fable.
3:22But I remember in December when we put out Opus 4-6, it was like this interesting time where everybody went home for the holidays, and a lot of people had that week off between Christmas and New Year's. And then they came back and were like, oh, I spent a lot of time, and I really get why Opus is good, and I'm going to do it. So I don't think Fable has had that opportunity yet. But despite that, though, I mean, this was a pretty big reaction from the Trump administration. And I think everyone here, especially if you were listening to the Alex Stamos session, understands what's going on. But this happened within a few days of the model release.
3:54If I can give Wired some credit, Wired just reported last night that it was due to Anthropik's relationship with or having given access to the model to SK Telecom that could have raised flags within the administration. how surprised were you by how immediate that that backlash was yeah I think the the the sort of reaction decision was surprising and when we were sort of immediately engaged with them to you know restore access as well um and so at the same time you know the one thing that we we think a lot about internally is you know there used to be a poster on the Facebook wall when we were so that was like every day feels like a week and I think that's becoming true in AI and I think a good thing to remind ourselves of in general in the industry is we're dealing with unprecedented times.
4:38We're dealing with new situations and they can develop really quickly as well. And so I think also developing the capabilities and the connections and to make sure those conversations can happen quickly is really, really important. I have a new motto suggestion for you for the wall. Move fast and jailbreak things. I don't think they're going to use that, Lauren. OK. You know, we just heard from Alex that some of these capabilities have been available on, you know, previous models or, you know, you can you can find bugs with other with other models that are available today. So why do you think Anthropic got singled out on this front?
5:14I don't know why like Fable specifically was singled out as, you know, again, Fable being the non-cyber intention model as well. I think that the thing that does change over time is there's, you know, the capabilities, if you knew what to prompt and knew what you were looking for, and then there's the capabilities, I think, you know, I'm not sure I resonate with the, like, juicing high school athletes metaphor from Alex, but like, you know, the uplift that you get, like, uplift is a thing that we think about a lot when we think about model safety. So if you look at our model cards, for example, one of the ways that we look at risk in the bio domain is comparing the uplift from sort of a lay person using the model versus, you know, an expert or a lay person just using the internet and seeing what the comparison there is.
5:58And so that is one trend that has been progressing as the models get more capable. And so maybe less why Fable is singled out and maybe more like what the overall trajectory is. Interesting. So we just had R. Karazian, the lead economist on Ramp on the podcast. And one of the interesting things that he said was, you know, we had during the Pentagon situation, there were these headlines, okay, this company will not use Anthropic anymore. But actually the data from Ramp shows that spending actually increased to Anthropic models. It was apparently a good publicity moment for the company. So I think it sort of plays into a debate that we have like here on the show about like how much of this is like real concern for the issues and how much of it is, you know, from Anthropic is marketing.
6:42And like we have somebody from Anthropic here who can actually shed some light on that. We just said Stamos talking a little bit about it from his perspective on the security side, but we're lucky to have you here today. So so is it material or is it marketing or some combination? I mean, I think it is one of the hardest things to like really deeply believe that something is true and not marketing and then due to not even just anthropic. I think like, you know, I think people are generally right to be skeptical of any company saying anything and you should like put it through your own filters as well.
7:14But for me personally, I was like, no, but it's real. And like we like both deeply care about safety and are like To the extent that we are being vocal about anything, it is to either sort of help paint the picture of what very likely is coming or we believe is coming or what we've already seen and spotted. For example, in the Mythos case, really just looking at vulnerability scanning and bug finding and doing it in partnership with companies that were in that kind of project glasswing initial announcement. And so the technology is awesome. awesome, not in the, like in the just, you know, it is doing really incredible things.
7:50And therefore, even calling out what we see as what is happening, I think can seem hypey. I wish I could press a button and I could make everybody believe that we are not being hypey. I realize that's not the reality that we operate in, but at least from my perspective, we try to call it like it is. My understanding is that within labs in particular, and research and development at Anthropic, that you're using the best AI models to actually build new products, to prototype new products, and sort of test out your thesis. So is this ban essentially now limiting your ability to do that within labs?
8:25Yeah, I mean, definitely Fable is the best model I've ever used. And it's not to say that work has stopped, but it's definitely less good than the other models that are using models that are less good than that as well. And it is also, I mean, maybe the reverse of the distrusting the first week, you know, response is what happens when you don't have Fable. And like, obviously the Twitter reaction for the people that had already kind of like gotten into the model and were using it, right, was strong. But I'd say even in my personal use, I'm like, oh, like I'm on Opus 4.8 and it's good. Like I'm still productive, I'm doing work.
8:59But, and we can go into sort of like how my work changed with sort of these like Fable or the models of that sort of family, but it is noticeable for sure. Yeah, I think we'd like to know that. I mean, the public had access to Fable for like a half a minute. Groups in this like Project Glasswing have had access to Mythos. We don't really know what the difference is between using a model that you can use today and using one of these super anthropic models. So what actually could you do differently with a Fable or a Mythos? I think for me, it's the sort of scope and scale of delegation. And all these things are really imperfect.
9:40Like, you know, people say like, oh, is this now a level five software engineer or level six? But anybody who's used these models extensively knows that they're still spiky in capabilities, right? In some ways, they're, you know, in many ways, they're a better engineer than me. And in other ways, like, you know, I was complaining today that it had missed a descender, like the G, that bottom part is called the descender. I'm like, how did you put that in the UI? And it's clipped in And of course, like there's vision capabilities that need to improve and there's debugging capabilities. And there's sometimes even just sort of human common sense that is, you know, we're way better at than the models are.
10:13But overall, I think the big shift for me working, and it was really interesting because it sort of coincided with me going back into a builder role. So I really got to see going from, you know, using these models as an executive, you're trying to do the most of them, but, you know, not going to have it right. all your email and I think strategy still needs to come from you and then you can use the models to sort of pressure test it. But going back into a builder role and going from, okay, I am delegating chunks like, please fix this bug or I'm thinking of implementing this feature, let's go back and forth to something that ends up being much more sort of, all right, I got this bug report from one of our users or I have this notion of something that I want to build, like, can you sketch out two or three ways in which we could do it?
10:57right like that seems plausible often I find actually sometimes it'll uh give me the the sort of explanation or proposal and be like okay that actually is over my head like you are clearly way smarter than me like explain it to me like I'm maybe not five but at least you know uh uh not you and it'll sometimes you know uh explain it that way but then go build it and get it right getting it right like you know at a very very very high rate and I think that starts really changing how you operate like I moved much more to before going to bed making sure that I had like queued up for Fable, like enough chunky work to last, I would call it the whole night and I'll check in later and it got it done in an hour.
11:33And it was like, they're just hanging out for the next seven hours, but like really like delegating like a, a much more of a goal than just a, so one, one task, for example, like give us an example of one task you would hand to it. Uh, I mean, here's a kind of crazy one, which is, uh, for the programmers in the audience, like I, uh, I had written one of our labs, uh, projects in, um, in Python. That's like the language I You know, all Instagram was all Python. And for like some not super exciting reasons, we actually needed it to be in TypeScript to deploy it. And I was like, all right, that's going to be like, you know, in Instagram, we for years talked about moving from Python to PHP or hack or the Facebook language after the acquisition.
12:10And at least when I was there, never did. But I basically, we have a feature called dynamic workflows where you can have it like also break down the task into like a lot of subtasks. And I trusted it to sort of not just do the individual action, but here's like a whole language conversion of millions, hundreds of thousands of lines of code at that point. Go off and do it. Go plan it. Go execute it. Go verify the work. Double verify the work. And then I came back to the work being complete. So that level of like, this is a big sort of chunky task. So that was, so you're basically saying it was faster.
12:41Did it in an hour? You're guessing compared to, with Fable, compared to what it would have been before. I think the main difference is in the past, it would be like, great, I did it. and you'd be like, well, did you, you kind of took a shortcut here, or this is not quite right, or I need to go verify it, or like, oh, you cut this corner, or. It's like the managing interns thing that everyone's been saying for the past year, which is very offensive to interns, by the way, but yes. Yeah, exactly. I don't know, have you managed an intern? Yeah, that's true. And it was, and you're saying it was more correct.
13:09It was, so it's faster, it's more accurate, more reliable, and then according to the U.S. administration, dangerous. I think the other piece is it has like a greater, theory of mind is the wrong word, but sort of like theory of project so that it's less, you know, oh, I'm going to make this change and it'll say, great, I'll make this change. But really, especially if you've done like software engineering at scale, the best engineers kind of keep in mind all the disparate parts of how this thing's, and they also see around the corners, like I can make this change, but if I don't do it in this way, then the next change is going to be incrementally harder.
13:41And I think that's like been a significant difference I've seen in that kind of class of models. So I think when we talk about Anthropic Labs, right, people think of Claude Code because it is really your breakout product. And it sounds like you've been tasked with basically figuring out what the next Claude Code is. Would you say that's an accurate description of what you're doing at Labs? And also, why does Anthropic need Labs? Labs, yeah. It's also maybe worth thinking about why we needed Labs in 2024 when I arrived and why we need labs today because I think that the answer kind of shifts. I started the original labs team with Ben Mann, who's one of the co-founders of Anthropic, in my third week at Anthropic, and it had been something that I'd been bubbling under.
14:23And at the time, the reason was really different. It was all of our product engineering team was 25 people, and we didn't have the models, really. We had Claude, when I joined, was Sonnet 3, like you, and Opus 3. Those were, for the time, good models, but you weren't going to, they were not even interns, right? They weren't even IC3 engineers. So if you have a team of only 25 or 30 engineers, they are working on like the next incremental thing. And we were feeling like the models are starting to get better, but we don't have any products that sort of show that off. Like a good litmus test for me is when we get ready to release a model, do we have either a product or a demo or some other illustration of something that is very different?
15:02And it gets harder over time. Like with Fable, you know, even like illustrating like that weekend task or, you know, this longer amount of work. So really, labs at the time was, let's make sure we don't, like, our products don't fall behind the model exponential that's happening. And yeah, so Cloud Code came out of that initial one because nobody in the rest of the product orgs, people were thinking about coding, but nobody was sort of had the, like, space to go and think about, well, what if we totally change the form factor and we embrace the fact that the models were going to evolve in this way?
15:32And a lot of the two most useful thought exercises we do in labs, one is visualize the gap between what the models can do today and how most people use it. It can be closed that gap. That's one. And the other one is imagine what the models are bad at now that they're actually going to be really good at in six months. And let's make sure we have a product ready for that by then. I think those are the two guiding questions for labs. And then also out of that first incarnation came computer use. Computer use was different, though, because when we built it, it was really bad. Like we tried a bunch of products with it and this was around, you know, Sonic 3.5 and be like, Claude, can you help me, you know, clean up my desktop?
16:07And it would like click the thing, delete the file. Like this is not safe for release. We're definitely not going to go and build this or to ship this. But we had that product so that every new model that we'd release, we'd first check it internally and say, the computer used to get better. And we'd tell the research team how it had gotten better or worse until the moment where we said it's good enough. we're actually going to put a product out around this. It also gives you this sort of beacon into the future that you can measure your future products against. But then compare it to now. So we have a thriving product team.
16:38There's Cowork. There's, you know, Claude Code has grown a lot. We have our platform. And now I think it's actually much less about none of these product teams are doing this sort of thinking. And I think it's much more that the models are advancing really quickly. and even our capability to interact with them needs to evolve. So one of the things we collaborated with labs and CloudCode that we shipped today is CloudCode Artifacts. Having CloudCode not just be able to type back to you, but also sort of draw a picture or give you an illustration. And that partially came from spending a lot of time in labs saying, just a text box and a big text response is not going to cut it anymore.
17:14Like when I mentioned that the models feel like they're way smarter than me when they talk to me, sometimes I'm like, can you draw me a picture? Because this is what I actually need to fully understand this. But it's really what we've been thinking about is, you know, yes, we have a lot more products. You know, we actually have a lot of consolidation to do in our products. Like, that's another initiative that we have. But within that, we still have an opportunity to make things much more accessible to a person that does not spend all of their time thinking about prompting and the exponential and the difference between high, low, and medium effort.
17:42Like, there's a lot we can still do there. Right. But, Mike, so there's a – it puts people using anthropic models in an interesting place, right? Right. You know, Cursor, I think, just sold for 60 billion to SpaceX. And someone put this meme on Twitter that, like, you know, Cursor would have sold for 300 billion if it wasn't for this guy. And it's a picture of Boris Churnity, the person who created Claude Code. And so for companies that are going to build on top of Anthropik technology, you know, they're going to wonder, do I want to partner with Anthropik or is Anthropik going to go ahead and build the product that I'm going to want to build, potentially even after partnering with them.
18:20Yeah. I mean, we'll take the agentic coding side, and I think the broader sort of aspect of being both a platform and a product I think is really interesting. When we take on projects, the goal is often to sort of push that area of the industry forward. So there were AI coding editors, and some of them were really good, but nobody was quite thinking about it in as sort of freeform a way as we got to think about it with Cloud Code. And now a lot more products have that flavor than I think would have otherwise. And so I think if we're ever, you can call me out on this, Alex. If we're ever entering an industry where you're like, all you're doing is the same thing everybody else is doing, but you've got the Anthropic brand.
19:00I feel like that's a bad use of our time and a bad use of our either labs or product team time. If we're going in somewhere, it should hopefully be to say, all right, we think that the direction of travel is this way. We can build a product of that. And then, by the way, there's no world, nor should there be a world where all the products are Anthropic products. That would be a bad world, right? So that is hopefully either creating new space for companies or sort of showing the way where other products can incorporate. Yeah. It would almost be like working for a tech company that has social messaging, video, if all.
19:30Right, Mike? Yeah. Okay. You did leave. Well, there was some question, for example, when Anthropic launched a product that was seen as competitive to Figma and you had been on the Figma board prior to that, and I think you stepped down. I did, yeah. Is that correct? Yeah, and so it's a good question that Alex has brought up. I think where Silicon Valley is known for this really healthy, vibrant, risk tolerant startup ecosystem. And when the big start coming in with tons of venture capital and, you know, a lot of resources, people say, well, wait, are they just are they essentially just going to steal my idea?
20:01Yeah, no, I think I think our dual existence and it's something that other companies have to navigate. We talked about Amazon a lot in the in the previous panel, like they have to navigate this role where they are both the infrastructure provider. They obviously have a very large e-commerce thing. They do video, but they also serve video. And then, you know, by and large, customers can live in that dual world of like, okay, I'm using their infrastructure, also knowing that they are also using their infrastructure to do that. And I think the, you can talk to our customers and see how well we're doing at this.
20:30Like the thing I always try to do is like at least approach it with a lot of transparency. So the cursor example is an interesting one where like Michael and I talked a lot over the, you know, time around here's where things are heading. And, you know, similarly with the other products that we think about, like, can we, I think it's a couple of things. It's transparency and then it's shared building blocks. Like, yeah, I think in general, and I actually don't think there's any cases where this is even true. Like we're trying to build on top of the same capabilities that are available elsewhere.
20:57The last time I was here in the Commonwealth Club on this stage was our healthcare day at the beginning of the year. And we didn't ship like Claude healthcare, only we have it like nobody else has it. We shipped a bunch of like plugins and skills and MCPs and like complimentary abilities. So that's how, I'm not claiming it's easy or that it's a straightforward thing, but it is how we're trying to navigate what is like admittedly a complicated sort of situation. Speaking of startups, Anthropic is still technically a startup, but you're worth a lot of money. I mean, what's the latest valuation? Is it?
21:28965. 965 billion dollars or something like that. You sold Instagram for a billion, right? Startup. In 2010. Right. Financials have changed quite a bit since then. And yet, Anthropic has positioned itself. It is a PBC and it's positioned itself as sort of a more ethical company around building AI. And I'm wondering if you could talk a little bit about how you see that positioning and Anthropics role in particular, changing the culture of the valley. Like, I know I think back to how Google in the beginning of the 2000s really changed the culture of Silicon Valley in so many ways. And how do you see Anthropics culture now dictating this next era?
22:08Yeah, that's a really interesting question. Maybe I'll start like insight. And I think there's an external component to I think the reason I joined in the first place, so I was winding down my second startup, and knew I wanted to go work at a frontier lab, because I'd started to use these models for coding. And they were bad at coding, but I could see that they were as bad as they were ever going to be at coding, they were going to improve. And I had started building on top of these API's. So the startup I was doing was called Artifact. And we did sort of AI powered sort of news recommendations.
22:35And she read a lot of big technology via Artifact back in the day. was one of the things we added. But not wired. You know, you guys had a really hard paywall, to be honest. Fair enough. We didn't do great on paywalls. Do you need a discount on my subscription because I can get one for you? Okay. It's actually really funny. Making deals. Yeah, making deals. It's the login cookies. It's really hard to keep people. I know, I know. Please escalate this to Condé now. I know. And email login's very hard to do in an app. But I was building on top of the APIs and be like, wow, okay, they're able to do really interesting things.
Read the full transcript
23:11But what ultimately made me go to Anthropic was they walk the walk and they really deeply believe in trying to make AI go well for humanity. And that's in the water internally and I think has been why I think the company has remained as cohesive as it has even as we've grown. And I think that it's a testament also to the co-founders there on how often they are talking about this as well. It was a surprise for me coming from a world where at Instagram we did a weekly all hands and we talked about product 95 % of the time. And maybe 5 % of the time we talked about something else that was going on in the world or around the company.
23:46Probably maybe underselling our like go to market. Maybe it was like 80, 20, but it was definitely a very, very heavy product. And I remember Anthropic about six months in myself and Kate Jensen, who's one of the leaders in the sales organization did a joint all hands where we talked about our like, you know, how we're doing product and go to market together. And people were like, this is so great. I finally understand our product strategy and like what we've been doing, it's like, oh, right, this is not, quote unquote, a product company. It is a very mission driven AI company with a very strong sense of why it exists in the world.
24:17I think in terms of the overall impact on the Valley, it remains to be seen. I think the positive signs that I've seen are interesting signs of things that I've seen are a renewed interest in philanthropy across the board. And I think that's something that has been written about. And I think it will be an interesting sort of outflow again who knows how all of this goes but like uh depending on how it goes it could mean a lot of interesting new sort of philanthropic um uh deployment and then i think the other piece is you know uh the conversation around how ai could or should go is one that is happening in real time with the technology versus retrospectively which i think has been the case for other technology waves and i think that is a good thing hi everyone alex cantowitz here i want tell you about a documentary I've made with Gravity to explore the future of AI agent security.
25:06To find out if we're truly ready for autonomous agents, I sat down with MIT professor Ramesh Raskar, former White House CIO Theresa Payton, Michelin's Group Chief Data and AI officer Ambika Rajagopal, and Sharon Guy, a former executive at Alibaba. They each offer unique insights into this evolving landscape. We conclude with Rory Blundell, CEO of Gravity, to discuss the path forward. With Gravity leading the way, join us on this journey. You can watch the full documentary at the link in the show notes.
25:50When you need to build up your team to handle the growing chaos at work, Use Indeed Sponsored Jobs. It gives your job post the boost it needs to be seen and helps reach people with the right skills, certifications, and more. Spend less time searching and more time actually interviewing candidates who check all your boxes. Listeners of this show will get a$75 sponsored job credit at Indeed.com slash podcast. That's Indeed.com slash podcast. Terms and conditions apply. Need a hiring hero? This is a job for Indeed Sponsored Jobs. No one goes to Hank's for spreadsheets. They go for a darn good pizza.
26:24Lately, though, the shop's been quiet, so Hank decides to bring back the$1 slice. He asks Copilot in Microsoft Excel to look at his sales and costs and help him see if he can afford it. Copilot shows Hank where the money's going and which little extras make the dollar slice work. Now Hank's has a line out the door. Hank makes the pizza. Copilot handles the spreadsheets. Learn more at m365copilot.com slash work. Mike, you talked a little bit about Anthropic has this gap that it sees between the capabilities of the models and where everybody is building products. And with labs, what you try to do is get ahead of that so you can show people what AI might be able to do now and six months from now.
27:06So please tell us. Please tell us what you're building, where you see the potential, and what people should be on the lookout for. Yeah, throw a roadmap. Oh, and if I can throw an and. And what you're building now, but also if you have a pie in the sky, like Elon Musk data centers in space type ambition, I want to hear about that too. Tell us everything. We have 13 minutes left. The rest of the minutes is the monologue in my product. I think maybe two themes I'm really excited about that we've been exploring a lot. The first one is giving Claude an environment where it has more agency and it also has more self-knowledge.
27:46and I'm going to unpack that because that's like a lot of AIE words, but I'll give you an example of where we are currently doing a bad job of this. Like if you are in a Cloud project and you make a file with Cloud, you're like, that's great. Can you add it to our project? Cloud will be, no, you have to go download the file and go drag and drop into this thing. And you're like, what? Until yesterday, I would have said the same thing about Cloud design and Cloud code, where if you're in Cloud code, you're like, cool. Like I need a design for this thing that we're building. Or you're in Cloud design and you make a mock-up and you want to go build it, be like, cool, here's a zip file.
28:18And you're like, what? And I think, so that's a little bit of interoperability. But in general, this theme of giving, if you give Claude a lot of notion of its environment, I was talking to actually a customer, like an API customer, and one of the things that they were experimenting with is actually even giving Claude like a secure version of their source code while it's running in the agent loop in their product so that if it hits an issue, it doesn't go like, I don't know, I hit an issue. It can be like, well, it's probably this thing, at least when it's talking to one of the sort of maintainers of the software.
28:46And so that overall theme, and of course you have to do it with safeguards and be really careful about what you unlock with it. It sounds kind of obvious, but it's actually night and day in terms of how expressive these products end up being able to be, right? And you can even see it going from maybe like core chat or classic chat in Cloud AI and something like co-work where it's got a little bit more agency and it has a runtime and it's able to sort of understand a little bit about its environment. But I think we are at like 10 % of the journey about where we could go. Actually, one of the reasons I think people got excited about things like OpenClaw is seeing how a harness that is modifiable and you can talk to it about things and you don't ever get the sense of like, oh, sorry, I can't do that.
29:26You're going to have to go to this setting screen and turn it on. It's just a thing it has access to and hopefully with the right gardening and permission. So that's theme one that I'm extremely excited about. And I think if we do it right, it should actually transform all of our products from head to toe. um the other piece um is uh and i'll maybe like share like the the the feel if not the internal product we're working on but like the the phrase i got as feedback was uh like i think closing the gap i mean i talked about closing the abdian capabilities in reality i think it's also closing the gap between how people understand their own work and then how the actual day-to-day is to do that work um i was talking to somebody uh internally who's on our privacy team and to move a ticket it from one queue through another one via the task tracker into another.
30:14It was eight different steps of copying and pasting, of manually moving, pretty annoying to have to do, probably error-prone. You have to keep spot-checking it. We helped her with one of our labs projects to basically make that not a pain. She's like, ah, this is the first time in my career, and she's been working for 30 years, what's in my head and what I I'm using is now this. It is now closed. I want to bring that feeling to everybody. Of course, Cloud unlocked a lot of non-technical people to be able to code, but we're still asking people to understand way too many concepts of what is the difference between my sandbox environment and production, or connected MCP as myself or others, or how should I store data?
30:58Of course, you can't abstract everything, but if you combine both of those themes, if you give Cloud a lot of self-knowledge and you're creating an environment where it can actually solve complex problems for people in repeatable ways, I think I get very, very excited about those things. And your moonshot. Not letting you off the hook. What's your moonshot? Moonshot? Yeah. Nothing in space, although I guess we're talking to SpaceX about spacey things. Wait, you're talking to SpaceX? Spacey things. I mean, those are the announcement for compute. It was exploring extraorbital. What was the phrase?
31:30Something about exploring post-orbital world things. Definitely not my department. But yeah, there's stuff in Moonshot. So the labs isn't working specifically with the team on compute? Right, exactly. Or ships? There's a separate compute team. Totally separate. Okay, so you're Moonshot. Do you personally believe in data centers in space? I had a conversation. I'm by far from a data center expert, but I talked to somebody who is a person who sends things to space, who is not Elon Musk. That's what you would say. Yeah.
32:03and they were really bullish and I was like trying to talk about why and it was basically like I guess effectively infinite power if you convert it well and you know infinite land and I was like okay you can buy that I mean and I think they feel good about the shielding you have to do again clearly not my area of expertise but after that talk I was like okay I see I see it you know even if it's going to take a few years at first I admittedly thought it was a crazy idea in general But now I'm like, oh, I actually really can understand why this might make sense. When you were talking earlier about the ways that the work in Claude is going to get compressed and all those steps, I couldn't help but think of tokens and how, you know, maybe it's good for your business model in the short term if people have to take so many steps and use so many tokens.
32:44But tokens have become this unit of economics that we're using to describe the industry now. And people are token maxing and now they're tokenmizing. And one, I want to see, I want to hear where you sit on that spectrum if you're a token maxer. And two, is there a near future in which the industry is not actually measured by tokens? You know, it goes the way of MIPS or dial-up or some other, you know, there's some other unit of measurement that actually defines the economics of this era. Yeah, I think both of those are really interesting questions. It was interesting earlier this year when you started hearing about, like, companies that have, like, dashboards showing, like, who used it the most.
33:22And we, of course, have internal metrics as well. And we found that there's not a lot of correlation between the person who's using the most tokens and the person that I... It's an interesting thought I should have made to do at your company. It's like write down your 10 most productive people that you think are most productive. And then get your top 10 token users and see how closely they correlate. At least for us, it wasn't that correlated. It seemed dangerous to sort of purely glorify the maximum usage. Obviously, it's very gameable. But even beyond that, I think it's, you know, yes, you can ask Claude to do 10 different variants on something.
33:52But if you thought about it deeply, maybe you would do choose two that you thought were most promising. And the third one, if you then had like some iteration on that as well. So I would not say like a token match. Actually, the tokeniest thing was that conversion thing I did, which is like a couple million tokens. There's like a lot of tokens that it took to convert the thing from Python to TypeScript. But I think people are being more thoughtful about these different pieces. And one of the things we look at whenever we look at a model launch is not just model intelligence, but we're also really thinking about model intelligence and effort and token efficiency as that combination.
34:25And I think that's a big lever we have to improve is how do we continue to be more and more token efficient for a given task so that you can also, hopefully you don't have to think very hard about this. We can do this automatically, but like we're able to tune the solution to the problem a little bit more. And then to your second question, yeah, I, you know, when I was still in the CPO seat, I was thinking a lot about sort of outcome based pricing. is something that would be really interesting to do if you could do it. And of course, if you talk to the Cierras and Finns of the world that have a really clear, like, we kept this, we were able to solve this customer request and not have it go escalated.
34:58That's really clear. It gets so much fuzzier on these tasks that we actually asked Claude these days. I had a strategy document. I used Claude to critique my strategy document. What was the outcome? It's like, well, I don't know. It's like, tell me how the strategy goes six months from now. It feels like it's going to be very hard to capture that as well. But I would like to see some more experimentation around, can you better capture what it's worth to the individual or the company? And then can we find the best way to do that as well? And I guess the most concrete thing we've moved towards that, and we have a product called Cloud Vantage Agents, where we'll run all of the infrastructure for you in terms of doing all of the agentic harness and calling the tools, et cetera.
35:37And you can either do it in sort of the normal mode, which is you give it tasks. We'll go through tokens. It'll tell you when it's done. or we have an outcome-based mode where you can say, here's what good looks like. Here's a rubric. Go and do it, and it'll go off and make it more outcome. So if everybody had moved on to that API, then I think maybe we could have a different sort of outcome-based pricing, but we'll see how that gets adopted. John, or the guys in the back, do we have the random image? Can we show the random image? If we can, great. I'm excited. The random image. Oh, here it is.
36:06Okay. It's just because we didn't have a good label for it, so we just called it the random image because it might come up at any point. But this is a chart from the Financial Times, speaking of utility, where it shows the amount of app releases that have come out, which are skyrocketing, and then apps with significant usage that seem to be going down and app reviews, which seem to be going down. So, Mike, I'd love to hear you respond to what we're seeing in the image here. Is it possible that, like, everybody's coding and releasing, but we're not really seeing a big boom in productivity? That's really, I mean, I think there's definitely a parallel in app usage in general.
36:40It'd be interesting to see if any of those app release became one of the apps with significant usage. All right, we can take it. We can take it down. Yeah, go ahead. I think it ties into something that I've been thinking a lot about, which obviously my background is in consumer. And I've been wondering what the consumer AI breakouts will end up being. And I don't know if we've seen a lot of them yet. And I think part of it is, you know, I don't know how far back that chart goes. But when we were releasing Instagram, it still felt a little bit Wild West in terms of the apps. Like people were excited about apps and like two kind of random people released an app.
37:13We were able to get to like number one in photos and video within three months. Right. I think that is much harder now when you think about how consolidated the top 10 is and how much time spent is spent on like the TikToks and reels of the world. It's a lot. Right. And so I think getting that breakthrough consumer experience, I think, is really, really hard. So I think that is as much a story about how sort of consolidated consumer products are these days, number one. Number two, how entrenched or how powerful it is to have that sort of data, data gravity, like the data gravity of something like your Google Docs are in your Google Docs.
37:51So even if somebody has a like 2x better AI powered, you know, doc editor, you're going to move all your stuff up? Maybe, probably not. So I think that it speaks to, you know, the things that are sticky. I think about it a lot. It's like the hard stuff is still hard. Like making something people want still really hard. You know, we have amazing bottles internally on top of it. Not all of our products work, right? And so I think that's a bullish sign for like product people like me because it means that I think we hopefully still add value. But I think that chart is maybe another place. It's just like it's harder in many ways than ever to break through, even if you can code more quickly.
38:25And what could we have done Instagram in, you know, a month instead of three or four, probably. But we got there after like a long winding turns and twists and turns process. You had, I think, 18 people at Instagram when you sold it. 13. With these tools, do you think you would have, how many people do you think you would have had? It's really interesting because of those 13. Just give us a number. Yeah. Like everyone's like, oh,$1 billion, one person startup. How close could you guys have gotten? I think we could have gotten there with like four to six you know okay yeah um or the thing that we would have done if we'd grown it we'd be able to do some things in more than a single track like Instagram was if you ever watched my five-year-old play soccer now by which I mean like the ball is there and every single person runs the ball like that was our product team it was like video go and everybody like goes and works on the one thing and like we'd be able to like play positions like Android we built in about a month for Instagram.
39:20We could have done it probably in a week with the models. And to build Android, we took everybody off iOS, and we all relearned to code Android OS, and then we went off and do that. And for that whole month, we were barely shipping updates on iOS. So I think you can be a lot more... Actually, a really good example. There's a labs project I have internally that helps accelerate how anthropic engineers code and do code review. And that project, I am maintaining an iOS and an Android version of. And I basically have the cloud that works on the iOS one basically ping the Android one and be like, hey, I implemented this.
39:55Sorry, Android users, it's still the second one, even in the NLL world. Sorry. And then the Android version is like, okay, I'm going to do this. Oh, that doesn't count because that feature doesn't make sense here. I'm going to drop it. And, of course, we wouldn't have been able to delegate all of that on Instagram. But we sure could have done a lot by having sort of platform parity. Like this dream of platform like close to parody is now actually quite doable. You're probably going to get calls now from the remaining six or seven people on your Instagram team going, was I, did I make the cut in the new era?
40:24Also, it sounds like you probably could bring Gotham back now if you really wanted to. I forget if we eventually, I mean, I think for April Fool's, maybe we brought it back one day. Yeah. Do we have time for one more question? Yeah. Last question. Okay, I mean, my last question for you is, you worked on a product that now as it has evolved, is in many ways ethically fraught because of some of the harms that people are concerned about with children. And when you talk about the fact that there hasn't really been a big breakout consumer app for AI, I think there has, and it's chatbots, right? And chatbots have also led to some real dangers and harms for young people.
40:57And so when you are building in labs, how are you thinking about the, you know, the potential harms and the risks that come with just making this technology that much better? Yeah. I mean, I think there are certainly products that we have either prototyped or conceptualized and been like this product. It sounds so hypey. I hate this. But like this product, if shipped, would be bad for the world or like would nudge people in the wrong direction. Or even if we did it right, the like wrong or like more morally fraught version of this would be actively, we think, bad. And so I think asking that question a lot internally makes a difference.
41:33And having it's a luxury to have core products and models that are doing really, really well. So we don't like that's in some ways an easy decision if we think it could get a lot of a lot of use. But yeah, I think going back to an earlier conversation, I think front-loading it is really valuable and really thinking through like It is now more normalized to have people at a company and definitely on topic does for like economists thinking about the impact of the thing That you're building on the world and that just was not the case and along the years on most of social media Mike, it's always great to be can we thank you again for bringing your insight today Let's hear from Mike and Lauren Thank you great job Thank you.
42:16This episode is brought to you by Google Chrome. You think you know a browser, but Gemini and Chrome, that's new. It can help you with practically anything on the web, like restoring a vintage motorcycle from a 50-page restoration block, or finally break down that long article you've had open for weeks. Gemini and Chrome is here for it. Ready to make anything online make sense? There's no place like Chrome. Check responses set up required, compatibility and availability varies 18+.
From the publisher
Mike Krieger is the head of Anthropic Labs and co-founder of Instagram. Krieger joins Big Technology Podcast live from the Big Technology AI Summit to discuss what it's like inside Anthropic the week the government forced the company to pull its frontier models, Fable and Mythos, off the market. Tune in to hear Krieger describe how working with Fable changed the way he builds — queuing up a full night of work before bed and waking to find it finished in an hour — why he insists Anthropic's safety warnings are material rather than marketing, and how Anthropic navigates being both a platform and a product as it competes with the companies building on top of it. Wired senior correspondent Lauren Goode joins as a co-interviewer. Hit play for a rare look inside the lab from the person building Anthropic's next breakout product.---
AI Agent documentary: https://www.gravitee.io/ai-agent-documentary
Enjoying Big Technology Podcast? Please rate us five stars ⭐⭐⭐⭐⭐ in your podcast app of choice.
Want a discount for Big Technology on Substack + Discord? Here’s 25% off for the first year: https://www.bigtechnology.com/subscribe?coupon=0843016b
Learn more about your ad choices. Visit megaphone.fm/adchoices


