In short
Podcast Notes: Azeem Azhar's Exponential View - Episode with Aaron Levie
Episode Overview Title: Inside Box’s AI Playbook with Founder & CEO Aaron Levie Description: Aaron Levie, CEO & co-founder of Box, discusses how an “AI-first” mindset is transforming Box and offers insights on building smarter, faster organizations.
Timestamps
- (00:00) Episode trailer
- (02:04) The "lump of labor fallacy" in sci-fi books
- (07:37) Individual productivity gains vs. team productivity
- (12:32) Box's Friday AI demos
- (21:23) Redefining management science with AI agents
- (26:37) Lessons from Ford's AI innovation
- (29:52) Leaders like Pichai and Nadella coding again?
- (35:16) Pricing models in a post-AI world
- (38:43) Impacts of cheaper AI tokens on usage
- (43:02) Addressing AI's verifiability challenge
- (48:24) Levie's personal use of AI
Key Discussions
- The Evolving Nature of Work
- Acceleration of Productivity: Levie emphasizes that AI acts as an accelerant in the workplace, leading to significantly faster completion of tasks.
- Roles Transformation: While few roles will disappear, the nature of activities within those roles will change dramatically.
- AI's Impact on Management
- Management Redefined: With AI agents taking on more responsibilities, the role of individual contributors (IC) may evolve into managing multiple agents.
- Implications for HR and IT: The need for a combined approach from HR and IT departments to navigate these changes effectively.
- AI-First Company Transformation
- Box's AI Priorities: Box is transitioning to an AI-first mindset, leveraging AI to improve operational efficiency.
- Friday AI Demos: Regular internal demonstrations encourage staff to explore and adopt AI tools.
- Challenges of Team Productivity
- Individual vs. Collective Gains: Levie notes a discrepancy where individual productivity improves, but collective team productivity does not correlate equally.
- Bottlenecks in Processes: As productivity rises within individual teams, overall efficiency may still lag due to existing bottlenecks in workflows.
- Pricing Models in AI
- Adapting to AI Outcomes: There's a shift from traditional per-seat pricing to models based on actual outcomes delivered by AI (e.g., by contract or negotiation).
- Unbounded Revenue Potential: AI agents provide opportunities for significant revenue growth through automated tasks.
- Verifiability and Trust in AI
- The Problem of Reliability: As AI models evolve, ensuring reliable outputs becomes essential, especially in critical applications.
- Two-pronged Solution: Improvements in model capabilities and vendor accountability will be key to addressing verifiability issues.
- Personal Adoption of AI
- Everyday Use: Levie shares how AI has become integral to his personal life, providing immediate information and enhancing learning opportunities with his children.
Key Takeaways
- AI as an Accelerant: Embrace AI not just as a tool but as an enabler to fundamentally change how organizations function and deliver.
- Management Evolution: The emergence of AI agents may redefine traditional management structures, requiring new skills and approaches from leaders.
- Continuous Learning: Organizations should promote learning and adaptation to AI tools to maximize productivity and innovation.
- Pricing Innovations: Companies need to rethink their pricing strategies to reflect the transformative potential of AI-driven outputs rather than traditional seat models.
Conclusion This episode provides deep insights into how AI is reshaping the corporate landscape, emphasizing the importance of an AI-first approach in enhancing productivity and redefining management roles. Aaron Levie's experience at Box serves as a practical case study in navigating these changes effectively.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00We're going to look at the last year and the next, let's say, two to three years as probably the most defining periods of workplace technology change that we'll ever live through. What would happen if many parts of your business, you could start to do things in an hour or two hours that would have otherwise taken a week? Which part of your business do you expect to disappear or become unrecognizable in the next 12 months? I think actually very few roles will sort of disappear kind of outright. I think it's what we will be doing in those roles that will fundamentally shift. We should have three, five, 10x more leverage as individuals inside of an organization.
0:34Everything that we do will be accelerated. How should we think about a world of resourcing when agents start to appear and become more reliable? In that world, the IC is now a manager and there's sort of no upper limit of how many agents they could be deploying. Then the question is, what is that manager of human managers of agents? What do they do? Should they also be managing agents? Open question. I mean, this is in the category of things that will change in a five-year period that might redefine 100 years of management science. I mean, it's super complicated because you have model releases every two weeks right now.
1:09The velocity is extremely fast. It's this incredibly exciting, stressful, crazy time. But for those that enjoy learning and changing and evolving how they work, this is a very exciting time to be diving into that.
1:29Today, I'm really delighted to have Aaron Levy, the CEO and founder of Box as my guest. Box is a company, a product that I first came across, I reckon, 20 years ago. Before, it had the.com domain. It was one of the first cloud storage companies in the world. Of course, now it does much more for enterprises. And Aaron has been leading a transformation of Box into an AI-first company, both for the company itself and for its customers. And Aaron, you've been really public about the benefits that AI will bring to all of us. If Fox is fast becoming an AI-first company, which part of your business do you expect to disappear or become unrecognizable in the next 12 months?
2:15I think what's going to happen, the part that will be unrecognizable, I want to first caveat that I'm probably going to be wrong on 90 % of anything that is sort of 12 months out. You know, anything in AI that you can't like see and feel right now is just so hard to predict. And it's actually, I think, relatively apropos of the name of your core content right now. It is this idea that everything that we do will be accelerated. And so I somewhat try and readjust when we get asked like, okay, what roles won't exist, for instance. I think actually very few roles will sort of disappear kind of outright.
2:48Now, there are some that we can get through in a couple categories. But for the most part, I think it's what we will be doing in those roles that will fundamentally shift. And then ultimately, the amount of output that we get from many of these roles will be the biggest difference. And so I think there's like two ways to think about AI. I'm sure there's more, but like one is when you look at work, if you kind of imagine like a mass of work, and then you kind of think about AI taking part of that mass, That's sort of like, I think, our default instinct because of, you know, kind of what literature says and sci-fi.
3:19Sci-fi doesn't sort of anticipate like dynamic economies. It usually says, OK, we're working in the factory, robots took the jobs, and now nobody's working. That's kind of the... It's the lump of labor fallacy, as the economists call it, right? That's exactly right. So like every sci-fi book ever written has a lump of labor fallacy, you know, built into it. I think there's a different way to think about AI, which is imagine work is actually on a timeline. if you kind of zoomed out over a 200-year period, let's say back 100 years, forward 100 years, and you just sort of look at progress in society, all we're really doing in work is sort of moving through some progress in society, right?
3:55We're trying to like cure cancer faster. We're trying to, you know, give people better transportation so they can get places. We're trying to entertain people in new ways. My view is AI is an accelerant to the timeline as opposed to a reduction of the kind of mass of work. And so if you start to think about it more as an acceleration on the timeline, then it gives you a little bit of some instruction of then what to go focus on with AI, which is do use AI to move faster as an organization. If as a byproduct of that, you save money, fantastic. Or if you can now do things that are now cheaper because you're moving faster, also fantastic.
4:32But use AI to actually move faster as an organization. So if you roll out a year from now, but then I think even probably, you know, it's almost easier to think about five years from now than one year. But at some kind of point in time in the not crazy amount of future, I would just argue that we should have, you know, three, five, 10x more leverage as individuals inside of an organization, which means that when we go and brainstorm a new marketing campaign to go do, instead of us having to have three brainstorms over a two week period and go in and, you know, talk to lots of other companies for best practices and benchmarks, that might be a 30-minute exercise, you know, just involving your marketing people and AI kind of, you know, coordinating on that.
5:14And then boom, you just, you know, shrank that process by a couple weeks into 10, 20, 30 minutes. And so imagine that across engineering, product development, marketing, you know, going and coming up with the sales plan for a customer. What would happen if many parts of your business, you could start to do things in an hour or two hours that would have otherwise, you know, taken a week. And what is that? Yeah. Sorry, pardon me. But there's this idea that embedded in that is that we all have big backlogs in our Kanban boards. And, you know, can you move through that backlog much, much more quickly, which, of course, talks about one of the behaviors that's important.
5:52Can you cultivate useful things to do further down the backlog? I think we'll get into that idea as well. So there is this notion of... Well, and to that on the lumber labor fallacy, like the campaign board might go off on a completely new direction as a result of this. I like as an example, I have personally and this is like to the chagrin of my my colleagues, I have lit up more projects internally because of my access to AI. Because what I do is I go experiment on some idea, then I then get convinced is much easier to go do now because AI kind of let me start to do the experiment on it. And then now all of a sudden, I've lit up a project that now people are working on.
6:34And that wasn't in their Kanban board a week ago. I think it's something you may have said on X. I think you said it makes it much easier to start projects. And I remember reading that X post thinking, yes, and who gets to finish them, Aaron? Right, no, 100%. I literally, so we have this project that we're doing right now on the tail end where basically I shaved about three days to a week off of the initial brainstorm process that would have taken. And I did that in a couple of minutes with AI. Then I rallied the team around it because the AI did the upfront part and it sort of showed what it might look like as a result.
7:09We still had another three weeks of work to go make it happen. I shaved off a week at the front end, but by shaving off the front end, I lowered the barrier to us sort of realizing what would be possible. But then I added three weeks of work to everybody else that wasn't on their plate. So AI is going to do that a lot more than we realize because what's going to happen is it's going to expand the aperture of now all the things we can go work on. I completely agree. And, you know, there's this idea of Armdahl's law, which is one of the famous computer science engineering laws, which says that, you know, if you have a process that's got a number of steps and you accelerate one part of the process, if that's only 10 % of the duration of the project, you've only shaved 10 % off the entire time.
7:53And so we have to turn that into reality. And I think there's this notion of what is the propagation curve of AI within an enterprise. So it feels like you've described individual productivity, your individual productivity, maybe it applies to my individual productivity as bosses who generate work for other people, accelerating. One of the things I found quite curious in talking to enterprises and CIOs around the world is that many of them report lots of individual improvements, right? Individual developers writing more code, whatever the metric happens to be, but they're not seeing the sum of the parts being necessarily greater than the individual contribution.
8:36There seems to be some mechanism that breaks down from individuals being able to do really well using AI and a team, a function or an enterprise being able to return the same result. Does that ring true in what you've seen within Box or your customers? I think embedded in there is an interesting kind of point about we will see stats of, okay, individual productivity goes up 10, 20, 30 % with AI. But like clearly at this stage in 2025, organization-wide productivity is not 25 or 30 % higher or already the economy would already be completely impacted. And so there's some kind of dilution that happens through each kind of turn of the crank of incorporating the AI productivity into the larger system.
9:18And similarly, there's, I don't know if you've ever read the book, The Goal, but it's all about basically process improvement. And the book basically goes through exactly what you just said, which is it sort of goes step by step and finds each of the individual bottlenecks in a particular kind of system. And until you sort of figure out where is the bottleneck as the next sort of gating factor in productivity, you don't get any benefit at the other end of the amount of productivity you get. So I can light up a project incredibly quickly. But if every other corresponding part of that process is not also benefiting from AI, then all I've done is in many cases actually added more work, which then won't inherently even show up in productivity metrics because I've actually now added even more tasks to the organization.
10:04And we haven't yet gotten efficient enough to then sort of see the productivity gains across the organization for that. So I think this is probably very much just a natural evolution of, I wasn't around for it, but I'm sure that like there was a multi, you know, probably decade plus evolution of like the first people with email in a corporation were like super productive, but then they were ultimately bottlenecked by everybody else that didn't have email. That was a probably a decade long process for everybody to finally get wired up where you could efficiently communicate with everybody to the point where that showed up as a productivity gain.
10:37And then reasonably quickly after that, we drowned in full inboxes because everyone had access to email. So you went through this sort of funny cycle. But you've raised something that I'm curious about. What is the balance between you as the CEO being good with these AI tools and coming up with new projects and needing your frontline workers to be really, really great at them? So sometimes we think of when it rolled out, the internet was very much in many companies sort of frontline driven and slowed by the center. If you look at things like ERP, that was driven from the top because frankly, who wants to live their life in an SAP interface and sort of forced down in the center?
11:17And you've got these two different models. To what extent do you need to enable the absolute frontline workers? And what's that balance between you as a boss leading them from the front? Yeah, so I benefit from just, I'm already pre-wired to just like love technology and always, you know, want to explore new things. That's why we created Box. And so I have a little bit of a natural advantage because like this is just this is like what I do for fun. These are like my hobbies usually are just like technology. Like it's not like I don't do I don't do fly fishing. So I already love technology. I love to play with technology.
11:49Most of my social friends are doing technology or AI. So I get it in every sort of part of my life. And so then the question is, how do you make sure that you're enabling all the folks that don't have that as sort of the same built in default way of way of operating? We spent a lot of time figuring out how do we enable kind of the front lines or the broader population of colleagues to become better and better at these tools. And it's interesting because this is a space that is moving so fast that, you know, most people are, you know, probably fatigued actually by the amount that's coming at them to try and figure out.
12:23And in particular, if you tried something with AI a year ago or six months ago and it didn't work, you might have written off that, well, okay, AI can't solve that particular thing where I want to review a contract or go and summarize 500 documents into a report. You might have been like, okay, AI doesn't do that. And actually, check back in today and it does. So the need to kind of stay current on everything that's happening is just unbelievable. So we're doing a lot of things like some of our practices are every Friday, we have an internal all hands, and we'll have somebody demonstrate how they're using AI.
12:59We're trying to make it more of a social collaborative exercise where you're demoing in front of your colleagues, how you're using AI to get some particular type of work done. We're working on more internal enablement, which would include kind of education. We're pushing everybody to read sub stacks and subscribe to newsletters and listen to podcasts so they can kind of stay as current as possible. I think about this as if you're working in technology, you know, in particular, but then probably the broader economy, certainly, you know, we're going to look at the last year in the next, let's say, two to three years as probably the most defining periods of workplace technology change that we'll ever live through.
13:36And so it's this, you know, incredibly exciting, stressful, crazy time. But for those that enjoy learning and changing and evolving how they work, this is a very exciting time to be diving into that. So there's a lot in that, a couple of different questions. The first one I wanted to pick up is give us a flavor of those Friday afternoon learnings. I mean, we do something similar with my team. In fact, starting in December 2022, we started to do a daily all hands, which is only a team of five, right? There's a smaller group, 15 minutes. What did you do the day before? And back in those days, of course, to show how random it was, it would be a bit of mid-journey.
14:17Well, I did this image in MidJourney. And now it's getting more and more sophisticated. You know, people are showing, you know, workflows that might involve eight or 10 agentic steps. They might involve decision-making by the agents and things that break rather randomly. Like, you don't know. It worked really well with the 2.0 version of Gemini Flash. And in the 2.5 version, it stopped working. And, you know, of course, but we're a small organization. We know each other very, very well. six people. Box is much, much larger. So what do those actually look like? And how do you get contributions?
14:54And is it still, oh, look, I've got a great prompt in Claude. Let me show you. Or is it we've built something in crew.ai that involves 500 steps? Yeah, the focus is we want to push the limits of our platform. So it's a narrowed scope, because it's just Box.ai functionality. All right, you've got agents and other things like that built into Box. Exactly. And so it's all going to be related to automating a document process, generating content, reviewing content, summarizing or adding intelligence on top of data as the benefit of the surface area is somewhat narrowed. But what happens is we'll have a sales rep come and it's all in Zoom.
15:35So come present for, you know, 500 plus people that tune in live and then thousands that will watch it. They'll present their workflow for, hey, here's how I take all of my transcripts from my customer conversations and I turn it into a summary. And then I sort of tune it for that customer's industry and I send off an email. or I take all of our sales material and I produce some best practice guide for my sales team or I'm in engineering and I want to take a chunk of code and have the AI look through it to improve it or do documentation. So these are the kind of use cases that are emerging organically and then we just have those basically boxers go and show that to everybody in the organization.
16:18And then for us, again, it's like very selfish because both we want to work in this new way so we can be helpful for our customers. But we also want to see then the upper limits of our own technology so then we know what to go improve on the product. If we can shorten the product feedback cycle that normally happens, which is you release something, you listen to customers, it takes a couple of months. If we can internally just learn those things very quickly, we can recycle that back into the product roadmap and then improve our own internal development velocity. So we're using it selfishly for our product teams to then also see, oh, wait a second, We ran into this issue where people want to share prompts, but we don't have a useful mechanism for sharing those prompts.
16:59Let's go build a prompt library, like those kind of things. Yeah, I could do with a prompt library. I mean, it sounds great. It sounds very, very inclusive. It sounds like there's a lot of activity, but activity doesn't necessarily beget results. So let's take that upper level to your board. They need clarity. They don't have a ton of time. So what is the single metric that cuts through all of this activity, all of this height, to really reveal benefits for Box from these activities that are showing the productivity? I think as an industry, we're very early in what is a normalized sort of metric to quantify AI-driven productivity.
17:37I think there's a range of approaches that you can take. We've actually explicitly said that right now we're still in the journey to figure out what metric we're going to land on. There's an easy one, which is sort of like, how much money did I save directly? Then there's another one, which is how much money in productivity or hours or system time did I have cost avoidance on? So that's like, if I had done this with human labor, it would have cost this much. And then a third easy one is just like, what are the hours, tokens, some metric of just sheer output coming from the AI that now you're utilizing?
18:13Right now, we're actually way more focused, and I'm very comfortable being more in the, we actually want like a thousand flowers to bloom mode. I don't know if I would feel comfortable if I were like at JP Morgan with that strategy, but being on the front lines of a software company that is fundamentally becoming an AI first company, I don't mind a little bit of a period of, let's just, let's let everybody go and test and push these systems. And I think the best scenario for us right now is this idea of reveal preference, right? Like, let's just like see what people actually do. And almost by definition, the market, you know, we set very high goals for anybody working in the company.
18:51And so it's quite easy just to then let the market decide how to use AI to accomplish those goals. And so if we're actually setting our goals properly, and if the managers or leaders know what AI is capable of, so if we know that AI is capable of making a product development process go 30 % faster, and we set that goal kind of built into that, then we don't need to tell people to use AI or how to use AI. By definition, the only way to hit that goal would be to use AI. So for us, at least, I'd rather ratchet up the goals. I'd rather accelerate the timelines, hold people to a high bar of how quickly things should happen.
19:27And then revealed preference will just be what tools do they use? How do they incorporate that in their workflows? We don't necessarily need to hand everybody the exact template to go follow. Do individual workers have a budget of tools that they can access? I mean, what we've done within my team is that everyone has a quite generous$1 ,000 a month budget for AI. So once you've got O1 Pro or whatever the chat GPT Pro thing is, that's 200 bucks out of the door, right? And then you need Claude and so on. But, you know, no one's really approaching that level right now. But the idea was to give people enough freedom to experiment.
20:03and then every three months you can do a triage and say, what worked, what didn't work, what don't we now need, what has fallen off the agenda, and what is working. But what is the right way to allow them to explore so you do have a thousand different flowers that are blooming, so you're exploring the possibility space for Box or any enterprise, but without losing sense of kind of financial control or security or data? Yeah, so I love that idea of a user budget or employee budget. And that's exactly probably what I'd recommend to anybody outside of our company. Right now, we have not kind of created any sort of stated limit.
20:42And mostly because I don't know what I would, I would have our organization do. And for example, so within Box, Box AI basically connects almost any leading AI model to your data. And so that means that as an employee, you can go into Box and let's say create a document and use Cloud 3.7 to generate, any of the content in that document. So you're effectively able to use any of the leading AI models within our company. And we don't have a budget limit only because I think it's very provocative and I think it's the right thing to do for us. We just said, go run wild. We want to learn what everybody's doing.
21:19We might now, who knows, if you could kind of like parallelize 100 coding agents all using Cloud 3.7 doing long running agentic coding tasks, maybe we'd be like, okay, sorry, you've just spent$200 ,000, you know, building some feature, like then that'll be a problem. We are not there yet. And so maybe we'll have to check in at some point if we've had to kind of evolve our practices on that. That's an important segue. Thank you. You're such a wonderful guest because we wanted to move towards what happens in a world of agents. And that has been the sort of the rhythm of 2025. I guess what happens with the world of AI agents is these things can spawn multiple instances.
21:59is they get to run much more autonomously because you move from the prompt respond, prompt respond, where the human is always in the loop. So that creates an issue around reliability and trust. But it also affects how you think about the design of the organization. I noticed that Moderna, which is the wonderful biotech firm, there was a story in the Wall Street Journal a few days ago saying they've merged their tech and HR under a single person so that you can align workforce planning with AI. Moderna had famously deployed more than 3 ,000 different GPT tasks in the previous couple of years. How should we think about a world of resourcing in general when agents start to appear and become more reliable?
22:49Yeah, I mean, this is definitely in the category of, again, things that will change in a five-year period that might redefine 100 years of management science, this is one of those. You know, if you think about like the Alfred Sloan typical management science approach of, okay, you know, we developed all these hierarchies and everybody's an expert and everybody has sort of division of labor and, you know, very clean systems. You know, this is sort of blowing that up a little bit because now you do have some fundamental questions of what happens when an IC or individual contributor effectively can become a manager of agents and there's sort of no upper limit of how many agents they could be deploying.
23:29In that world, the IC is now a manager. Then the question is, what is that manager of human managers of agents? What do they do? Should they also be managing agents? Open question. The other question is, how much is this the responsibility of a reinvented HR department? How much of it is a responsibility of a reinvented IT department where you can either have HR sort of thinking about, I have a new type of labor force a la agents, and I have to think about them in coordination and concert with people? Or is it the IT department saying, hey, we're like the technical experts. We're going to always be on the front lines of what's happening in the AI space.
24:10And we're going to bring that to the organization or the lines of business or to HR and say, hey, I have a new type of labor I can supply to the organization. Here's the latest set of offerings that we can bring to the business. I'm probably partial to the latter because the technology is moving so fast that I think you're going to need to have deep technical expertise to understand where this is all going and how to incorporate it into the rest of the business workflows. Because these AI systems will not show up in any kind of physical form for the foreseeable future. you know, like we can get into robots or whatever, but like that means that they're only going to be deployed from software and IT is going to manage that software.
24:51So, so I think that this puts IT in this unbelievably interesting role, which is you're now kind of responsible for the digital workforce of the company, which means that, that you need to then also be incredibly astute at the business processes of the company. You need to be incredibly astute at, at basically, you know, if I could go and add a workforce to the sales team to go in and do lead generation or lead review, what part of the sales productivity metric would that move? That's not usually a classic IT conversation that they would normally go in and be that partner on. In 5 % or 10 % of companies, definitely.
Read the full transcript
25:30But for 90 % of the world, what usually happens is the sales leader says, hey, IT, I need a new tool for helping my sales reps go and do lead gen. But in this moment, IT should probably be coming to the sales leader and saying, hey, I can deploy this kind of digital labor at this part of your set of workflows. So the premium on understanding the business, the business model, the business processes is massive right now in IT. And I think this is going to be a major moment for companies to kind of get around that. Wow. I mean, it's super complicated because you have model releases every two weeks right now.
26:03So prior to ChatGPT, the big five model companies released a new model every six months. Now it's every two weeks. Within that, you have their specific product releases. So OpenAI announced a new coding agent just before we went on to do this recording. You also then have the growing plethora of applications that are wrappers, as they used to be called, right, that are built on top of these. So the velocity is extremely fast. And IT has, I suppose, in recent years, also got a sense of being the chief technical authority of the company, right? What can we use? What has got the right trade-offs between the investment we put in, the result we get set against the risk?
26:52Because IT has a risk component to it. So there's an entirely new discipline. So I think internally, it ends up having to be the CEO. That I totally agree with. But very quickly, I've met the CEOs of these big companies like they're not going to spend more than two hours a day personally trying to read all the AI news. So then the question is, who are they going to go and say, I need you to be my expert to help me figure out what we're going to go and deploy. And I guess my point being, because of how technically complicated and quickly this space is moving, I personally wouldn't want to rely on a non-technical function to kind of, you know, guide me through that.
27:30the kind of vocabulary that we've had to develop just to even make it through this conversation is like, it's already two years of us having to be wired into this to be like, okay, you know, Cloud 3.7 Sonnet is better at coding than O1. And think about the productivity, you know, difference. If you deploy one AI coding tool versus another, that is not something you want to leave to chance. So that's why eventually I still think that it becomes, you know, IT or engineering or whatever we want to call it, but it's a technical function that will have to basically be that counterpart to the CEO, to be very clear.
28:07Or maybe there'll be a lot of organizations say, okay, HR, you should go define what this looks like, but the CTO, the CIO, the CDO, the chief AI officer, that's the right hand to kind of getting through that journey. And when you think about most large companies and most of your customers have already been in existence, they have systems of record, They have established processes. There seems to be a question between, are we just able to throw AI systems, whether it's LLMs or it's LLMs based within agents, on top of our existing decade-long relationships with our core sort of enterprise systems of record?
28:47Or do we need to start to rebuild these, right, from, you know, function at a time? How do you help CEOs think about that particular question? Yeah, I mean, this is super interesting. And I would say that my thinking on this continues to evolve. And so, you know, in five years from now, I reserve the right to be totally on a different, you know, have a different view. Right now, I think that for existing categories of technology, a lot of the incumbents are very good. They're with it in terms of this latest tech breakthrough. This is not an environment where I don't meet a lot of incumbent CEOs of software companies that have their head in the sand on this movement.
29:25If you compare that to, let's take a snapshot of 2007, let's just say. And if you were to go to ask the incumbent software companies, how much are you going to invest in cloud? So you go talk to the CEO of SAP, the CEO of, at the time, Larry Ellison, etc. And you say, how much of your business is going to be cloud first right now? There was this kind of culture of resistance. Larry was actually a big thought leader on cloud. There was a little bit of a time where they took longer than they probably should have to move to the cloud. But he was obviously within the 90s. He was a big believer in network computing.
30:00So credit to him. But the company still took a while to shift. If you ask the equivalent incumbents today and you say, hey, Benioff or Bill McDermott, how much are you going to do AI? Like these companies are all being completely pivoted to being AI first platforms. So the good news for an existing large enterprise who is using one of these vendors, your vendors are 100 % motivated to becoming as AI-centric as possible and ensuring that behind the scenes, they're making their data models improved. they're making their architectures improved to be able to support AI agents. And just in the past, just to give you an example, just in the past month and a half or two months, we've announced agent interoperability with Salesforce, with ServiceNow, with Google.
30:46We just announced something with Microsoft. And so the good news, everybody has gotten the bug. Every platform is going to be an agent-first kind of software platform. So that's existing software. The neat thing, though, is AI is also offering the ability to have new categories emerge very regularly. that are not just CRN or ERP or ITSM. These are categories of software where it's kind of quasi-professional services. It's a new sort of category of digital experiences where you say, now I want software to translate my marketing asset into 20 languages. Or I want software to obviously, you know, be an AI agent for developing code.
31:28These are going to come from new companies, by and large. There might be incumbents that try and insert themselves into that space, but these are going to be lots of new startups. So those will be areas where you can kind of reinvent the process fundamentally. But I'm bullish and optimistic on how the incumbent companies sort of work through this. I know that that might come as bad news for some ruptured startups where they would prefer that Salesforce wasn't with it. But I think Salesforce is going to still maintain their dominance on CRM plus agents. I think ServiceNow will maintain their dominance on ITSM plus agents.
32:00There's still lots of startup opportunity, to be very clear. But I wouldn't anticipate those guys going away in the process. Yeah, I mean, I think it does feel like they have leaned in much, much more quickly. And I think with changes at the FTC, the M &A window reopens. Yes, that's a great point. I mean, Moveworks being acquired by ServiceNow, like that may have not happened before. Then maybe there's one other byproduct. This era of enterprise software, for the most part, you still have leadership teams or even founders running companies where they themselves were the disruptor. And so everybody's like very clear that like you do not want to get disrupted.
32:37That's scaling up to the biggest companies. One of the fun stories of software right now is Sergey Brin goes in every day and is working on AI. Like this is like, think about the era that we're in where Sergey is just like coding and doing AI training right now. Because not only is this probably, I mean, I shouldn't speak for Sergey, but probably the most exciting, you know, moment that they've had, you know, as a technology company. But also he probably understands what's at stake to make sure that Google makes it through this journey and Sundar, you know, as well. So you have these kind of like leadership, Sundar fully gets it, Sergei, you know, onboard driving things.
33:16It's been rumored that Bill Gates is, you know, back doing AI stuff at Microsoft here and there. So what an incredible moment that founders are coming back to their incumbent companies because of how important this period is. You know, absolutely. And even non-founders. I mean, I heard, I spoke to someone who knows Satya Nadella very well. And he said that Satya is coding and he keeps hitting the rate limits with his coding assistant, which I think you're the CEO of the company. Somebody turn a switch. On that question, though, agents, the world of agents and these software platforms, we've come from a world of per-desk pricing.
33:49And these disruptors, like Benioff, they disrupted a world where you paid an enormous license fee in maintenance. That's what on-prem software was like in 98, 99. So how do we think about agents and what this software, this business model turns into? We spend$100 billion a year, roughly, on enterprise software. and, you know, I think people have started to say this is going to be the business of selling outcomes and not just tools. So maybe we should even start with your experience at Box, right? If you're going to be selling outcomes rather than tools, how are you going to charge your customers?
34:27Yeah, so I think it's probably going to be a hybrid business model ultimately because I don't think AI disrupts still that sort of desk worker dynamic. So I think we're going to have, it's going to be seats plus. So basically, you know, think about it as a seat model for the end user that's interacting with the system. And then you've got this kind of like potentially exponential in some cases, but certainly augmented business model of basically, you know, agents doing additional work. So now all of a sudden, and this is like why I think it's the most exciting time to be building software. if you were a startup and you were pitching, I'm just going to make up the example, but you were pitching legal document or contract management software 10 years ago and you go to a company and they would say, okay, well, I have 15 lawyers, so I'll buy 15 seats of your software.
35:18And that software startup is like, okay, obviously it's not that big of a TAM because it's only the number of lawyers. So, you know, we'll probably have to price more than like, you know, the office subscription, but like obviously there's an upper limit of what you can charge for software. So maybe it's like 200 bucks a seat, you know, something astronomically high relative to productivity software, but like 200 bucks a seat or 300 bucks a seat. That's your upper limit. You have 10 seats that you can sell to that company for a couple hundred bucks a seat. Now you're that exact same company in 2025.
35:46What you're going to do is you're going to go to that company. You're going to say, we have software that not only makes your lawyers more productive, we actually start to automate some of the work that they otherwise do or never got around to or have to kind of outsource. and we're going to charge at the rate of$5 a contract or$10 a negotiation or$100 a legal review, some metric that an agent is going to go off and do. And that metric has no upper bound for an organization. It's not capped by any inherent thing other than that company's own volume of customers or sort of business that they want to go through.
36:23So now that same software company, you could underwrite 5 or 10x the amount of revenue that they could generate versus what they could have generated before. To make this not even a hypothetical, think about Cursor. If I told you in 2010, there's an IDE company that's a fork of VS code that's going to get to$300 million in revenue in two years, you would be like, none of the words would, you'd be like checking which words. That wouldn't make any sense. Yeah, which word did I miss here? Like, it doesn't make any sense. Like, even at$20 a user for an IDE, like how many people could there even be?
36:58But it's because agents have an unbounded amount of utilization that they can go and drive. And so I don't think it disrupts the seat model. It expands on top of the seat model of software. Yeah, I mean, if the TAM, the total addressable market, is growing so rapidly, it almost doesn't matter what the seat model is because there's so much upside. There's a few things just to unpack here. So one is that, of course, the underlying cost of delivering this input is declining really rapidly. So we saw the cost of a GPT-4 equivalent token decline about 200 times in an 18-month period to June, July last year, and they continue to decline.
37:39So that moment is really, really noticeable and that trend continues. The second is the utilization of those tokens, especially with inference time or test time scaling, which is what the reasoning models do, especially with agents, dramatically, dramatically increases. As a little aside, I was sitting with one of my colleagues, and we just thought we would map our company use of AI tokens. A lot of it comes through APIs. And he used 400 ,000 tokens in 15 minutes, just doing some little thing. And so you get these two variables that are inputs, like the cost is coming down, but the amount we use is going up.
38:16It's very clearly got positive elasticity of demand. OpenRouter, which is an aggregator of LLMs, has seen a 45x-fold increase in per token usage over the last 12 months, which is pretty dramatic. And on the other hand, agents are getting more reliable. So Meta, which measures this stuff, says that on certain coding tasks, an agent can do 50 minutes of work unattended 50 % of the time, but that is doubling every seven months. so you're not far away from the five-hour task, but it's only at a 50 % reliability. So we're trying to make sense of all of that, right, in terms of what does the shape of the market actually look like.
39:02If I'm paying for outcomes, I want a certainty that it's been delivered. So if I'm the agent provider, but my agent is only 60 % reliable, I'm going to have to run this agent several times to give you a guaranteed verifiable outcome, which raises my cost. and my time. So when I start to package all of this, I'm thinking, does this actually slow down that future that you and I have talked about? Or is the rate of technical progress so fast that this becomes a kind of attractable problem to get to the future that you described? Yeah, I'm going to go with the latter. I think you can outrun it with just the efficiency gains of the models.
39:41Here's a fun example. We just upgraded one of our default kind of use case models from 4.0 to 4.1. And we saw about a 15-point improvement basically overnight because all we had to do was switch the model, maybe tune the prompt a little bit. 15-point improvement in basically favorability score of answers. You know, I'm going to make up the numbers, but we went from like mid-70s PDFs to 90, you know, kind of more or less on PDFs. So asking a question, getting a summary, kind of adding expertise. And so basically in less than a couple quarters, going from 4.0 as our default to 4.1 as our default, we got that level of boost.
40:20I haven't checked in on the latest on 4.1 costs, but like was not, you know, didn't kind of astronomically change the financial variables of this part of the product. So I think what's going to happen is you're going to see improvements to these models just continue at the exact same rate that we've seen. The cost curve come down at basically exactly the same rate. We've got this great dynamic, which is you're now seeing TPUs sort of bring down costs on the Google side. That's going to obviously continue to drive competition for the NVIDIAs of the world to have to continue to improve their pricing and their kind of gains.
40:55So I think I basically, asterisk that I could totally be wrong, I think costs are not going to be the underlying problem with AI. The things today where to your exact point, you've got to run something through a model 20 times, And either in a year from now, you'll run it through a model five times because the model will be 4x better, or it'll be 25 % the cost because some other efficiency gain has happened that has caused the amount of times you run it through the model to be just much cheaper light. You can solve it through either variable, either the token goes down or the quality of the token goes up.
41:26One or the other, you're going to end up being able to outrun this. I suppose where we've seen the tremendous growth on the agents has been in coding. And coding as a domain, of course, is very, very verifiable. There's a did it work or did it not work? Lots of things you've discussed in your business, unstructured data in contracts, in sales proposals, in market research. A summary of that is not necessarily right or wrong. There's a verifiability problem. But verifiability of a softer question is just a more expensive verifiability. but it does seem to me that that needs to be tackled because a lot of the value in the world is not the lines of code.
42:07And it's quite difficult. We've discovered this, that, you know, we run tests. I know the Box does this as well, where, you know, you have lots of the LLMs and you're trying to get it to do some work with unstructured text. And we just discover that seven times out of 15, O3 is best. Four times out of 15, Gemini 2.5 is best, except sometimes when it's called 3.7. And then occasionally a deep seek outperforms on the remaining tests. It's really hard to make a choice. Now, we're not putting them in mission-critical roles right now. There's always a human in the loop. Is it the underlying capabilities of the models improving on this exponential curve that solves that problem?
42:45Or is it what the vendor offers to their customers in making sense of it or guarantees or assurances? I'm talking my own book for a second, but this is where I think having the software layer be different from the model layer ends up being useful because we don't have any, I mean, we have got great relationships and friendships and partnerships, but we don't have like a technical bias toward one of the AI vendors. So we will go where the model is best able to solve a particular problem in that domain. And so you can easily imagine an agent in the future that says, okay, for legal contracts, I like Claude 3.7.
43:21For financial documents, I like Gemini 2.5. For, you know, large code bases, I like Grok 3.5 or whatever. And it will route, you know, based on whatever data type it sees or the length of the data type or some other nuance. It's like not that hard. Like, we can technically do that today. We don't, it's not that elaborate because we haven't found the need. But like, you would want to rely on your vendor to do that. So a la Box. But the good news is that these are increasingly tractable problems. And if you had said this two years ago, I would be pretty nervous because, you know, think about the state of the world of like ChatGPT 3.5.
43:58You know, you're like rolling the dice whether the answer is accurate or not. And the context window, it's bizarre to even think back this much. The context window at the launch of ChatGPT 3.5, I think, I could be wrong, somewhere between 4 ,000 and 8 ,000 tokens. That's right. That's like, you know, a couple pages of a document. And now we have million token plus models that you just give the whole document to the model and it comes back. In two and a half years, to see a 250x improvement in how much context you can give the model is truly insane. There's not a technology that's ever improved at that rate.
44:36So I'm increasingly confident that any of these types of problems are going to be tractable. And then there's a whole category of unstructured work that to your point is not, it doesn't get compiled like code, but almost anything in the realm that we're seeing today of 2-5, 3-7, 4-1, 3-5, which is funny because like, you know, anybody listening to this will know every single one of those numbers, what it meant. And to the outside world, it's like the most insane sounding thing. But there are many problems with any of those class of models will do 10 times better than any human for that task. Like if I go into, I love all my doctors.
45:14They're amazing. But respectfully to them, there are many times where I've been to a doctor two times in a row over a two-month period. And the notes that I get from the second meeting are like 40 % accurate relative to what I said in the first session. And so God bless all of the doctors in the world. But like making our doctors have to take copious notes to transcribe an interaction that then is going to miss like many key bullet points that actually I shared, you know, in that process. My God, I would love everybody to be required to have a transcription of every conversation with a doctor, run that thing through Gemini 2.5 Pro, and that is the note that's in the doctor EHR system.
45:55So I think for a large body of work that we do today, every day, these things will already exceed what the humans are doing. Now, maybe they'll interpret the wrong insight as the key insight, but their ability to capture, synthesize, diagnose, summarize, transcribe is absolutely above human capability, you know, at this point. Yeah, well, I mean, very, very true. And we have that experience within all of our meetings. You know, we're always using one or two different transcription methods to help us actually speed up how those meetings go. And of course, in talking of speeding up, I can't believe we're already top of time.
46:33And normally, I summarize some of the things that we have described and discussed, I will have to do that in the show notes, because it's been such a rich set of guidance for people in the enterprise and bosses and workers trying to figure out how to make sense of AI. I want to just leave with one question for you, which is, what is the one thing that you are now doing with AI as we're in the middle of 2025 that has really started to change your life outside of work that is surprising that you didn't think you'd be doing a year ago? I think I've reached the point in my personal use of AI that is very much akin to, it's very early in this, but it's very much akin to probably what sci-fi had intended.
47:19If you think about the Jetsons or any, you know, author would have kind of predicted, I have a little command center where I'll be talking to my six-year-old about some historical thing or something interesting. he'll ask a question. And it's the kind of thing that normally you just punt on and you'd be like, I don't know, man, like, I don't I don't know why that happened. I don't know how that works. And we just like 100 % of the time pull up Gemini or ChatGVT and just ask the question and go back and forth with it. And it's like, it's this remarkable thing where just anything you want to know is now at your fingertips as this personal assistant that's just always with you.
47:59And so that's how it's changed my personal life to the point where I think we'll just get to a point in a couple of years where we won't remember what it was like to not be able to have every piece of information at our fingertips. It would just be like, wait a second. So you just had to like guess how much Tylenol to give a five-year-old, you know, when they're sick, like what were you supposed to do back before you had these assistants where you just ask any question? I think that summarizes exactly where we are. It's starting to become so deeply normalized for us. We are still at that moment of wonder, though, where actually we recognize that we have it and we used to not have any of this.
48:34Aaron, thank you so much for making the time this morning. Thanks, man. Good to see you. Okay, cheers. Bye-bye.
From the publisher
Aaron Levie, CEO & co-founder of Box, joins Azeem Azhar to explore how an “AI-first” mindset is reshaping every layer of Box – from product road-maps to pricing – and what that teaches the rest of us about building faster, smarter organisations.
Timestamps:
(00:00) Episode trailer
(02:04) The "lump of labor fallacy" in sci-fi books
(07:37) When individual productivity gains don’t translate to teams
(12:32) Box’s Friday AI demos
(21:23) How agents might redefine 100 years of management science
(26:37) A lesson on AI innovation from the early days of Ford
(29:52) Sundar Pichai, Satya Nadella, and Sergey Brin are coding again?
(35:16) Pricing in a post-AI agent world
(38:43) Cheaper tokens, heavier usage: AI’s margin math
(43:02) Solving AI’s verifiability problem
(48:24) How Aaron uses AI in his personal life
Aaron's links:
- Box: https://www.box.com/
- LinkedIn: https://www.linkedin.com/in/boxaaron/
- X/Twitter: https://x.com/levie
Azeem’s links:
- Substack: https://www.exponentialview.co/
- Website: https://www.azeemazhar.com/
- LinkedIn: https://www.linkedin.com/in/azhar
- X/Twitter: https://x.com/azeem
This conversation was recorded for “Friday with Azeem Azhar”, live every Friday at 9 am PT / 12 pm ET. Catch it via Exponential View on Substack.
Produced by supermix.io and EPIIPLUS1 Ltd
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
