Box's Aaron Levie: On Reinventing Yourself in the AI Age and Enterprise Diffusion

15 Sep 2026 · 1 h 5 min · 31 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Reinventing Box/enterprise software for the AI era by focusing on the “bridge” between LLM capabilities and real enterprise workflows; why application-layer diffusion will be massive; how agents should be deployed on top of systems of record; and how to evaluate/tune models for document-centric tasks.

Guest

Aaron Levie, founder and CEO of Box (enterprise cloud content management and collaboration). Background: previously built Box around secure cloud storage/sharing of unstructured enterprise documents; has worked on AI since ~2015 (including early OCR/classification/security work) and led a company pivot after the ChatGPT moment to build an AI stack and “Box agent.”

Key claims

LLM wrappers alone aren’t enough; enterprises need workflow-specific “bridge” software plus change management and domain expertise. Model providers will face a strategic tension: move up the stack vs leave room for applications. Model token subsidization is likely temporary; value should spread across the stack. “Work slop” acceptability will evolve; trust signals will remain messy.

Notable examples

Box agents answering questions over “hundreds of billions” of files; extracting metadata/structured data from contracts and research; long-running onboarding/security/governance agents. Box “agentic harness” evals (complex work eval; holdback eval) and model performance correlating with coding benchmarks (Fable 5.1, GPT-56 class; Gemini sometimes better).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Career Tips for Students

0:00 to 0:51

Learn the importance of Twitter for career advancement in college.

“I want to send an emergency alert to like everybody who's like a sophomore or junior in college and just be like, follow these 20 accounts on Twitter.”

The Evolution of Application Companies

1:11 to 3:48

Discover the changing landscape of application companies amidst AI advancements.

“I am wondering though, I know this is a different podcast, but when do you talk about root canals?”

Bridging the Gap Between AI and Workflows

3:48 to 6:46

Understand the challenges and opportunities in integrating AI into enterprise workflows.

“It's a little bit secret, but we pay very close attention.”

The Role of Domain Expertise in AI

6:46 to 9:15

Explore why domain expertise is essential for successful AI deployment in businesses.

“On the other hand, you also want to be able to have an ecosystem so people trust you.”

Box's Journey and AI Reinvention

9:15 to 12:54

Learn about Box's history and its evolution with AI technologies.

“And then you kind of alluded to this point, Fox kind of guarding the head and house.”

AI in Enterprise: Opportunities and Challenges

12:54 to 14:00

Discuss the potential and challenges of implementing AI in enterprise settings.

“and it was a very simple idea, but we kind of just kind of cracked a nut or struck a nerve.”

Challenges in Image Data Processing

14:00 to 15:10

Learn about the difficulties and costs associated with image data processing in enterprises.

“And you had to have a model for every single use case required a, you know, kind of a hyper-trained model for each workflow that you wanted to do.”

Box's AI Stack and Agent Development

15:10 to 16:50

Discover how Box pivoted to develop an AI stack for processing enterprise files.

“Like everything was kind of exactly academically what you should do.”

Automating Document Analysis and Workflows

16:50 to 18:20

Explore how AI agents can automate the analysis of documents and workflows.

“Your tagline, your business lives in content.”

The Evolving Perception of AI Content Generation

18:20 to 21:10

Understand the tension between AI-generated content and traditional standards of quality.

“Like that's the dream state of most of these enterprise workflows is what if we can onboard a client faster?”
Show all 31 chapters

Trust and Quality in AI-Generated Work

21:10 to 23:10

Examine the issues of trust and perceived quality in AI-generated documents.

“Because I myself am doing the same thing.”

Box's Approach to Harnessing AI

23:10 to 24:10

Learn how Box is utilizing AI to enhance its file management and search capabilities.

“Yeah, no, but it is psychologically kind of like weird because I'll read these like, you know, X articles and I'm like now like doing two X the amount of work to read these.”

Evaluating AI Model Performance

24:10 to 28:04

Delve into how Box evaluates AI models and their performance in real-world applications.

“Yeah, like we always talk to labs like, hey, we'll give you as much data as you want about kind of how the system works.”

The Race of AI Models

28:04 to 29:30

Explore the competitive landscape of AI models and customer preferences.

“There's obviously rumors about other models.”

Adoption of Open Weight Models

29:30 to 30:58

Discuss the current adoption rates and future potential of open weight models.

“So probably higher than people think, lower than what enterprises actually want, and much, much, much, much, much lower than what it'll be in five years.”

Customization and Memory in AI

30:58 to 32:45

Delve into the challenges of memory and personalization in AI systems.

“this that I totally subscribe to is Jesse at Decagon do you probably read that post of like of like this paradox of like we're gonna have like you're gonna see clothes just you know continue you go exponential.”

Continual Learning and Data Privacy

32:45 to 37:09

Examine the implications of continual learning and data access control in AI.

“weights themselves aren't fundamentally changing yeah it seems like i mean if i listen to my friends at the labs like continual learning this idea of like the model's weight should adapt as it gets to know you.”

BoxLabs and Applied AI Research

37:09 to 38:21

Learn about the research focus of BoxLabs on AI and accuracy improvement.

“And there's not a lot of sort of, you know, kind of church and state problems for, for drug discovery, you know, workflows.”

The Role of Systems of Record

38:21 to 40:38

Discuss the importance of systems of record in relation to AI agents.

“So if you're a bank, you have a bunch of loan documents coming in.”

Innovative Use Cases for AI Agents

40:38 to 42:01

Explore innovative customer-driven use cases for AI agents in enterprises.

“And you're like, oh, we could just like have our agent go do that for you.”

The Role of Agents in Governance

42:01 to 43:11

Explore how AI agents can enhance governance in enterprises by proactively identifying compliance issues.

“So like, I think that it was at least half serious.”

The Value of Systems of Record

43:11 to 44:35

Discuss the importance of systems of record and how they can create value through data.

“even remotely a hard debate is you have to just make sure as a system of record that you can find a way where commercially it sort of makes sense on the other side and is sort of valuable and interesting.”

The Future of Product UI with AI

44:35 to 46:33

Analyze the potential future of product user interfaces and the implications of AI integration.

“If I just had a way of just like always understanding like, okay, the CIO is doing this thing and I need to reach out or whatever, I'll like take my money.”

Applied AI Use Cases in Enterprises

46:33 to 47:46

Examine various applied AI use cases in different fields and their future implications.

“I think you're going to have these in every field.”

AI Diffusion in the Workforce

47:46 to 54:33

Discuss the differing rates of AI diffusion across different job roles and sectors.

“And ultimately, like in five years from now, I would bet like 90 % of all tokens in the enterprise are things that a user never kicked off and they just see a result.”

Advice for Founders in the AI Space

54:33 to 56:00

Gain insights from Aaron Levie on how to engage in and navigate the AI landscape as a founder.

“Maybe let's start with founder advice and company building advice.”

The Importance of Being Wired In

56:00 to 56:50

Learn why staying updated on AI through social media is crucial for career development.

“Cause it is, it is the global town square for AI.”

Exploring New AI Products

56:50 to 57:40

Discover insights on the latest AI personal assistant technologies.

“I have not done the Instinct invite code yet, simply because I have a backlog of like three other personal assistant products.”

Building AI into Company Practices

57:40 to 59:50

Understand how to integrate AI into company workflows for maximum impact.

“Um, so we still have like, there's still some blocking and tackling at the infrastructure level for these things to be totally awesome.”

Challenges and Opportunities for Founders

59:50 to 1:03:40

Explore the competitive landscape for new startups and the evolving role of AI.

“So, you know, we had this thing like two weeks ago where I was like, can you just get everybody in a room, you know, on this particular team and just like, like show them how this one person is using AI.”

The Future of AI in Business

1:03:40 to 1:04:40

Learn why the fastest route to enterprise adoption will define success in AI.

“And we weren't going kind of nuts with that.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I want to send an emergency alert to like everybody who's like a sophomore or junior in college and just be like, follow these 20 accounts on Twitter. Yeah, I agree. And also join Twitter because like this will just help your career. You're either like a year ahead or a year behind simply based on your feed. And my feed is like so, so wired in. I'll still talk to 20 year olds that are like, yeah, like, you know, I see some articles and I'm like, what do you mean? How do you see articles? Like, I don't even know what that means. Like, do you just like get lucky that somebody emailed you an article?

0:29You just have to be wired in. I enjoy it. It's a lot of fun. I play with everything.

0:50Today, I'm excited to welcome Aaron Levy, founder and CEO of Vox. Vox is building a collaboration platform and content management platform for the enterprise that has now taken on really new life with AI. And I'm excited to chat with you today about Box and AI, about your general thoughts on AI because you are such a thought leader and about how founders can reinvent themselves for this AI wave. So thanks for joining us. Thanks for having me. I'm a big fan of the podcast. I religiously watch every episode. And so great work. I am wondering though, I know this is a different podcast, but when do you talk about root canals?

1:22Oh my gosh. We're going to call in Doug Leone. Have we figured out how to weave that in yet? I mean, I'm happy to talk about different kind of pain and suffering that we should just do this pod with a with a live root canal Let's do it. I see how it goes. Actually, it would be like hot ones But you get a root canal and you have to talk about your strategy while you're in a dentist chair And Doug is just on the other end of it. Amazing amazing Doug is so pleased with himself Okay, I was seen every tweet. Oh, yeah. Okay. Oh, yeah. He's very pleased with himself. Okay, good Okay, let's start with, I'm curious your take on this, application companies are the hottest Neolabs.

2:00Agree or disagree? Two years ago, I think it would have not made that much sense as like, what does that mean? But very clearly, I think this is what's playing out in the market. And it's all working out mostly because of open source. But what's pretty amazing is right now, the whole concept of being a sort of LLM wrapper or model wrappers actually working out. Because what I think people underappreciated was that in the real world, in the enterprise, what you need is some bridge from the model's capability to the actual workflow that the enterprise has. And that bridge, basically, probably a trillion dollars has been bet on basically one of two outcomes.

2:45Either that bridge is very kind of limited or that bridge is actually very vast and needs to be able to take on a lot of depth within organizations. And that basically looks like, are you only long the model itself and superintelligence, or are you long this sort of application tier, or maybe previously would have been just pure Neolab. But I think it's very clearly playing out that actually there's a lot of gap between the model and the workflow. And as you bridge that gap, over time, you get to a point where you realize, oh, I should also do the model. And then you have enough data, you have enough sort of domain expertise where that becomes its own sort of flywheel.

3:25So I'm very bullish on this idea. And what's cool is it's opening up like multiple layers of sort of opportunity for startups because you could either be the actual applied company itself, i.e. the Neolab, or you could be the infrastructure provider to the Neolab. And you have like multiple layers of of going and attacking that space. But I think a huge update for the market and everybody's kind of views of it. Box Labs, let's go. We already have it. It's a little bit secret, but we pay very close attention. Right now, the main focus is let's make the agent really, really good on any model. But over time, obviously, you would peel off certain use cases, either on a per customer basis or kind of across the whole data set.

4:06Yeah. I mean, it seems like there's two forces that are happening. One is people don't want the fox to be guarding the hen house. They don't want the seller of the token to be the one that is also metering and gating what is the best token for each use case. And then the second is, there's actually a lot of work to do on that bridge to cross and there's real research involved in it. And it's pretty bespoke to the exact workflow and the exact end customer that you have. Yeah, I think what I tend to see happen in the Valley is, and for very good reason, and it's actually why these companies have been so successful, is like everybody's so kind of research pilled, which is again, totally awesome, big fan.

4:41It leads to all these breakthroughs. But the sort of sense that, okay, the model and the intelligence in the model is kind of the only form factor that matters. And then you go to the real world and you sort of see how intelligence actually gets rolled out in people's workflows. And the model could be the most intelligent, super intelligence in the world, but that workflow still requires you to connect up to other data systems, still requires these moments where there's a human in a loop interaction. There's delays in the workflow. And so the sort of agent has to sit idle. There's change management of the actual business process.

5:15There's legacy systems that haven't been modernized. So you kind of go through these five or 10 things that are way sort of, you know, much more operational, much more blocking and tackling than just the pure super intelligence of the model. And the last thing that I think a classic sort of research organization wants to go do is go attack every single one of those things. And this is sort of not a there's like this is not a sort of a one off in history. Like we've always had this relationship between infrastructure and application. Like, you know, obviously, AWS or GCP or Azure have created, you know, trillions of dollars in value of kind of market cap of infrastructure.

5:49But guess what? There's also trillions of dollars of value in software that only exists because of that infrastructure. And if you were to go back, you know, 10 years ago, and you were to look at what was happening in the data space, and you looked at what what, you know, GCP, early kind of versions of GCP was building or AWS is building, I guarantee you would not have predicted Snowflake or Databricks existing, you would have been like, the infrastructure just already does that. Why would you pay another$10 billion of revenue to all these other products that are just making it so you can work with your data?

6:16The same thing is going to be true for intelligence, which is the models will be insanely valuable. but the application of bringing those models into real workflows in banking and life sciences and healthcare and government, that's just going to be a lot of software. Now, the sort of challenge for the next, let's say, two to five years is how much do the model providers need to move up that stack and also try and go and compete at that layer? Or do they either leave open intentionally or accidentally that entire space for the application sort of ecosystem? And actually, to some extent, is like a big question strategically for them because on one hand, you want to be closer to the customer.

6:50On the other hand, you also want to be able to have an ecosystem so people trust you. So that's going to be like a really interesting tension over the coming years. How do you think they'll play out? I feel like that's the question we spend every day wrestling with. Yes. You can, you know, I don't want to speak for Sequoia, but you can sense some of the existential dread of a VC right now is like, should I just put another billion dollars into anthropic or, you know, should I attempt to sort of see what, what the applied layer is going to, is going to play out. So nobody's, you know, that envious of, uh, of your guys' position having to, to figure that out.

7:20But, but obviously it's even harder for the entrepreneur. Um, uh, but you know, I, I am, uh, you know, I'm pretty long, obviously the application layer, like I'm also equally very biased. Like I'm, I'm, I have a very concentrated bet, uh, with very limited diversification on, on it working out that you still want to buy technology that that sort of understands your workflow and can get to the the core enterprise data but i don't know that that there's i just don't see a different event happening than the than all of history i mean you know we only have like 50 or 60 years of of computer software history um but basically when you go to that law firm and you go to that you know pharma company and you go to that bank they need something that bridges the core technology to their workflow and their business process.

8:05And AI has not meaningfully sort of changed the need or the shape of what that looks like. And the best sort of manifestation of this is sort of this chatbot versus kind of agentic workflow demarcation. So the chatbot can be totally universal and can be totally horizontal. But all of a sudden, you kind of look at that and you say, well, my workflow kind of needs something to kind of ping me at the right time in the process, or it needs access to a certain kind of data that just the chatbot can't natively get access to. So somebody has to go into that organization and get it set up. And somebody has to go and provide domain expertise to this model so it really understands our particular business process.

8:45And so then unless you really just underwrite the big labs at, honestly, I'm not exaggerating, like 100 ,000 employees, if you don't underwrite that, then the diffusion kind of economy is going to be massive. because every single one of those companies, whether that's a 50-person firm or certainly a multi-hundred-thousand-person firm is going to need an army of people to go in and help them with that transformation, that change management. And so this is, you know, like whether it's the FDE phenomenon or just, again, understand that domain expertise, that's going to be a very big deal. And then you kind of alluded to this point, Fox kind of guarding the head and house.

9:22And that might actually be singularly the biggest sort of reason this has to happen, which is even under all sort of like complete benevolence and like nobody's actually doing something in any kind of like, you know, sneaky way, it just stands to reason that if I'm going to sort of give a task to an agentic system, I just want that task to be cost optimized, like with accuracy as kind of holding constant. And so who can do that? It's the company that doesn't care among, you know, 10 different models, which model is performing that task. It's like definitionally, you would want that to be the company that does not have a sort of a preference.

9:58And then the force you have going against that is that the model companies can subsidize their models or offer them at different rates. I don't know how long that lasts, though, because when these companies become public, I think they will be held to basically the same laws of capitalism that everybody else is. So the subsidization is working very well up to a certain threshold of spend. and we might have exceeded that spend when you're at like tens of billions of dollars of kind of capital. And in fact, if anything, it might even be worse because eventually you have to go in and sort of pay for your training runs as well.

10:30So like the subsidization of tokens is a, I think it just has to be a temporary phenomenon. You mean from a gross margin perspective or from like an antitrust perspective? Entirely gross margin. Like if I - But they have such high gross margins on their prints right now. But then what are they subsidizing? the like then then it's actually then they're then they're charging actually like you know decent rates yeah but they can afford to price the api higher than their own first party products right so that part totally fair you know there's an interesting kind of you know dimension which is well like if api is the high margin thing that is sort of paying for the subdivision but you've moved all your customers over to the applied product and there's no api revenue so like there's like an equilibrium you have to strike with this um and uh and so then all the while if you have some sort of either non-economic actors or people just with a totally different game theory in this, you know, Meta being one, maybe SpaceX being one, China certainly being a giant one, even NVIDIA being one, like that changes the calculus as well, which is those four kind of cohorts don't necessarily need to make money on inference in sort of at least the same margin structure at like what an anthropic or an open AI needs.

11:39So they would might be fine to bring down inference to 10 % margin because they just want to basically pay for massive compute clusters. And then as long as that happens, and as long as there's not insane proprietary sort of closely held secrets, then no matter what, you're going to have kind of cost per token go down on a like-for-like basis. All of which means more value accrues to the application layer, which is, I think, a very long-winded way of saying, I think there's just value in kind of everybody in the stack. I just don't know that the only thing I probably wouldn't bet on is just, okay, one or two labs get 95 % of the value creation.

12:17I think there's just going to be a much more dynamic environment. And honestly, if I were one of the two or three biggest labs, I think I'd prefer this outcome too. Because back to your antitrust point, like at some point you'll just be nationalized if you're the only thing that sort of exists as intelligence. So you kind of want a little bit of healthy competition in this ecosystem anyway. Yep, totally. Let's transition to talking about box. I want to come back to talking about tokens and Jevon's paradox and China and all this stuff. But let's talk about Box for a second. I love that topic too.

12:44Give folks, I mean, I'm guessing a lot of people that listen to this podcast use Box, but give people a brief kind of explanation of the history of Box and how you're reinventing yourself with AI. We started the company as a way to be able to kind of securely store and share data in the cloud. and it was a very simple idea, but we kind of just kind of cracked a nut or struck a nerve. And for Doug out there, we were able to sort of scale up quickly. We pivoted rapidly in the enterprise and the idea was enterprises would be moving from on-premises systems to the cloud and they would need a better way to be able to sort of store, share, collaborate, manage all of this unstructured data, their corporate documents, their financial documents, their marketing assets, their research materials in the cloud securely.

13:32So that was the kind of company. And we had been sort of flirting with AI kind of products and sort of experiences really since like 2015. If you remember like the first kind of rise of the, you know, maybe the, at least in modern times, the AI winter that happened in like the 2015 to 2018 period, which is like, we think it's going to happen now and and it didn't. And that was a period where we were like, okay, these very, you know, sort of early AI models were showing us signs that, okay, if you, you know, looked at an image and you could classify the image, that's pretty useful if you're in an enterprise, because now maybe you'd take all of your image data and sort of label it, or maybe you'd OCR something and you'd be able to kind of pull out, you know, kind of, you know, sort of the text in there.

14:16That's enormously helpful. The problem was, is insanely expensive. And you had to have a model for every single use case required a, you know, kind of a hyper-trained model for each workflow that you wanted to do. So we kind of shelved it, you know, a few years later, you know, started paying attention to the GPTs. We had some hackathons where people were like, oh, we could do like, you know, type ahead in one of our kind of, you know, note-taking products. And that was, that was sort of early kind of versions. We did some, some early work in, in sort of text detection and classification, which helped with security use cases.

14:47Then obviously chat GPT moment sort of hits. And, And that was the big sort of head-exploding moment. If for no other reason why, then they kind of figured out a form factor that opened up everybody's mind to, oh, these could be these interactive systems that you just ask a question, get an answer back of sort of kind of increasing complexity and length. So we looked at that very quickly. We jumped all in. We sort of did the whole company pivot. Like everything was kind of exactly academically what you should do. We had a team carved out. We put the best people on the team. And we met every day, looked at the updates, and then slowly but surely sort of built out what today is kind of our AI stack and then basically the box agent.

15:29And for us, you can imagine the use case is very straightforward. We sit on hundreds of billions of files. Every single one of those files contains critical information for an enterprise. That could be their contracts, their research files, their marketing assets, their loan documents, like all of this critical information. The problem is they rarely know what's actually inside of it. So unless you literally look at the document and kind of search and find it, you just don't know what's inside of it. So now agents can go and basically be farmed out to go and answer questions about that data. They can pre-process it and extract metadata from those documents and turn it into structured data.

16:07You can use agents to automate sort of steps and workflows. So we built a platform that basically lets you deploy agents against all of that unstructured data. And that's been the kind of core focus. Are you using agents to create new content? We are. There's a couple modalities where that shows up. One is we have, again, an online sort of collaborative product that an agent can just like generate any amount of content in it. And then we've done most of the, I think, probably more exciting work with OpenAI and Anthropic on just how do you do like advanced document creation, PowerPoint creation.

16:39We've decided that their tech is, you know, at this point going to always be frontier. So we have an agent that goes and interacts with those systems to produce a high quality PowerPoint, et cetera. Awesome. Your tagline, your business lives in content. Unleash it with AI. Yes. What are the hero home run use cases for how people are unleashing it today? Yeah. And then if you had to fast forward a few years, what do you think people will be doing with AI in your product in a few years? Yeah. So probably the easiest hero for, again, more of a traditional enterprise to think about is just as simple as you have a million contracts.

17:12Why don't you find out what's inside them? Or you have a million research documents, be able to go and pull out all of the critical structure data, put that into a database, and then be able to query, analyze, automate workflows around that. So that's kind of the thing that just knocks it out of the park every single time, because it's been a longstanding problem that people have never been able to go and sort of apply human sort of labor to, because it's just too expensive to read every contract, every research document. Maybe you could do it if you had like a loan document process, but most other data just never gets read at that scale.

17:44And then I think the stuff that we're probably, you know, as much if not more excited by is really the equivalent of what we see with, let's say, coding agents or other complex agents, which is you have long running agents that are just executing your entire kind of workflow or process. And this would be in the form of you go to a bank and you're onboarding at bank and they've basically like automated every step that is possible to automate and then sort of jumps out to a person in the steps in the process for extra review or extra verification. But now instead of that sort of one or two week back and forth, it just happens in like an hour.

18:20Like that's the dream state of most of these enterprise workflows is what if we can onboard a client faster? What if we could discover kind of critical data inside of our research much more quickly? What if we can alert to a security event much more quickly? So to do that, you need these sort of background agents or workflows that are sort of pre-established for those processes. Totally. It's the year of the long running agent. It is, it is. Yeah, yeah. I'm curious, like you made the analogy to cloud code. It seems to me that in the coding domain, using AI is like not only accepted, it's embraced.

18:51Yes. In the content domain, which is I think where a lot of the content in box sits. Yep. Using AI to produce content at least, it's just like there's almost this allergic reaction to it. like all the pangram stuff on Twitter. There's the, you know, it's like this concept of work slop. I'm curious what you think about work slop and like, will this still be a thing in a few years? I'm going to sort of separate the box corporate hat and just now kind of maybe riff as a consumer of, gosh, I wish there was a better term, but work slop as, you know, inside of an enterprise context. I get board decks that are entirely written by AI these days and it kills me.

19:29So here's the difference, I think, on the acceptability. So there's probably like more symbolism to this actually topic than just like the slop element. But like actually like diffusion of AI in general sort of almost ties to this code. Like other than, you know, the top engineers that we hang out with it, like are like they have deep taste in the code. And the judgment is incredible. And it is them as much an art as it is a science. So take that group aside. For most of the world, code is a utility. It is just trying to accomplish something. You're just trying to automate something. You're just trying to put a sort of interface up there that somebody presses a button and moves to the next step.

20:14So for most of the world, the value creation of code has been to automate things and to be able to have it as a utility. So at the end of the day, like, and we'll probably still use the term slop for a while because there's taste in kind of front end design and there's taste in sort of systems and you don't want to have vulnerabilities in your code. So that's going to exist for a while. But at the end of the day, if you can tell an agent, like, please go and generate my entire backend system or my front end system. Like, it's just, it's not only acceptable, it's preferable because it's just like, that was the thing that was blocking us from moving forward.

20:50So we need to go do that. At least the way society functions and the way the world works and our brains work at the moment, maybe this changes. You know, when you get a presentation from somebody, there's still this association which is like I'm trying to decide if I can trust that person to go and execute on that thing or deliver that result or understand that topic. And so when you see work slop, you're like, I'm losing my ability to sort of know for a fact that like how much of the thought process was them versus how much was the AI. How much did I even care about that? Because I myself am doing the same thing.

21:25So like we have this weird like it's this very weird sort of like collective issue that we have, which is like I'm doing work slop for some of my brainstorms and decisions. But when I get it from somebody else, I'm like, hmm, should I trust you? and I don't know I mean it just might be a thing that as a society we have to kind of keep cranking through over the next kind of three to five years and end up at the other side I hate to use these totally busted analogies but obviously you don't care when you see somebody's financial model, you're like yeah that was generated clearly by like a macro that was like you did not personally go and compute all of that but you're showing it to me and we're talking about it so why can't the same exist for a strategy deck or whatnot but i think right now we're going through this evolution of like what is the person's role what is the content a proxy for is it supposed to be a proxy for how much that person knows is it a proxy for what we think they can go and execute on i think i think we're just in this very messy period where we have to kind of figure that out totally um did you read the stan druckenmiller well street i i read the uh i I read the discussion about it.

22:35Yeah, the reaction to it. But actually, I didn't read it. But was it, like, very sloppy? I didn't think it was sloppy. I loved it. And so to me, it was just a nice counter example of I have this, like, allergic reaction. How many it's not X, it's Ys were in there? I don't think there were any. Okay, okay, okay. But it does show up as 100 % AI in Pangram. Okay. And dashes? I think there are dashes. Okay, you can do that. Yeah, yeah. But it was a nice counter example to me because normally I read something that's clearly written by AI and I just have this allergic reaction. Whereas with the Stan piece I didn't.

23:06And I'm not sure how much of that was just, you know, it's Stan, therefore I trust in Stan versus. Yeah, no, but it is psychologically kind of like weird because I'll read these like, you know, X articles and I'm like now like doing two X the amount of work to read these. I'm reading it one for the substance. I'm also reading it two for the calculation of like, did the person write it or am I just literally reading like a Claude prompt? And then like, I'm like, I'm literally like my mental processing is now like should i now does that up weight or low or like lower my my sort of judgment of the person or the post and i think we're yeah we're in for some weird times because of this i would hate to be a college professor i would just i would totally quit because you're just like are you like i don't i don't know anymore what you did i i like what it does seem like the calculator is the closest analogy though yeah except it's just like that was like more like finite in terms and you still had to piece together so many more things yeah like we like the and some of these analogies are breaking down of like it's just a task because it's like well at some point like this thing is doing like at least 10 tasks at once yeah um but uh but yeah let's talk about harnesses how does the uh how does the box great transition to harnesses okay speaking of calculators let's talk about harness okay this is back to the kind of neolab kind of applied applied layer uh there's there's a bunch of things that that we know about our file system permission structures our search engine that, you know, certainly, and by all means, we actually would love all the labs to train on our understanding of this because it would only make external agents as good as possible.

24:40Yeah, like we always talk to labs like, hey, we'll give you as much data as you want about kind of how the system works. But let's just say like that aside, we have a lot of depth of understanding of what do people do in Box? How do they search Box? How do they decide when they look through 10 files, which is the one to go pick? What is their internal kind of calculus or heuristic on sort of figuring out the most relevant document to look at? So we know all that. And we basically have built an agentic harness that attempts to kind of understand that set of domain understanding about our system.

25:13It obviously has access to our search system, our file system. It has a bunch of kind of mechanisms for just pulling out just the text of a document, just pulling out chunks from the document, doing embeddings on the document on a fly. So there's a set of kind of tools that it can use. And effectively, it's a harness for asking questions of a large data set. So in my box account, I have I don't even know the latest number, but on the order of tens of millions of files. Just because it's like it's every everything that has ever kind of accumulated over over 20 years. but I can now ask any question of all that data set using the box agent and it goes around.

Read the full transcript

25:50It does multiple searches in one. It re-ranks it. It then very quickly sort of pulls out the most relevant information. Then it'll, in some cases, read the full document. It does all the steps. And then we compare that against like, well, what if we just gave, you know, Claude our API or gave OpenAI our API and we see like meaningfully better results on accuracy and latency because, again, we kind of know exactly how to tune it for our workflows. So that's effectively the harness that we built out. And then what evals matter the most to you? So I have a couple just funny personal ones that I just keep track of, of my own use cases.

26:23But we have, I don't know, hundreds of different tests that we do on every single model. We actually have two evals at the moment. One is we put out a thing called the complex work eval, which is a set of domain-specific work in life sciences, financial services, public sector, tech, et cetera. And it's kind of exactly what you'd think of as a document-centric eval. So given these five documents and this set of problems, what would your answers be? And we test every single model against those with our agent. And then we have a holdback eval, which is actually the first one is holdback also. But the second one is just like then our box instance and how box employees use their data.

27:01And then we eval every model again on that. So we're able to kind of roughly keep track of all of the incremental progress. We see when things move by half a point in terms of model improvement. And then we roll out sort of default models based on different kind of cost and accuracy thresholds. And then we let customers also choose any model they want from effectively our model garden. What's your current view of the race and where all the horses are in terms of model performance on your use case? They more or less closely correlate code with one exception, which is actually in some of our use cases, Gemini is disproportionately better than what you would see from coding.

27:40And it might be, you know, sort of just better tool use, you know, given kind of, you know, the Gemini ecosystem and what they need to build for. It solves, you know, kind of a strong set of sort of general knowledge work use cases as well. But I think mostly correlating to kind of code. So Fable 5.1 was clearly kind of state of the art and the best model that we've seen. There's obviously rumors about other models. And so we'll see how the kind of race kind of continues on this front. But basically, by and large, like when you look at GDPVal, Mercore has their Apex eval. These things will all generally follow the coding models.

28:23And so I think it's we're just neck and neck on like Grog, Muse, the Fable class and, you know, GPT-56 slash whatever the building next. Like these are just like it's a total race right now. And your customers typically express a preference on which model they want to use or do they just use your default? So they, you know, kind of by volume, they use our default because it's just easy and it works extremely well. And it's tuned for there's a few ways our kind of agent manifests. So like the way that you'd most commonly experience it as an end user is you would just be searching and kind of, you know, asking questions of your data.

28:59But by volume, the volume of tokens tends to go through more of our workflow agents or data extraction. That's where actually you have customers actually doing evals. And they're basically saying, OK, I want to really make sure that at this cost profile I can get 98 % accuracy on data extraction. And that's a place where like we'll have an FDE that goes in and helps you, you know, helps understand your data environment, tests against five different models. And then you're just basically at the mercy of the eval. Yeah. What are you seeing in terms of the adoption of open weight models in your customer base?

29:30So probably higher than people think, lower than what enterprises actually want, and much, much, much, much, much lower than what it'll be in five years. So some mix of that would be the kind of message. And it's primarily cost that's driving that decision? I have to probably attribute 30 plus percent to just kind of the sexiness of like I want to try GLM. Yeah, I think there's that. Like I've heard CIOs of Fortune 500 companies say we're playing with open source here. And I look at that and be like, well, I know for a fact like Gemini or Muse would have been just fine at that particular cost profile that you're trying to do.

30:08Or probably even like 5.6 Luna or Terra or whatever, whichever one had the crazy discounting they just did. like it probably would have been totally fine but you want to be able to be like okay i'm a little hedged like it's cool like like we're that phase still um over time i think it stands to reason that that you'll see meaningful different costs because you'll be able to peel off workloads that that just only make sense at at sort of you know grinding down to the cost of inference in which case open open weights will have this sort of economic advantage right now there's you know this challenge of like sometimes it's more token inefficient you know sometimes like randomly like I've heard stories like randomly it'll just like speak Chinese like like mid mid chain so you're like okay well that'll be weird for a bank um so it's funny so like we need to like probably work on some of those things but long term I think it has to be you know the case that that you're gonna you're gonna peel off those workloads you know one of the more interesting posts I think on this that I totally subscribe to is Jesse at Decagon do you probably read that post of like of like this paradox of like we're gonna have like you're gonna see clothes just you know continue you go exponential.

31:12But what's going to happen is each use case that kind of matures, you can peel off to open source. And once you have kind of stability in that use case, it starts to make sense to veer it toward an open weights model, assuming one of two things is true. One, that it's actually literally cheaper or two, having some post training gets you X percent more performance. And so I think you will just be in a reality where we will, and this is going to be very confusing probably for like the press more than people in the valley because you'll be like wait a second like the revenue of anthropic opening eye etc are like off the charts but somehow open weights is like also growing exponentially you know like how is this like how is the pie is growing so fast yeah and it's like the pie is growing so fast yeah but what's happening is is actually there's there's an interesting duality it's it's it's not even just like like rising tide lifts all boats it's like no no like we either use fable or or five six for orchestration and then we farm out all these long tail tasks to a cheaper model or the opposite is true like you use you have some kind of orchestration agent that like by default does the cheaper stuff but occasionally sort of see some something that is just way too hard and then it pops it out to to one of these heavier models and so you might have blended 50 spend on each but 10 times the amount of tokens you know on the open weights model and so like everybody's kind of winning but there's an there's an interplay between why they're winning i'm curious about how you think about memory and uh customization or personalization and where that's going to go because it seems like today the dominant architecture is kind of like rag based system still like you can get fancy on the rag but it's still context look up where the weights themselves aren't fundamentally changing yeah it seems like i mean if i listen to my friends at the labs like continual learning this idea of like the model's weight should adapt as it gets to know you.

32:58Yeah. We had Ngram on the podcast. I don't know if you know Dan. I just got introduced. I listen to the podcast. Amazing. I would have loved to been the fourth person in the room. Amazing. Yeah. I think they're working with customers to help basically bake in some of the context into the weights themselves. What direction do you think it's going to go? You know, you're catching me at a time right before I'm actually doing my call with Dan. So I wish I could have talked to him first and then I'll have like a way more eloquent answer. I'm extremely fascinated by the approach. I have no reason for not wanting it to work and exist.

33:34We live in a world at Box where we see the high degree of complexity on permissions and access controls and data that tends to be the rub on a lot of these types of approaches. And I'm going to put Engram aside because I'm sure they've already thought this through. So I'm going to talk more generic philosophically. I think sometimes you will talk to a researcher that sort of imagines the world working the way they work, which is like I'm a researcher. I have access to everything. And so if I had a model that was trained just on my world, this would be amazing. And then you're like, let me introduce you to a lawyer.

34:10and the lawyer like has this tiny little you know access point of just like the five projects they're working on because somebody right you know one door over is working on the competitive project to to another company in the space that they can't have any sort of overlap with with what they see or what they know and there can't be a single document that passes between those two walls and they have to be these kind of hard barriers so you know so like sure like you could still you could still now train a model just for that one user but like what happens if every single day they get added or removed from something that that adds important context to sort of what they need to understand and and and you know again like i think there's gonna be probably breakthroughs in continual learning that sort of all resolve this but like this is why previously it was just like you know there was no way you could pull this off five years ago because it would be insanely expensive impossible to kind of wrap your head around around how those access controls are supposed to work.

35:05But, you know, obviously with like, as the cost curve goes down, as open weights, you know, get, you know, cheaper, smaller, faster, better. I think this becomes super interesting. One thing on the podcast that I found very fascinating, and I just need like a T chart, honestly, is just like, you know, like, what is the decision point of what goes in context and what goes in the weights? Like, you have to be a little bit thoughtful about like, where is the massive performance gain that you get by baking in the weights? And there's probably some like incredible like like calculation of like like when the rate of change of the data is not you know so far but the upside of the of the weights you know you know dramatically change the accuracy of the model like you know you'd have to kind of land on some sort of you know rubric like that i mean if you could wave a magic wand it almost seems like i think carpathia said this in some prior interviews like if you could almost remove all the memorized information from the models and just have it encapsulate the specific reasoning capabilities, like the ethos of how we do things, for example, at Sequoia.

36:05And then you have all the actual content and a lookup system that almost feels like the, if you could wave a magic wand, that's what the system would look like. And so that one's super interesting. I think the question will be like, how much are enterprises different at that level versus it's actually their literal IP that is what makes them different. Like how many different types of styles of execution are there in the world versus no, it's just like the depth of knowledge about that particular legal case. And how do I apply it to this other project I'm working on? That's where so much of the value sits.

36:38So, so, but again, if you can just like wait till my zoom call with Dan and then I'll really know the answer, but like, I'm a fan because no matter what there's going to be like, like I've jumped right into like the individual, you know, but like no matter what, like at a firm level, there's probably ways to take this approach. Like I'm a big sort of fan of what trajectory or applied computer doing or prime intellect. Cause, cause I think there's like, there's, there's no question that if you're Eli Lilly, you want a model for how you do drug discovery. And that probably does need to go like farther or deeper or more sort of specific than what you're getting off the shelf.

37:10And there's not a lot of sort of, you know, kind of church and state problems for, for drug discovery, you know, workflows. They probably want as much of that information available to as many people as possible. So I think it's going to be like domain specific. And, you know, you're just going to have different outcomes based on which vertical or, you know, type of use case and where the firewalls need to be in that process. Makes sense. Okay. So you hinted at the beginning that there's a BoxLabs. What type of work is BoxLabs doing? So it's the equivalent of BoxLabs. Like, I don't know if we've used capital L yet, but basically, you know, it's our applied sort of AI team.

37:42And what research areas are most interesting to your team right now? Yeah. So the biggest areas that the most sort of research kind of, you know, of the continuum of engineers, there's some cluster that is sort of more on the research bent. And of that cluster, the things that we've spent time on still, again, is at the kind of applied layer. But it's a lot around how do you take agents and make them, you know, another 10 points of accuracy improvement given X problem. So what is the, you know, how do you build a map of the problem set, you know, with a given set of data to sort best execute on that task.

38:18So we spend a lot of time on that style of work. We have a team, for instance, working on how do you do effectively at the agent level, at the harness level, some form of kind of auto research on sort of hill climbing on accuracy of answering questions or sets of problems on given a kind of a set of client data. So if you're a bank, you have a bunch of loan documents coming in. And these are like 100 page documents, like whether you're getting 70 % accuracy with an off the shelf model or like 97 % is like basically, obviously a world of difference in can you actually go and automate that process.

38:54So somehow you have to hill climb from the base model to the 97%. And there's a lot of work going into the system to basically pull that off. Maybe zooming out, what do you think of as, and this can be a box question or a non-box specific question, the role of systems of record in a world with agents. And I'm sure you saw some the Twitter discourse on like, you know, every software company is trying to sell me their own agent right now. I don't want another agent from them. I want their system record to work well with my agent. And so like, how do you think about that? Hashtag Claude Force. So - Catchy name, by the way.

39:25Very catchy. I mean, literally, it's like one of those things where like the first three minutes, you're like, man, that seems funny. And then four minutes later, especially when you see the stock, you're like, ah, brilliant move. Like, this is great. We're doing this. And then when you, I think somebody, somebody actually said this the best. They were like, when they heard Matthew McConaughey say it out loud, that was like really the, that sealed the deal. That was the aha moment. That was the aha moment. And it's like, man, he can sell software. Like, it's actually incredible. Like his voice is so good for selling systems of record and agents.

39:56If you're in like, like our sort of, you know, contemporary group of like, you built a SaaS platform and you have some set of data and workflow that your customers kind of operate in, there's effectively two things you just have to do. And I think anybody attempting to do one over the other is just going to lose. You have to build an agent that is insanely great at your product. Like that agent has to be, you have to provably be like 10 or 20 points better than an off the shelf agent using your system. Not because it's hobbled the other side. It's just like you are so eval maxed and you're like so tuned to your particular workflow that you can improve your system, you have to have that.

40:39And you probably, because you understand your domain, unless you're like totally asleep at the wheel, you probably have use cases that no one has sort of thought to build products around because you talk to customers every day and you see what they run into. And you're like, oh, we could just like have our agent go do that for you. Like I've had at least a dozen. I mean, so I probably talked to a couple hundred customers a year in variety capacities i've had at least two dozen times where the customer has a use case that is like a breakthrough moment for me of like shit that would be actually totally insane like what's example um unfortunately since you put me on the spot i don't know if my example will pay off the the level of excitement i just had uh because it'll be on the spot no no i it's just like i think like the thing i was thinking of is just like i think it's gonna be a womp womp for the podcast but um uh there was there was basically a customer had this idea of they wanted an agent in the background trying to sort of figure out when documents sort of met their governance policies and like does something need to go into some kind of archive or something into some kind of legal hold or whatnot oh cool and and see exactly see that's exactly the voice that was exactly the voice i was worried about yes yeah no you couldn't even you couldn't even pull it off okay so so but in our world, this is awesome because you're like, oh, like, no, because think about it.

41:58Every company has a head of governance. I meant Cole sincerely. Okay. I know. I believe you. I mean, listen, you do enterprise. So like, I think that it was at least half serious. Imagine you're an enterprise. You have a head of compliance and a head of governance. Okay. They can only be like overseeing the whole sort of enterprise. They've never been able to be everywhere at once. Now imagine if they could sit next to the employee and basically be able to be like, oh, you're about to go do something that breaks our governance policy. So, so like the idea was like, oh, what if there was just an ongoing agent that just like automatically was just like, nah, that's going to break your governance policy instead of the user having to like try and predict or understand this stuff.

42:36So anyway, those are the kinds of things where if you have an agent within your product, you're going to be able to identify sort of sooner and better than, than the rest of the market. And, or just like do things that maybe would be impossible to do off platform. On the other hand, it's just like, obviously you have to go headless. You literally have to make sure that your APIs are exposed to Cloud and ChattoBT and all the different platforms. And you have to make sure that you have either a direct way into deterministic APIs so that those agents can use your APIs, make calls via MCP or whatever, or at least make your agent be headless and be exposed in those systems.

43:10And then the only reason maybe this is like even remotely a hard debate is you have to just make sure as a system of record that you can find a way where commercially it sort of makes sense on the other side and is sort of valuable and interesting. And the reason why I think a lot of people got that wrong that weren't in these companies was just underestimating the amount of sort of new use cases that are just total upside. They're like complete white space opportunities for these systems of record. So like in the Salesforce example, I use Salesforce more today, probably by an order of magnitude than I ever have because I MCP into it via Claude or Chachapiti.

43:46And so I just am always asking questions about the data inside of our CRM system. And do you think that means the systems of record become more toll booth businesses then to make sure that they're capturing the opportunity? I don't love that term because like no one's had a good experience at a toll booth. I love toll booth. Yeah, exactly. You love it more than governance agents. So I would say that because they sort of have a depth of purpose of organizing the workflow, managing the data, securing the data, providing guardrails, then it's really just, yeah, you'd have to have some kind of volume oriented business model on that other side.

44:21And I just think there's like, if you're solving real problems for customers, it'll just like make money. I've like, this is so cheesy, but like I've told like LinkedIn product managers, I'd probably pay 10x more for LinkedIn if I just could MCP into it. If I just had a way of just like always understanding like, okay, the CIO is doing this thing and I need to reach out or whatever, I'll like take my money. So these systems actually have a tremendous amount of value based on the data that they have. And customers will absolutely sort of find some way to reward you for that value creation if you're doing a job.

44:55Super interesting. Maybe related, let's talk about product UI. Yep. Big, you know, generic chat bot, chat box agent. Is that going to be the dominant UI for how people are using AI in the future, especially when it comes to the application layer? This is why I think the applied layer has so much room to run is because probably the sort of universal chat system that you ask a question to, you get an answer back or does sort of some work in the background. I think it's going to be obviously like that's going to be a mainstay that that that UI will always exist. It'll be incredibly powerful. The horizontal products will have it.

45:34The vertical products will have it. Everybody will have it. It's just like your product has a search box. Obviously, it does. Yeah. So that's always going to be here for these sort of one-off asks of an agent or go find this thing or answer this question or produce something for me on demand. But most of the enterprise is sort of made up of these processes and workflows that are kind of just happening behind the scenes. Sometimes they're happening with computers and computers are running these things. Or sometimes they're happening with other people that are doing these things. Or sometimes they should be happening with people, but you could never afford to have them happen with people.

46:04So they just didn't happen. And so that sort of is a slightly different kind of metaphor than a chatbot where you ask a question and it comes back with an answer. That's like, okay, I kind of want agents in the background to do things for me. Read every contract. Look at every log. Triage every security incident. And then instead of me chatting, maybe I'll chat as a means of doing kind of a catch-up. I want a dashboard. I want a workflow. I want a queue. I want a task list. so then the the the challenge becomes well does the horizontal product kind of take on every one of those those components and and manifest every one of those experiences in one in which case i think you'll start to be like man that thing is like really like a that's pretty heavy and like and then we'll start to like be like oh this is no longer this simple easy delightful thing anymore so then the vertical players actually like like they actually sort of understand the process and can manifest all the right buttons and tabs and the names of the things for that particular workflow.

47:06So I think it puts, I think as you have agents that are doing more work in the background, doing more, you know, async work that is just like I farmed out a bunch of agents to review things as they happen or whatnot, that leans more toward the applied companies that can understand those workflows, that can understand those processes. I think you're going to have these in every field. You're going to, certainly we already know that, you know, how they're going of looking legal with Harvey Laguerre, et cetera. We are seeing them start to emerge in areas like security. We've seen them start to emerge in the long running kind of coding agents with cognition and factory.

47:39So I think that that will be, you know, one of the bigger kind of applied AI sort of use cases. And ultimately, like in five years from now, I would bet like 90 % of all tokens in the enterprise are things that a user never kicked off and they just see a result. They just, they see a task show up and they have to go review it. And it's just like, it's just happening. Yeah. Yeah. Makes sense. Okay. Let's talk about AI diffusion. Coding agents, it was like, boom, January 1st, 2026 happened. And the fastest diffusion of anything into the economy we've ever seen has happened. The diffusion of the rest of the AI magic into the rest of our jobs seems like it's been a lot slower.

48:19What are your thoughts on that? And where are the areas where you think we're going to see faster diffusion and how is that going to happen? Yeah. So you always have to kind of compare and contrast coding versus everything else to really understand the dissimilarity. So in coding, and this is back to the kind of utility point on slop, like the utility of code is almost 100 % represented by the amount of text that you can generate. Obviously an insane amount of value went into the text, but like knowledge and expertise and meetings and everything. But like ultimately like the text is the thing that produces the program that is actually the thing that you're trying to do.

49:00You know, if you could have the world's greatest programmer like and never had to sleep, never had to eat, they could intuit what to build and they could just sit on a computer all day long, your value creation would be 100 % correlated with how many hours they could sit at that computer and type code. Like lines of code, ideally good code, is the most thing that will be correlated to whether you produced software that people wanted. So it's all text. The models are hyper-trained on these. Everybody in AI labs treat coding as a competitive benchmark to constantly try and exceed. They get to do their own evals on it every single day because they are the ones coding the models themselves.

49:40and it's the most technical audience of all time where when they deploy an agentic system and they run into either a bug or a problem or like some MCP server comes back with like connection invalid. They know how to triage the problem. They don't call IT. They just like, oh yeah, no, I didn't open up that port. Sorry, I'll go fix it. So that's like five things. Oh, and maybe like the sixth, like it's just like a very, very high paying vertical that like it's just like automatically valuable if you could get 10 % or 20 % productivity gain, let alone 5x productivity gain. So, so, so take those five or six things that coding has as, as sort of beneficial properties to sort of automation, then compare that to, you know, you know, every other form of knowledge work.

50:24And you'd probably have like a histogram and I don't know if anybody's published, maybe you can like the similarity to coding and like, like what, what are the domains that like, like start to sort of, sort of, you know, look closer, like, like less and less like coding as you, as you kind of scale out. And it's like, okay, well, lo and behold, legal is kind of interesting because, because like there's a lot of value creation to somebody sitting at a computer, reviewing legal documents, writing legal documents, like processing large amounts of information. Okay. So that's kind of blowing up. And then you kind of go down the list.

50:54Now let's take something like, like a sales rep. Okay. So much farther down the list in terms of, of sort of likeness, the sales reps value creation is basically convincing an external customer to buy software or technology or a caterpillar truck, you know, from them. That is the value creation to the economy of the sales rep. And so let's say we brought the world's best automation to them. Like, first of all, again, they'd have to like figure out how to technically wire it up. They'd have to make sure they give it all their data, all these kinds of things. But no matter what, they're still rate limited and constrained by like, did the customer respond to them?

51:30Do they want to meet? Can they meet next Tuesday or can they meet today? Like, does the customer have budget? All of these other things. So that's maybe the entire continuum right there is like one is like of knowledge work. Like, obviously, this is not even touching the sort of working with with Adam. So that's a continuum, which is on one end, you have somebody rate limited by so many external factors. On the other end, you have somebody who could sit at a computer all day long and just type type text. And that is your ability to basically automate things is that continuum. So for the real world, We have to basically bring intelligence to these workflows in ways that are that sort of somewhat feel like the shape of their workflow, somewhat feel like the shape of their work, and then find a way to deliver the change management, deliver the implementation, get data into a format and into an environment that actually works with these systems like Asterix.

52:19Like one of the other big things is like if you go to most engineers in 2026, like maybe like minus two months ago, given the given the latest phenomenon. But like like the codes in GitHub, you just like connected to GitHub. Like like remember, there's this period where like you're when you launched a coding agent, like there was no like sign up or register. It was just like, give us your GitHub. That doesn't exist in knowledge work. There's no like give us your GitHub for knowledge work. There's like give us your box. Well, box customers have a much easier time with all of this. Unfortunately, we're only 1.3 billion in revenue run rate.

52:50So that means there's a lot of people not using Box. And so what are they using? Their data is in on-premises systems, legacy file shares, legacy infrastructure, enterprise environments that don't talk to agents particularly well. So just think about that sort of distinction between implementing coding agents versus everything else. Even something, again, you'll sort of fall asleep about is access controls in the enterprise are totally different. I actually totally forgot that point about coding. In coding, you get access to basically most of the stuff ever relevant for your job. In knowledge work, you're like, hey, Sally, can you open up that sort of file share for me?

53:27Can you open up that sort of project? Because I didn't get access to it. How do you make sure the agent has access to those set of things? All of that work has to get done. So the thing I think we have to prepare for is two things. One, Silicon Valley has to prepare for diffusion taking a lot longer than they think. Or then I'll say we, but I really think they because I know how long it'll take. And then the second thing is, is the good news is this is sort of all correlated to the applied layer value creation. Like because the companies that will just have the patience, the sort of the full sort of domain expertise, the just the sheer work ethic, because it's not like everybody's just from the floodgates.

54:09Like you have to go in and just, you know, pound pavement and get out there. That will be the applied layer. So I think this actually represents, you know, a trillion dollars of applied layer AI value is actually how do you go get the technology to the lawyer or to the sales rep or to the life sciences researcher or to the person that runs the customer support team? That's all sort of opportunity right now that exists. Awesome. I'm going to close by asking some advice for other founders. Maybe let's start with founder advice and company building advice. On the founder side, it seems like you are in every AI cap table, you know, every cool new company, like, you know, you know, and Graham, you know, how did you kind of get yourself in the middle of the AI conversation?

54:50There's probably two parts. One was just very well primed for it. Like working with unstructured data for 20 years, you just like can instantly see the benefit of agents on that. So, like, it honestly, like took longer than I would have wanted that we got to have this conversation because we tried to have this conversation, you know, eight, nine, ten years ago. and now it's finally happening. So first of all, just super well primed. Our product sort of shape and what people do with our product already lends itself extremely well to agents. So like obviously we had to bet the company on that and then lo and behold, we had the positive feedback loop of customers actually saying, yeah, that would actually be very powerful if I could go read every document and answer any questions.

55:28So that's the first and certainly biggest by kind of a factor of 10. And then the other is just like, I'm extremely fascinated by the technology and it's just fun. It's like - Where do you learn about it? Um, your podcast, uh, Dworkesh's podcast, um, uh, the, I mean, uh, probably unfortunately for brain cells, like it's probably 95 % Twitter. Um, I have a routine where like at the end of each night, I just go through the feed and it's just like, I do, I like, I, I look like some sad meme probably of just like, I'm just scrolling and scrolling and scrolling and like attempting to like triangulate all the information.

56:04It's amazing. Cause it is, it is the global town square for AI. It is. And I want to send an emergency alert to like everybody who's like a sophomore or junior in college and just be like, follow these 20 accounts on Twitter. Yeah. And also join Twitter because like this will just help your career. You're either like a year ahead or a year behind simply based on your feed. And my feed is like so, so wired in. This is the number one advice I give to people when they're asking how to get current on AI. It's like follow these 100 accounts. It's like, but I mean, virtually first step, join Twitter.

56:33Yeah. Like, I'll still talk to 20-year-olds that are like, yeah, like, you know, I see some articles. And I'm like, what do you mean? How do you see articles? Like, I don't even know what that means. Like, do you just, like, get lucky that somebody emailed you an article? Like, just join Twitter. Like, what are you talking about? So, I would, you know, you just have to be wired in. I enjoy it. It's a lot of fun. I play with everything. What's your favorite new AI product? Please say Instinct. So, okay. Full disclaimer. I have not done the Instinct invite code yet, simply because I have a backlog of like three other personal assistant products.

57:09You got to try Instinct. I know. I know. Everybody's rate. I'm very excited. I'm on. And I'm not an investor. Oh, you're not. Okay. So this is like totally genuine. Okay. So I 100 % will have, I don't know when this is going to run, but I'm sure I will have played with it by the time it runs. I'm in pre-release in a couple right now that I'm spending some time with. I think the personal assistant stuff is super exciting. at least in some of my use cases, it's still showing some of the limits of like browser use as an example, like, man, we still have some work to do there. Like probably the entire internet needs just like a CLI for their product.

57:40Um, so we still have like, there's still some blocking and tackling at the infrastructure level for these things to be totally awesome. Um, and then, and then I kind of look like your average, probably, you know, AI pilled knowledge worker on, you know, every day I'm asking one of five different AI systems, you know, 20 to 30 questions, on just like doing research, looking for talent, looking for what is a competitor doing? What's happening in this market? How do you expand there? So to kind of that. What about on the company building side? How do you, you know, 20 year old company at this point, what's your advice for other people that are trying to reinvent their companies to make sure they have, you know, max adoption of AI, not just at the individual level, but also at the kind of company level?

58:21How do you make your business legible for AI? Yeah. Again, some of this we have as a byproduct of how we've always thought about information systems in the company. So like, no exaggeration, if you have a question that you would like to ask about the business that has ever been documented in a form of unstructured data. So a meeting note, a project plan, a roadmap, a presentation, a financial document, a financial planning session. It's 100 % in box. So like we benefit from a data architecture that already is insanely sort of tuned for because this is how we've run the company. We didn't let anybody use anything else.

59:00And so the data is very easy for us to be able to work with at scale. And then, of course, we have Salesforce and all the other kind of core canonical systems. We've been able to kind of have, I think, pretty good data hygiene. So agents running on top of that makes it a little bit easier to be kind of AI first in how we operate. Maybe a couple best practices or things that we've seen. First of all, trying to figure out where the highest leverage impact workflows are going to be and trying to kind of target those. So our CIO is very AI-pilled. We have a little bit of like a center of excellence on AI.

59:31We've hired some internal AI FDEs to kind of help with these processes. Do you have leaderboards? We don't token max. we do actually have like I mean like I guess like literally we have a list of people buy a number of tokens but but it's usually to like go and inspect like okay do we do we think that's useful or is there a learning there that we should take back to some other function so like probably our top you know in the top three all AI users at box like probably one of them is wasting you know half the tokens and then two of them are like oh shit whatever they're doing we need to go like do like an internal training session for everybody else.

1:00:06So, you know, we had this thing like two weeks ago where I was like, can you just get everybody in a room, you know, on this particular team and just like, like show them how this one person is using AI. And then like, you know, six hours later, they were like in a room and the guy was doing like a full demo of what he was doing. Shout out to Mick. And so like, that's the kind of stuff that we're trying to do, which is just like, how do we show everybody what it looks like to work in this way? But, you know, for as fast as we're moving and we are shipping at some parts of the stack two or three times more sort of, you know, actual customer facing products.

1:00:41So, like, I don't care about how much code, but like, did we actually deliver more functionality that customers asking for? Some parts of the stack, we're doing that. Then you'll go talk to a friend in Anthropic and they're just like, oh, my God, we still have we still have ways to go. So like, you know, the particular meta kind of constantly changes, which is like, you know, two years ago, you'd be like, you're just using a plugin in your IDE. And you're like, man, that's we're not going to be that like we're not ready to fully be first. And then you're like, OK, everybody roll out, roll out cursor.

1:01:09And then you're everybody rolls out cursor. And then you're like and then that finally happens. And then you're like, you know, you're walking around. You're like, you're not only working from Slack, just at mentioning bots doing your code. Like, what are you doing? Like, it's like we're constantly just changing what are the workflow paradigms on this. Do you guys have like a Slack coworker agent, for lack of a better term? We have a few that have that shape. I'm pretty excited about like Cloud Tag as like a form factor. You know, you still have to, again, kind of get the team construct right and the data right.

1:01:39But we have a variety of ways that people, you know, kind of work with agents in Slack. I don't know if it's as sort of Slack-pilled as Benioff would like us to be or as like arcanthropic or opening eyes but like we're heading in that direction yeah i think that history books will be written about you know the art of business in this time like i would love to read the new the art of war yeah with like everything that is happening here because i think it's pretty extraordinary stuff we're seeing what do you think it takes to win an ai versus pre-ai and like what does it feel like to be a founder right now versus when you started books i am both jealous of and also not jealous of like like the young founders you meet because you're just like like ought to be young again and and the whole world is your oyster and you can go in you know any direction and the leverage you have you'll meet with them and they're and you're like oh my god like i saw a product a week and a half ago and they did a demo of the product and i was like i was like absolutely this would have been a 40 person project five years ago like or you know especially when we were starting out easily 40 person project and it was it was two people and you're just like how do you have so many tabs that work and they all seem to have stuff behind the tabs that all seem very functional.

1:02:47Like this is not fake. Um, and, and so I'm very jealous of that. Like, it's like incredible because you could just like, you can start your company from scratch with that as the design principle. Now we will, we will get there because we're just going to muscle through it. There's a couple of things that we just can't do, which is like, we're not like, we're very uncomfortable at the idea of like sort of, you know, removing the code review and like some of these things that get talked about because our customers can't possibly entrust us with their data security and compliance if we don't take that seriously.

1:03:12So we're always going to have a little bit of a discount on the productivity because of where we are in the stack, what we do as a business. But so jealous of being able to kind of be fresh in that. And then on the other hand, it's like, man, at the same time, for every great idea, it's like instantly five competitors. We didn't have that problem. We had a good kind of couple years where we just could like, we could just grind on our product and our sort of experience. And it wasn't like every three days you were like, oh, like Sequoia funded this thing and And benchmark funded this thing. And we weren't going kind of nuts with that.

1:03:46Now, we had our own version of that. So at the time, I probably was going nuts. But in retrospect, it was not worthy of going nuts. Now it's like, oh, man, this is a real race in every one of these markets. Totally. So I think I'm probably pretty consensus on this, which is in a world where AI builds things so much faster, then probably the shift moves to whoever can actually get it to the customer is in the best position. so um you know it's it's fun talking to founders that are pretty kind of pilled on that um you know scott or matan are like they they get the mandate they're just like this thing is going to be an enterprise diffusion play and and so you have to get it to the enterprise um and and anybody who kind of mistakes the the mandate right now you're just going to lose it's just like it's game over sorry like there is quite literally trillion to trillions up for grab at the applied layer and And the companies that know how to build the teams and get to the enterprise will be the ones that win.

1:04:44It's just like obviously guaranteed. Well said. Aaron, this is a very fun conversation. Thank you so much for joining. Thanks for having me.

1:05:09Thank you.

From the publisher

Starting a company is hard. Reinventing your company for AI as a public company with quarterly earnings results is even harder. Aaron Levie has pulled off the transition with Box and offers hard-won advice for founders. The cofounder and CEO of Box argues the value isn't only in the model; it's in the bridge from a model's raw capability to the actual workflow inside a bank, a law firm, or a pharma company. That's the case for the application layer, and Box is building it: an agent harness tuned so tightly to its own file system, permissions, and search that it beats handing the raw API to Claude or ChatGPT on both accuracy and latency. Aaron explains why token subsidies from the labs can't last, why you want a model-agnostic company routing your tokens rather than the one selling them, and why coding diffused fast while the rest of knowledge work won't. (There's no "give us your GitHub" for a sales rep.) His prediction: within five years, 90% of enterprise tokens go to work no human user ever initiated.

More from Training Data

All 110 episodes
Box's Aaron Levie: On Reinventing Yourself in the AI Age and Enterprise DiffusionTraining Data · 1 h 5 min
Listen in VO