In short
The episode explains why enterprise AI spending often fails to produce durable business value, even when individuals adopt coding assistants. It argues enterprise AI breaks down across three layers: (1) model/benchmark mismatch with real enterprise needs and unclear “north star” evaluation targets, (2) system integration gaps (authentication, APIs, schemas, context engineering, fragmented multimodal data), and (3) human/organizational adoption failures (lack of leadership ROI alignment, no change management, treating AI as magic). It highlights that only about 6% of large enterprises succeed by moving from pilots to production with meaningful workflow integration.
Guests
Emily Hsu, head of enterprise AI at Scale AI; previously a Google researcher/engineering lead for Vertex GenAI tuning/evaluation and an early Google Brain member; also a Vanderbilt professor. Kevin Ball (KBall), VP Engineering at Mento; independent engineering coach; co-founded/CTO of two companies; organizes AI in Action via Latent Space.
Key claims/examples
Frontier benchmarks miss operational correctness; enterprise data mess can’t be “sprinkled AI” onto. Example: Mayo Clinic patients bring hundreds of pages of non-standard PDF records despite FHIR. Success patterns: fix data foundation (including entity resolution), front-load leadership sponsorship/change management, and combine internal domain expertise with external AI specialists. Evaluation: test agent behavior across environment/tooling state, happy/edge/negative cases, and robustness; use anonymized/expired data proxies when real enterprise data can’t be shared.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe AI Spending Gap Overview
0:00 to 0:45
Learn about the disparity between AI investment and its realized value.
“It is widely reported that a gap has emerged between enterprise spending on AI and the durable value captured from that spend.”
Introduction to Emily Hsu
0:45 to 1:57
Discover Emily Hsu's background and her role at Scale AI.
“gives the company a rare view of why enterprise AI may be stalling.”
Emily's Career Journey
1:57 to 4:23
Explore Emily's journey from academia to AI leadership at Scale AI.
“Thank you so much for inviting me, Kevin.”
Overview of Scale AI's Services
4:23 to 6:13
Understand the broad range of services Scale AI offers beyond data labeling.
“When I was doing the research for this call, I initially was like, oh, I've heard of scale.”
Challenges in Enterprise AI Adoption
6:13 to 6:41
Examine the unique challenges faced when adopting AI in large enterprises.
“I think we're all trying to figure out where are the value layers in this new AI world.”
Two-Sided AI Adoption Problems
6:41 to 12:20
Learn about the dual nature of challenges in AI adoption, focusing on technology and people.
“Like what is different about adopting AI in a large enterprise and what are the classic failure modes?”
Complexity of Enterprise AI Needs
12:20 to 14:00
Delve into the complexity of AI needs across various enterprise sectors.
“So tech layer, there's some gaps, absolutely.”
Challenges of Digitalization in Healthcare
14:00 to 15:32
Discusses the complexities of digitizing medical records and patient data.
“But then let's look at in the past two decades, when people are talking about digitalization, like everything from paper to digital form, digital asset, there's a lot of technical depth actually in that wave.”
Challenges of Digitalization in Healthcare
15:40 to 16:47
Discusses the complexities of digitizing medical records and patient data.
“Like every shortcoming that we all have that a human has just said, oh, you know, that's not perfect, but I'll paper over it.”
Challenges of Digitalization in Healthcare
17:07 to 18:04
Discusses the complexities of digitizing medical records and patient data.
“Start building for free today at xweather.com.”
Show all 22 chapters
Challenges of Digitalization in Healthcare
18:06 to 18:39
Discusses the complexities of digitizing medical records and patient data.
“which means GitHub Actions is running more than ever.”
Leadership and AI Integration
18:59 to 28:00
Explores the necessity of leadership commitment for successful AI adoption.
“If the leadership does not recognize the value of the ROI, and then when we actually get into all these barriers, then people could easily sort of, will not have a clear actually success criteria.”
Challenges in AI Adoption
28:00 to 29:09
Learn about the critical human factors in AI adoption and success.
“Yeah, so that's one of the first things we noticed The second one, as we also chatted about, is the human part of it.”
Buy vs. Build in AI Solutions
29:09 to 32:09
Explore the complexities of whether to buy or build AI solutions in enterprises.
“Versus the successful mode of like, hey, this is something new and tricky, and we're going to work together and have change management and co-develop.”
Evaluating AI Models
32:09 to 38:48
Understand the evaluation criteria and methods for AI models in workflows.
“So one of the things you mentioned there that I'd love to dig into, because I think you have done a fair amount of work in this space, is around evaluation.”
Creating Realistic Evaluation Environments
38:48 to 42:05
Discover strategies for establishing realistic and reproducible data environments for AI evaluations.
“You start data vendors and we're one of them, right?”
The Role of Forward Deployed Engineers in AI
42:05 to 43:34
Learn how forward deployed engineers contribute to AI deployment and co-evolution.
“But really then is the details are implemented either by the company that it's sold to or by forward deployed engineers.”
Co-Evolution of AI Agents and Human Interaction
43:34 to 45:30
Explore how AI agents learn and adapt within enterprise environments alongside human users.
“Because the agents become more capable, but then what people are trying to do with them evolves and kind of what that looks like.”
Liability and Responsibility in AI Decisions
45:30 to 48:26
Understand the importance of human oversight and responsibility in AI decision-making.
“I think there are two dimensions of things.”
Human-AI Cognitive Interface and Trust
48:26 to 50:38
Discuss the cognitive interface between AI and humans and how trust can be established.
“Well, and that fits with what we were talking about in terms of our lack of a truly high fidelity eval environment is you got to put it into the actual production environment and see how it behaves.”
Improving AI Confidence and Decision-Making
50:38 to 52:55
Examine how to enhance AI models' confidence levels for better decision-making.
“And that kind of like calibration capability is something that people start to look very much deeper into and it's still kind of like open research at this moment.”
Keys to Successful Enterprise AI Adoption
52:55 to 54:13
Learn the critical factors for successful AI adoption within enterprises.
“If there's a single point, it actually, I think, is boiled into get the right pilot use case.”
Transcript
Automatic transcript. May contain errors.0:00It is widely reported that a gap has emerged between enterprise spending on AI and the durable value captured from that spend. Individual employees have enthusiastically adopted coding assistants and chatbots, yet those gains do not seem to be transforming businesses at an organizational level. One of the most important questions in the tech industry today is understanding why AI is not yet delivering returns that match the investment, and what separates the small number of enterprises succeeding from the many that are not. Scale AI is known for supplying the human-labeled data behind many frontier models.
0:39It now also builds AI applications and agents for large enterprises. That combination of working alongside frontier labs and inside enterprise deployments gives the company a rare view of why enterprise AI may be stalling. Emily Hsu is the head of enterprise AI at Scale AI and previously spent over a decade at Google. In this episode, Emily joins Kevin Ball to discuss the three layers where enterprise AI breaks down, why frontier model benchmarks miss what enterprises actually need, the data foundation problem, how the most successful companies combine internal domain expertise with outside AI specialists, and more.
1:22Kevin Ball, or KBall, is the Vice President of Engineering at Mento and an independent coach for engineers and engineering leaders. He co-founded and served as CTO for two companies, founded the San Diego JavaScript Meetup, and organizes the AI in Action discussion group through Latent Space. Check out the show notes to follow KBall on Twitter or LinkedIn, or visit his website, kball.llc.
1:57Emily, welcome to the show. Thank you so much for inviting me, Kevin. It's my pleasure to be here. Yeah, I'm really excited to dig in with you on this. Let's start with a bit of your background. So can you kind of give us a little bit of your career to date and how you ended up at Scale? Yeah, I absolutely love to share that. A lot of people were saying like my career is kind of a reverse of many people. So currently, I'm the head of enterprise AI and Scale AI. It's kind of a company, a startup company, but it's not like as small as a startup anymore, that people kind of note from the perspective of supplying the data to many of the frontier labs and being an active participation to the Gen AI large-language model to where we are.
2:39So it's very interesting work from both understanding the customer requirements on the enterprise side, as well as driving the innovations to research for pushing enterprise AI to land with revenues. Then before that, I was actually working for a really large company, Google, one of the hyperscalers, and also the Frontier Labs, two-in-one. I worked there for 11 years, been very hands-on engineers. I'm also early member of Google Brain team. So I was a researcher, hands-on model training, publishing papers. And I was engineering lead for several of the cloud AI product, Vertex product, tuning, Gen AI tuning, evaluation agent engine with my team members who really start with the concept, the idea and see the market needs for it.
3:25and then put our hands down there, build up the product and bring it to the customer. So it's a really fruitful experience working with the team, put up a team to bring this to the market to see what the enterprise adoption for this platform products are. So that's my experience at Google for about 11 years. But actually, before I joined Google, I was a professor. I got my tenure from Vanderbilt University. So my research is kind of a combination of like applying a lot of the theoretical analysis from optimization and statistical learning in understanding sort of networking system, a network and large-scale distribute system, and with the application to several domains, separate physical system, healthcare system.
4:08After I got tenure from Vanderbilt, I spent my sabbatical years at Google and decided to stay with the industry. So that's kind of like me. Yeah. That's awesome. Well, I'd love to dig into a few different pieces of that, but let's start also with a quick overview of scale because you mentioned scale. When I was doing the research for this call, I initially was like, oh, I've heard of scale. They do data labeling, right? But y 'all do a lot more than data labeling now, don't you? What is the span of what scale is tackling? Absolutely. This is a great question and segue. So people really know about scale.
4:40And this is also how scale actually incorporated as a company is really to bring data to machine learning labs and the machine learning companies as a data provider, starting from providing data for image, like vision perception models. And later on, the market is actually expanding from the vision models and then getting to language models. So this is where we actually get into partnership with OpenAI in very early days, providing human annotations, human feedback data to make GPT where they are now. So that's a core business. But meanwhile, at scale, and also I think the broader market also realized that beyond the foundation model capabilities, the landing of the value of all these models really depends on building reliable applications and agents that powers the applications.
5:30So that's actually another big chunk of business for scale. And within that sort of the application business, we'll have two business units, depending on who are the customers we're speaking to. So I'm currently leading the enterprise AI side where I'm in the enterprise business unit, where our customers are private sectors, 500 sort of enterprises. And we'll also have another business unit where it's public sector facing US and the global governments. But the goal is kind of like really building applications and solutions on top of leveraging the intellectual capabilities for models to serve the business needs.
6:11Yeah. Well, and that makes a ton of sense. I think we're all trying to figure out where are the value layers in this new AI world. And it definitely looks like it keeps moving up the stack. Well, I'd love to really dive into the enterprise side. I think there's a lot of interesting things there. And, you know, I work at a relatively small startup. The adoption challenges in a startup world are very different than they are in the enterprise world. I think you've actually written a bunch about the different failure modes of enterprise AI and things like that. So let's maybe start with that. Like what is different about adopting AI in a large enterprise and what are the classic failure modes?
6:48Yeah, absolutely. So this is very insightful question. And there's actually many kind of directions people are thinking about it. Right. So maybe what we can do is I would think about like to start with is it's a problem of two sides. On one side is on the machine learning AI capability side. Like, you know, people usually say, think about from sort of AI perspective. On the other side is actually from the platform system perspective. So in order to have system solution application that truly serves the enterprise customer needs, both sides have to meet with each other and be successful. And then on top of the technology layer, that is also adoption layer, which is the management side.
7:34It also touches the people side of the enterprise usage. So it involves education, leadership, routing, and all the stuff. So I just want to sort of maybe anchor this discussion on this people versus technology. And within technology, there's the AI part of it. There's also the system part of it. So maybe we can start with on the AI side of it a little bit. So if you think about AI side, the market, the community have a lot of exposure to what foundation model Frontier Labs are doing. I come from Google. Actually, I'm part of the Gemini team. I've been going through Gemini 1.0 all the way to 2.5.
8:13So if I look from the insider lens when we develop Gemini, right? So if you look at it, the development team actually have very limited exposure to what enterprise needs are. And this is exactly what I feel, you know, even from Cloud Ad, enterprise part of business of Google. Fundamentally, every model iteration, which we call data flywheel, understanding where the model gaps are and how to contribute data, train it. the whole engineering practice is pretty much owned by the engineering team who actually do not have the business exposure. And that is reflected as, right? So people are very benchmark driven.
8:56They have very focused goals, say, hey, now we actually have different capability pillars, math, reasoning, code, conversational, multi-turn safety. And then if you look at that, The enterprise part is becoming a very narrow pillar in all the frontier labs leaderboard progressively into 2025. So which means the model knows a lot of the human intellectual capability and knowledge, many of them is reflected by our flagship benchmark humanity last exam. But in terms of concretely, what is the professional requirement? What is the operational workflow? What are the policies that is actually super critical in the enterprise domain?
9:43There's no evaluation criteria. There's no intuition knowledge within the people who are actually in the mobile. So then from AI perspective, the fundamental capability has a gap from what it means. Well, it sounds like not only a capability gap, but even a goal or target gap. Exactly. It's not clear where it needs to evolve. Absolutely. So we need to set up the North Star. And the enterprise, it's a single word, but it's not as simple as math and code. Math and code are the true kind of like, we have a lot of breakthrough in terms of intellectual capability advancement in the past. But if you think about math and code are the things that majority of the people who hands on the code truly understand.
10:32right there are also there are also domains in which it's if not trivial much easier to define what correct looks like so you can build an optimization function and have an iteration loop and things like that absolutely right this is where the reinforced learning with verifiable reward is actually bringing a lot of capability advance in the last i think last year let's make the time flies. Yeah. So this is - I know, right? A year these days is like five years. Exactly. I think it was starting from the early of 2025, where we're thinking about reasoning models. So now I'm thinking about this like one year and a half.
11:11Yeah. And you're absolutely right. Having a precise definition on what is correct, what is right, actually bring a lot of benefit in the post-training process with R. But then in the enterprise domain, and back to what are discussing about. If you think about it, it's a single word, but it's even hard to define what it is. There's so many verticals, right? And then let's say healthcare. Now within healthcare, to think about there's primary providers, medical centers, there's insurance companies, there's pharmaceutical life science companies. Even with health providers, there's big academic medical centers, there's community healthcare service.
11:51So they have very diverse needs regarding what they're going to use AI for and what they regard as acceptable AI service are. So that's kind of like the complexity of the enterprise world. And the complexity is not very well captured and presented to the lab's developers. That's one of the root cores for where the capability gaps are. Yeah, that makes a lot of sense. So tech layer, there's some gaps, absolutely. But let's not leave it there. Let's look at the other layers too. The other layers, yeah. So that's kind of like starting from the foundation model, but then we also talk about the system side, right?
12:34So not just the models. When you develop an application, it doesn't matter big or small solutions or reusable applications agents, you need to have it to be integrated with the enterprise context environment. You need to understand what data are, how to get yourself authenticated. When we say yourself is the agent themselves, to get authenticated to the right data. And also to understand for all the enterprise services, their APIs, their data schemas behind it, what do they mean, right? With different table names, with different value names, how do we actually understand the context? A lot of cases people talk about context engineering.
13:18How we actually learn those tribe knowledge within the enterprise and bring that knowledge to be part of the context into the agent. And that is actually a non-trivial problem, right? So because it's not like all the information is right there presented to you. the information. Absolutely. Like I'm dealing with exactly this in a startup environment where we have like, you know, a handful of systems and a few different databases. You look at an enterprise that's been around for decades and has hundreds of thousands of people, like this is a mammoth problem. Totally. And then you hit another point in terms of the pain point, what we see is when we think about enterprise and then we talk about like, this is a new era of AI waves, But then let's look at in the past two decades, when people are talking about digitalization, like everything from paper to digital form, digital asset, there's a lot of technical depth actually in that wave.
14:19Because during that wave, the entity who actually are the information receivers are human. Humans are extremely flexible, and they are able to tolerate, like, you know, give me a PDF, give me a handwritten note. It's not like perfectly digitized in a way that is to the high standard, but I can still get it. This is actually like the problem we get, for example, from the engagement with Mayo Clinic, right? We think right now all the electronic back-to-record system are digitized, right? it confirms a uniform standard called FHIR. But the reality is there are always patients who are not actually within the network.
15:02They go through another system. They come to you with very complicated medical history. They're not going to give you the standardized, normalized electronic medical records. They will give you hundreds of pages of their historical medical records printed in PDF. So how people can read it, but then in order to actually ingest the data from different modality with the technical depth from the last two decades, fragmented, received, processed by human beings in the past, now you have to plug in an agent that is powered by foundation of large launched models, then the gap is actually emerged from there.
15:44Yeah. Yeah, absolutely. Like every shortcoming that we all have that a human has just said, oh, you know, that's not perfect, but I'll paper over it. Now we have to go back and revisit and say, well, can an agent paper over it? Maybe not. What do we do with that? Absolutely. That's kind of like one of the things we experience, how we actually integrate an application, and agent into enterprise environment really address all these gaps from fragmented information, tribe knowledge. So all the things become the barrier actually prevent accelerated deployment of AI solutions in enterprise. You're building agents that can write code, summarize documents, and automate workflows, but they're missing one thing, awareness of the world around them.
16:30X-Weather combines enterprise-grade weather intelligence with agent-ready APIs, natural language capabilities, and an MCP server built for tools like CLAWD, Codex, Copilot, and modern IDEs, so your agents can adapt workflows, automate responses, and make better decisions based on real-world conditions. Backed by Visilla, whose instruments fly on NASA missions to Mars, X-Weather delivers trusted data and unique insights that go beyond conditions to actual impact, from real-time lightning strikes to road surface forecasts. Start with 15 ,000 free API calls every month and pay only for what you use as you grow.
17:03Your full weather stack for developers by developers. Start building for free today at xweather.com. Your customer lives in the real world and it's messy. Intermittent connections, mid-onboarding drop-offs, edge cases on devices you've never tested. Mobile apps reflect reality in a way no other surface does. Yet, from an engineering perspective, they're the hardest to understand. It's common for mobile engineers to see green backend dashboards, normal error rates, no crashes, but inevitably somewhere there's a frustrated user watching your app spin. After a few seconds, they'll lose patience, close the app, and turn their attention somewhere else.
17:41They may never come back. The worst part? Most observability tools never see any of it. BitRift, on the other hand, captures 100 % of mobile data, unsampled and in real time, so it's immediately queryable by engineers and AI agents. It's mobile observability built for the real world. Try BitDrift today. bitdrift.io slash signup. This episode of Software Engineering Daily is brought to you by Warp Build. AI is writing more code than ever, which means GitHub Actions is running more than ever. Your GitHub Actions bill is now a function of how much AI code you generate. And every engineer knows the feeling.
18:15You push a commit, and then you wait. Warp Build makes GitHub Actions twice as fast at half the cost with a one-line change to your workflow. Linux, macOS, and Windows Runners in Warp Builds Cloud or your own, enterprise-ready, SOC 2 Type 2 attested, and trusted by teams like Sky from Comcast, Bitcoin, and Braintrust AI. Get started with$50 in free credits at warpbuild.com. Well, and this kind of then ties into this last piece, which is the human layer. Like what needs to adapt in terms of how we're engaging in work? Absolutely. So this is like at the lower layer of the technical issues, but the solution to the technical problems really lies in the human's hands.
18:59And this is where, from organization perspective, the leadership commitment, the leadership, because all these motions, all these AI motions, shows us it's not going to be successful if there's not a clear incentive on what is the benefit, what is the return, is the ROI. If the leadership does not recognize the value of the ROI, and then when we actually get into all these barriers, then people could easily sort of, will not have a clear actually success criteria. People could easily pause at impressive demo and not pushing forward because there's no incentive to show the fundamental return and the impact, economical impact to the solution, right?
19:46So when there is a, it's just drop it in the middle of nowhere and then not get into productionization. So from this perspective, the most important thing is the alignment at a high level on a leadership level. what is the we're aligned on the north dot successful metric what is the expected output of it and then driven by that we would have a clear technical solutions technical roadmap to say now in order to achieve this high level north dot success metric what other things we need to do what is the path you know what's the low-hunging fruit we need to do in the first step to get into the system foundation to be developed?
20:33And what's the next step to get to the user engagement? How we actually collect the feedback continuously from the usage to improve the system from the agents. So those are the things that actually defining the driving force to be successful. Yeah. Having leadership agreeing on where to go definitely seems important. Another thing that I'm pondering about, like one of the things that I've seen and I've heard a lot of people struggling with is, at least in the coding space, which is the one that I and I think much of our audience are most familiar with, these tools are really focused on individual productivity.
21:08Yes. But if you just drop a bunch of individual productivity improvements into an organization without rethinking about your organization, you often end up just bottlenecked on the next layer that didn't speed up the same amount. And you don't see nearly the level of organizational improvement that you expect based on the individual. How is that showing up in these large enterprises in terms of change management or reorganization or rethinking of how we organize ourselves? Yeah, this is such a great point. It just kind of really resonates what we see in the field as well. If you go to a lot of the enterprises and you ask them the AI adoptions, the answer could really go two directions depending on how people think about AI adoptions.
21:51If we say the people who are using cloud code, you know, GPT, and you're using all these individual AI tools, your adoption is amazingly high. Everybody is using it, right? But they're not using it in the context of like exactly what you captured in contributing in a coordinated enterprise workflow. They just use it to improve individual productivities. And then to do that, there's actually two major problems, kind of very related. The first one is there could be a lot of IP leakage. Because when you interact with an individual tool, all the queries you have, all the questions you have, there are a lot of enterprise confidential information.
22:36And even if it's not proprietary data, there's also the way you ask questions and then ask follow-up questions, review a lot of the judgment and the implicit knowledge within people's mine. And those information is actually leaked into the frontier labs. So that's one side of the problem. But on the other side of it is actually these are very valuable enterprise digital assets, the IPs. And when people are actually getting used to just interacting with the individual productivity tools, all this digital asset, all this implicit knowledge are actually not captured and shared across the organization.
23:24This is exactly the point you're talking about. In many cases, when you work on a project and then when you have experience and the context, the skills, all this expert knowledge from an enterprise perspective, this information is to be consolidated across multiple team members and then into a coherent knowledge base so that the next generation of the people who are working on the project will be inheriting the knowledge, the experience being accumulated from the previous conversation. So those are the things that are actually missing by building this. So that kind of brings us to the flip side.
24:06We've talked about where this breaks down, where the gaps are in different ways. What does success look like? I know you all recently published this report and you said, I think it was 6 % of large enterprises are really right now succeeding at integrating AI and seeing the benefits. So what are the patterns? What are the ways that they get around these gaps? How does that end up looking? Yeah, so when we have the 6 % sort of the reports, the first thing we think about is how we define success, right? So in this domain, defining success is actually not easy as well because as we said, if you ask them about AI adoptions, they could say some shallow adoptions.
24:45and that's not success, right? So the first thing what we did is we really think about what is the success criterias are. And then for that, we basically have a pretty rigorous way to say that. It's kind of like how people among all the projects people have been engaged in, how many of them actually go from pilot into production and then meaningfully integrated with the workflow. and meaningfully integrated with the current enterprise workflow that is measured by very concrete usage metrics. So those are kind of like the things we're looking at how we think about the success in a very concise way.
25:27We actually look at the characteristics to see what are the characteristics the 6 % winners actually share. We are not able to attribute to say, hey, these are the factors that are causal factors. If you did this, you'll be successful. We're kind of like doing machine learning. We try to capture what are the shared attributes about all this 6%. Have you ever been able to make a true causal argument in a human system? Like that's very, very hard. But yeah, I think that is useful, right? The pattern matching is what we're all looking for. It's like, okay, what are the patterns that I can try this?
26:02And maybe that will help move my organization closer. Yes, that's exactly the sentiment that I can, you know, what we're doing. So there are a lot of findings in the report, but the three things kind of really strongly reflected from all the statistics and the study we have. The first one actually is not surprising, is the data foundation. And this is actually we've talked about how data is fragmented, multimodal data, connecting data actually to understand the data access policy. What are the governance or how we actually collect the feedbacks from the field? All this data infrastructure foundation.
26:39is actually a prior, kind of like very strongly related with the success of AI, search deployment success. No surprise, right? It's not surprising, but I feel like this is a thing that comes up and there's memes that have gone around, right? Of like, oh, we have this huge data mess. We could solve that. Or we could just sprinkle AI on top and hopefully it will work. And it all comes back to, no, you actually have to solve your data mess. It's a chicken egg problem, right? If you don't have the data, the AI will not do anything. But then when the AI is actually getting to the enterprise in the right way, then AI could be helping you with, for example, one of the key technologies we're actually providing to the market is, say, entity resolution.
27:20So you have different tables. The keys, the way you name corporate entities are different in different tables. In the past, this is a very traditional machine learning problem that has been there for a long time. But now with the large language model, with a lot of priorly trained work knowledge, they can do this job really better. So now if you bring AI to solve the problem, you have a much better normalized, standardized linked data asset. Then the next layer of AI agent is able to build on top of it and then discover and unleash more capabilities of solutions on top of it. So it's an iterative process, but you need to have the foundation together first.
Read the full transcript
28:01Yeah. Yeah, so that's one of the first things we noticed The second one, as we also chatted about, is the human part of it. Usually, if there's early stage investment in terms of senior leadership sponsorship, and then they are committed in change of management, they provide resources in employee engagement training. Because personal usage versus enterprise workflow AI, a lot of cases have different steps, different UIs, if they actually would get into the co-developed product, co-developed motion, and then envision when AI is co-pilot, work with them in a workflow environment, then the success of adoption is actually much increased versus initially we don't actually imagine what a change of management look like when the AI agents are being deployed.
28:59So if the enterprise front load, all the organization work, that is actually something we noticed among the 6 % winners. No, I think that's really important because, you know, I talked to a lot of people in a lot of different environments here, and it feels like one of the most common failure modes I hear about is essentially treating this as a technology and sort of magic adoption thing. Everybody must use AI. It's a magic technology. It will make everything better. Just do it. You figure it out. Versus the successful mode of like, hey, this is something new and tricky, and we're going to work together and have change management and co-develop.
29:38and really refine how we're working with this. Absolutely. And this is actually a great segue to the third point we actually discovered, which is like on the market right now, if you talk to enterprise customers, right? There's always these questions about buy versus build. And many cases, if they ask you, do you have a product that we can buy immediately, right? And on the other side, if you get deeper with them, they're thinking, oh, we have a pretty strong engineering team. Can we build it together, right? Build internally. But the answer to the question is actually not that simple, as you pointed out, right?
30:13So it's really we need to bring the expertise from both sides. If you just get a product, even with like a customization consulting team to get to the enterprise side, we still face the problem how to integrate a product with the workflow, with the enterprise context, right? And then how the product is actually going to be used. What's the CUJ look like? What are the touch point for customization? So that's kind of like one. On the other side, if they think, hey, we have the internal team, let's just build it. It's really going to fit very closely to our internal usage pattern, the workflow pattern.
30:58Then what is missing is the internal team. Of course, the Urali is very strong, genealogy with many years of seasoned experience. They usually don't have the broader understanding about the AI technologies that actually one of the benefits when we work very closely with the frontier labs, we actually have a best practice in terms of how we do evaluation of the foundation models. What are the best patterns in developing agents around on top of these models? And what are the typical failure modes as well? So when we actually see the 6 % winners, they strategically combine the internal expertise, which is that domain knowledge, in terms of when this is useful, how to use it, and what outcomes are actually going to change the workflow and decisions.
31:48with the specialized partners who actually bring in the expertise regarding AI models, AI application development, and system development. So when the expertise from both sides are combined together, this actually creates the motion and the foundation to be successful. Yeah. So one of the things you mentioned there that I'd love to dig into, because I think you have done a fair amount of work in this space, is around evaluation. right and i think one of the things that i definitely see hiring engineers and looking for folks to work on ai products and things around this is one of the biggest gaps is people who who can grapple with non-determinism and those who struggle with it right and the fact that in a non-deterministic environment you can't do unit tests you have to have some other way of validating quality and understanding what's within your sort of trained scape scope or the training corpus or the nice fit of the model and what isn't and all that sort of thing.
32:43So how do you all approach evaluation? And particularly, I think there's a lot of people thinking about evals at the prompt level. How do you think about evals at an agentic or workflow level? Wow, this is my favorite topic. And then it's a list of questions that worth the time to dive in a bit. So in a very simplistic term, right, if we think about evaluation, there are two things about evaluation. The first one is what needs to be evaluated, right? And then as you correct point out, like in the foundation model development, what we evaluate is the model, whether the model is going to work, right?
33:22And then that's the subject to be. And then in order to evaluate, there are two things surrounding it we need to do. One is how the model behaves and the expected input. This is actually super important. When we do model development, and you test it under a very broad spectrum of queries. And the queries, of course, will be sliced into different pillars, some are coding, some are conversation, and some are adversarial for risk assessment. And then the last part of it is when the output, doesn't matter from the model agent, is produced, how are we actually going to tell is it good or is it actually less desired?
34:01And in the undeterministic case, as people recognize, there's another aspect of the robustness. Can you actually reliably, robustly to give the consistent output when either you repeat your task or you have small perturbation in the input? So these are kind of the three ingredients of it in the context of agent evaluation. So the first one is actually the agent self, also the environment you're actually operating in, right? So it's not just like what you develop the agent intelligence self and build on top of the model. It is also what is the backend environment? What are the tools that this agency is connecting to?
34:48What are the service APIs that the agent has access to? And more importantly, what is the backend state? How many databases you are actually connected to? What are the file systems? What are the knowledge base that enterprise? So this actually defines the environment during evaluation, right? Because if you change that, the evaluation result could change, and that's the robustness question we need to get into. So people need to understand what are the environment, what is the entity we are evaluating. And then the second part of it is, as I said, it's kind of software testing that you need to define your test cases to cover the happy path, the typical, the edge cases.
35:30So we need to strategically define what is the set of tasks that cover different dimensions of the capability and the functionality of the expected behavior of your agent, right? So just give you an example. Let's say we are doing patient safety triage agent. And then the goal of the agency is going to look at a report to say, is this an event I should report? And then what is the policy term that I need to ground my decision to? So this is the agent task. And now when you get into the evaluation part, the high level mindset is actually very similar to what software unit testing is doing. Then you need to think about how do you design your unit test to cover different aspects of it.
36:26Then you need to think about there's the positive case. The event is true. But then you also need to cover different policies. like, you know, and the worst scenario, because if you look at the event safety guidance, there could be 1 ,000, well, that's a little bit too high, like 100 different criterias. Then you need to make sure your query actually hit the different categories of the events so that you don't have some missing edge cases where things you don't have the confidence to speak about. And of course, you also need to think about the negative case and the hard negative case. You need to think about what if there's ambiguity where you're missing evidence?
37:11What if there are terms that is not very well specified in the guidance? So you need to think about all these things, cases, and those are the evaluation, data set, curation. I totally agree. I want to dig in a little bit more about the environmental piece, because I think looking at building out these types of eval frameworks in different environments, one of the biggest challenges that I have encountered is how do you get a realistic, but also like reproducible data environment, right? You don't necessarily want to run this against your prod environment, or maybe you do, maybe that's right. But like, what does that whole setup look like?
37:47I think this is the open question, to be very honest. And this is actually a research direction we're working very hard. And I'll actually have some result that to be shared with everybody coming up and hopefully we'll talk to you again when it's out. But you're actually asking a real hard question. The problem right now is if you look at across the industry at this moment, there's very limited in a conservative, but it's not known, evaluation environment that is backed up by real data. Because the real data from enterprise, it involves a lot of the privacy, legal usage concerns. But if you don't use real data, then all the dependencies, the very complex dependencies within the data will not be truthfully reflected in the environment you are developing.
38:38Right? Yes, exactly. This is exactly the problem I feel like we're all grappling with right now. We're all grappling with. Yeah. The solution I can share a little bit of like how people approach it. You start data vendors and we're one of them, right? So we're actually looking at data not actively being used, but similar in nature, in similar semantic domain usage domain, but from the past, expired data. And we do DID, we do sort of anonymization over the data so that the sensitive information will not be leaked, but the semantic relationship will still be there. So there are some lot of research in that domain.
39:15People are looking for proxies, basically. surrogate and proxies to build a similar environment. But I agree with you, this sim to real is always a gap. A simulation to real environment is always a gap in the machine learning research deployment challenges. Yeah. I am curious, actually. So you mentioned this is an area that you all are providing. What are the services that SCALE is providing here in the enterprise space? Because we've talked about some of the needs and the gaps, and you've sort of alluded to, we do co-development, we're working with people, You provide some of these like simulated data environments.
39:50Like what actually does the, if I was reading it correctly, the Gen.AI platform or the enterprise platform on scale provide for folks? Yeah, absolutely. So at very, very high level, we do three things, right? We provide platform, which we call skill generative platform. That's the foundation allowing all the system integration work, all the trustworthy agent execution, and then allowing the agent execution, all the traces to be logged and then for evaluation for auditing purposes that the platform we're providing, which provide the system infrastructure foundation to build a reliable AI solutions.
40:27So that's one pillar. Another pillar is actually our full deployment motion. So this is like, we're not just giving a platform. What we do is we actually have a team. It's a very well-constructed team, which have strategists, purely people coming from management consulting like McKinsey, BCG. they would work with the customer to understand their business North metric, and then look at the use cases that allows us to produce the impactful outcome to meet that North Star metrics. And meanwhile, we have the technical people who works very closely in the initial scoping process to say, what is technically feasible within the timeframe we have?
41:09And this we co-define, we discover the use case, co-define the use case to be worked on. Then we get into the delivery motion where the forward deployment people, the people with applied and as well as system background work together to really wire all the fibers, code all the AI solutions, perform the evaluations, understand the worst gaps are, and heal climb on the quality. So that's kind of the full deployment service motion we provide to enterprises. Yeah, it's really interesting because I feel like this is a pattern that is starting to, like if you look at our industry as a whole, we're evolving from essentially deploying software that was relatively repetitive, but in that enterprises were integrating then trying to find ways to integrate into their processes.
41:58And we're moving to a world where we're deploying a set of capabilities, which has software, But really then is the details are implemented either by the company that it's sold to or by forward deployed engineers. Like that's probably the most rapid growing software field right now is forward deployed engineers. Exactly. And then actually just on top of what you're saying, here's another dimension. When AI capabilities being deployed, it's fundamental as a learnable system, right? So it's not just like a one-shot thing. what a good for deployment engine team, like, you know, scale AI, we're not just saying here's the solution.
42:38We actually think more long-term as we previously alluded to, how we actually help you to collect all the digital asset that reflected from the usage pattern through the traces. How would actually help you to identify the quality gaps when we talk about the evaluations, right? Where things fall and then provide you the research innovation, which was the last piece I was going to talk about, to improve it progressively over time so that agent becomes smarter over the time of people using it, because it would capture human's knowledge, human skills, and then use it in the future. It would incorporate it into the memory of the agents and then help its usage moving forward.
43:28That's interesting. Can we talk more about that? Because I think there's like, coming back to our co-development, there's like a co-evolution that needs to happen there as well, right? Because the agents become more capable, but then what people are trying to do with them evolves and kind of what that looks like. So how do you think about setting that up? Yes, I think that's actually a very essential piece when we get agents into the enterprise environment because it's a living thing. And don't think like in the previous era, when we say service, you know, a SaaS kind of era, you connect it, but what you get are just data into it.
44:00But then once the agent, agents actually serve as the copilot that works with you, then agent would actually learn from like, you know, Chris talked about, it will learn your institutional best practices as your institutional knowledge base, right? In a collaborative environment, let's say we have a group of people who are working on a project in like three months. Then when we actually work together, on the same agentic system, then agent is able to actually accumulate our usage questions we had, working solutions that fails, would accumulate this project-oriented information. And then over time, it would make our project execution a lot more productive.
44:44And meanwhile, on an individual basis, you could also remember what you did yesterday, what is your usage pattern, and so on and so forth. So the agent would have components built in to learn from the interactions and become smarter and interact with people on the project's execution, workflow execution in a smoother way is actually one of the differentiations that it would bring to the customer enterprise. Fascinating. So how do you think about the evolution of human in the loop here? And what can get fully absorbed and automated versus where you have people checking in versus where you have this very co-pilot-esque approach?
45:29Absolutely. I think there are two dimensions of things. The first one is, you know, when AI gets into the enterprise, it plays a very active role from any aspect of it. But there's one thing is, who's going to be liable for the output, for the decision, right? Someone needs to be held responsible for it. You can't say this is the AI did it, But in my humble opinion, I think the human is going to play a long-lasting role in being responsible for the AI decisions and what the work from the AI is. Technically, how that's going to happen, I'll get to that in a minute. But in the long run, human stays.
46:13Human will be liable for it. What AI does would be helping AI in making decisions. in various ways. But towards the end of the day, we need to have the people who are responsible for the AI decisions in the enterprise environment. So that's kind of like a very long-term, sort of like the foundation of human, stay human in the loop. But with that being said, there's also a tactical direction. Like, you know what human does, right? So I believe that part is actually what we see is going to have a path towards initially when the trust between, like when we talk about evaluation, right? There's offline evaluation.
46:57There's also online evaluation. Initially, before an agent is being used in the production environment for a long time to gain sufficient trust, there's always going to be a human in the loop component where human will be triage. Ambiguous cases, human will be validating the high-stake operations where the system state got permutated. So right now, a lot of the writing, almost in all the applications, if you see something's been written, and then there's always asked you to confirm, right? So all this risky high-stake operations, humans need to confirm that. And then even for the insight derived by the AI agents, human needs to review it to make sure it really aligns with the business policy and the business standard, right?
47:48But over the time when the trust is built and then the products is matured in a way that we are going to significantly reduce the cognitive load of human being to a point some of the very confident answers, confident operations from the agents should be just automated. But that is a tactic roadmap to initially we have a higher human oversight to pay to get reliable, trustworthy deployment. But this cost is going to reduce gradually when the mutual trust is established. Yeah. Well, and that fits with what we were talking about in terms of our lack of a truly high fidelity eval environment is you got to put it into the actual production environment and see how it behaves.
48:38And while you're doing that sort of live action evaluation, you don't want it changing key data or making big decisions. I think there's something interesting on the sort of long-term vision there. So you highlighted liability and responsibility, which I think is true. Like that's an important piece, right? I think it's really important that whatever, and we're already seeing this in the coding world, right? People will be like, oh, the agent made this decision. Well, who said it was okay? Who committed it? Like you, you still made a decision there, even if your decision was to delegate your thinking.
49:09But I feel like, like in that framing, it feels very burdensome, right? It's like the agent's doing the work, but you have the liability and the responsibility. What's the positive side of that? What is the, what we're going to be doing as we automate more and more in terms of like this direction, decision-making exploration. Like, I think there's something on that side of the world as well. Yeah, absolutely. to your point, I feel like people are still trying to figure out what is the cognitive interface between AI and human, right? Fundamentally, AI is actually expanding human's intelligence and capability, which leads to higher productivity, as well as higher quality decisions, right?
50:00But then to your point is where, what is the interface? Where it actually helps instead of asking me all these trivial questions right now, like everybody's in cloud code or like, yes, confirm is confirmed. After a while, I got fatigue, right? And then not going to be making the right decisions anymore. So this is actually also from a research perspective, actually people are actively looking into in a sense that when the AI knows what is the point where we truly need human input, right? When he's sure about what he's doing, when the AI is not sure about what he's doing. And that kind of like calibration capability is something that people start to look very much deeper into and it's still kind of like open research at this moment.
50:49Yeah, there's something interesting around that in the, I remember digging in at some point, you know, large language models, your default interface doesn't actually expose the confidence at all. Exactly. But you can, especially if you're doing an open model, right? You can look at what was the distribution of the logits on there and see like which one, you know, this was very sure versus this was actually a pretty even spread, could have ended up in a bunch of places. So there's something fascinating at exposing that, you know? Exactly. There are research about like how the logits and then, but the logits at token level.
51:24And then now you think about like a long sequence, right? Right. Yeah. What does this look like at an agent level or a task level or something? Yeah, yeah, yeah. Exactly. And it's calibrated, right? So if the agent said 60 % confident and that actually empirically calibrate with the statistics on that. So this is open research at this moment. And then people are still thinking of like talking about sometimes the agents are over confident on many things. Yeah. That's actually a really fascinating area. people should really spend more time looking to one, especially when we're getting deeper into productionization and the human agent collaboration.
52:00Yeah. Yeah. I've been thinking about this a lot with regards to some of these like entity extraction and context extraction and things like that. Like, can we add a confidence layer, right? Okay. We are 80 % certain this is all one entity, but we need some more evidence. How could we validate things like that? Yes, yes, completely. Awesome. Well, so we're getting close to the end of our time together. Is there anything we have not talked about yet that you think would be useful or important for us to discuss? Yeah, this is a really fruitful and productive conversation. Really enjoyed it. Yeah, I think we covered everything.
52:36Yeah. What do you think? Yeah. I mean, I think we covered a lot of ground. A lot of many, many ground. Yeah. Awesome. Well, I think let's maybe do one thing back. Like if you were to look, you just did all this research looking at the different failure modes inside of the enterprise and what these 6 % companies are doing really well. What would you say is if someone was working inside of one of those enterprises, but was not maybe in a leadership role, or maybe they are in a leadership role, but you know, a software engineer listening to this or a line manager or something like that, like what's the leverage point you would push them towards to increase their chance of success in terms of enterprise adoption?
53:15Yeah. If there's a single point, it actually, I think, is boiled into get the right pilot use case. And there is a lot of things within the right, right? The right thing needs to be truly valuable. People have the incentive to make it work. It's not a toy use case, right? And then with that, then you could actually organizing. You could even like a software engineer, they could convince their leadership because it aligns with their business metric, right? So there's a lot of things within the right part. The second part is the right use case needs to be ready in a sense that you actually have the right data.
53:55You have the right resources and assets to make it correct. And the third one on the right part is actually you can think about whether the technology and expertise within a team outside is actually would enable this to successfully execute it and land in the room. So it is basically find the sweet spot to start with and the success would compound from the first successful pilot to production launch.
From the publisher
It is widely reported that a gap has emerged between enterprise spending on AI and the durable value captured from that spend. Individual employees have enthusiastically adopted coding assistants and chatbots, yet those gains do not seem to be transforming businesses at an organizational level. One of the most important questions in the tech industry today is understanding why AI is not yet delivering returns that match the investment, and what separates the small number of enterprises succeeding from the many that are not.
Scale AI is known for supplying the human-labeled data behind many frontier models. It now also builds AI applications and agents for large enterprises. That combination of working alongside frontier labs and inside enterprise deployments gives the company a rare view of why enterprise AI may be stalling.
Emily Xue is the Head of Enterprise AI at Scale AI, and previously spent over a decade at Google. In this episode, Emily joins Kevin Ball to discuss the three layers where enterprise AI breaks down, why frontier model benchmarks miss what enterprises actually need, the data foundation problem, how the most successful companies combine internal domain expertise with outside AI specialists, and more.
Sponsorship inquiries:
sponsor@softwareengineeringdaily.com
The post The Gap Between AI Spending and AI Value appeared first on Software Engineering Daily.
