In short
Podcast Episode Notes: Dev Interrupted - "Why enterprise AI lives or dies on applied research"
Episode Overview
- Hosts: Andrew Zigler, Ben Lloyd Pearson
- Guest: Elizabeth Lingg, Director of Applied Research at Contextual AI
- Focus: Discussing the challenges of transitioning AI models from research to reliable products and the importance of applied research in enterprise AI.
---
Key Topics Discussed
Challenges in Transitioning AI from Research to Application
- Transforming AI research into practical products is complex.
- Importance of applying AI research in real-world scenarios to ensure customer satisfaction and effectiveness.
Measuring AI Impact
- Inner Loop Metrics: Metrics like accuracy that focus on the performance of AI models.
- Outer Loop Metrics: Metrics correlating to customer satisfaction and user experience.
- The concept of a "vibe check," which emphasizes the subjective experience of using AI products.
The Necessity of Specialized AI
- Specialized AI is crucial for enterprise applications over generalized models.
- Organizations must avoid biases by employing diverse and multiple metrics to assess AI performance.
Collaboration Between Research and Engineering
- A framework is needed for research and engineering teams to work together effectively.
- Emphasis on communication to foster a shared understanding of how AI tools can solve specific problems.
Insights from Industry Conference
- Discussion of the Engineering Leadership Conference (ELC) where AI was a central theme.
- Conversation with engineering leaders about their challenges, tools adopted, and the metrics they focus on, specifically around PR throughput.
---
Key Takeaways
- Metrics are Not the Solution: Metrics should serve as proof of improvement rather than the goal itself. A focus solely on metrics can create unhealthy team dynamics.
- AI as a Force Multiplier: AI should enhance existing processes rather than complicate or replace them. It amplifies the strengths or weaknesses present in an organization.
- Diverse Perspectives in Measurement: It's essential to incorporate different viewpoints when defining success and accuracy in AI to avoid biases.
- Grounded Research vs. Sycophancy: There's a need to balance user satisfaction with the accuracy of AI output to avoid reinforcing incorrect responses.
- Continuous Learning Culture: Leaders should encourage engineers to experiment and learn from each other, promoting cross-functional skills and understanding.
---
Practical Advice for Engineering Leaders
- Start with small-scale experiments to evaluate AI capabilities without overwhelming teams.
- Foster a culture of collaboration where engineers and researchers share insights and skills.
- Use multiple metrics to gauge performance, integrating both quantitative and qualitative data.
- Ensure that improvements in AI performance are aligned with user needs and actual product usage.
---
Additional Resources
- Contextual AI: [Contextual.ai Website](https://contextual.ai/)
- Elizabeth Lingg: [LinkedIn Profile](https://www.linkedin.com/in/elizabeth-lingg-79b85624/)
- Related Articles:
- [Throwing AI at Developers Won’t Fix Their Problems](https://www.aviator.co/blog/throwing-ai-at-developers-wont-fix-their-problems/#)
- [Why Language Models Hallucinate](https://openai.com/index/why-language-models-hallucinate/)
- [Cursed Programming Languages](https://ghuntley.com/cursed/)
---
Conclusion The episode emphasizes the critical role applied research plays in the successful deployment of AI technologies in enterprises. Collaboration, diverse metrics, and a focus on user experience are key to translating research into effective practices.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:06Welcome to Dev Interrupted. I'm your host, Andrew Ziegler. And I'm your host, Ben Lloyd Pearson. We're here on site at the Engineering Leadership Conference here in San Francisco, meeting with over 120 engineering leaders, including our friends at Expedia, showcasing their squads framework that they built on the Linear B platform. We'll have that episode coming out here in a few weeks for all of our listeners. It's a really great one, I think. You know, we've been sitting down with a lot of engineering leaders here in the dome at this event. A lot of episodes getting recorded that you all are going to hear in the upcoming weeks.
0:39So, yeah, stay tuned. Lots of good content upcoming. Been a really great event, Andrew. And, you know, I really feel like the whole tone of this event was set like the very first morning, like the first conversation I had. I had these two engineering managers walked up, had barely finished my first cup of coffee and had a really wonderful conversation with them. Naturally, in 2025, anytime you talk to somebody in tech, it always ends up coming down to AI. Like eventually you're going to start talking about AI. And, you know, they told me they've done what everyone's doing these days. Like they've adopted all the tools.
1:15They've got Copilot. They've got Cursor, Cloud Code, Windsurf. Like they had a laundry list. They just kind of listed off. And I was like, well, you guys are doing it all. And they even told me they had like bought some like dashboard from one of the vendors here to like measure it all. And, you know, it's paying a lot of money for it. Got all these pretty visualizations. but they were like almost distraught over it. And, you know, they're like, there's this AI everywhere. Like we have it all over our booth even, you know, but it's all over all the booths. The event itself makes AI like really a central component of the entire experience.
1:51And, you know, they were kind of distraught about how like they had all these tools, they had all these metrics, these dashboards, but they still felt stuck. Like they weren't actually sure if they were making the right decisions and if productivity was getting better and if their developers were actually being more effective, being more efficient, producing higher quality. In fact, they were actually starting to experience some backlash from some of their developers. And particularly over one of the metrics that their executives had started to really get stuck on was PR throughput. They saw that dashboard and they were like, ooh, can we make that number go up?
2:29And naturally the developers are like, why do we need more PRs? Like, is that really the thing that we should incentivize? And even the engineering managers are, you know, they're a little skeptical of that metric as well. And, you know, personally, me, like I immediately just sort of cringed a little bit because I once worked on an engineering team that we had two North Star metrics. It was the number of commits that we created and the number of lines of code that we changed. So, you know, naturally, we just wrote some scripts to make those two metrics go up. And sure enough, these two engineering managers, they even described a recent conversation they had with their engineering team where it was getting a little fraught.
3:09And the engineers were wondering why they were so focused on these metrics. And the engineering manager just had to say, look, if our executives are so focused on this one metric, PR throughput, and you have these AI tools, you guys know how to make that metric. You can use AI to make more PRs. Like, look, you can just do it. I know it's not the thing that's going to be best for all of us, but you can just do it. There's nothing stopping you. And I felt so bad for them, honestly. Like, that's not what you want to do to your team. Like, that's not the kind of relationship you want to have. That's not healthy.
3:44That's not going to be productive. But it's a common refrain we hear. A lot of these organizations, they really never make it past that measurement phase of the developer productivity journey. and they like buy this dashboard, they get a framework and the execs might think that's enough. But the reality is that you've really just barely even made the first step at that point. And if you don't take action, you know, if you don't realize, if you don't take any steps to improve upon the things that you're measuring, then you're never going to realize the full potential of AI. And I think that's a theme that we're seeing like time and time again here, you know, not only at this event, but it is coming up a lot here.
4:25and elsewhere. And, you know, we're obviously here with Linear B, you know, we have the Dev Interrupted Dumb. Linear B is also here with their booth showing off a lot of the cool new features. But what I really love is the, you know, particularly with Linear B is they're not here to just like show you how to measure AI and be done with it. Like we're really focused on how to get better with AI. So, you know, as a part of this, Linear B has announced a bunch of new features around it. So like we've got the AI code reviews that are doing really awesome, especially in head-to-head comparisons against all the tools that are out there like Copilot.
5:00Because you think about code reviews are one of the most common bottlenecks in the SDLC. So really giving practical improvements to dev teams. We also announced the new MCP server that just generates helpful artifacts to help you make decisions. It's a new way to consume data within teams. And it's almost like an evolution going beyond metrics and dashboards. Now you have like this companion that helps you make sense of your reality. But then things like developer surveys to help uncover pain points in the AI journey. And then, you know, of course, AI dashboards. Like after you've made improvement, you know, you do need something that proves that you've made those improvements.
5:41So, you know, rather than being like metrics being the first step, I think that's one of the things that, you know, us here at Dev Interrupted, but also Linear B, I think really makes a point is that metrics aren't the answer. They're the proof after you found the answer. It's been a really fun event so far. And we've had a lot of conversations like that. But I think if I had to really like summarize where I feel like the zeitgeist is today, and that is, you know, there's so much hype in the AI space. I think we are beginning to transition out of things being purely forward looking like everyone's going beyond just like, what is what will the world look like two years from now?
6:19There's still lots of that, of course. But there's a lot more now of here's what we have done and are doing and here are the practical changes that are happening now. And here's how that's going to change the world in two years or three years or five years, you know. So it's really starting like AI is really starting to show up in the real world, not just be some like futuristic thing that we're all waiting to show up. You know, it's a really powerful anecdote you shared from those engineering managers who came up and talked with you because across the conference, I think a prevailing theme has been that AI is a force multiplier for the things that are happening within your organization.
6:58So if you value things like DevOps and really strong process and quality, then AI can help magnify that. We just heard an amazing talk from Natalie Gleam, the chief engineering officer at Duolingo, who talked about how they harness AI as a force multiplier in order to create more content quickly without reducing quality. And that's the real key, is that balancing that quantity with the quality is how you achieve really great success within your team. And you can actually use it as like a force of change within your engineering org. And that's like a prevailing thing I'm seeing and hearing from a lot of the talks here.
7:36I think every single backdrop has the letters AI on it somewhere. Everyone here is talking about it and what it means to build incrementally. I think that's a really good call out that you made that there's less of the, oh, the nebulous look forward. What does it look like in X years? And it's more about what does it look like right now within our organizations? We've all been tangling with it for a year. What's happened after all of that? And are we upside down? Are we actually better off? And I think those are real conversations happening within all of these engineering orgs that are here. So it's amazing to all come together and share those insights with each other.
8:11Yeah. Yeah. So, of course, not everything is happening here in San Francisco at ELCA. There is the rest of the world happening. And there are other stories out there that, of course, we want to cover while we're here, too. So let's get into some of the news stories, Angie. So what's our first story that we had today? Yeah, so along the same theme, we read an article from Ankit Jain. This is at Aviator about throwing AI at developers and about how it doesn't really fix their problems. And this is a prevailing idea of something you just talked about, something that is really encountering a lot of engineers within their organization right now, that bottlenecks and problems within your org, they don't go away when you add AI.
8:47In fact, AI will magnify those problems. And so this article breaks down kind of a flipping the script on how you think about some of the ways that you would utilize AI. So like one that stood out to me is that complexity can never be removed. It can only be abstracted. So when you use AI to make a process more simple, it doesn't mean that all the complexity is no longer happening. It's just happening in a place where maybe you're not seeing it as much or you don't have as much touch on it. And that's what saves you time. but it's not good to save you time if you lose the quality and the exact like what you're trying to achieve with that actual process and so uh it really dives into like why it's important to understand under the hood what's happening within your ai workflows and goes back to things we talked about on dev interrupted like the power of evals making sure things stay on the rails i think it was really interesting to see how it kind of like laid these top level points for why it's a force multiplier.
9:48Yeah, there was a line in this that really stuck with me that has actually come up with a few times, actually just in these couple of days we've been here. And that was, don't just pick a tool, pick a problem. You should always approach AI from solving something that your organization is facing. If you buy a tool and then search for a problem, you're probably going to misapply AI. or even worse, you may find yourself going in the wrong direction using AI. And in fact, I spoke with John Amaral while we were here, and that episode will be coming out in coming weeks. And he brought up a very, very similar point.
10:26And his company is doing some really, really cool things with AI and actually productizing it. And he made that exact same point that when they were using AI to build a product, he was very focused on finding the problem and then figuring out how AI can be applied to that problem. So it's definitely worth checking this article out. What's next? I think we have an article from OpenAI about hallucinations. Oh, yeah. So this has been a really interesting one to follow. OpenAI released a blog post building on a research paper that they did about why language models hallucinate in the first place. And if you've been working with large language models like many of us have, you know that the hallucinations are actually more like a feature, right?
11:07When we call things hallucinations, the idea is that it's a plausible guess within the realm of what could be true, but is actually not grounded in reality. And what we deem as a hallucination from the machine is actually a common occurrence that we do all the time as language interfacing things, right? And so the way that we build these systems and then evaluate them actually has a lot of similarities with how we build our own knowledge and how we test that knowledge. So this paper from OpenAI dives into the provenance of hallucinations and why they occur. And it really dives into that there are actually more things that are reinforced through post-training as part of the process of creating a model and bringing it to market.
11:50And it's a thing that happens because of how the models are incentivized. And there was a really great post that we saw from recent guest, Dr. Tatiana Mamout, who came on DevInsrupted earlier this year to talk about the sycophancy within AI where, you know, we all know this experience where you're told you're absolutely right when you're absolutely wrong. And this behavior, it stems from, of course, the way that the models are post-trained and evaluated after they've been fine-tuned, right? And so she calls out rightly that it's just like on the SATs, how students are penalized for leaving a question blank.
12:27These evals put these AIs on rails to where they're always pressured to make some sort of answer, even if it's not true. And this is what ultimately ends up kind of bubbling up these hallucinations in these experiences that we have, because we tell the LLM to be right, to be certain, to be confident. And that's the same thing that it's tested in, in that same language in a post-eval, right? And then it's put to task. It has to make a guess. If it leaves a blanket, it knows it got it wrong. If it makes a guess, at least it's maybe one out of a four chance. And so this is just like how humans work.
13:00And so ultimately, the incentives for what will make AI more truthful are the same incentives that make humans better capable of synthesizing thoughts. And this means that don't just rely on binary results. Yes, no. Is it true? Is it false? Have them write short responses, essays, avoid the weasel words, and actually figure out what it is that you're trying to do and say. And that's the mark of a highly intelligent person who can convey their thoughts. And that's the road we have to go down with LLMs to get them there too. Yeah. One of the examples they included was if you ask a model to tell you or to say when someone's birthday is and it doesn't know, and it might take a guess, and maybe it guesses the month right, but it gets the day wrong, and that gives it a partially correct answer.
13:45It would rather do that than say, I don't know, because that is 100 % wrong. So that incentive model encourages it to get that partially correct answer. And it's going to say that partially correct answer just as if it was a 100 % correct answer. And as you mentioned, part of the issue with hallucinations is that variability is actually what makes probabilistic tools like LLMs valuable because it gives you some level of adaptability. I think a lot of people who are working on improving these models do believe that hallucinations, if they're not completely fixed over time, they will get much better.
14:24You know, we'll get better at solving it, at reducing the impact of them. Something else to consider is, you know, if you're operating in a space where this does cause problems for you, you can also find ways to implement deterministic checks to sort of mitigate the risks for that. Yeah. So if you want to understand a little bit about why hallucinations exist and why it's a problem and what the companies that are building this stuff are doing, this is definitely an article that you should check out. Absolutely. All right. So our last one is filed firmly into things that I don't think any of us knew that we needed, but here it is today.
14:58And now we all get to live with it. So what's our last story for the day, Andrew? This is my favorite one. So we've covered Joffrey Huntley and all of his work, Vibe Coding and how he uses things like Cursor and Cloud Code and then AMP to build things that he never possibly could have built before. And the way that he thinks about using these tools is so optimistic and just so cavalier that I'm just always constantly plugged in to what he's working on. I think if there was a Vibe Coding Olympics, I think Huntley would be like the best Vibe coder, like maybe in the world right now. Yeah, he has been at the top of some charts, maybe to his chagrin.
15:40But I will say that the work that he does is really interesting. And we're going to cover one of them right now. So he's been working all this year. Well, Claude has been working mostly to create a new programming language, a C compatible standard lib using only Gen Z slang. So what this means is that he used C to write basically a compiler in a new language. And then once he wrote enough of that language in that new language, he then used this cursed language is what he calls it to write itself. And in this process, what he did is he unveiled that you can really kind of do anything now within software engineering.
16:21And he really kind of breaks down this falsehood that a software engineer skills will atrophy in this world of AI because he's somebody who never had the curiosity before to even want to build a compiler or understand how they worked. But because he was enabled by these tools, he was able to work with Claude on a loop, running a very simple prompt to incrementally create a Golang-like programming language using only Gen Z slang. And the lexical structure of this is amazing. He released this website that we'll include in our newsletter, so you can go and check out how you could use this language to write your own functions.
16:55And what ends up happening is every single declarative word that you would imagine writing a program with has been literally replaced with Gen Z slang. So I'm going to call out some of them that are really funny to me. For example, you declare a function with slay. So you would say like slay add as like a way of declaring a function. Another one that's really funny is when you return, you don't say return, you say damn. That's pretty good. And then, of course, you have, of course, working with things like true and false. True is based. And cringed is false, of course. And my favorite, my absolute favorite is doing a comment block.
17:33You know, every language has it different. Maybe it's a double backslash. Maybe it's a backslash. And then an asperic. No. In cursed language, a line comment is frfr. So for those listening that don't know what that means, it means it's literally saying for real, for real. and if you want to do a block comment well you open with no cap and you end with on god it's hilarious and it's something that i didn't think the world needed until i saw this website and honestly i'm gonna probably sit down in the saddle and try to write something with this language he calls it cursed it is literally cursed it is the cursed gen z language and he also made the joke that he cursed at uh claude a whole bunch while i can i can imagine what do you think about this one well of course we had to consult with our resident gen z slang consultant yes producer adam he's the only way we could make sense of this although it is it once you know the slang it is a surprisingly readable language which is really phenomenal i'm gonna list off some of my favorites so to import a package yeet that's brilliant like i feel like that should be what javascript changes their import i feel like yeet should be like export it'd be like i'm done with to eat but i like it a constant is facts that's a great one that is a good one uh to break ghosted an int is a normie man there's there's just so many like you really have to go check out the cursed web website to go to go see all the the incredible ones and if you haven't seen huntley doing this vibe cody he actually does it while he sleeps which is the best part yeah it live streams on youtube he's got all this agents just like hacking away as he sleeps um he's australian so it's like during our daytime he's he's at home sleeping and like i'll just have it like casually playing while i'm working some days and like i'm like man this guy's working 24 7 you know like it's kind of incredible it's amazing and then he and then you know he'll wake up and he'll see the work that it did and he'll review it and it's amazing actually if you go to the repository where he's been building this language the amount of incremental work and all of the artifacts of production that have been created by him in the lom through this whole process i have extracted so many techniques that i use all the time from how huntly does this and the last thing i'll leave you all with is on the website we'll include there's a whole bunch of leak code examples so you can actually go read full leak code functions of common things people are asked to do in programming languages like determining the maximum depth of a binary tree using only the cursed language so definitely be sure to check this out and we would love to see if you write anything in this yeah yeah and i and i want to just point out that i think what this really illustrates is the brilliance of when you construct this in an ai agentic system with all the context it needs with the structure with the tooling that it needs how simple the the capability actually becomes because this is not like a thing that's going to get deployed into some high production system that is like well okay maybe maybe we'll see who knows so he sets this whole up and then gives it like the most simple prompt, which is the most incredible part of this.
20:36You know, it's for three months, he says his agents had the simple goal, produce me a Gen Z compiler, you can implement anything you like. That's it. That's all he told it to do. And here we are three months later. Like that's really a pretty extraordinary thing. And I think this will be a case study in how vibe coding this type of work happens in the future. So it's gonna be a case study. It's gonna be a full conversation. We're definitely have to get him in the seat here. to talk about it. Yeah, yeah, yeah. Stay tuned. I have high hopes that he'll be a guest on our show soon enough and we'll definitely get to pick his brain on how all this works.
21:09So amazing. Well, you know, those are our stories for the week here at ELC. There's just so many insights that we're extracting. We're going to keep talking about them week after week as we drop the amazing episodes that we've been recording here on site. So stick around. Are you investing in AI but struggling to see the real impact on your engineering team's productivity? Well, you're not alone. In a free 35-minute workshop that I'll be hosting with Linear B, we're going to show you how to translate AI metrics into business ROI, just like Expedia and Adobe have. And you'll learn a simple framework for understanding where AI is helping, where it's hurting, and where to focus your next investment.
21:47And besides, you're going to get a nice takeaway report on AI productivity as well. So don't miss out. The workshop's coming up on September 17th or 18th. Grab your slot, and we'll see you there. Today, we're diving into a really cool topic. We're exploring the world of specialized intelligence and the challenges of bridging research and product development. And we're doing this with Elizabeth Ling. She's the Director of Applied Research at Contextual AI. That's the team behind RAG. And ultimately, Elizabeth's impressive background spans major players like Microsoft and Apple. She brings a unique scientific perspective to making AI work for real-world applications.
22:29Elizabeth, welcome to our show. Thank you very much. I'm so excited to chat. We're excited to dive into your research background. And I wanted to start just by talking a little bit about your career because it's taken you through some impressive places that I just named. And it all started with you for a love of learning. What does it look like as an applied researcher on a software team? I would say that I've always been passionate about math AI. Even as a kid, I read 2001, A Space Odyssey, and I was so fascinated by the book and whether, you know, Hal had a software hardware malfunction. Did he achieve singularity when he ejected the person from the spacecraft?
23:11What happens with intelligence, with AI? So I've always loved AI and the research side for the longest time, math, statistics, data, data science. It's always been a passion. But then how do you actually bring that into a product? How do you build something, something that works for, you know, real users, real customer data? So that basically is where my career has taken me to actually apply the research, apply the ML model and build real and exciting products. So your journey starts by becoming a researcher, someone who's obsessed and great at science and math. And that lets you go deep on like an academic type of route of learning how to conduct research and collect great data.
23:55And then from there, how did you look to pivot those research skills to become an engineering leader? Sure. So I think just the fact that I joined industry after grad school at Stanford, I ended up, you know, working on different aspects of the code. So not just, say, clustering and similarity algorithms, but also information retrieval, building my own Lucene frameworks, and calling different APIs, external APIs, and so forth, which brought me into a lot of this software side as well. Of course, for anyone with a computer science degree, you do learn the fundamentals in different areas. But even though my job title was on the ML side, I'm still touching all these different aspects of the system and I'm working with the DevOps team and I'm working to train a model on a server on a specific compute.
24:50So I think what happened is once I joined the real world, I have to build these products. I have to make sure the customers are happy. Then I start worrying about things like latency, right? So even 16 years ago after grad school, latency was a thing. So I was working on governance, risk and compliance and fraud detection and using AI for that. But even for that, you know, if it's too slow, the customer is going to be unhappy. So that's that's what brought me into the software side of things. And then honestly, it's been the best fit for me. I realized that even though my greatest passion is on the research, AI and math side, it's really good to have those solid computer science fundamentals as well and at least have an understanding of code.
25:32Of course, you can vibe code these days, but you have to have the understanding of the system. Having that full understanding, I think, is really valuable. And what I've noticed is the industry has shifted. Like when I first joined, you either pitch, you're an AI researcher or you're a software engineer. You're a data scientist, you're a software engineer. But these days, it's very common to have the full stack, MLE, applied scientist, applied researcher, which works out great for me. I think it is important to have that full perspective. Yeah, that's a great call out that you've made there about the diversification of the skills and the baseline that folks are expected to operate at.
26:10We're hearing that from everybody. We talked with Lee Robinson, VP product at FurSell, echoed the same thing happening in web development where you have like the full collapse of front end and back end into this full stack engineer that's a baseline now. That's table stakes as someone working in that environment. And you see the same thing happening. We talked with a guest from Google DeepMind who talked about how most engineers and how engineering work. It's a collapse of back-end, front-end. It's not that specialists go away. It's just that the expectation level that everyone's supposed to be operating at shifts.
26:41And there was something else that you said in there, too, that I really like. I want to double-click on. And you said, when I entered the real world, when you went from doing the theoretical work to working in the practical application world, what was a big challenge you faced when you made that transition? I think anyone who goes from the academic world to a more of a production environment to learn about all the basically all the failure cases that exist and, you know, inner team dependencies. And to your point, I mean, yes, specialists still exist. They're great. I love specialists. It hasn't been obviously my trend in my career per se, but it is the baseline you do.
27:23Even if you're a specialist, you do have to have that baseline understanding. And so I think what was shocking to me is just at first, once I was getting to know how to build software and cloud software and applications and so forth, and eventually SaaS is like, okay, there's still much that goes into it. You know, I always thought in academia, the research is very hard. It is a very challenging aspect to it. But it's surprising to see how much you have to adapt that research to really put into production, work with multiple teams and multiple disciplines. Yeah. And especially when you work at like a large company like Apple and Microsoft that has a bunch of teams, has a bunch of disciplines, it's a huge spread out organization.
28:08What was it like shifting from working in that kind of environment to being an applied researcher and working in a more startup like environment? So my trend over my career has actually been either startups or sort of the startups within the large companies. So I have been a part of several early stage startups or startups that are quickly acquired by a big guy kind of thing. And I think what I love about it is for both of those situations is the ability to really contribute to the research community to be scrappy and move quickly. For example, at Microsoft, I was most recently working on the co-pilot team and leading the language understanding for the office co-pilot.
Read the full transcript
28:53And in that role, you know, we started with a hackathon. And then eight months later, we had to deliver entirely new product, which you think would be easy for the large company. But, you know, large companies also have their downside. They have their processes. So it was really exciting getting there and getting to that large number of customers and large sales number and moving to GA in just eight bucks. So I always love both within the startup and the large company, trying something new, working on risky teams. And then I would say at startups, though, it's the best growth and learning. I see so many people, especially in the AI community, attracted to startups because you wear so many hats, you touch so many things, you learn a lot, you challenge yourself.
29:37And then certainly at large companies on the right team, there's a lot of learning and growth as well. And when you're on those teams and you're working on applied research within an organization like that, how do you measure your impact and how productive y 'all are as a team together? So that's an interesting question is always measurement and metrics, which, as one knows, can be biased, right? So you may measure accuracy, but you have one perspective. Someone on the other side of the world or coming from a different culture may have a different definition of accuracy or golden labels. Or what is the representative data set?
30:14Is it the data on the Internet? Probably not. We've read the Stochastic Parrots paper. But bottom line is, I would say we look at we do try to quantify measurements typically like accuracy and so forth. or say for retrieval, you know, DCG, MRR, recall, precision, all that kind of thing. And then we try to, with like large language models, we measure groundedness, equivalents. Contextual has a really cool methodology called LMUnit, which I could tell you more about, which runs software engineering unit tests to verify criteria. So we have all these, at Microsoft and at startups, we do have these quantitative and at large companies like Apple, we do have quantitative measures, which we use.
31:01And we try to diversify those metrics. That's always my goal, to avoid bias. So get perspectives from different people, aggregate different metrics together, try to balance out the bias of the data set, have a holdout data set. That's one way to measure success. But the biggest measurement of customer success is always, I would say, always customer utilization, customer satisfaction. interaction so one thing we've worked on doing is correlating our what we call our interleap metrics with outer loop metrics so how many queries per day how often are people using the app you can correlate that with for example you can see if there's a correlation with some of our internal metrics in terms of of recall for retrieval can you apply a linear regression and do you see that there's a correlation there what features should be weighted higher or lower in order to determine and how the inner loop influences the outer loop.
31:58So we do some of that quantitative work. And then what's interesting with what I've seen at certain companies is there's a lot of like vibe check and user experience. So even Apple, I think is really good at that. Like Apple is not historically been like the, you know, AI data company, but for example, they're really good at like building these really cool engaging products and really beautiful products and making data beautiful and making AI, like the excitement of using AI when I worked on the Apple health team, for example, on your phone, on your watch. So I would say there's both the quantitative and the qualitative metric, and they both matter.
32:39And for both of those, you have to be very careful to avoid bias. It's really helpful to understand how you walk through the process of evaluating your own work and looking at it when it's actually applied at scale and inside of applications and those things that you look for, the signals from the end users on is this good, is this bad, but also just internally your methodology for making sure everything stays grounded because that's why you're there. For those listening, you know, they might see a lot of things related with this kind of mentality with their own kind of mini startup groups and cultures within their orgs right now because everyone's experimenting with this kind of technology and they're trying to be grounded and methodological about how they approach it, right?
33:21And I'm wondering from your perspective, You know, how could a leader foster that kind of culture of grounded research and innovation at their own company? So definitely at Contextual, for example, we're very much focused on groundedness and accuracy. So the reason why is for a lot of enterprise use cases, it's very important that you're grounded in the retrieve knowledge, in the specific knowledge, and there's not hallucinations based on the model space knowledge. So that is one thing you can foster is like if the company's goal is enterprise customers who need to be grounded, who need to be accurate.
33:57That's just such a big need, I think, right now in the community and in AI. Then that's something you should focus on because that's what's important to your users. On the contrary, if you're like a consumer company like Artificial General Intelligence, you know, the OpenAI, Anthropic, etc., who's looking at the consumer market versus enterprise, you might be more lenient on creativity versus accuracy, just not important to your users. So you do always have to be like customer first and user base in mind. But even then, and this is something I learned across my career at different companies, you also just have to try to seek the viewpoints and the external metrics and try to understand from like diverse team members how to avoid bias.
34:45so like large companies what's great is they have rai but even they're even you know tools around this like um there's models which will tell you if something is offensive or they'll be sure to be careful about rai the llama guard type model especially yeah yeah so and then the the problem shield and so forth so there's there's things you can do whether you have the resources of a large company and have a whole team around it, or you're using one of the open source models. But I would say that's important to just have, like, try to have a broader perspective or make use of tools that can kind of give you a new sense of awareness.
35:24And then I would say definitely use multiple metrics. That's another thing. And then the final thing is just beyond all that, don't forget the vibe check. Don't forget the actual user experience. What is it really like to use it. Even having humans read to emit tell you what they think, that's invaluable. I like what you're calling out at the end of the vibe check aspect. It actually reminds me of the recent memo that came out from OpenAI about their own model releases around sycophancy, because they use that language vibe check as part of their review process that maybe didn't ultimately pass the right test.
36:02If you remember the memo, I'm not sure if you're familiar with it. It was very much about how the LLM responses you would get would say what you wanted to hear as opposed to being grounded in truth or having a reasonable response, having those safeguards. Are you familiar with that memo that they dropped? That's what I think of when I hear you describe that topic. Yes, yes. But this is, I mean, human preference-based alignment is a whole research area. In fact, we have, in my company, we have a lot of researchers who are collaborating with Stanford. That's one of the dangers with RLHL and reinforcement learning and any kind of human preference alignment through thumbs up, thumbs down.
36:42People have their own bias. They want to be told what they want to hear, but that's also dangerous, right? So that's why you do have to balance it out. You don't want to just rely on that sycophancy. You want to maybe have some vibe check to counter that. That's one thing. Or you could use another metric. Like I I mentioned we have a model that measures groundedness to say how accurate our data, how grounded the responses are in the retrieve knowledge, how accurate they are. So those are different methodologies you can use to balance it out. But that's why you need those multiple metrics, because you don't want to just optimize for making the user happy.
37:21I know before I mentioned that was the ultimate metric, but it's not their happiness is just not being told what they want to hear. Like enterprise users also want something accurate. And even consumer users, like there's a good number of consumer users who will be angry if they're told something inaccurate. So you just have to have a broader perspective on that. Yeah, it absolutely depends on the reason to which you turn to the tool, right? It goes back to what you've rightfully called out about the usage of generalized models versus specialized models, understanding what and why you need to be using the tool that you are for the stakes of your company or your product or what you're putting out there.
38:00So in your kind of environment, you know, you spend a lot of time thinking about the specialized intelligence and how do we create it in a way that is grounded in the facts and truths of our org and our live data and the things that matter to that agent doing that workflow. What are some things that you see working in specialized intelligence that just off-the-shelf models just can't do that maybe folks are missing if they never look deeper? I would say what folks are missing with off-the-shelf models is even with them, there's just a lot of effort, even so, to make a specialized use case work.
38:39Like if you look at some of our customers working on technical documentation, code gen for not just general code gen, but code gen for their libraries, their specialized use case, hardware configurations, customer support for electrical engineers, like really experts or another, you know, customers in the tax and legal space, like experts in different regions. For example, there's different tax laws. And then so out of bots, you could get a certain thing. But a lot of organizations, they have their own internal data. They have their own internal knowledge base. Right. And so I think what's missing is it's not so easy to just throw all that into context or set up a rank platform and use the foundation model.
39:29It takes a lot of, you know, human effort and work. Like you have to start data cleaning, re-rank for a generation. And so a lot of what the value proposition of what my company is doing is making that a lot easier to spin up these specialized agents very quickly by having all these components, document, understanding, parsing, re-ranking, generation, retrieval models, etc. and the ability to specialize them and also evaluate for groundedness and so forth. And so having that all together and having that acceleration within a product and a platform is definitely something that I think is much needed because otherwise you set up one use case, you spend a long time, and then this new use case comes back to ground zero.
40:15And I know there's a lot of companies doing that. A lot of companies are just, they have this customer and then they have a bunch of folks like set it up make it work and then they're back to ground zero for customer number two number three it's still a good business model but our goal is is is to make that much easier and so i think what's missed there is not having the state-of-the-art components and modular components that you could use and quickly stitch something together and i think there's a big market for that it makes a lot of sense and i i think of i put myself in the shoes of like our listeners too And many of them find themselves, you know, a leader at an engineering organization.
40:54They're experimenting with AI workflows, trying to help their engineers have a better developer experience, be more productive, remove friction. And Elizabeth, I know that you also, your background includes working, you know, very closely with engineering teams and engineering managers. I'm wondering what advice you would give to those folks for how they should go about exploring this problem. Because for many of them, the accessibility of starting with a specialized model approach, no matter how much institutional knowledge they're sitting on, it might be very daunting. But I'm just wondering if you see opportunities for them to ultimately work with off-the-shelf models along with their own data or eventually moving towards the specialized intelligence approach you've said.
41:41How do they really get started with the kind of tools that are available to them these days, you think? so i mean there's definitely some options out there and then there's for example on our platform we have a certain free trial period for use of certain apis they go play around with different options like contextuals see how they how they like it but i would say um they can also just experiment with that versus out of the box i would say honestly um out of the box again it's only going to get you so far and you're going to probably have to dedicate a lot of resources to that. So that's an option.
42:21That is an option, but it depends on the size of your company, the number of use cases. But you can consider working with a company like Contextual. Again, you just have to weigh the pros and cons. For some people too, if it's a very simple use case, maybe accuracy isn't that important. Maybe something out of the box will work great. And then in terms of working collaboratively as an AI leader with engineering leaders, yes, I've definitely done that. And then I've also worked even on the software engineering side and big data myself, which I think is very valuable, again, to have that full stack, full perspective.
42:59But I think it's very important to try to communicate and understand the other perspective as much as possible and kind of explain where you're coming from and work together with the engineering leader to say, hey, this solution has these pros and cons, this cost associated. this has this, here's the first sequence of experiments we should try. And let's work together with the software engineering team and with the applied researcher, applied ML team, and come to a conclusion on what kind of roadmap. And then definitely starting with those initial small-scoped experiments. In my opinion, that's the best way to go.
43:38So starting with small, repeatable experiments that have a reason, you see, was it good? Was it bad you do the vibe check along the way and ultimately you know you're working with a team of engineers and and you said you you spoke to you know you understand this very well as an engineering manager someone who is an engineering leader um that you know ultimately everyone just wants the right good code ship good software right and and and they want to try to use new tools to do that better faster and safer so it's about trying to roll that out and kind of like pilot programs and and and don't obviously overload your teams with tools that you know they're all trying at the same time because then it's hard to measure what's really working and what's not right right and it does especially these days with a fast pace it does get a little challenging because you know there's the people from the research perspective like myself who more typically slam more the side of let's experiment let's try this new technology let's try this new model.
44:37But those folks need to be also grounded, you can say, in the fact that, hey, no, this model uses a crazy amount of memory and GPUs and the code is strung together. So that's why the partnership with engineering and then even for myself, again, because I do have some perspective of both kinds of things, trying to help bridge the gap and collaborate and come to a common conclusion is helpful. It can be hard too, because on the research team, it's more about innovation, prototyping, not so much about writing production-ready code. So that's why really clearly partnering with people who love to write that robust, product-ready, well-tested code is critical.
45:24And while also just being flexible to StopPoint, I think that's what's helped me too, is just having people on my team try, maybe experiment with some of the things that the other team is doing or at least observe what they're doing. Oh, this is what it's like to train a model. This is why, you know, this case doesn't work and you can't just fix it. Or, oh, this is what it's like to write some production code. Let me try to write some code in the system. And I found actually what I had, because I've had people on my team who are like very strong engineers or very strong researchers and then more people in the middle.
45:57But when I had those people who were kind of like on the opposite sides of things, especially those with like a great growth mindset who want to learn, I found that when they kind of experience the other side of things, it also helps them to understand. And the only trick there is it has to be a balance because, for example, in my current team, I work with our platform software engineering team and they're like really brilliant, brilliant engineers. And so sometimes my team, we do have engineering knowledge so we can debug something. But you have to think, OK, how much are we going to and then how much are we going to ask them for their help?
46:34And then my strategy on that is at least in good faith, put some time in, look at the log, send over a log, share what you've tried, shared what you think the next steps are. And then also the other team will feel, hey, this person actually tried to look into it. Though again, sometimes time doesn't permit that. It's like, hey, we need this now. There's an outage, like help us. And then on the opposite side, like, hey, there's an accuracy issue in the model. We need it fixed now kind of thing. That's a good call out about the idea of like, as an engineering practitioner, you know, you can practice, you can do it, you can apply it, you can try it yourself before you try to pull in other folks.
47:11And what you do in that case is you're creating this culture of learning where folks in the organization are learning things outside of their immediate team scope and their responsibilities. And I think that's ultimately one of the biggest benefits of having like an applied research team and folks who specialize in that is they're able to connect all of those dots between the engineers and the silos in which they operate. Right. And in doing so, you build a culture of learning. And I'm curious from your perspective, what are opportunities you see now for engineering, maybe like new engineers or people who are trying to level up to become those engineering leaders, managers, someone like yourself?
47:48What kind of skills do you think are most important these days? Yeah, so just to follow up on what you just said first, I think that is absolutely critical to have that growth mindset and that openness to help where needed and work together. But I also just wanted to say that that has its pitfalls, too. That's my natural style. And I think collaborative teams do the best. But it's also you have to be careful because I have had people on my team working on projects completely inside their area of expertise. And it can be like, OK, how long is this person going to spend on this for the good of the team and the company?
48:25There's different reasons for that. So you have to you have to be careful there. Now, in terms of growth, being an applied researcher and having the perspective of both software engineering and ML or an MLE and how you reach a leadership level there, I would say, honestly, when you're really new, just be really good at your job, do your best. and then as you start to feel more comfortable think about how you can start solving problems for more people bring people together whatever your talent is so your talent could be you're really great at architecture so you really take the lead there or your talent could be you're really great at collaboration so you focus on bringing everyone together getting everyone the common ground in order to solve a problem or your talent could be you're really good at like you have strong skills in debugging.
49:16So then you could take the fact that you have those skills, plus you have ML background, and then be able to debug across the full stack and then maybe teach others. And then I think for any job, it's all like sort of a matter of like, do the job before you're promoted. But you should talk to your manager probably about how you could do that or your mentor. Usually, if they're amenable to that, like they want to help you grow to what you want to do. If not, probably you should start looking for somewhere where you can eventually grow. But one thing I do is I'll give people side projects. So a lot of people might want to get into a specific research topic.
49:54And I'll be like, okay, well, just focus on delivering. But then 25 % of your time, see how you can apply this here. And then often that can help them level up. And then another thing is I feel like in startups is you really make your own opportunities. opportunities. So startups, they're growing like crazy. There's opportunities available. So try not to think in terms of restrictions, like I don't have this title, so I can't do this. Just think, okay, I'm good in this area. I think this can help. But within that, I mean, talk to your manager, talk to your team. Don't just go and like, well, I'm just going to come up with this new project and not get it approved.
50:31It can't be like that. But if it's a side project, most likely it will, I would say. Also, I think just being open-minded for those opportunities that come to you and just thinking if it's a good opportunity for growth within your company or externally. Yeah, it's like understanding your impact and what your impact can be. I love how you called out how in a startup, you know, you make your own opportunities. I think that's completely true. Folks who experience a lot of success and growth within a startup are the ones that can market themselves and the impact they have on the day to day really well.
51:04And that And that requires obviously being deeply familiar with the problem that your company is solving and being really close to those critical conversations that are happening every day, whether you're a part of them or not, like just understanding how it's shaping. And then thinking about how can my impact tie back to that? And in doing so in a way where you don't tank your current expectations for what you are there to do, because it's great to be curious. It's great to want to grow. But if you're there to perform a specific function, you have to, you know, first and foremost, make sure you fulfill your function because your company is relying on you in that way.
51:39So it's like a good bit of nuanced advice, but it really calls out how you should be focused on understanding, you know, what's the core thing that my company is trying to do right now? And what is available in my skill set, in my wheelhouse, to get us a little bit closer to that goal? and then once you do it and you start working on it it's about explaining to people about like you know this is what i tried this is what worked this is what didn't this got me a little closer we're all stumbling in the dark trying to figure out you know where the light switches so we can see everything is this that this is me trying to get that direction and so it's kind of it's it's helpful to understand it from your perspective too as someone who bridges both like the theoretical and the practical world um i can't think of a better perspective to draw upon to kind of connect those dots for us.
52:24So I appreciate that. And Jan, I kind of want to round things out too, because we've covered a lot of really great ground here about grounded research and about generalized and specialized intelligence, but also about your career and how you've moved through these types of roles and in these kinds of environments to ultimately take the theoretical research-minded approach and apply it into solutions, software, and applications that are shaping our world. But before we wrap up, where can our audience go to learn a little bit more about contextual AI and what you do, Elizabeth? So if you can definitely go to our company website, we also have pretty active social media on LinkedIn, Twitter as well.
53:05Just search for contextual AI. And then because, as we mentioned, the leaders of the company are part of the original RAG team. Our CEO is one of the co-authors and a professor at Stanford. he was interviewed recently. So yeah, feel free to check out anything about our company or send me a LinkedIn message. I do get quite a lot, but I'll try to reply if you have a question. Yes, yes. Please reach out to Elizabeth on LinkedIn. You know, you should definitely give us a shout out if you listen to this episode there. Dev Interrupted posts every week. We're going to be sharing clips of this conversation with Elizabeth there as well.
53:44And we'd love to hear your thoughts on what we discussed today. so for those listening thanks for joining us this far be sure to rate and subscribe to the podcast if you haven't already and we'll see you next week thanks for joining us on Dev Interrupted thank you
From the publisher
What does it take to transform a brilliant AI model from a research paper into a product customers can rely on? We're joined by Elizabeth Lingg, Director of Applied Research at Contextual AI (the team behind RAG), to explore the immense challenge of bridging the gap between the lab and the real world. Drawing on her impressive career at Microsoft, Apple, and in the startup scene, Elizabeth details her journey from academic researcher to an industry leader shipping production AI.
Elizabeth shares her expert approach to measuring AI impact, emphasizing the need to correlate "inner loop" metrics like accuracy with "outer loop" metrics like customer satisfaction and the crucial "vibe check." Learn why specialized, grounded AI is essential for the enterprise and how using multiple, diverse metrics is the key to avoiding model bias and sycophancy. She provides a framework for how research and engineering teams can collaborate effectively to turn innovative ideas into robust products.
Check out:
Follow the hosts:
Follow today's guest(s):
- Learn more about Contextual AI: Contextual.ai Website
- Follow Contextual AI on Social Media: LinkedIn | X (formerly Twitter)
- Connect with Elizabeth: LinkedIn
Referenced in today's show:
- Throwing AI at Developers Won’t Fix Their Problems
- Why language models hallucinate
- i ran Claude in a loop for three months, and it created a genz programming language called cursed
Support the show:
- Subscribe to our Substack
- Leave us a review
- Subscribe on YouTube
- Follow us on Twitter or LinkedIn
Offers:
