Can Your AI Strategy Be Future-Proof? | Galileo’s Vikram Chatterji

18 Feb 2025 · 41 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Dev Interrupted Podcast Episode Notes

Episode Title

Can Your AI Strategy Be Future-Proof?

Guest

Vikram Chatterji, Co-founder and CEO of Galileo Hosts: Ben Lloyd Pearson and Andrew Zigler

Episode Summary In this episode of Dev Interrupted, hosts Ben and Andrew discuss the evolving landscape of artificial intelligence (AI) and its implications for software engineering teams. They delve into the challenges of adopting AI strategically, including balancing innovation with practicality. Vikram Chatterji joins the conversation to provide insights on how engineering leaders can future-proof their AI strategies amidst rapid technological advancements.

---

Key Discussions

  1. Current AI Challenges
  2. AI's Limitations: Discussion on a study illustrating how AI fails to recognize obvious patterns in data (e.g., the "gorilla in the room").
  3. Importance of Visualization: The need for human intuition in data exploration versus the AI's rigid processing of data.
  1. Banks Competing for Tech Talent
  2. Tech Layoffs: Many tech companies are undergoing layoffs, prompting skilled developers to seek opportunities in traditionally non-tech sectors such as banking.
  3. Modernization Efforts: Financial institutions are investing in tech stacks and improving work conditions (flexibility, compensation) to attract talent.
  1. Focus as a Core Skill
  2. Shift from AI to Focus: As AI becomes ubiquitous, the ability to focus on tasks that AI cannot do becomes crucial for engineers.
  3. Creating a Culture of Understanding: Leaders should foster environments where engineers grasp the nuances of AI inputs and outputs.
  1. Future-Proofing AI Strategies
  2. Evaluation Frameworks: Importance of establishing clear criteria for assessing AI initiatives based on business needs.
  3. Rapid Prototyping: Engineering leaders should experiment quickly but also measure impacts to build organizational trust in AI technologies.

---

Key Takeaways

For Engineering Leaders

  • Understand Your Culture: Tailor AI adoption strategies to fit the organization's risk tolerance and operational dynamics.
  • Crawl, Walk, Run Approach: Start with small-scale AI projects, evaluate their effectiveness, and gradually expand based on initial learnings.
  • Build Trust: Ensure that AI tools and outputs are transparent and reliable to gain acceptance among stakeholders.

For Developers

  • Continuous Learning: Stay updated on AI technologies and community best practices to remain competitive.
  • Experimentation Mindset: Engage in building small AI applications to understand practical challenges and solutions.

Vikram's Advice on AI Implementation

  • Focus on practical use cases before diving into broad AI strategies.
  • Prototype rapidly with a structured evaluation framework to gauge success and iterate effectively.

---

Additional Resources

  • Galileo AI: [Galileo's Website](https://galileo.ai)
  • Follow the Hosts:
  • [Ben Lloyd Pearson LinkedIn](https://www.linkedin.com/in/benlloydpearson/)
  • [Andrew Zigler LinkedIn](https://www.linkedin.com/in/andrewzigler/)
  • Follow Vikram Chatterji: [LinkedIn](https://www.linkedin.com/in/vikram-chatterji/)

---

References Mentioned

  • [Your AI Can’t See Gorillas Article](https://chiraaggohel.com/posts/llms-eda/)
  • [Hiring Trends in Financial Services](https://leaddev.com/hiring/how-banks-caught-up-in-the-battle-for-developer-talent)
  • [Focus Will Be the Skill of the Future](https://www.carette.xyz/posts/focus_will_be_the_skill_of_the_future/)

---

Conclusion The episode emphasizes the duality of opportunity and caution as organizations navigate the complexities of AI integration. By establishing frameworks for evaluation and focusing on the human aspect of tech, leaders can position their teams for success in an AI-driven future.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:07Welcome to Dev Interrupted. I'm your host, Ben Lloyd Pearson. And I'm your host, Andrew Ziegler. In today's news, we're talking about how banks are capturing tech talent, why your AI can't see this gorilla, and how the skill of the future is focused. Wait a second, did you just say that AI can't see gorillas? Yeah, well, you ever hear that saying about not seeing the forest for the trees? Yeah. Well, it turns out if you're an AI, you might not even see this literal gorilla drawn in a scatter plot. My favorite article from this week was one called Your AI Can't See Gorillas. And it explores a study that demonstrated how when you're given a hypothesis to prove or disprove from data up front, it challenges the way that you can explore that data freely.

0:55And basically, there's a literal gorilla depicted in a group of data that when depicted on a scatterplot, it shows that gorilla. But if you don't actually plot the data on a graph and look at it and explore it, you're never going to see that gorilla. If you just go in and look at the columns and the rows and what's in the cells, you're never going to see it. For me, it really highlighted the importance of intuition and curiosity and how we as humans explore data. Because AIs, they totally fail at this right now. Even in the article when the author would take a screenshot of the scatterplot that the LLM is so helpfully rendering for it that shows the gorilla plane as day.

1:38It still doesn't want to see the gorilla because it thinks it's looking at a chart. So for me, this is spelling strawberry all over again. And it really emphasizes why it's important to keep a human in the loop too. Yeah. So I guess the generative AI comes from a universe where scatter plots are never used to visualize things or to make pictures, you know, but so, all right. So if I understand this, if you plot the data on a graph visually, the human is almost certainly going to immediately understand that it's a gorilla right now. Literally waving. We'll make sure we share the link so our readers, our listeners can see it too.

2:14It's quite hilarious. Literally waving, but it can't see it, no matter what you ask it. Yeah, and even humans, I mean, if a human just looked at the data and never visualized it, they would completely miss the plot here as well. So in other words, get off the command line and look at some images for a moment, you know? Yeah, and sometimes it's even applying common sense. And I think the biggest takeaway here too is it really highlights the difference in which you and the LLM are processing data. This is strawberry all over again, literally, because we don't struggle with spelling strawberry or knowing how many R's are in it because we look at each individual letter.

2:51But as we all know, LLMs don't work that way. Similarly, in this situation, we learn from data and graphs by looking at them with our eyes. but an LLM understands them by looking at the actual data. Yeah, and for our listeners who aren't aware of the strawberry phenomenon, there was a while where ChadGPT would gaslight you to convince you that there were, what was it, two Rs in strawberry? Oh, it would commonly change. It would really never get it right. Yeah, it'd never give you the correct answer no matter how you ask the question. And I believe they had to manually fix that one use case to get it to work.

3:27Yeah, the strawberry news was tragic. and very viral. Perception is a hell of a drug, especially when you combine it with pattern matching like us humans commonly do. So were any of the GPT models they tested actually able to detect the existence of this gorilla? Not successfully, no. Even when given the picture. However, it was noted that Claude was able to kind of start to pick out that there's some sort of depiction in the image. It doesn't go as far as to understand that there's a gorilla, you know waving at it in the plot but it does start to understand that there might be visually something depicted it goes one step further they even say that because there's an image depicted in the data it's highly likely that that data is generated or not authentic in some way and so then claude warns the reader about the data they're using and its authenticity which is a fun twist for the story to end on yeah it's almost like it detect it detects that like something is waving at it, but it doesn't think that it's important to know what that thing is.

4:30More importantly, it knows that in its world that scatterplots don't wave at you and scatterplots don't have pictures in them. So the fact that this one does both something suspicious. Yeah. Now I want to switch gears to this, this story that you brought up about banks, you know, in financial institutions trying to catch up and the battle for getting, for hiring developer talent. So I read this article and, you know, I really broke down how, financial services industry as a whole is trying to seize an opportunity to hire more top technical talent from all these big companies that are out there undergoing layoffs.

5:05And I think these layoffs have become such a seasonal occurrence that many of these big tech companies are kind of losing their allure. It's like a certain number of people are going to get laid off every January. That's not a great environment to work in. And traditionally, I think a lot of developers, particularly from these big tech companies, they gravitate towards those companies because they just have like, they want to work there. Like they pay really high salaries, they're great institutions, and they've not really viewed financial institutions in the same light. But the reality is a lot of these organizations are now modernizing their tech stacks, they're increasing budgets into development, giving them access to all the latest and greatest tools like AI and machine learning and all the cloud technologies that we know and love.

5:53And compensation is also going up with this, as well as flexible work arrangements and the benefits that they offer to be more in line with the big tech companies. And there's actually been – there's an element of geography too. U.S. banking hubs are a lot more dispersed. So I think as a result, it's easier for them to get into the remote work movement and to find developers that maybe would work remotely for these big tech companies but live in a city where these banks have a headquarters. But the main reason I brought this up, because it reminds me a lot, actually, of a conversation I had with Louis Vega from Bloomberg LP back in May of last year.

6:32And he builds all these really cool internal tools for Bloomberg. I really loved learning about the engineering culture they built because, you know, one of the coolest aspects was how he was describing how they create these like unique brands for each of their projects. So like a log manager that you call logger, like without an E, and then you go to a designer and get a logo for it, gets added to this internal marketplace for all the other developers of the company. Like they really kind of celebrate building cool things and sharing them internally. So if you haven't listened to that episode, I definitely encourage you to check it out after you read this article or before you read it.

7:09Yeah, definitely a good one to check out. I also really resonate with what you're saying about what you learned from this article. I do think it says a lot about how tech is being more distributed everywhere and technology is not something relegated to the tech world per se. The entire world's undergoing transformations with technology and AI is pulling that even faster forward now. So the result is there's so many opportunities to even be a tech person within a non-tech company. That's how I actually started in technology. By being that technical expert for non-technical folks that are trying to execute something at scale, you can actually find that your insights can go really, really far in places that are maybe non-traditional to where you work from.

7:56And this also reminds me of even some of the insights that I learned from today's guest, Vikram Chatterjee of Galileo, who's coming up here in just a few moments. And in order to implement AI at the scale you're talking about in a bank, you have to prototype rapidly, but you have to measure that impact and ultimately build trust in your organization, which is probably going to be very rigid against that kind of rapid change. But before you can get there, you have to attract that top talent and catch up people-wise. So this news article makes perfect sense to me. It even relates to what we saw recently with Goldman Sachs acquiring that AI transformation head from a tech company.

8:37I think this is a trend we're going to see over and over again this year. Yeah. And one thing I've actually seen some specific examples of is when a lot of these more traditional organizations launch a more coordinated software development arm of their company, they're often set up as almost like an independent entity that gets a lot of freedom that you don't typically see within a financial sector company. So that allows them to move a lot more quickly to use new technologies. So yeah, you can't, you know, banks don't write them off. If you're one of the unfortunate people out there who's been laid off, or even if you're not like, you know, go check them out.

9:12Maybe they're a cool place to work. But, you know, speaking of your skills and where your skills can take you, there's another article I read this week that talks about how the skill of the future is not AI as everyone's talking about, but it's focus. And this one, I loved this article. it really highlighted how leaders need to create a culture where engineers can validate and understand AI's inputs and outputs and seek alternate solutions and understand ultimately what's happening underneath. That way you can engage in a stronger understanding of how the AI is working, but also then harness that understanding to get really good uninterrupted work done.

9:51In sum, basically use AI and understand it, but then harness that free time and focus on the stuff the AI can't do. And it really talks about how you build a team culture, the harness that unlock. Awesome. Yeah. And it kind of reiterates a point that I've been making for a while is that, you know, a lot of these GPT models are, they're very accurate, but they lack precision. So what I mean by that, and a good example of this is if you go out and ask any of these GPT models to create a picture of a waving gorilla, for example, they're probably we going to create something that, you know, if I took the output and handed it to our producer, Adam, and asked him what it is, he would probably immediately say it's a waving gorilla.

10:34Like they do that quite successfully. But if you look closer, you'll probably see a lot of like micro errors that it makes, like, you know, to borrow a phrase from Westworld, the answer's in the hands. Like you can tell that it's fake or it's AI generated because the hands look weird. And in that environment, it's actually very difficult for humans to find errors, right? If everything on the surface looks correct, but it's actually filled with a bunch of problems, you need focus to suss out those problems. And it goes to show like we've been seeing research and anecdotes of the difference in how junior developers and senior developers adopt AI.

11:09One thing that's come up again and again is that junior developers are a lot more likely to just blindly accept AI output and then try to fix any problems after the fact. Whereas a senior dev might treat it more like a sparring partner where they ask it some questions, get some ideas from it, and then the developer themselves goes out and creates the code themselves, you know, asking for advice where they need it. In that environment, you don't have to catch as many of those precision problems. Well, some great stories we had today. Andrew, why don't you let us know about who we're having on the show today?

11:42Yes, we've talked about how AI is changing the way engineers work, but also the risks of blind dependence. And if we treat AI as an oracle instead of a tool, then we lose the ability to question and refine and truly understand the solutions it generates. And focus is the real skill that separates great engineers from the rest. So in today's conversation, we're talking with Vikram Chatterjee, the co-founder and CEO of Galileo, about how to future-proof your AI engineering investments. Galileo is building the system of record for AI models that help teams open the hood and understand what's happening with their LLMs underneath.

12:23Stick around. Habits are a powerful thing. Maybe you've read one of the many books about the habits of highly successful people out there. Well, Linear B is out with their own book, The Eight Habits of highly productive engineering teams. This practical guide offers advice and templates to help you establish durable data-driven habits. It covers things like setting actionable team goals, coaching developers to level up their skills, using monthly metrics check-ins to unblock friction, and run more efficient and effective sprint retrospectives. This guide has something for everyone within your engineering team, so check out the link in the show notes for the eight habits of highly productive engineering teams.

13:08Hey, everyone. Welcome back to Dev Interrupted. I'm your host, Andrew Ziegler, developer advocate at Linear B, and joining me today is Vikram Chatterjee, co-founder and CEO of Galileo. Vikram has been on the front lines of the AI revolution for many years, from leading product management at Google during the birth of Transformers to building tools that help engineering teams confidently evaluate and deploy AI systems. Here's the crux of today's conversation. Engineering leaders are stuck between a rock and a hard place. They know they need to experiment with AI to stay competitive, but they're under immense pressure to justify those costs.

13:50All the while, AI evolves so rapidly that today's wrong move could cost them tomorrow's opportunity. Vikram, welcome to the show. Thank you, Andrew. Super excited to be here. Likewise, let's jump right in. Starting with the biggest challenge currently facing engineering leaders, AI being that moving target and experimenting feeling risky, it's a big barrier to adoption and the fear of making the wrong investment is ever present. How do you think leaders can take the first steps without putting themselves or their teams in a bad position? I've always thought about AI as it's another tool in your arsenal, right?

14:30even before generative AI became a very big thing with classical machine learning and with NLP, it was never about like, hey, you have to use this thing. It was more about what's your use case. And based on that, is this a good fit for your use case? Now, I guess the difference is what we've seen with a lot of engineering leaders that we talk to, there's a lot of top-down pressure to just use AI for the sake of it. To get to your point about the rock, between a rock and a hard place. It's very important for engineering leaders to think about what are the heuristics, right, that is going to help me figure out, you know, is this something that we should be even going ahead with?

15:08And that includes things like, what does the business need? Because what I've seen is folks just do a massive hackathon within their org, and you're going to get a hundred ideas just given how open-ended AI can be right now, right, in terms of whether it's generation of text, whether it's generation of images or completion of a task with agents, you get a hundred use cases. But it's very, very important for them to then go back to like, you know, product and business owners to figure out which one of them should they prioritize. And also on the back of that, which one of them can you actually get out the door very quickly?

15:44And that's kind of where the operational rigor has to kick in of like trying out X number of ideas very, very fast. So you have to have that machinery in place and the ability to push back and say that, hey, maybe I don't need to use AI at all. But when I do, here's how I need to do it. And at that point, we can talk about this more, but we have to think it through in terms of if I succeed, what does that mean for me in terms of the number of engineers I need for this, the evals, the cost of productionizing this thing at scale. So there's a bunch of things that people have to think about and a lot of trade-offs at the onset itself.

16:16So it sounds like in balance with that proliferation of ideas, you really need an evaluation framework or a way to understand and extract scenarios that are maybe have a higher ROI or a higher impact on your business and focus on those. Because I think that's kind of part of the problem, too, is you're drowning in possible solutions and everyone can come up with maybe a way to integrate it in some way. But is that the most effective way? Is that where we should focus our attention? and the more attention you put on something, it can skew than how the rest of your organization is using and thinking about AI.

16:50So those decisions, especially early on, they're really impactful. How would you advise or what habits do you think make for somebody to be able to evaluate and de-risk that experimentation? Like what are those tools that those kinds of people are always using again and again to do that well? Yeah, it's a great question. I will say it depends a lot on the organization and you have to just know your organization well. So if you're a big bank, right, the amount of harm that can happen if you put something out there in the consumer world with AI and it misfires is very, very large. It can literally derail your entire bank's reputation.

17:30For a commoditized entity like a bank, you're going to be out of business very quickly. On the other hand, if you're, let's say, a DoorDash or an Instacart, the bar is probably a bit lower because not to say that they have a low bar in general, but if something goes wrong with their chatbot or something like that, it's not going to be the end of the world because they're not dealing with people's money. They're dealing with hungry people, which is bad, but they're not dealing with people's money. Maybe you didn't get my burger, but, you know, my bank account didn't have like an unauthorized transaction or something.

17:59They're totally different stakes. They're totally different stakes. And so what I've seen as a result of that is when you talk to these large enterprises, they're very, very excited about generative AI and they're very excited to add agents and everything else. But they're taking a very, very careful crawl, walk, run approach to it, which I think is good. What I'm seeing with the other companies where they are earlier stages, they're tech first. Let's say, as an example, like a DoorDash or like an Instacart or like an Airbnb or a Twilio, they're taking a much more experimental approach to this.

18:29right like let's try things out let's see how it goes let's see how we feel it's very much on the lines of you know build fast break things learn quickly and that's really led to them kind of figuring things out as they go and it is also useful when as this industry is moving super quickly so based on the organization i've seen like the barrier to entry from a fault tolerance perspective is different now within that if you're an engineering leader at a faster moving company if you're all about experimentation and going fast then the question becomes how do you plan well and there is a certain crawl walk around there as well to be honest andrew because what we've seen there is you have to think about you know is it going to be 15 use cases 10 use cases you have to have some forecasting there because once based on the forecasting you have to do a couple of things you first think about the compute costs and then staff yourself with enough i don't know a100s and have that like available for you so that everyone can just build otherwise everyone is going to come back to you and say like, hey, where's my GPU at?

19:29So you have to have like X amount of compute and give that out very judiciously. You then have to think about like, how can you optimize the compute and, you know, start to invest in tooling at that layer to minimize the cost of compute. And then comes the eval piece, which is kind of what Galileo does, where you have to think about how do you create the right kind of guardrails in place for like an AI, CICD process, and then basically go to the teams and say, awesome, all right, you want to you want to launch things, you want to experiment with things. Here's the stack that you can use. Go knock yourself out, right?

20:01And some people will use the entire stack, some won't, but you have to create that enablement almost within the team before they can like go crazy. That's a big unlock. So let me try to unpack this playbook because I think there's a lot of really interesting tips in here. One of them being first and foremost, understanding your company, your company culture, their risk tolerance and what they're doing, and understanding that there's a big difference between a traditional enterprise and a digital native company experimenting with AI right now, especially with different levels of risk within what they're working on.

20:35So it's about understanding your own company, your own environment, the level of risk and the tolerance for experimentation. But then getting into that, you know, crawl, walk, run loop that's going to get you up and get you moving on this process. It's about creating the actual resources to enable those teams to make effective tools. And so if you're going to create a way for them to grow within your company, you have to think about resourcing them and prioritizing them based upon, again, you know, maybe the profile of your company culture. So it really does start with understanding your culture.

21:11Exactly. And you hit on a good point. It's the culture, it's the people. Again, it comes down to like also the business that you're running, whether you've an Instacart or DoorDash, again, going back to that example, they have many different ways of instituting general AI, but maybe there's a different company that really doesn't have those many use cases and that's okay. You kind of have to look at all of those different angles and then figure out how you want to act and how quickly you want to act. But it does stem with that. As an engineering leader, I would say like step one is always just that.

21:42Right. And part of company culture and it's kind of what I want to focus on next is also, you know fears and anxieties around making the right or wrong decisions or creating tools that are doing jobs that traditionally people did within their org and understanding how people need to reprioritize their time to better use these tools and in doing all of this i think we're all trying to make systems that are future-proof we don't want to rebuild these ai workflows again and again and again we want to make it once and evaluate it and iterate on it So that's part of justifying to the ROI and going back to creating the resources within your team for that to grow.

22:22So how can an engineering leader, perhaps someone who's in a company culture that's maybe more in the digital native side, have an appetite for risk and experimentation? Maybe they even have a little bit of resources. How can they shift those conversations with non-technical leadership from immediate gains to creating future-proof solutions that are going to help the company in maybe like a year or five years from now? It comes down to building trust. I think the first thing becomes like as an engineering leader, I've talked to a lot of different leaders in the space and it all comes down to number one, first having a gut instinct reaction, a gut instinct on their own side, personally.

23:03Like some kind of a trust that You know, these use cases are great. And this use case is actually going to be a good one to start with. The second is staff that. Just go very, very fast and build out a prototype and start to see, does that actually add value? And then shop that around with others in your business. They could be leaders in the business, depending on how large the company is. And you can now with Gen.AI, the unlock is that you can build a prototype pretty quickly. And so you can at least start to get a sense for like, what is the appetite here? And that's typically what I've seen happen.

23:34and then from there you can start to figure out what the KPIs are, right? Because the main thing is you want to be able to ship something pretty quickly. You want to go to production quickly with the right checks and balances in place. So you start with at least one to two use cases as quickly as you possibly can with the right checks and balances and guardrails in place and then start to see how that looks. Roll it out to 1%, 5%, 10%. Start to see what you learn and then with that playbook you can go very fast with everything else. This is exactly what we've seen with the largest banks in the world, as well as the digital natives.

24:09It's just that the digital natives are just moving much, much faster than the largest banks. But this playbook is kind of similar on both sides. It makes a lot of sense in terms of how those teams can get started. It's also about evaluating, going back to the evaluative frameworks from earlier and having measurements in place to look at the results week after week. And that creates within the company culture, a healthy socialization and understanding of the tools that people are building and we're using. And that's what cultivates the trust. And it's also something really hard to find in an LLM based world.

24:44You know, LLMs are stochastic. They don't want to repeat themselves. So when you put them into an environment where you want a repeatable workflow, where you have an AI agent evaluating things on the fly, there's so much to consider. And that's kind of part of, you know, the dauntingness of that I was alluding to earlier. In our initial conversation, you know, there was something that you mentioned to me that's really resonated with me since. And as you mentioned how AI agents in the future, or even now, they can act as smart routers or kind of like load balancers for workflows and help optimize them over time.

25:19And that was really fascinating to me. I'm wondering if maybe you could dive into that a little more just on a technical level now that we've talked about introducing these products and or these projects within your company. if you are somebody who's revolutionizing a workflow right now with AI, how should they best kind of think about it? What I'll say is in terms of agents in particular, how that works. So for context for folks, Galileo is a leading provider of evaluation tooling for any AI developer out there that's building an AI app, right? And what that means is an eval basically includes not just the metrics, but also like the data set that you're working with.

Read the full transcript

25:57and it includes like a workflow so you can really understand what your failure modes are, build out evals around that and then use that in your CICD process. So you're not, you know exactly if you're shipping a good product and once you ship, you can check for regression test. That's the net net of how evals work. Now with agents, what's interesting is we've kind of moved from people building like just chatbots. That was like, it almost felt like that's just as far as the imagination could go in terms of use cases. Now it's like you can complete any task. What's the task that your product does.

26:28So, you know, we started seeing this with operators launch as well, which is a good example of an agent where all of a sudden I started to see booking.com's folks talking about like, hey, now you can book a hotel room with just natural language. And then box.com's CEOs started talking about how you can add files and folders. So it just unlocked all these use cases. So that's kind of what we're seeing with agents coming. Are these apps perfect? No. And here's why. Because what's happening with agents is it's essentially a way to say like, hey, LLM, instead of just generating something, why don't you just act towards choosing the right kind of tool?

27:05Or here's a function, I've wrapped it in a certain manner, do X for me, right? So making it do specific things. The tool piece of this especially is interesting because now the API can act as, as you mentioned, as a smart router to figure out, you know, which tool should I use for finishing this task? Which is, I think, the biggest unlock right now because earlier you would just code all of that. You would literally deterministically say, this is exactly what you need to do. Versus now it's almost like, I just want this done. You figure it out. You figure it out for me. Which is also why there's almost like this leader worker relationship almost with an agent that's happening.

27:41And so that's been a very big unlock because now these agents can just find out the right tool and go ahead. However, it does lead to different kinds of failure modes, right? Did it choose the right tool or not? Did the right tool get called in the right way or not? Even if it did choose the right tool, what happened after that? Can I see the entire flow of how all of that worked out? At the end of the day, did it complete the task? What do you mean by complete the task? How do you measure quality of task completion? Did it plan the whole task properly or not? So there's a bunch of these different kinds of qualitative stuff along the way that you need.

28:19And they need to get evaluated in the evaluation system that we've been talking about. So all of those things need to get looked at. Because when you're in this environment where it's sitting as a router or it's making smart decisions, it's no longer like a chatbot where it's query response, query response. Instead, it's request or demand or you need it to do something and it's then a sequence of actions. and those actions lack visibility to you. You're not staring at it, making the decisions or the actions or making the API calls. You just see what it tells you at the end. So that's kind of where, that goes back to building trust as well because you need to understand what's happening in those stages.

29:03Yeah, exactly. It's funny how similar this is to working with a human being. Yeah, exactly. That's what's coming to mind for me too. I really liked what you said about leaders and contributors, about it kind of being this balance of that you as an engineer are overseeing the output or the work of the agent and you're responsible for its success just like how a manager is responsible for their IC's you know success you have to give it the right tools the right environment and the right context so it's a it's I think it's a big unlock are you kind of seeing that from folks that are starting to engage in these workflows we are because that's exactly the kind of question that they're asking around like, hey, how do I make sure that everything worked fine?

29:44And also if it did work fine and the answer is correct, can I see what the route was that it took? Because maybe it's just making unnecessary API calls that it doesn't have to in the first place. Can I optimize this even further? So there's a big question around what are the failure modes, but also can I have a visualization into like how the entire AI app like actually made its decisions and how it was planning so I can maybe tweak things here and there. The system itself, I think the folks at Databricks, Matei Zaharia coined the term compound systems for what we at Galileo basically called your AI app.

30:19That compound system is becoming more complex because of these agentic scenarios where essentially people are adding function calls. So it's fascinating because now it's even closer to classical software engineering and it's all about like how good are your functions and how well are you managing it all? And those engineers come to us and they ask us about like, great, this is fun, but how do I, what is a unit test for me now? And what is a regression test for me now? And that's where evals come in. Right, I'd like to understand a little more about how does Galileo open that box so that you look in to understand how those tools are working and how somebody would use something like that to build confidence.

31:00Because what you're describing is like a whole new category to me of how we think about tools. And when you're in an environment where you're defining a novel category, that's really, you have to do a lot of like definition setting and understanding. So we're all for the first time opening the box and looking inside on a workflows, workflows that we're building now for the first time. What does Galileo provide? What are the things that people should be looking for? What Galileo tries to do is our end goal is to help you build high quality AI apps fast. That's our end goal. So we win if you're building those apps 10x faster and those apps are 10x better.

31:40That's our goal. Now, in order to do that, that includes what I think of as the visualization layer, meaning like you could just see your traces and spans. I feel like that's the easy part. It is that it's good for them, but it's highly commoditizable, right? Anybody can build that thing. So we obsess about the user experience at that layer. But then beyond that, what we've been seeing is we initially gave the developers the ability to just build their own metrics as well and score their agents. And what we saw was developers struggled with that because they kind of had an idea about like, hey, I need to see if it's planned this thing out well enough.

32:14But then in order to build that metric, they almost had to build a very complex prompt. They had to figure out the instructions. They had to figure out how do you optimize this? Do I use a GPT-4 model for this, something else? So that's the layer where we basically realized that, wait, this is actually ripe for a lot of research. So we have a fairly large staff of AI researchers that are constantly working on our Luna evaluation metrics, where it's not just the prompt instructions and things of that nature, but we also focus a lot on how can we optimize the cost and latency of these different metrics.

32:46So what does this mean for the user? What this means is I've built my agent and I'm building my agent. You could just use Galileo's TypeScript SDK or JavaScript SDK to be able to start logging your application. and then on the other side without any ground truth needed you basically magically see not just the visualization layer but you also start to see two things one you'll see these very very highly explainable metrics show up with an explanation for exactly what went wrong and why the second thing that you see are automatic insights as well which are fairly easy to understand because what we want to do essentially as a developer what they want as well is i just I'm trying to build this agent out.

33:25I cobbled together a few things. I'm trying to run a quick experiment. Which part of this complex compound system should I focus on? Should I focus on this API, that one, the prompt, or something else? And so we basically dumb it all down for them as much as we can and tell them, we have all your logs. We also have these metrics that we've built out based on all of this. Here's what we think you should do. So it's almost like a co-pilot for your AI application development. That's the journey we want to go with them on as they're going from like building towards scaling. That's very fascinating to me because when someone's going to look at all of these different variables, what Galileo is doing is it's helping you isolate those variables and find the ones that are going to have the biggest impact for you to focus on, which kind of resonates with like the whole top line objective of folks having to, or being in a position where they are experimenting with AI or building to workflows as they're trying to evaluate and justify where they should be spending their time because it's moving so quickly that their time always needs to be spent on the most impactful part of the project or anything else you're working on is likely just adding risk to the project because you're doing things that are probably going to be outdated by the time they're really in practice or really in use yeah and that requires like a whole mental i think flip about how we build um and how we look at them.

34:44So it helps you isolate the variables and make those tools a little more natural to develop in a classical way where you can evaluate and understand. Yeah. Yeah. No, I agree with you. That's exactly right. And when someone is maybe building and managing a bunch of these AI bots or tools or workflows, you know, are you seeing that it's like people are like almost like doing performance evaluations on their agents? Like maybe like how somebody would do for maybe someone's getting ramped up, right as a BDR or as a customer success manager. And there's, you know, basic tenets that you want them to do across the board and you're evaluating their interactions with customers or whatnot.

35:21And because you understand what is good and what is bad in the environment of your company's culture. Are you saying that that's how people are using and evolving as they build these tools? Yeah, it's similar. It's very similar because to take your example for a second, like you, let's say you hire an SDR, the SDR is mostly focused on outbounding. And then you got to see the quality of that outbound, who they outbounded to, what the result of that was, did they book a call or not, and all sorts. There's an entire funnel there. Imagine all that's being done by an AI agent. That AI agent's basically going to have to call, do a bunch of different kinds of API calls to make sure all that happens.

36:00And now the question becomes, great, if it's an AI agent that's not sleeping at night, then how do I make sure that there's some sense of potential failure modes that can happen? You would probably do this with the SDR as well, a human SDR, but you'd probably want to have like some kind of inspection, right? You have an expectation setting, but I want you to book 10 calls this week. That's my expectation. And then with a human, you have some level of inspection around, great, like how many outbound calls do you do today? And then I'm going to start to look at a funnel. And similar, you come up with that mental model of what those potential failure modes would be as the person was building the app.

36:36And then you have to build out the guardrails accordingly. And then what happens is as you go through the motion of interacting with that human SDR or the agent, or the AI agent, you start to figure out more failure modes because you go deeper and deeper and deeper. You're like, oh shit, this can go wrong. That can go wrong. And then on the fly, you're going to have to create more evals and more metrics. And you also start generating more data around like, ah, you know, like when it's trying to reach out to this specific kind of person with this specific queue, that's when it's failing. So maybe I should isolate this data a little bit.

37:08That's when you start to create like this, this data set that you want to test against and these metrics. So that's the flow that people have been going down of like exploring and creating these models and making it part of their process. What really stands out to me about that is it can get very complex and it sounds like a whole new set of skills even. From your perspective, what do you think is the most powerful habit or skill that an engineer or an engineering leader should be picking up right now to stay ahead of this kind of curve? They have to build. I think the best engineering leaders that I've seen, they're doing two things right now.

37:43They're building, they're keeping themselves upskilled. I think everyone has to do that. I certainly do that all the time. Everyone has to do that. And if you're expecting your team to build out these kinds of AI apps, you have to get very, very familiar with it. That's one very large part of it. And the second piece is just being in touch with the community, like learning from each other. Because I've noticed that organizations that are moving really fast are the ones where they have inch leaders who are also talking to other inch leaders about what they have learned really, really quickly. Because everything's moving so quickly, you can't wait to make those mistakes yourselves, right?

38:17So I'm seeing folks where there's like a good cross-pollination amongst leaders that they're moving much, much faster. It could be that, you know, an inch leader might tell somebody else that, hey, you should get Galileo because without evals, it's going to be really bad. and the other one is going to probably learn that the hard way. It has happened with a lot of engineering leaders before. So I feel like those two things are very, very important. Keep being at the forefront of learning by doing because you're an engineering leader, test off those skills and you can actually build simple apps on the side.

38:46And the second one is just being a big part of the community somehow, being in touch. Build, read, and communicate. Three core defining traits. And that's a really powerful takeaway. for me what really stands out about that is that people who put in place small incremental changes over time you know they get ahead faster and faster in an incremental world like ai somebody who's building that tool yesterday is going to be much more ahead of you tomorrow if you didn't build it so those engineers that are out there building and talking about what they're building and they're reading about what other people are building those are the ones that are getting ahead and are staying on the forefront.

39:28And that's a really great takeaway. And Vikram, this has been an incredible conversation. There's so much insight packed into what we talked about today. I want to thank you for sharing and giving our listeners some actionable takeaways. Before we wrap up, where can our audience go to learn more about Galileo and the work that you're doing? Yeah, for sure. So we are at Galileo.ai. You can also check us out on LinkedIn and on Twitter. we post a lot of content on our website and we put that in our blogs and the website. There's an entire research section where we publish our papers and everything else that we've worked on for AI evals.

40:05We've been around for four years. So there's a rich history of a large body of work there. So they can check it out there. We're also hiring right now for a lot of engineers who have built out their own AI apps. Excited to chat with anybody who's interested in being at the forefront of helping builders build. That's fantastic to hear. and we'll definitely, I'll include those notes in the show notes on Substack. So if you listen to this and you're interested in getting involved with Galileo or following up on anything that we talked about today, you know, please be sure to subscribe, share our episode.

40:35You've made it this far. And if you check out our Substack, there's even more insights from today's discussion. And we'd also love to hear from you on socials. So don't be a stranger. And that's it for this week's Dev Interrupted. We'll see you next time.

40:55We'll be right back.

From the publisher

Ben and Andrew open the show by dissecting why AI can't see gorillas, how big banks are stepping up to attract tech talent, and why focus is becoming the must-have resource for devs.

Then, Vikram Chatterji, co-founder and CEO of Galileo, joins Andrew for a discussion on how engineering leaders can future-proof their AI strategy and navigate an emerging dilemma: the pressure to adopt AI to stay competitive, while justifying AI spend and avoiding risky investments.

To accomplish this, Vikram emphasizes the importance of establishing clear evaluation frameworks, prioritizing AI use cases based on business needs and understanding your company's unique cultural context when deploying AI.

Check out:

Follow the hosts:

Follow today's guest:

Referenced in today's show:

Support the show:

Offers:

More from Dev Interrupted

All 208 episodes
Can Your AI Strategy Be Future-Proof?Dev Interrupted · 41 min
Listen in VO