From DevOps ‘Heart Attacks’ to AI-Powered Diagnostics With Traversal’s AI Agents

24 Jun 2025 · 41 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: From DevOps ‘Heart Attacks’ to AI-Powered Diagnostics With Traversal’s AI Agents

Overview In this episode of the *Training Data* podcast, co-founders Anish Agarwal and Raj Agrawal of Traversal discuss how their AI agents are revolutionizing the way enterprises handle system failures. Their technology enables rapid root cause analysis, drastically reducing the time taken to identify issues from hours to just a few minutes. This conversation dives deep into the implications of AI in DevOps, troubleshooting, and the ongoing evolution of software engineering.

Key Concepts AI in DevOps and SRE

  • DevOps and Site Reliability Engineering (SRE): Essential roles in maintaining the performance and reliability of software systems.
  • Traditional Incident Response: Often includes chaotic communications via platforms like Slack and requires significant manpower to troubleshoot.

The Role of AI Agents

  • Root Cause Analysis (RCA): The process of identifying the fundamental reason for system failures.
  • Traversal's AI Agents: Designed to automate RCA by analyzing complex dependency maps and logs to pinpoint issues quickly.

Metrics and Observability

  • Golden Signals: Four key metrics (latency, traffic, errors, saturation) used to monitor system health.
  • MELT Data: Stands for Metrics, Events, Logs, and Traces, a framework for observability.

Key Takeaways Current Challenges in DevOps

  • Traditional processes are cumbersome, often involving many engineers and prolonged investigation periods.
  • High severity incidents can lead to significant downtime, impacting business operations.

How AI Transforms Troubleshooting

  • AI agents can perform RCA in 2-4 minutes, contrasting with hours spent by human teams.
  • The integration of AI aims to let engineers focus on creative and strategic planning rather than routine troubleshooting.

Predictions for the Future of DevOps

  • DevOps roles will evolve to be more strategic as AI handles more tactical troubleshooting.
  • The future workforce will require both SRE principles and fluency in AI, as systems become increasingly complex.

Observability Landscape

  • Current observability tools are fragmented, often requiring multiple platforms to get a complete view.
  • Traversal's approach aims to unify data sources and streamline the troubleshooting process.

The Impact of AI-Generated Code

  • As AI becomes more integrated into coding practices, the challenge of debugging systems that are not entirely human-written will grow.
  • AI-powered debugging tools are essential for effectively managing these systems.

Product Insights Traversal's AI Agent

  • The agent works by orchestrating various tools to fetch logs and metrics, analyze data, and conduct RCA.
  • A key architectural decision was to use a read-only access model to avoid complicating enterprise data structures.

Accuracy and Measurement

  • The effectiveness of the AI agent is measured through live incident evaluations and post-mortem analyses.
  • Traversal claims to achieve over 90% accuracy in identifying root causes when the data supports it.

Cultural and Organizational Insights Team Composition

  • Traversal's team is predominantly composed of engineers, with a focus on creating a culture of rapid iteration and experimentation.
  • The shift from traditional software engineering to AI-focused roles emphasizes adaptability and ongoing learning.

Observability Teams of the Future

  • Expected to be more hybrid in nature, blending traditional engineering skills with AI fluency.
  • Understanding of AI reliability will become crucial as systems become increasingly automated.

Conclusion The discussion highlights the transformative potential of AI in the DevOps landscape, particularly regarding system reliability and troubleshooting efficiency. As AI tools like Traversal's gain traction, the industry may see a shift in the skill sets required for engineers, emphasizing the need for a blend of technical knowledge and AI capabilities.

Recommended Readings

  • The Bitter Lesson by Rich Sutton: A brief article on themes in AI and learning.

Final Thoughts The future of software engineering will see more integration of AI, potentially allowing engineers to focus on high-level strategic planning rather than getting bogged down in routine incident responses. As tools evolve, the way teams operate and troubleshoot issues will also need to adapt, signaling a significant shift in the DevOps landscape.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00If you're in product, if you're in design, if you're in core engineering, you constantly have to make bets as to where it is going to be six months from now and then you're willing to re -evaluate everything six months from now. The good news is that it's only going to get better. So it's not like your product's going to get worse six months from now. And so for example, one really interesting thing bet that is paid off is Raj and Raj's were very pressured about the fact that reasoning models are going to get better. They've made this bet in September, just seeing where the world is going. And we architected our system such that the reasoning models would get to shine.

0:35And that has really played our dividends right now in the way our entire architecture is set up.

0:57Today we're excited to welcome Anish and Raj, co -founders of Traversal. Traversal is building AI agents to transform the world of DevOps and site reliability engineering. Today, companies have war rooms and armies of dev ops and SREs, troubleshooting their production failures and fixing them in the code base. Every minute of production downtime is costly and sometimes life or death. Keeping a company's application running is very hard and valuable work, and it's a problem that is only getting worse with the advent of AI -generated code and vibe coding. Traverse will believe that agents are the scalable solution to the problem, and that seen exciting results are ready from the troubleshooting agents that they've deployed into production.

1:33We're excited to have them share more about that vision and the results so far. Anishan Raj, welcome to training data. So wonderful to have you today. We have been working together for a year, and we can't wait to share with our audience how AI, NAI agents can transform the world's of satellite reliability, engineering, and DevOps. And somehow we need to weave negronies in the conversation some place later on, but we'll get to that. Let's start with some quick hot takes. Will DevOps or SRE, S -Win -O -M today, even exist five years from now? Yeah, it's a great question. I think it will, but I think it'll look fundamentally different from the way it looks now.

2:17And I think in this world of DevOps and SRE, I think healthcare analogies generally fall quite naturally, I find. And in this world, we think about healthcare and the mass law hierarchy of needs there. Imagine the stage one is where, let's say you're having a heart attack. You have to solve that right now. Nothing else matters. Nothing matters five minutes from now other than solving the heart attack that you're having. Right, and to me, that's analogous to dealing with the high severity incident. And then stage two is where, let's say you're dealing with some sort of chronic issue that you face, you know, you have spray knuckle or whatever the, whatever illness that might be affecting you.

2:56And it's very hard to plan three months in advance because that's something that's just affecting you every day. And I think that's the analogous thing you do in DevOps is dealing with streams of alerts, streams of checking whether deployment was safe or not safe. And then stage three is where I think of it as like life hacking, where you're thinking about how do I optimize my sleep in my nutrients so that I have a high quality life that's you know, and fulfilling life, right? And I think that's the equivalent of like planning out what your next five years of infrastructure look like and how you're investing the right places.

3:32Now, I think of what unfortunately people in the DevOps SRE or Encore engineering space live like right now, it's like having a heart attack twice a week and dealing with a debilitating chronic condition that you have to handle every day. SM listening to you, I'm just trying to picture all the people who I know in DevOps and SRE in scrubs or wearing CGI monitors to life hack, like, you know, their infrastructure. And I think people don't realize that when you make this connection to healthcare, you realize how what we've gotten used to and what life really should be like in this world, as we think about the healthcare of large -scale software systems.

4:10And so I think if traversal does job right, then that first stage and second stage of needs, which is there's really high severity incidents and that constant pain by death by a thousand cuts should be something that AI and AI agents take care of. And DevOps and SRE get to deal with the creative fun parts of what is it my infrastructure should look like for the next day for the next five years and it becomes a much more fulfilling job. And many engineering teams are nowadays adopting autonomous coding things like cursor, wind serve. Do you think that's going to have profound impact on how people maintain the reliability of their infrastructure or will it be the case that AI will have a major role to play in fulfilling that vision, moving people away from like being the intensive care unit surgeons to being more thoughtful planners.

5:04So I think it's going to be, there's a short term answer and there's a long term answer and I think it becomes in the short term at least it's a tale of two worlds I think. So, there's this because of everything happening in the world of cursor and windsurf and so and so forth. I think the idea of vibe coding obviously has become very popular, right? And we can write a prompt or a few prompts and have something stood up that people can use and play with. But I think the analogy here is of fast fashion where you try something on, you like it, that you don't like it, you discard it, start where the next thing, right?

5:40And I think in that world actually, reliability doesn't really matter, because you don't really need to take care of what you've created, there's no craft to it, you've created it through it, right? And I think, but there's another world, which is where you start applying these AI -powered software engineering techniques to mission critical systems, in payments, in financial institutions, in security, infrastructure, streaming. And in that world, I think it's gonna lead to a major issue, it probably already is, and it's only gonna get worse, in my opinion. Because everyone, we've seen this with the enterprises we work with, everyone is using AI -powered software engineering tools to guide them as they write code.

6:19And what you actually find is that it actually even passes their local unit test. So, in that local piece of code, everything looks perfect. Right? The problem is that in lot -scale enterprises, things break when different pieces of your system interact in a way you just didn't realize, or you didn't couldn't foresee. And when that happens, because all of this code is being written by an AI system, it's very hard to debug it because you just don't have the context anymore. You didn't write it anymore. And I think in that world, unless we find ways of using AI systems to do software maintenance, we're going to be throttled.

6:48And either people will disallow AI software engineering tools to be used because there's too much downtime or something of that form. And I think in that world, we're going to need new tools, a new software to help maintenance of of such systems. I think that's what Krivos and Kulip can play a part. You definitely have your way with words and analogies, so I love the tale of two worlds. You either have Louis Vuitton, high -fashion, or you have fast -fashion machine or or or or something like that. But I think we jumped into into the deep end, maybe like a bit too quickly. Let's go back a few steps and maybe explain for our audience what what actually root cause analysis stand for?

7:34Like what does it do? Tell us a bit more about this wonderful world of troubleshooting and root cause analysis. Yeah, so I think, if I just think about the tail the life cycle of an incident, right? The story of an incident. It always, surprising looks the same across companies. A customer will log on to a platform. Things are not working exactly right. Either it's already been recorded or they make it complain to customer support. customer support looks at and says, okay, it's not a user error, it's an actual issue. It escalates in this game of telephone up like five labbers of different sophistication of engineering organizations within an enterprise.

8:11At some point it hits the DevOps SRE team who look at the issue and decide, is this worthy of an incident based on the impact of it, based on the severity of it, the immediacy of it? And if they decide it is worthy of an incident, then suddenly chaos in series, right? you'll have a slack incident channel that's created. There's 30 to 50 people in the channel. Everyone's kind of implicitly blaming each other for what happened and a bit of a who -done -it type thing. And almost without fail is the same thing that happens, which is there'll be this 10X engineer who's been at the company for - Who's the Agatha Christie?

8:44Yeah, I'm the Agatha Christie. I'll show a lot of poems of the situation. And like, aha, they'll figure out exactly what happened, right? and not clear how they got there, but they get there. You roll back the, you come up with some sort of hot fix and then everyone kind of, it looks for a longer term fix over time, right? And that's typically how it all plays out. And that whole flow is what we call the lifecycle of incident and root causing analysis in particular, right? And I think what I find incredible is that all of observability, right? These incredible companies have been created. The other second largest software spend typically companies have after your cloud spend and yet we still have the state of root cause analysis.

9:30But anyways, that's my take on what RCA looks like nowadays. And tell us maybe like a bit more color on where does root cause analysis live with relations to the plethora of observability tools that companies use today. a wise ad that if this still the case, if this is the number two spend in technology, wise ad that you have 50 people in a Slack channel with a Houdani kind of plot. It speaks to the importance of the problem, right? And what observability is done is the best they could have done given the technology that was available to even like six months ago, right? Which is, so I think what observability is, it's fundamentally about the creation of telemetry data.

10:13It's called melt data, so like metrics and events and logs and traces. So melt for short. And it's about the creation of this data, storing it, and then providing a nice visualization layer on top of it so people can create the right slices to create little eyeballs in different parts of their system. Right? And I think that's what observatory NSAT, which is a storage and visualization layer, because that's all I could have been done. And the complex workflow troubleshooting remains super manual. And that's just because that's all we could have done so far. And I think the promise of what these AI agent companies can do is that they automate complex workflows that we do in top of software.

10:54And to me, this problem of troubleshooting would cause analysis is one of the most complex workflows that any study humans do in software. And I think now there's an opportunity to kind of move up in terms of sophistication where we use AI systems to actually automate this workflow of us are just stopping at the storage and visualization now. Can you talk about your product? What is the agent or the agents that you're building to kind of take on this problem? And how does it work? Yeah. And sort of a double agent? So maybe stepping back, what is an agent? And it's really an LLM orchestration of tools.

11:27And in our world, those tools might be data fetching tools, how to get logs, how to get metrics back, data processing, how to format it in a way that the agent can really process or statistical tools like running a normally detection. And I think really what we're trying to look for here is you need to define the tools rich enough so that an RCA can be expressed as some combination or some sequence of these tool calls. And that's really where a lot of the complexities and the multi agents come in is how do you piece together these tools to solve these complex tasks. And I think at least in our world, the data is so big that it needs to be sequential, because a single trace might not even fit in LLM context.

12:13So you need to know how to slowly just build this context of the LLM to get to the root cause. And that's why it's really challenging of a problem. Are you mirroring the cognitive architecture of a human troubleshooter? Is it, you know, it's a map of operations where a similar to what a human would be doing? Or is it very different for an agent? Yeah, I think that's where at least when we were first thinking of the problem, we tried to mimic how an SRE would debug. And it was very manual, very sequential, where an SRE typically might look at a piece of evidence and then, okay, figure out what's the next piece of evidence to look.

12:51But a lot of times they make these hops with system knowledge, which the agent might not know. And that's really where we used a lot more scale to figure out how to bypass some of this issue of system knowledge. So this agent will basically look in a more systematic fashion of, there's this Google SRE handbook, which tells you, here's a key golden signals, you should be looking at latency, error rate. So it looks at this way of kind of health monitoring to piece together how to sequentially search this rich space. and it's less about having holes in the reasoning where a human can just immediately get maybe to that hop.

13:33The agent is making these sequential flows that get to the answer. Are there particular conditions or environments where this agentic approach works better or worse? Yeah, I think fundamentally it's about data access. So what we have found, and surprisingly that's what we have found is that we sometimes can add less value in a series A startup versus a large enterprise. Because when you become a large enterprise, your observability is quite mature. Everything is being instrumented in a way where the fundamental data is in place, but your teams are very fragmented. There's no one team or no one person has enough contacts to piece together or everything about how to debug.

14:17That's why these 30 or 50 people in a Slack incident room. And so what we fundamentally found is that the place where we add the most value is when the reasoning steps for this agent to go from a high level trigger to the root cause, those steps can be found in the data. Right? It's too much data for any one human to keep in their head. When that data is fundamentally not there, then it's when the agent suffers, and that tends to not be the case of the enterprise is what we're found. Very interesting and really somewhat counterintuitive because it's usually the case with startups that the way you choose your design partners or customers you start at smaller companies or mid -market and then you try to graduate your way up to enterprise.

15:02And if we're going to use like the L1 to L5 like analogy from self -driving cars, how close Are you with having agents correctly get to the root cause of incidents or maybe more broadly like what? What counts as success like are we trying to get to a hundred percent? Correct resolution or what even counts as correct resolution like in this context? Yeah, that's a great question. So I think there's really two cases here There's the first case where the root cause fundamentally belongs in the data. So it's in some log. It's in some PR And in that case, I would say with traversoral, we're at L4 and I won't say L5 because for us, we might be able to flag that problematic PR or that smoking gun log, but then there's still that last mile of the fix and the remediation and in cases where the fix is not so localized to a specific file or code change where you need to do basically a bigger system -wide change.

16:07That's we haven't gotten there, but I think that's where it's really exciting to see all the developments with code agents because that's where we can get to that L5. I think for places where the root cause isn't in the data, that's where we're more at a L2. I think traversal finds a lot of the important symptoms that really help people debug. But sometimes it's just not in that data. And this human kind of needs to make those additional hops to get to the actual root cause. But still, the symptoms really help figure out how to make those additional hops. And sometimes, right for us, when we notice that, we can tell customers, well, maybe you should instrument in this way to make the system more observable.

16:51Across the AI landscape right now, there's really interesting kind of AI native companies being formed like yourselves. There's also the incumbents that are, you know, in many cases not asleep at the at the whale. How do you think about the incumbent risk here and why your customers are choosing to go with you? Yeah, I think it's a great question. I think the heart of it is, observability is expensive, right? It's such an expensive product that people pay for. As a result, it's extremely fragmented. Right? You go to any big enterprise. They're using data dog, they're using Splunk, they're using DietRace, they're using Elastic, using Grafana, using ServiceNow, I mean, using everything.

17:28And all of them are encroaching in this game of attrition. And the problem is that if you just think about the pricing models in this world, it's all based on the metadata of the store. And so the result company A is not incentivized to give you better insights from anything that store in company B. And to debug something, if you just look at any of the SRE on college NIAs, they're calling upon all five, six tools that they have access to. And is that fragmentation of this historical industry, I think, which is going to ideally lead to companies such as ourselves that are somewhat agnostic to whether the data is stored at least for now, it gives us a chance.

18:04Do you have customers deployed on traversal today? And what have you found is the difference between what you kind of expected that you've been coming from an academic background versus what you're actually finding in real -world environments? Yeah, it's been quite a journey. I think when we started, as one should typically do, we started with like very small companies and build something that worked for them. And that typically took the form of, we went through the last hundred incidents that they had, right, in the Slack channels, and kind of try to use our brains to figure out, what is the meta workflow of how they always debug an incident in that particular company, right?

18:39And in that situation, a very popular framework for agents is something called the React Framework, right? And that, as a general idea, we could somehow imbue the meta workflow into a React agent system, and it was able to really, a really good job. And then at some point, we started working towards larger companies that had an actual like a lot scale observability system, thousands of microservices that kind of situation. And I think I remember very clearly this one week where it was the first time we dealt with that kind of scale. And the second we tried to apply our system on just some historical incident that happened, in our accuracy with a zero percent.

19:19And we just, it would not move. Whatever we did to the prom, so whatever we did to anything, it was stubbornly a zero percent. And that was a rude awakening. I remember Raj and I had a negroni.

19:36And I think at that point, we kind of had to think through ways of, we made some interesting decisions which played out in our favor, which is we said, okay, nothing about a specific company will be hard coded into the prompts, right? And nothing about a workflow that humans do there will be hard coded into the agent like workflow. And that complexity has to go somewhere, right? And eventually where it went to is computation, which in our world takes the form of spending tokens in the problem, like using it at inference time. And once we were able to kind of find an architecture that exploited inference time compute, which is something now everyone is finding to be important.

20:19Accuracy starts shooting up. And what we found then is for if the fundamental answer lies in the data, we get to the answer more than 90 % of the time. And we get it within two to four minutes. Right. And which is amazing because now you just look at these slack channels. Humans are spending most of the time just verifying the answer to what we find versus actually root causing it. And just the time, the month to month time resolution has dropped, I'd say, as is the number of people in average in this LAC channel, which is two things that I think any enterprise cares about. And how do you measure accuracy like in this context?

20:51Like, for example, there are a lot of companies that focus on LLME valves. Like, does such a thing exist in the world of like SRE and RUCAWS analysis, or like, you know, how do you know that you are correctly identifying a RUCAWS? The gold standards, honestly, trying it on live incidents. I think when you onboard to a customer, oftentimes incidents can happen like two to three times a week. And in those scenarios, you get the best feedback. Maybe it takes a couple of hours for that incident to complete. You look at the post -mortem and you can really evaluate it. So I think that's one definitive source.

21:28There's other subtasks we do for evaluation, such as when people are trying to search for specific information from their observability. There's smaller chunks of tasks you need to do to do RCA. So we evaluate on those tasks as well, which are just higher volume, but ultimately, real live incidents are the best way for us to evaluate. Awesome. It sounds like you know you had this, uh, a rude awakening at some point where you deployed the product and it was stuck at 0 % accuracy before you kind of went back to the drawing board and really rethought like the entire architecture. So maybe can you give us like a very quick tour under the covers of how the product works today and kind of what is the magic that enables?

22:13Precise worth cause analysis. Well, one important decision we made was we only require read only access to the data And I think that was a decision basically based on enterprises not wanting to just have yet another tool to generate more data. So how do we actually do it? Well, I think there's two phases. There's an offline phase and an online phase. So during this on offline phase, we're really trying to learn this rich dependency map. How do different functions, how do different logs relate to each other? And one way to do that is through LLMs. So LLMs go traverse, really understand semantically, how these logs, different tags within the logs all relate to each other, and then we also use statistics.

22:58and statistics comes in when, for example, there's natural variation in time series, and that turns out to leave traces of causality, which basically a niche that I worked on in grad school of how do you pull causal relationships out of this data, and that's really key to build this rich map. And we also use self -play to basically prioritize certain paths that are very promising for RCA. So now, once we've constructed this rich dependency map during the online phase when an actual incident comes to us, what this agent is doing is it's using that real -time information and this dependency map to basically figure out what hops to make to do the Rucos analysis.

23:41How long does it take for the offline part to become effective, such that in other words, between deploying your solution and the first incident that you can actually troubleshoot how long is that gestation period? It kind of depends on if they want to troubleshoot a live incident or a historical incident. In general, I would say we take about five to ten hours to kind of look through all of their code base, look through their observability, really have that system understanding. For larger customers, it can take a day, but generally five to ten hours. So you've mentioned reasoning and inference time compute a couple times in there.

24:24Are you using kind of foundation models and fine tuning them for your purposes? Are you building a lot of that kind of architecture yourself? Maybe just talk a little bit about the architectural decisions you've made? Yeah, one really interesting thing we have learned is that if you work with enterprises, they typically have a existing relationship with an alumni provider. So they might have an enterprise control of the OpenAI and Thropic. And if you try to bring your own model or your own fine tune model to them, you're going to be stuck in security hell for a barrier. And so you have to kind of tie your hands where you say you have to be able to use whichever model I give you.

25:02Typically, OpenAI is a pretty safe bet and Thropic as well. And so then most of the complexity is really about how do you get the right set of tools that this LLM has access to to orchestrate the RCA itself. And the other thing you can do is fine tune within an company's environment. So, you know, let's say they pointed the R Azure OpenAI instance, then you can fine tune that by every time an incident happens, you see how far you got, you saw whatever the pin would cause buzz, you can see how far away you were, and then the system can fine tune on making sure that gaps get small and small over time.

25:38Right, so that's generally how it we've seen it's played out. And you also mentioned that you guys had part of the architecture here is based on years of your academic research. Can you share a bit more about kind of this interplay between your PhD dissertations and translating that now into a company? Yeah, actually, so one thing that's kind of interesting is at least during grad school I worked closely with the Broad Institute to basically understand what these CRISPR interventions, you have these gene regulatory networks, you do a CRISPR intervention, and you're trying to understand what is the effect of this drug or this knockout experiment on how these genes express.

Read the full transcript

26:24And the techniques we developed there for learning that causal structure between genes turns out to be pretty related to this problem we're facing with production systems where if you think of the nodes of swapping them with genes as microservices and then you're learning okay what happens when I make a PR change or I break this part of the system how does that percolate it becomes almost the identical problem and I think that was honestly by sheer luck we didn't know that was gonna be the case and only until we got into the weeds of the problem, did we realize, like, oh wow, we got really lucky that our grad school research played out well here.

27:05What has been some other surprises on this journey over the past year, year and a half? Well, for one, I think, which is part of why it's been such a joyous thing for us is just how it's like the industrial age of AI. And so I think all of the most interesting innovation, and I feel it's happening in small research -focused startups, such as the ones that we are part of, or have the privilege to be a part of. And so I think just having seen the best of research in some of the best universities and now actually seeing how it gets played out in a company, it just feels like this is where all the magical special work is happening.

27:43And I think that was surprising to me coming from a world of academia, becoming a professor. I thought that's where all the innovation will be. And so that's been quite surprising. Honestly, I think the other thing that's been quite surprising is just how hungry enterprises are for using JNVI to solve real problems. And you can show that the pace at which you can move with them, what typically would take a year or two years to close, can be rapidly done in a couple of months. And so I think the hunger of the market and also the pace of innovation in industry has been quite surprising. Speaking of the pace of innovation, things are moving so quickly and you have been working working on this product for slightly more than a year now.

28:24Have you been in situations where you have to go back and rework something? Oh, because now there's MCP or something new on the market that previously you had to engineer yourself, but now it's readily available like all the shelf, so to speak, or in some ways, like how do you future -proof your architecture and what you're working on? Honestly, I think that is going to be a constant challenge for companies of this generation. If you're in product, if you're in design, if you're in core engineering, you constantly have to make bets as to where, as you're going to be six months from now, and then you're willing to reevaluate everything six months from now.

29:01The good news is that it's only going to get better. So it's not like you're products are going to get worse six months from now. And so for example, one really interesting thing bet that is paid off is, Raj and Raj's were very pressured about the fact that reasoning models are going to get better. They made this bet in September, just seeing where the world is going. We architected our system such that the reasoning models would get to shine, and that has really played our dividends right now in the way our entire architecture is set up. I think you have to keep making these kinds of six -month bets about where AI is going to be and surf it just right.

29:40I think the companies are able to string together a few of these bets just right, other ones are going to win. I think that's why in my opinion, if you asked me, would you rather be an AI team taking this problem more? And it was our ability team on bias, but I'd say our other ability, an AI team, because you just have to, you get more fluent in making these kinds of bets about what things are going. What is your team composition? Like how many are researchers? How many come from the domain? Like, I'm curious about this, what the shape of an AI native agent company will look like more broadly?

30:10Yeah, I mean, right now I would say we're 90 % engineers. I mean, a lot of people, you know, there's a few of us with PhDs who have done PhDs in machine learning. A large fraction come from like traditional software engineering backgrounds who have been very interested in Genai and have used it in their own life. But I think for some of these things, at least the barrier is lower and you can kind of get everyone's trying to learn how to make an agent. And I think it's not like the old days where you might have needed a PhD to write out these gradient updates on this crazy LM. I think it's a lot more democratized and that reflects in the way we've made the team composition.

30:57We also have people working with infrastructure. So how do we scale these agents? And I think that's a really interesting problem of, to Anisha's point of inference time compute, is how do you really get this AI agent to just do a lot more work than a human can possibly do? And that involves for us, we make thousands of network calls per investigation. So that's reflected in hiring the best infrastructure engineers. And then I think the other part is we have product managers and some of those folks are coming in the next month. So we're really excited just to bring all of these different people with different backgrounds together to solve this hard problem.

31:40Yeah, I think one thing that unifies everyone of the company is an excitement for a gender V .I. And I think because of that, they're willing to learn and adapt and throw away things that didn't work, which is kind of being stuck in your way. Because I think those are the people that are not, that's the type of culture that is not going to survive in this world. Because it's just moving too quickly. And if I think about what it's interesting, I think about an AI researcher versus an AI engineer versus an engineer, like how do I organize that? I think there's nothing to do with whether you have a PhD or any sort of credential.

32:12It's about how experimental are you in the way you think? Like how willing are you to make little bets and hypotheses about how things can be and then quickly put that to the test, right? I think that mentality of quick iteration on a with AI and data is I think what makes someone a good AI research or an AI engineer. And so it's more a mindset than a credential. Speaking of the composition of your team, I find it very interesting that Nita, one of the two of you or your two out of co -founders are actually observability industry insiders. And your entire company is very heavy on AI and engineering prowess.

32:53And honestly, quite short on observability, like domain knowledge, and yet, you are making waves in this domain and successfully resolving these very complex issues, but otherwise take little armies of people scurrying around and dumpster diving to go figure out what is the smoking gun log. It's almost begs the question of, What do you think a future observability or SRE team at one of your customers might look like? Today, you mentioned that there are people who have all the insider knowledge, like all the tribal knowledge, the intuition, they can make these hoops from here to there, without data, is that still available skill sets, like five, ten years from now, or would the observability team look very, very different?

33:48A couple of answers. One is, I think, there will be a lot of work to be done about the reliability of AI systems themselves. And I think the way you reason about reliability of AI systems will require, obviously, for this principle, a good SRE principles, but also some sort of fluency with AI as well, because they'll break and interesting unique ways that we're just not used to. And so I think modern SRE teams will have to kind of be fluent in both in some ways, which is how do LLMs and AI systems break and also how regular systems break. And marrying those two things with what I think will make you really good DevOps in SRE engineer.

34:28And also even simple things like, I think one very interesting thing that's gonna happen is that the data layer of observability is gonna fundamentally change as well. like the way a log looks like is going to look fundamentally different. And a part of being a good SRRU DevOps person is knowing how to write a good log, right? So that over time when instant happens you have the right instrumentation. There's an art to instrumenting your system in the right way. We're not flooding too many logs, but just the right amount, right? And I think the way we will write logs will look fundamentally different as well, because it's no longer meant to be scrolled by a human, but meant to be consumed by an AI system.

35:03That's a very interesting point. How do you rate cursor today on variability to write good logs? I think it does a pretty good job, but I mean I think it's better than a human. I think the way it can form it. Human just make typos, and they don't, I mean I'm really bad at spelling. I think cursor does a great job there, but I think a lot of it is still mimicking how traditional logging looks. I think to Anisha's point, you need less of that, and you really, for example, want to Barry is much information and the message field because that's the meat of what is going on. And now when the LLMC's at it can kind of have a better understanding of how to make those hops.

35:43But a lot of times people don't like to add so much content because a human can't read such a long Aristach trace in a log. We have shorter context windows than the Nizela Lums. I think also then connecting with the business logic, right, a lot of times because typically it may not be engineering system is fine, but it depends on what finding even means when engineering system, it has to connect to some sort of business logic for it to be healthy or unhealthy. And I think logging in a way that connects these two things in the right way is an art still. And I think that will be tough for cursor or any other AI system to do unless they have full understanding of your business logic, which might happen, I don't know.

36:21As you dream about the future of the software engineering market and the role that observability and traversal will play in that future. You started this podcast episode talking about, people aren't vibes coding up payment system or banking system today. Do you think that if everything that traversals building goes right, if observability changes in the way you think it might, that engineers at the largest banks and payments companies and healthcare institutions will be actually vibes coding their software programs? To some extent, yes. And I think what we will have is just a much better sense of, like, did it fulfill some function?

37:02And that will typically take the form of unit test in some form, right? And just much better testing. And so as long as it fulfills some test, that code is fine. And who cares what was written that code? I think that will somehow, in my opinion, is how software engineering is going to evolve, which will be much more about, did it fulfill some function versus the way it's written? which will be good, but as we talked about earlier, I think also bad because when systems fail, it's because they interact in ways you didn't expect. A lot of times, you are failing because this third party that you call upon is not in the SLA.

37:33And that's going to happen a lot more, I think. And it's going to be a lot harder to debug because you don't actually know what was the content of a particular code file. And so I think that's the problem that we're going to face. And that's why I think products like traversal will have a big part to play, is solving those kinds of issues, which is a lot more subtle. Because I think that's where humans shine, right? Is right now, humans mostly write all of their code, so they have that rich system knowledge. So when they're seeing an incident, they can short circuit a lot of paths, because they already kind of have seen something in the past, or they kind of already understand this error message they wrote.

38:11But if an AI generated that, they don't have that structural advantage, and then you really need a AI troubleshooter. Okay, we're going to close out with some rapid buyer questions. You guys ready? Yeah, so. One sentence or one word answers. Maybe first, what application category do you think will break out for AI in the next 12 months? Well, try to say this in two sentences. Which is, I think, anywhere where you're using reasoning models better, our way applications are going to shine. And I think diagnostics is one of them, healthcare being, I think, a prime example of that. I know the startup or founder of a two admire.

38:46I would say Demis Hossibis, hopefully I didn't butcher the name, but I think just an amazing researcher, so much foundational work, and just always trying to solve the hardest problems. Recommended pieces of content, how to make yourself smart on AI. My favorite article which you can read in five minutes is the better lesson by Rich Sutton, who won the Nobel Prize last year for his work on reinforcement learning. Yeah. Any other AI agent companies that you admire? You're one of this very first cohort of truly agent native companies. Any others? Outside from the obvious, like, Glean and, I mean, I'm just a perplexities one for sure.

39:29I think is just doing a fantastic job. Favorite new AI applications that you're using in your personal life? I honestly don't use anything other than Chad GPD. Unfortunately, the same for me. It's already, it's my best friend at this point. Wow, I love it. Maybe one last question. Will we be vibes coding, banking apps, payment systems, healthcare systems in five years' time? Do we have something to send you? Wonderful. Thank you so much for joining us today. We really love this conversation and congratulations on what you've built so far at Jversal. Thank you. Thanks so much.

From the publisher

Anish Agarwal and Raj Agrawal, co-founders of Traversal, are transforming how enterprises handle critical system failures. Their AI agents can perform root cause analysis in 2-4 minutes instead of the hours typically spent by teams of engineers scrambling in Slack channels. Drawing from their academic research in causal inference and gene regulatory networks, they’ve built agents that systematically traverse complex dependency maps to identify the smoking gun logs and problematic code changes. As AI-generated code becomes more prevalent, Traversal addresses a growing challenge: debugging systems where humans didn’t write the original code, making AI-powered troubleshooting essential for maintaining reliable software at scale.

Hosted by Sonya Huang and Bogomil Balkansky, Sequoia Capital

Mentioned in this episode:

SRE: Site reliability engineering. The function within engineering teams that monitors and improves the availability and performance of software systems and services.

Golden signals: four key metrics used by Site Reliability Engineers (SREs) to monitor the health and performance of IT systems: latency, traffic, errors and saturation.

MELT data:  Metrics, events, log, and traces. A framework for observability.

The Bitter Lesson: Another mention of Nobel Prize  winner Rich Sutton’s influential post.

More from Training Data

All 110 episodes
From DevOps ‘Heart Attacks’ to AI-Powered Diagnostics With Traversal’s AI AgentsTraining Data · 41 min
Listen in VO