In short
Podcast Summary: Engineering in the Age of Agents with Yechezkel Rabinovich
Podcast Information
- Podcast Title: Software Engineering Daily
- Episode Title: Engineering in the Age of Agents with Yechezkel Rabinovich
- Episode Description: Discussion on modern software platforms, microservices, observability, and AI's impact on software engineering.
Key Guests
- Yechezkel Rabinovich (Chez): CTO and co-founder of GroundCover, an observability platform.
- Kevin Ball (KBall): VP of Engineering at Mento and independent coach for engineering leaders.
Main Themes
- The Challenge of Modern Software Platforms
- Distributed Systems: Modern platforms consist of numerous microservices, third-party APIs, and cloud resources, complicating system observability and troubleshooting.
- Operational Risks: Lack of clear visibility into systems can lead to increased operational risks and inefficient troubleshooting.
- GroundCover's Approach to Observability
- eBPF Sensors: GroundCover leverages eBPF sensors to capture critical system logs, metrics, and traces directly from the kernel.
- Bring Your Own Cloud Model: Ensures data privacy and security by keeping all data within the user's cloud environment, thus offering cost efficiency while maintaining control over sensitive information.
- The Impact of AI on Software Development
- Accelerated Code Generation: AI-generated code increases the speed of development, raising challenges in code review and validation.
- AI in Code Reviews: Mentioned the use of AI tools for generating and reviewing code to ensure quality and compliance with standards.
- Modern Observability Needs
- Complex Integrations: The interaction between different components in modern systems can be unknown until visualized through observability tools.
- Centralized Data Visibility: The need for all observability data (logs, metrics, traces) to be in one place for effective analysis and correlation.
- Future of Observability and AI Integration
- Root Cause Analysis: Aiming for deeper understanding and automated assistance in incident investigations using AI.
- Enhanced Testing and Guardrails: The importance of creating deterministic systems that can validate AI-generated code and its implications on architecture.
Technology and Concepts Discussed
- eBPF (Extended Berkeley Packet Filter): A technology that allows for efficient and safe monitoring of system calls and application behavior without altering code.
- Observability Pipeline: A framework to process and interpret observability data, emphasizing the cleaning and structuring of logs before analysis.
Future Directions for GroundCover
- LLM (Large Language Models) Observability: Tracking and understanding API calls made to LLMs and correlating them with workloads.
- Using LLM for Observability: Developing systems to clean and structure observability data using AI to facilitate better inquiries and troubleshooting.
Conclusion This episode emphasizes the evolving landscape of software engineering, where AI and observability are crucial for managing increasingly complex systems. GroundCover's innovative approach aims to provide engineers with the tools to maintain visibility, security, and efficiency in their software operations. The discussions also highlight the need for better practices, tools, and understanding of underlying architecture in the face of rapid technological change.
---
Key Takeaways
- Observability is critical in modern distributed systems to mitigate risks and enhance troubleshooting.
- GroundCover's eBPF-based platform provides comprehensive insights while ensuring data privacy.
- AI is reshaping code generation and review processes, necessitating new approaches to testing and validation.
- The future of observability will involve deeper integration with AI, focusing on providing actionable insights for incident management.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Modern software platforms are increasingly composed of diverse microservices, third-party APIs, and cloud resources. The distributed nature of these systems makes it difficult for engineers to gain a clear view of how their systems behave, which can slow down troubleshooting and increase operational risk. GroundCover is an observability platform that uses eBPF sensors to capture logs, metrics, and traces directly from the kernel. Critically, GroundCover runs on a bring-your-own-cloud model, so all data remains within the user's own environment, which gives increased privacy, security, and cost efficiency.
0:38The company is also focused on adapting to how AI-generated code is changing observability. Code can now be produced at superhuman speed, which increases the challenges for reviewing code before it enters production. This means that observability is likely to play a growing role in code validation and providing guardrails. Yehezkel Rabinovich, or Chez, is the CTO and co-founder of GroundCover. He joins the podcast with Kevin Ball to discuss his journey from kernel engineering to building an EBPF-powered observability company. The conversation explores the power of EBPF, the realities of observability in modern systems, the impact of AI on software development and security, and where the future of root cause analysis is headed.
1:27Kevin Ball, or KBall, is the Vice President of Engineering at Mento and an independent coach for engineers and engineering leaders. He co-founded and served as CTO for two companies, founded the San Diego JavaScript Meetup, and organizes the AI in Action discussion group through Latent Space. Check out the show notes to follow KBall on Twitter or LinkedIn, or visit his website, kball.llc.
1:59Hey, Chez, welcome to the show. Hey, thank you. Nice to be here. Yeah, I'm excited to have this conversation. Let's maybe start a little bit about you. So can you give our listeners just a little bit of your background and then how you got to GroundCover, where we are today? Yeah, sure. More than a decade in software engineering, specialized in Linux and distributed systems. So mainly worked on kernel modules until I got sick of it and then kind of got in love with eBPF. Then we founded Roundcover, which I'm the CTO and co-founder. And that's what I'm doing for the last four years. eBPF is really cool.
2:36We did another episode about that and I hadn't been exposed before. My previous Linux background had been 15, 20 years ago. And then coming in and being like, wait, you mean to integrate with the kernel? I don't have to go through this arduous patch process and kernel process. It was mind blowing. So actually let's maybe even start a little bit there. So how are you utilizing eBPF for ground cover? Yeah. So as you said, after a few years with kernel models, you kind of fall in love with the power of extending the kernel, which is very cool, but it's also very, very hard. The development cycle is so slow because any mistake you do basically can crash the system.
3:16So after a few years as a R &D manager and leader, me and Shaha, my co-founder, kind realize there is a big problem with instrumentations. We all know the SDK instrumentation. You have all the classic vendors. You have Auto, the open source standard. But still, you have to instrument your application, right? And it sounds very easy. You have a lot of tutorials. But in real life, it's very, very hard. Most companies have a lot of different runtimes, different versions. And you need to keep on track on all those SDKs. So why won't we do it with eBPF, which basically let us instrument the application from the kernel side without any risk for the application itself.
4:01You run in a sandbox, you have the kernel verifier that will not load you if you're doing something wrong or that can potentially harm the application. But you still get 95 % of the value. So we can inspect any syscall. We can intercept HTTP requests, SQLs. Redis calls, whatever you're doing, we can probably see it. So maybe leads to GroundCover, which is an observability company. I would say modern observability company. We utilize eBPF alongside with classic instrumentation, but our main sensor is based on eBPF. So you deploy the sensor in less than a minute, you get traces, logs, metrics, everything in one place without changing your code.
4:47The other cool thing that we do is our backend is based on bring your own cloud, which means your data stay inside your cloud account. Very nice in terms of privacy and security and also allow us to reduce costs because we're not charging by volume and your customers usually happy with it. Yeah, for sure. Well, one of the things that you alluded to there is an area I think we can dive into a lot. So you said a modern observability company. And I think a lot is changing right now in terms of how we build applications, how we deploy applications, how we need to be observing them. So what is needed for modern software development from the observability side?
5:29I think nowadays, average platform is so complex with integrations, third-party cloud resources, different SaaS vendors that you're using for feature flag or hosting. And the average engineer kind of struggles to even know what components the platform is relying on. So that's where classic instrumentation fail, right? Because it's like the unknown unknown, right? You don't know what you don't know. And this is a nice experience we see with customers that deploy the sensor for the first time. All of a sudden, they see those links between their application to third-party. It could be even something very mild or very small, like fetching an avatar from a third-party website that they didn't even know or things like that.
6:21So I think nowadays, the basic of modern observability is to have all the information in one place. You'd be surprised how many of our customers before they used Brankova used five, six, seven different tools. So their signals were across different platforms and just the correlation between that, it's so hard. So I think having all the data in one place, this is the very bare minimum of observability platform in these days where most companies have 100, 200, 300, microservices and the number is just increasing. Yeah. Well, and the fact that you're able to gather all that data without the engineers having to add the instrumentation to their code means, yeah, you're able to capture those unknown unknowns because if an engineer didn't think of it, if they had to instrument it, it wouldn't be there.
7:13It's 100 % visibility on what your application is doing. It's very nice and very mind-blowing for the first time. Yeah. Well, and I think there's another big trend which makes this very interesting, which is It just feels like with, in particular, like the way that software is changing within the AI development world, like the volume of change going on is so high that being able to even keep track of like what's going on has just gotten harder. You can see it even from how the development lifecycle changed. You know, a few years ago, not that many. It would take you a lot of time to write the code.
7:48Something that product could explain in a few words, it just takes time. Engineers optimize for typing speed, right? There are competitions around how fast can you type. Currently, that barrier basically does not exist. Any software engineer can basically write superhuman code speed, velocity, and this is no longer a barrier. You basically can print unlimited lines of code in minutes. Are they doing what you think they're doing? maybe that the architecture that you planned to be, maybe. And, you know, how do you code review it? I am feeling that pain tremendously right now. So what's your answer?
8:35How do you code review it? With an AI, of course. No, I'm half joking, but to be honest, like obviously at GroundCover, we use AI to write code. We use AI to do code reviews. Because even one of our engineers just wrote a utility that used AI to create the pull request with ground cover flavor. So it collects the information from the ticket, maybe the Figma that correlates to the ticket, and basically generates a PR with AI and that the engineer could just edit the last mile. So obviously, the short answer is, of course, we use AI for code reviewing. But at the end of the day, engineers still need to be accountable on the software that we ship.
9:24So I personally think that testing should be very, very mindful. Because I've seen tests that are actually making sure there are bugs because the tests are also written by AI. The code review in the AI era for me starts with, first of all, let's look at the test. That's been true forever. But with AI, it's just sometimes impossible to read all the code. And maybe sometimes it's unneeded. There are different software and different requirements. For instance, me personally, I just wrote a very simple library that can pass metrics QL to logical representation. We integrate that in the platform. Apparently, this does not exist.
10:19I was shocked that this does not exist in TypeScript. And I wrote it. And I don't know TypeScript. Never code in TypeScript. So I manually crafted all the scenarios, half manually, right? I described all the scenarios and then pressed tab. But I manually crafted the scenarios that I think logically would challenge the platform, the lib, and then follow on the implementation and seems reasonable. And I've never read the code. But for this kind of library, it doesn't really matter, right? It's simple input of string. The output is very simple. There's no side effects for using that library. It's not something in the hot path of the platform.
11:05So that makes sense. But on the other hand, if we implement a new parser inside the sensor, which runs high throughput, zero allocations, code itself matters. You have to know how much do you allocate? Do you have any memory leaks? So I think the question is very dependent on the context of what are you building. And that's something that we at GroundCover started to differentiate. Like, let's think, what are we checking now? Does the code matter? Maybe it's not. Maybe we can replace it in a week and it doesn't matter. Yeah. I mean, I think to some extent what you're describing there reminds me of a metaphor I've used before, which is increasingly the code gen is essentially like a compiler, right?
11:56Like when was the last time you read the binary that a compiler for your code generated? You didn't, probably. But you did check, did it do what I expected? Did it pass the tests? Does it behave as I anticipate? And for a lot of code now, that's essentially what it is. Does it pass my tests? you described a few other things that might matter like performance functional pieces like how do you validate those in a world in which you know agents are building all of your code yeah i think that the difference between transpiling maybe compiling c c code to assembly is you don't miss out on the architecture when you when you change c language to to assemble but But when you only test for input output, you do make sure that that piece of software does what it needs to do, but you can't make sure it does it in the way you want it to be.
13:00And it's also not deterministic as compilers, which tend to be very deterministic for the most cases. And I think it is different in that, what about the architecture? How does it fit with future features that you want to integrate? I think when you talk about architecture, the comparison between classic compilers doesn't represent it well enough. Yeah, well, that's definitely true. Or at least you need a new source code, right? You need technical documentation or tests or some sort of validation. But yeah, no, you're right. Not all the metaphors align. I think this does, though, get you into this sort of interesting thing with what you all are doing, which is tests are one form of validation.
13:46Reading the code is another form of validation. Another form of validation is like, what do the logs say? Is it allocating memory? Is it performing? Is it doing all of these different things? Like having some sort of feedback loop between the actual running code and whatever is writing it, whether it's an engineer or an agent, is extremely important. Yeah, the problem with LLM is that it can lie, right? It's statistics. So I don't trust logs that are being generated by LLM because, you know, it just can make up logs. Maybe it refactored that code and forgot to rename that variable. Maybe it thought it would write it and it eventually did not.
14:30So I personally less rely on logs and also markdowns or cursor rules. I less tend to use it because I think if you want to embrace AI with the options that it will fail, you have to put some guardrails that are very deterministic. So you mentioned tests. Tests are brilliant for that. Tests usually don't lie. I've seen cases where, you know, AI kind of inject every scenario in that code base to possibly pass the test, but it's very rare. It's very rare. And I anticipate we're going to see it less. But I think linters should be very fashionable now. I think we should have more complex rules to make sure the complexity of functions is something that we can live with.
15:31Or even conventions are something that we want to enforce because eventually you are going to read that code. There is a chance you're going to read that code. we have to make that assumption because end of the day when you wake up 3 a.m something is wrong you can tell anyone that you're trying to craft a prompt to fix it but it's your responsibility that you're not sharing with that ai bot right and we're still we're still relying on humans at the end of the day for for the future but for the near future it still looks like it so i I personally think EBPF comes very nice with this because EBPF will tell you the truth.
16:14This is the HTTP request. It's not dependent on your code. This is what happened. Exactly. We even start to instrument testing environment, getting the traces back to the AI and saying, look, this is what happened. What do you have to say about that? But you have to create guardrails that rely on a solid ground truth. And it's a good observation for why your observability should be separate from your code base itself, because you don't want the LLM to write lies into it. That's really cool. Now, a challenge I've seen before with observability is it can be very verbose, right? There's a lot of stuff that happens in a modern system.
16:55And you get, you know, thousands and thousands of lines of logs or what have you from relatively simple interactions. That's one challenge on the pure data storage and transfer side. But if we're feeding these back into LLMs, there's a context management challenge too. How do you all think about that? Yeah, I've seen a lot of people saying, hey, let's just send all those logs to OpenAI. And they will tell us what's happened. And then you realize in that five minutes, there were 20 million logs. You're right. And when we introduce AI-based code, it actually increases the data that we're sending.
17:34So that even just makes things worse. What we're heavily trying to do at GroundCover is being able to summarize data on a stream aggregation fashion to basically represent brands or patterns. So for instance, logs are the simple use case of, you know, you have 50 log lines. And if we can nail the patterns, we can actually convert them to kind of time series. Time series is very compact. Any AI agent happily look at a graph and say, oh, this is interesting. And then you can narrow it down to a service or a timeframe that will allow you to dive deeper. So this is one way. This is maybe the simplest signal to summarize.
18:27But we're also doing it for APIs. We create baselines on interactions. This is a bit more tricky because the dimensions on what is the same communication pattern is a very complex issue, especially when you look at it from a network perspective. Just to give you an example, think about a query param that generates 1 million different routes. You need to understand that those are the same API and this is a variable. That's a very simple example, but we are trying to compact all those signals to baselines. And then this is something that you can speed the AI for trends. and then narrow it down until you can find the raw data, but still manageable to send to an agent.
19:24Another interesting idea is to look for other signals, for instance, change management. So think about image change, something happened. This is probably a good place to start looking around that time and you basically narrow the time and with that, you narrow the context needed. So if I understand a little bit, and just to kind of replay back to make sure I'm on the same page. So you've got kind of looking for a set of patterns, the ability to build a baseline, the ability to kind of pull in any sort of external information, like when an image changed to sort of show, here's a relevant area and a high level description.
20:03and then you sort of expose some sort of explorability so that the agent that you've passed that off to can say, okay, great, this looks like it's a problem. Let me, give me more data for this spot. Yeah, yeah, exactly. And you can also go wrong with this process and then go back because sometimes every engineer knows that, you know, you see some suspicious log and you think you found it. And two hours later, someone's coming, someone comes to the office and say, oh no, this is not the issue. I know this. And the same thought process will happen to AI. It's an iterative process, but to make it efficient, you have to start with some kind of patterns or time series represented signals.
20:51And then do you expose those to your LLM via an MCP server or how do you approach that? Yeah, so GroundCover was one of the first observability vendors that exposed MCP. It actually, we started before MCP had the OAuth authentication, and we were almost ready to release, and then this got announced, and then we recreated the entire authentication mechanism. So that was fun. Yeah, and while we worked on the MCP, we learned those kind of stuff. At the beginning, we were very, very naive. We thought, hey, we have the Swagger, the OpenAPI, very simple. We're just going to convert it to MCP, which basically it's another HTTP server.
21:33But eventually we realized agents are not that good with reading OpenAPI docs and using it. Especially if you have a very large API, right, with a lot of different things. Yeah. Yeah. And you can imagine that our APIs are very sophisticated in terms of you can use a bunch of operators and conditions with recursive groups and some kind of operators between those groups and things that the UI is doing constantly. And even human, when you want to write a query language, you pour a lot of intuition about how you're using those conditions. And we found out that the agent didn't really like it. We had to limit the amount of options in order for the agent to make a reasonable call.
22:21So we kind of encourage the agent to use simpler APIs and let it kind of narrow the search and only then expose more complex APIs. But you have to keep the API in some way a bit more relative to what you would do with an SDK where you want the developer to have all the options in the world to find what's relevant for that scenario. Yeah, no, I've definitely seen something similar. If you have too many options available, it just gets confused and starts throwing random stuff at you. Yeah, and also I must admit that at the beginning, I thought we can also feed LLM with the open API and tell it to create an MCP that works well for it.
23:09But that didn't work well as well. It didn't understand how an agent will effectively consume those APIs. We were very disappointed with that process. We had to go back and manually craft the endpoints and think about the use cases. And we were worried about, are we leading the agents to do things that we think are the right thing, but maybe it would prefer to do something else. But end of the day, we saw 100 % better results when we closed some APIs, limited the number of results, we forced the agent to get up to 20 results, for instance. Otherwise, we just got 2 ,000 log lines and kind of got stuck with random nonsense.
24:01Yeah, that's interesting. So if I'm hearing correctly, some of what you did is you kind of, one, applied your expert judgment. Here's a set of things that probably will be helpful. We're just going to expose those. And two, kind of gave it like this progressive disclosure where it's like, here's where you start. Okay, now we're going to expose a little bit more. Now we're going to expose a little bit more along the way. Yeah. We also make some parameters required, although they're not required from the official SDK. For instance, it meant to tell the agent, if you are looking for traces and you don't know what cluster you're looking at, something is off.
24:44You need to do something else. Like if you're that clueless, you're using the wrong API. So it has to go through a certain way of thinking because it understands it has to know first what cluster are they looking at. So now it can think, how do I know what cluster do I want to check those traces? And then that led to some kind of flow where eventually it got the right cluster. And some APIs where the output was less verbose and more high level, we allowed more primitives set of variables. So you could ask questions like, what change happened in my entire production? Or what incidents do I have?
25:37and then get the labels of those incidents and then think where do you want to check, right? So some kind of leading indicators to where should I look for the heavy stuff? Where should I look for the actual raw logs, actual raw traces, which could be very overwhelming. It could be, and that kind of API could return easily 10, 20 megabytes of response. That's fascinating. So essentially, you're making the API much more restrictive than you would for a human because you're saying, hey, there's a right way and a wrong way to do this. And if you're calling this without these variables, you probably have not thought about or figured out enough to get a useful response here.
26:23Yeah. And when you think about it, this is very human nature. So imagine you're standing behind a junior software developer looking on some kind of an incident in production. And all of a sudden, you realize that person just read all those log lines randomly. You're like, hey, stop for a moment. What are you doing? Let's figure it out. High level, where is the issue? Let's filter all those logs. like let's focus on what we know probably going to lead us to some kind of realization on what happened so we tried to put that notion into the flow of the apis we're not it's not 100 successful still it can still get lost but it helps to keep the the wild investigations in somewhat direction of narrowing that.
27:26That makes a ton of sense. Now, we've talked a little bit about identifying patterns, doing that very deterministically. I was curious if you had explored any sort of statistical pattern recognition or LLM-based pattern recognition internal that you then can expose up to an end user agent or anything like that. We play with it, and we're still playing with it because I think this is obviously the future. We have a feature that is copy to agent, which you can see in modern tools today have this kind of copy to agent when you have this prompt that can guide the AI agent to what are you doing. We started doing that for more than just a section.
28:12So imagine that I can represent flow you did in the platform visually and explain it to an agent saying, you went from traces to that workload page and then you clicked on logs. You filtered those, you added those keys and you landed here and then take a screenshot or something like that and give it to an AI. This is like an alternative universe for NCP. Instead of letting it use APIs to your platform, it kind of let it experience the UI in a way, in a markdown way or in a screenshot way. And the results were actually surprisingly good because you would think that, you know, we put so much effort within UX, right?
28:59To make the app human-friendly or help you slice and dice a lot of data. And then you throw everything until, okay, the AI agent will just use MCP or APIs. but when you just if you just take a screenshot of your application and send it to an ai agent you'll be amazed how much does it understand from that scenario so we're definitely looking at this kind of stuff nothing yet production ready but we are playing with it it is interesting right like these things are trained on human thought processes and they can only incorporate so much information. And so all of this thinking we've gone into, how do we help a person reach the right things they're looking for can be applicable.
29:46Switching threads a little bit and looking at another hot topic in this sort of AI expansion space is privacy, security, those sorts of things. As you're building an observability solution for this modern era, What types of privacy and security needs are there that are maybe more relevant today than they were a few years ago? Yeah, so logs and traces have the potential to contain PII at a minimum. We all try not to do those kind of things, but this happens. like someone just put an object inside a logger and you're all of a sudden you printed some PII into your logs and now you need to understand how you delete it or how you contain it.
30:39GroundCover is built on bring your own cloud which I think is becoming very very popular with the LLMA era you know you have non-humans sniffing into your code or your data or your most private data. So all the data contained inside your environment. So imagine if you have a very sophisticated agent that run in some kind of SaaS, you don't know where it is, and it learned from your data, analyze it, summarize it. Now you're more exposed to prompt injections and even just bugs, right? Like even just someone didn't think that will happen. So, BrownCover is built to use your LLM inside your account.
Read the full transcript
31:30So, the risk is a lot smaller. You basically have all the data with the agent inside your cloud provider. So, at least you know what are the physical perimeters with data being served and saved. I think this is the future for agents in general. You don't want a very sophisticated agent to have your data externally as we do with humans, right? When we onboard new employees, when we onboard contractors, we want them to use our laptops. We want them to stay in our offices or in our network environment, I think AI agents just make it the potential, the risk is just a lot bigger. It does feel like that's kind of the next generation of SaaS is everything can be done within your cloud, wherever it is, at least if you're on the big cloud providers, but increasingly, whatever cloud you're in.
32:34Yeah, and I think if we're already at it, I think Kubernetes changed the world in standardizing where do you deploy code. So now Kubernetes is commoditized like an RDS. It's easy to assume you're going to have a managed Postgres instance and compute and maybe even object storage. And to be honest, that's 90 % of what you need to build a very complex platform. So this is what the GrandCover is actually doing. Use very simple services to run that production and we make sure we can run it everywhere. We have obviously the big cloud providers. We also have on-prem and AirGap. Once we rely on Kubernetes, a lot of things got way simpler.
33:21Well, now that we're sort of talking about the way things are evolving, what do you see as, like what needs to happen in the observability space over the next five or 10 years? You know, where are we not yet solving the problems that are really there? The holy grail is root cause analysis. Everyone wants agents to tell us what went wrong and how to fix it, right? You see it everywhere. Everyone is targeting root cause analysis. And I think this is the very expected future. We're still not there. It's still very complex. So you can see root cause analysis with very basic scenarios, right? You have that kind of log and error log and you can explain it.
34:07you can correlate it with infrastructure changes that's very basic and this is already happening but what about like a two-hour research you want a co-pilot you want someone to be there with you and do the research with you while you are investigating an incident in production and we all been through this incident where you are six hours eight hours into the incident and trying to figure out not only what happened, and you also want a way to remediate. How do I recover from it? And for that, we need a deep understanding of the architecture, a deep understanding of the changes, and also a lot of world knowledge on engineering.
34:50And that's still not there. And I think this is definitely what we are targeting, but it's going to take some time to get to it. What would you say are kind of the underlying components that would go into making that possible? Good question. We need to have this model of how production looks like. We need to have the notion of how we learn production, right? How we learn the architecture. This is still something that is usually represented by a simple graph. but think about how long it takes to an engineer to understand the entire architecture it takes years in in a modern company and for some companies it's never been one person that can understand everything so we need a way to create that knowledge and also keep it up to date because things change so fast, it has to get the live data from production to understand how it's built and also the consequences of the architecture.
35:58So I think, for instance, eBPF is a very good way to understand the behavior of applications, the interaction between applications, the dependency between applications. So you can definitely build that graph represented with some kind of degrees of what's the side effect of this component failing. We also can inspect disks and network calls. So that's a very good start. But you also want to understand the interactions of people with those components. You want to understand the teams and the responsibility of the teams on those components. So, we're thinking of a lot of things happening on Slack, right?
36:47We need to understand the interaction there, a lot of planning on new features. So, okay, we need JIRA or Linear to also understand the planning for the future. Maybe someone already wrote, this is going to break, we need that planned. And we need to understand that as well. Like, what's the future? A lot of organizational knowledge on DR &D and product should be there as well. That's a lot of knowledge. That's interesting. So I was thinking about, like, you brought up Kubernetes and how that has really shifted deployment. And I think one of the things that it does is it takes at least one part of your stack and makes it kind of declarative, right?
37:27You can see some piece of your architecture and pull that knowledge out. you know, Terraform is similar, like any of these sort of declarative, like infrastructure as code types of solutions, give you the ability to analyze that part of architecture. And then as you highlight, like EPPF gives you live behavior data. So that's another piece. The organizational side is an interesting one, right? Like how are these systems used? What's going on with them? I wonder if there's other places where, like Kubernetes, there's some sort of declarative structure we can put in place that would let us short-circuit that a little bit and just get a sense of like, oh, here's how info flows in this system.
38:07I think you can learn a lot from code reviewing. If you read on code reviews, you probably get a lot of insight on where the weak spots of the components. You basically want to understand why things are like they are. Because you can always jump to conclusions, right? We maxed out the database connections and that's why it broke, but there is something more than that, right? There is a reason why we limited that number of connections. Maybe the developer kind of created this kind of new query that consumed a lot of this connection and we actually rotled that to protect that database from being over over-consuming CPU or something like that in order to serve other applications.
38:59So a lot of going on between the APIs and the applications and the Kubernetes, I think we need to find a way to understand why things are built in the way they built. What was the architecture decisions? Did we make it intentionally? Maybe we didn't. Maybe we just didn't know this is the default behavior. Is it default behavior? There is a lot of things to know that are between the code and the application and documentation of those services that we still need to figure out how to get that information. But I think what's comforted me is that at the end of the day, engineers solve those problems.
39:41So they go to Slack, they search everywhere, they're reading GitHub issues, they open AWS documentation, understand that's the default behavior. You're at the opening ground cover, see how many connections they have currently. What's the sort of throughput at that moment? Is it different than it was yesterday? So the knowledge is there. We just need a way to represent it and also to make it efficient because humans usually have intuition on issues. We need to understand how that works. Yeah. No, I think that the modeling piece is really interesting, right? In some ways, even you mentioned with log data, right?
40:20You're finding that time series is a very effective way to model things for an LLM to utilize and to use in different ways. If LLMs become the glue that is running through these different inferences, it matters. How are we representing this data to them? They're often quite linear thinkers from what I can see using them. And so if there's nonlinearities to it, that needs to be represented somewhere else in how you're then presenting that info to them. Yeah, it's fascinating. I think we also need to decide how much are we going to help it because most use cases, most incidents can be represented as 20, 30 questions that you have to answer.
41:01Most incidents are missing resources, noisy neighbors, bad version, exceeding quotas, some kind of infrastructure failure. We can help with crafting, as you said, 50, 60, 100 scenarios that AI can go through linearly and then start to think on its own. Basically, that's what we would do as an engineer, right? We're going to have those 10, 20 playbooks. If it doesn't fit none of those, you know you're in trouble, but you still have some kind of learning that you did. because you know, it's not that, it's not that, it's not that. Now you're looking for something weird. Something weird means you're going to call X, Y, or Z and ask for help.
41:49And I think even if we just do that, right, we can do that before an incident maybe because we can have more questions answered or even just present a report saying, here is a custom dashboard I created for those scenarios. This is what I checked, blah, blah, blah, blah. And that's what I know. I'm stopping now and now you. help me help you. That's a big milestone. That's a big milestone if you can have a coworker that is actually a bot that help you actually investigate and taking branches. I'm thinking of mentioning wrong cover agent, go look for storage issues in the cluster. Coming back after a minute saying, you know what, I found this, I don't think it's relevant with a graph.
42:35That's a big thing. No, that would be super nice. And to your point, I think a lot of the folks who are most effectively using AI for whatever purpose right now, they're building out playbooks, they're building out scenarios where someone has thought through it, maybe using an LLM, maybe on their own, they've sort of described what they think should happen. And that can be then reused and built upon. And so yeah, you could have these 60, 100, 200 scenarios, some of them might be purely deterministic, you can just go and look and see like, hey, hey, did this happen or not? Others might require an agent to do some decision-making.
43:10But that, yeah, that would be a really nice roll-up. Well, we're working on it. You are. Okay, so that actually feeds into this question, right? So what is coming next from ground cover or coming soon? What are the dimensions on which you guys are pushing this forward? Yeah, so there are two main verticals. One is observing LLM and the other one is LLM for observability. We just released the LLM observability piece. We use eBPF. We talked about it. It's very easy for the sensors to pick up LLM calls. And we now can see every API call to LLM inside your production or dev or staging. So we can now track token consumptions and comparison between models and correlate that to workloads and even showing you what caused that API call.
44:06So this is a track of LLM observability. We're just getting started. And this is a huge milestone for us because until now, we focused on classic APIs. And now we're trying to get the Bedrock API core, Azure OpenAI and all that. And it's very interesting and very fascinating. We're hearing customers requests for rendering images, sending to LLN and a lot of wild stuff. So this is one vertical that we're going to invest on it. The other one is using LLN for observability. We are taking a very different approach. we started by making sure the data the observability most data you would think is json logs no that's not true most data is just random printing of logs with weird formats and multi-lines and a lot of sorry to say that but garbage data and no matter how much your agent is smart if the data is garbage garbage in, garbage up.
45:14So we are actually now focusing on using LLM to create observability pipelines to clean the data. So we offer to pass those log lines into a much more meaningful format. We suggest passing specific fields that we think are meaningful to create time series from it and represent it in a much more sophisticated way that will allow you, the user, to create more advanced queries. So that's the first step. The second step will be, after the data is organized and clean, we're going to use, obviously, LLM to analyze the data. We already started doing that with patterns and, as we talked before, the baselines of the signals.
46:02But the target is to have that co-pilot running with you when you investigate or even just doing the research. It doesn't have to be just on crisis. I found myself using CloudCode for understanding repositories, not even just writing code, just asking questions about that repository. Sometimes it's just making things up, but sometimes I can actually get very good answers from it. And that saves me a lot of time. So imagine you can have it on your observability data. You can ask questions like, I'm looking for a cross-AZ consumer that is inflicting$10 ,000 a month because of cross-AZ network. That easily can be a two-hour research for an engineer and maybe two minutes for an AI bot that you can just mention on Slack and get dashboard answering your question.
46:59So yeah, that's the goal. No, I love it. And I think the dashboards and being able to surface up the right dashboards at the right time is a really nice thing. I read an article about looking for more AI-enabled heads-up displays rather than co-pilots, where it's just going to show you what you need to see at that time so it's easy to understand what's going on in your system. Yeah, I love it. Awesome. Well, we're getting close to the end of our time. Is there anything we haven't talked about today that you think would be important to leave folks with? I'm going back to the beginning. We want to leverage AI to write a lot of code.
47:36We need to find smarter ways to make sure, in a very deterministic way, on what the code actually does or how it behaves, what is the architecture that it represents. And no matter what we're going to end up, it has to be deterministic. Otherwise, we're just lying to ourselves with another layer that is not deterministic, that is judging another layer. So I don't know who's listening, but we need better linters. I feel like this is the answer. We need new tools. I feel like this is the era for having another stage for the CI to make sure things we care are being forced on the AI agents that write the code.
48:25So tests are very good. Linters are a good start, but I feel like something smarter needs to be created. Yeah. No, there's something interesting there. I almost wonder, like I observe a wide difference in how well LLMs write particular software languages. And I wonder if there's a designed for LLM generation software language that needs to happen where you can kind of enforce those architectural pieces at the link level. Interesting. Probably there are a lot of developers that will code in just English, right? because if you think about it, code language is built for being compact, but for specific tasks, right?
49:05Very efficient for specific tasks. But now with AI, I feel like English is probably a good way to represent an idea. But they're not only compact, they're also, to your point, deterministic, right? This thing means exactly one thing that can be reproducibly created. English is terrible for that. Yeah. I agree. So I love English as the first specification level, but you need something that you can, as you highlight, deterministically verify, evaluate, enforce restrictions on, something that where it's very, you can draw very clean lines. Yeah, sounds interesting. Well, this has been an absolute pleasure.
49:47Thank you, Chez, for the time today. And we'll call that a wrap. Thank you.
49:56Thank you.
From the publisher
Modern software platforms are increasingly composed of diverse microservices, third-party APIs, and cloud resources. The distributed nature of these systems makes it difficult for engineers to gain a clear view of how their systems behave, which can slow down troubleshooting and increase operational risk. groundcover is an observability platform that uses eBPF sensors to capture
The post Engineering in the Age of Agents with Yechezkel Rabinovich appeared first on Software Engineering Daily.
