In short
The TWIML AI Podcast Episode Notes
Episode Overview Title: An Agentic Mixture of Experts for DevOps with Sunil Mallya - #708 Host: Sam Charrington Guest: Sunil Mallya, CTO and Co-founder of Flip AI Description: The episode explores Flip AI's incident debugging system for DevOps, utilizing a custom mixture of experts (MoE) large language model (LLM) trained on a novel observability dataset. The conversation covers topics such as challenges in integrating time-series data with LLMs, the system's agent-based design, and the chaos gym for testing robustness.
Key Concepts and Discussions
Flip AI’s Observability AI
- Objective: To streamline observability data for developers, focusing on root cause analysis (RCA) of incidents in software systems.
- Target Audience: IT personnel and developers dealing with observability challenges.
- Unique Approach:
- Combines metrics, events, logs, traces, and code (termed "CoMELT") for comprehensive incident analysis.
- Employs a custom LLM trained on 100 billion tokens of DevOps data for domain-specific insights.
Agentic Design Philosophy
- Agent-Based Design: Clearly defined roles and boundaries for system components to ensure reliability.
- Chaos Gym: A reinforcement learning environment to simulate failures and test system robustness.
- Multi-Decoder Architecture: Different decoders for different modalities (e.g., logs vs. time-series data) to optimize performance and accuracy.
Training and Integration Challenges
- Data Types:
- Utilizes various data forms: metrics, traces, logs, events, and code for training the LLM.
- Emphasizes the importance of curating high-quality training data to avoid the limitations of generic datasets.
- Integration of Time-Series Data:
- Recognizes that LLMs struggle with numeric data; thus, traditional models may be integrated to enhance performance.
- Uses a mixture of experts approach to handle different types of data for optimal performance.
Practical Considerations for Deployment
- Scalability Concerns: Discusses the reality of deploying systems in diverse environments (e.g., on-premises vs. cloud).
- Computational Profiles: Focus on maintaining a balance between model size and computational resources to ensure efficiency and cost-effectiveness.
- Testing Framework: Emphasizes the significance of both unit tests and integration tests to ensure that the system functions as expected.
The Role of Language Models (LLMs)
- Model Selection: Strategy of selecting models based on their suitability for task performance and computational efficiency.
- Fine-Tuning vs. Generic Use: Focus on fine-tuning models for domain-specific tasks to enhance accuracy and reliability.
- Agentic Workflows: Discusses the future of LLMs in terms of developing robust agent-based systems that can reason and make decisions akin to human experts.
Innovations and Future Directions
- Mixture of Experts (MoE):
- Exploring the potential of integrating various models to enhance task performance.
- Discusses the challenges and benefits of implementing an end-to-end MoE architecture.
Conclusion
- The conversation with Sunil Mallya provides valuable insights into the intersection of observability, machine learning, and DevOps, shedding light on practical applications and future developments in AI-driven software management.
Key Takeaways
- Clear Roles and Interfaces: Essential for reliable agentic workflows.
- Data Curated for Domain-Specific Needs: Critical for effective LLM training.
- Integration of Traditional Models: May be needed to address numeric data challenges faced by LLMs.
- Testing and Validation: Both unit tests and integration tests are crucial for ensuring robust system performance.
- Future of Agentic Systems: Emphasis on optimizing the use of LLMs while integrating various models for enhanced decision-making capabilities.
For complete show notes, visit
[TWIML AI Podcast Episode #708](https://twimlai.com/go/708).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00One of the emerging patterns is defining clear roles and boundaries and interfaces. That part is what's lacking today in most agentic sort of workflows or orchestration systems. And that's what we said is like, no, we got to get this right. Because it can't be like out of 10 runs, we saw one magic. It has to work nine out of 10 times, right? Like, or we want to get 10 to 10 out of 10, but like, that's at least a start. We cannot be one out of 10. And so it was really that effort that we put into, okay, what are the fundamental pieces?
0:49All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. And today I'm joined by Sunil Malia. Sunil is CTO and co-founder of Flip AI. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Sunil, welcome to the podcast. Thanks, Sam. Good to see you again after many years. Good to see you for sure. It has been a while. There are going to be a few folks listening who were at our TwiMLCon event back in 2019. and they will remember that you put on an amazing DeepRacer demo slash contest for us at the conference.
1:38That was a lot of fun. And you did that because you were on that DeepRacer team at AWS. Tell us a little bit about what you've been up to since then. Yeah, DeepRacer was a crazy ride and still seems to be going strong. um so it's amazing they continue to host contests around the world around deep racer yeah like 120 000 people that apparently last year uh you know participated that's crazy wow wow yeah yeah uh since then um sort of uh uh you know ventured into nlp uh which was sort of gaining a lot of traction with all of sort of, you know, the fine tuning or building those foundations was sort of emerging with UML Fit and, you know, BERT sort of coming into that picture.
2:31And that sort of got me thinking that, hey, this is sort of a rocket ship that's, you know, building up. And I was so wrong. It was more like a Voyager or, you know, sort of outer space mission. Like, you know, nobody could have predicted. So it was sort of a lucky break to be on that ship, I would say. And you did that by switching over from DeepRacer to the Comprehend team at AWS? Yeah, so I took over the Comprehend team and then eventually laid the foundation for what's now Bedrock. So it was quite a crazy journey to sort of, you know, BERT was considered LLM in our world back then. The term didn't exist, but it was like, you still had to put all that compute to train that.
3:26And then suddenly, you know, you can probably train BERT on your laptop now with, you know, five minutes of compute time.
3:39That's awesome. And since then, you've gone on to found a company, co-found a company. What's Flip AI up to? Yeah, Flip AI, we are Observability AI. So the thesis around this is, you know, the majority of our sort of team has been building software at scale, maintaining five nines of availability. And one of the challenges always was the operations part of like, how do you run a service that's always up. And that contributes to hair loss, loss of sleep, and many other side effects. So we sort of like, hey, this is a pain we know really well. LLMs are going to... Coding is an obvious sort of...
4:30I would say developers love coding. It's what comes after it is what they hate. So let's go after that and solve that sort of the genesis of the company. I always need to ask when I hear AI observability, do you think of yourself as primarily observability for AI or AI for observability? Yeah, it's, you know, the whole ML ops and AI ops and AI for, it's all convoluted. That isn't a standard way to describe, but we're basically taking the pain from the developers, which is taking all the observability data, which is metrics, traces, logs, events, and making meaning out of that. So when you get that page that something is broken, we tell you exactly what is broken and why is it broken.
5:26So that's the definition, I would say. I guess maybe another way to ask the question is like the target developer profile. Are you going after someone who would be using a weights and biases or are you going after someone who would be using a Splunk or a honeycomb? Yeah, it's the latter. So it's the developer, sorry, ecosystem. So IT folks, like applying AI to like the traditional DevOps IT observability problems. Exactly. Okay. Yes. Very cool. Very cool. Awesome. And so like what, like how do you see your unique approach to that, given that there are the honeycombs and the splunks of the world that are already out there, data dogs and who knows?
6:17Very true. I think what's interesting is majority of the companies don't use a single tool. So your data is spread across different tools. So they end up using all the names you described. And ultimately, when something is broken, people have to go look at all of these sources and sort of stitch that story together as to little breadcrumbs all around. And that's a pretty tedious process. and what we are able to do is sort of be that intelligence layer on top of all of these tools and do the querying for you, do the data wrangling and understanding and reasoning, okay, this is broken because these two other things are broken and they're putting pressure.
7:01So, you know, it's typically when something breaks, it's not necessarily the service that you're sort of debugging is, it's something that is downstream, five levels down that is broken. And it's extremely hard to find, even with the existing tools. So that's what we make simple. And on top of that, what we've done is we've built our own LLM from the ground up and trained it on 100 billion tokens of DevOps data. So it's very domain specific. So it doesn't need to know what Napoleon does or hasn't done. It's very focused on just DevOps and really trying to solve that pain. So that's the unique approach.
7:52And we're able to deploy in your VPC, on-prem, wherever the entire FlipStack deploys, which is really important for a customer profile. being enterprises, you want all of the data governance because this data can have sensitive information. So we give our customers all the necessary guardrails in terms of controlling the data flow and governance. What does training data look like for a DevOps focused LLM? Are you just like throwing log files at an LLM? And then like, what's the value of an LLM that can, you predict the next made-up number in a log file. Right. We actually went after training. We've done different kinds of training.
8:44But one of the, I like to say, machines talk to each other with APIs and they express pain in logs. And that's essentially what we model. That's essentially what we are modeling is that pain that the machines are expressing. So it's sort of English, but it's not really. It's its own language. So it's important, even with the latest models and releases, they'll still fumble. They don't understand that. They're not necessarily... It's sort of like Pig Latin, sort of like... It's a half language, sort of. So you sort of need to fundamentally understand that language to be able to interpret that, which is why we train it ground up.
9:36But it's also not just logs, right? Like there's metrics, like you're tracking metrics, you have trace, which is graph data. But it's also got code, honestly, because the pieces of code in log, there's exceptions. You mean stack traces when things blow up? Correct, right? So you sort of have to understand. So we sort of coined this term. We call it co-melt. So we just code and melt data. Melt. There's like metrics and what's that acronym again? Metrics, events, logs, and traces. Traces. Traces, yeah. So we add a co with code. And so that's what we train ground up. So our training data is all of these modalities, so to speak.
10:23And a lot of this is available on the internet, but it's more generic. We've taken an approach of curating data. And I come from the old school ML, been doing ML for, I don't know, 15, 16 years. So back when, well, not back when, I still label data. So it's been a practice that I haven't lost touch with. So it's important to curate your data, expert labeling, make sure that data. So we sort of use whatever is available on the internet as this way of pre-training, sort of understanding the domain. And then you sort of start specializing by using data sets that are highly curated, labeled by experts.
11:16but sometimes like yeah i want you to dig into that more because my first thought when you said training in llm on this co-melt data was like i was trying to think through like how and why that would be valuable because i would think that uh in order to do what i would imagine you'd want to do, like in the observability domain, would require not unsupervised learning of the structure of a log file, but like supervised, like when this happens, that's this class of problem. Correct. So talk a little bit about how that end-to-end comes together. Yeah. So the first phase of sort of training you can think about the pre-training is just understanding what sort of just regular pre-training, right?
12:10Like what's the next word or predicting the masked word and so on. But then, as I said, like sort of, you know, I sort of jokingly said about the pain being in logs, but it's actually also in metrics and other places. And you sort of have to look at both of them, right? Like this graph is showing a certain data and then you've got logs that are showing. So they have to be, they're often telling you slightly different sides of the story, but you got to use both to complete that story. And so we use like, you know, joint training of this data to be able to make continuous meaning out of that. Yeah.
12:52And another sort of thing that is unique, what we do is, I like to say, pre-training is like graduating from high school and then supervised fine tuning is graduating from college. But you can't put the smartest person who's graduated from college and handle production incidents because you need sort of real life scars. And we sort of induce that to our LLMs by putting them in like a training gym. And this is sort of my reinforcement learning background coming in. So we have this training gym where we actually bring up applications and we use another LLM to break these applications. So we actually sort of simulate code injection or fault injections into the infrastructure.
13:41And because we know what we've created, we sort of use that to, oh, did you get it right? Did you predict like, or did you predict like, was that the issue? And so we can then use reinforcement learning to sort of help you, help guide the model in making the right decisions. And that's super important because there's only so much data you're going to find on the internet or you can label, you need an automated way to sort of scale your training. This is our chaos gym that we've built for the models to get as good a zero shot as possible. Architecturally, as well, what we've done is we recognize that each modality needs to be treated differently because, say, code and logs are predominantly...
14:33You can still use the same. Time series, LLMs are really bad at time series. They just don't understand numbers. And when you look at RCA's and reports, there's a lot of time and, well, this happened at this time, this number went up, this number went down. You got to do a lot of that. So we sort of came up with, okay, you can't have a single model. How do you sort of build this mixture of expert? and we sort of went with this hybrid approach of like, well, time series is not going to be, it has to be its own little component, but then attached to the mixture of experts. So time series is not a, say, a traditional transformer, but the rest of the parts are transformers.
15:21So we end up like, I don't know, I call it like a single, like it's a multi-decoder approach. So we have different decoders of the data or the interpretation to be able to make the most meaning out of it. So yeah, it's taken a lot of experimentation over the last two and a half years to get to where we are. Interesting, interesting. When you talk about integrating time series and LLMs, there are folks that are trying to do time series with Transformers. And, you know, depending on who you talk to, like the reports of success are either high or low, but that notwithstanding, it sounds like you're like incorporating in more traditional things.
16:16Like that's the impression of like a Remo, that kind of stuff, or like what? No, not quite a Remo, like a lot more advanced, but like it's using them as an input to inform the LLM. So I would say you're sort of using like a collection of traditional models to give more meaningful input to the LLM rather than a raw time series. So, you know, one of the fundamental problems I feel with like, let's say LLM is doing math, like that's a really popular sort of topic. And, you know, you have benchmarks and, you know, more like benchmarking, I would like to say.
17:00And one of the things is like, no, like the LLMs don't understand numbers together, right? Like, you know, five is a different number. Five, four is a different number. And it's 54, you know, it's collective. And that sort of doesn't quite exist because the tokenization just fundamentally doesn't understand. And I honestly feel like I have this theory, I call it the sunk cost fallacy of LLMs, where LLMs are so good and have done so close to breakthroughs with numbers that we're like, hey, we're not going to go back and fundamentally rethink that we need a new tokenizer, something that understands numbers fundamentally.
17:46so we can actually build this the right way, you know, with the right building blocks. Because we feel like we're so close. So everybody is like, maybe if I just do this one thing. Put my head down and push harder. Right. Like, maybe if I add the period or exclamation at the end of my prompt, maybe it gets everything correct, right? Like, and I think time and again, like the paper is coming out. I mean, even Apple had a recent paper on debugging GSM 8K with perturbations, right? You see every LLM is like, oh my God, the results, the variance of the results are just super widespread. So I think you just need to rethink what are LLMs good at.
18:33So we sort of take the traditional ones and convert them into what LLMs are good at. it's translating that into actual text that are more meaningful. Like if you actually, so it's actually getting close to the dimension that LLMs can understand and use that in all of that. So that's one aspect. The other aspect is when you want to deal with numbers, you've got to like, we have existing things that can give you a finite answer. So why are you going and suddenly changing your stack and adding non-determinism? And I think what's great is we, you know, tool use or function calling. So a lot of what we do is Flip ends up generating its own DSL with the LLMs.
19:23And the DSL does have things like, well, here, go call this function or use this tool to do the math, do the aspects. So that gives us really good results in terms of not screwing up the numbers because you have a very sort of number heavy output at the end where the database connections went up by this much, which put pressure on this tier where you started seeing higher latency and which ultimately caused X to happen. So that entire stitch is a lot of numbers. And by doing what I mentioned, two approaches that I mentioned, we end up getting very high accuracy here. You know, when you were speaking earlier about the relationship between logs and graphs and telling a story, you know, you went on to talk about time series data, which is like what underlies those graphs.
20:28But I'm also wondering if you have experimented with using VLMs, vision language models, to take the graphs themselves as data. Do you see any promise in that? Yeah, we did that. We experimented with that. I think we got some decent results. but ultimately one of the challenges with VLMs is more like the data that how you curate is the resolution of the images. What happens is you can get like, so suddenly like as you sort of zoom out, suddenly that appears to be a peak when in reality you got to take the interpretation of the, well, it's only going from 0.1 to 0.15 versus it's going from 0.1 to 9, right?
21:17Like it can still, So that sort of focusing on that little piece of information, which is slightly different to me, and it's very sparse. Unlike VLMs and text that where you see where this is more of a description, et cetera, we didn't quite get it right, I would say. but there's always this promise of where you get like sort of infinite resolution, so to speak, with actually having the raw data, which is much easier and more reliable in terms of operating on. So hence, we sort of paused the whole VLM approach. However, the VLM approach, I do think, could be really interesting in finding visual patterns that are much harder to find, you know, because it's, again, it's a much higher level representation of data, condensing.
22:21When you operate with raw data, you're not going to have, you have way too many dimensions. Again, it acts against you, especially when you want to find patterns. Yeah, it's an interesting give and take between you know using inspiration from what you know how the human would do it and you know doing it the way computers are best at doing it like a human would drown in the raw data that's why we have the charts and the charts very much help us drill in on on what's actually happening but you know maybe that's not the right approach for a computer bingo like that was certainly our intuition in terms of like oh we should be using vlms and what we found was um you know for a small subset they were good but like as you start like actually interpreting the data the scales and uh sort of like uh think about like visually well that as you compress data and you decompress like the timeline starts getting a little confusing.
23:23You mentioned the DSL and tool use as part of the way the tool operates, the product operates. When I hear tool use, I start to think of agentic behaviors, agentic workflows. Do you think of the system in an agentic way or as an agentic tool? Absolutely. And less in the marketing sense in that everybody's thing is an agent now. Correct. Yeah. Hi. I'd like to think I was calling them agentic before. Again, it's a very reinforcement learning term for me. So to me, it's very goal-oriented, right? Ultimately, what we do is, well, we got to find these RCA. So that's the goal. We got to go find the RCA, right?
24:17Like root cause analysis of an incident as a root cause analysis. So that's the ultimate goal as we've found, okay, why this happened? And along the way, we're sort of collecting, you know, observations from the observability, you know, data and processing that and then guiding what is the next step we need to take, which is very much a reinforcement thing. agents changing states through their observations and getting rewards from the environment to guide them and get to the role, right? Sorry, goal. So that to me was sort of agentic in nature, but we didn't necessarily think of this as like the agents as they appear today.
25:11It was more like, okay, we have a bunch of tasks that need to be done. We got to break the tasks down. Now, we need to make sure the hops between these tasks or the hand-holding between these tasks need to happen in a more reliable way. So in many ways, we actually came up with this framework. We call us agents, which are the lowest level of abstraction. We have actors and actors are this next level of abstraction, which are sort of cooperative agents. So the actors are able to take a collection of agents and perform a task. And then you have the director, which is the orchestration layer as to, hey, here I'm going to assemble these actors to do something.
26:04So in a more practical sense, we can think of the directors as like, well, you got an AWS, sorry, actors is like an AWS actor, a Kafka actor. So those are experts in these things. And then they have lower level agents like, well, I'm going to co-query Splunk. I'm going to go look at some other piece of data and so on. So really building repeatable patterns, right, and well-defined interfaces. So it's like, I'd say like we sort of heavily applied software engineering principles into how to build these, say, quote unquote agents and have them to operate together. And so is the director, is that more the infrastructure-oriented piece?
27:05It's more of the planner. Like, think about something. Yeah, it's okay. So because what happens is you're going to discover information along the way. So you need to backtrack and perhaps change your hypothesis or explore a different hypothesis that you have as to what went wrong. So you need something that orchestrates saying, well, it doesn't look like AWS had an issue. Maybe we should look at Redis or Kafka or other piece in your infrastructure. I guess where I was going with infrastructure was like you are describing this like multi-collaborative agentic thing. And part of the question that jumped up for me is like, are those things all running in one process or on one box or something?
27:55Or are they running in different places and what tells what to run where? Is that even an interesting and important factor here? Yeah, they do run in different processes because you want to parallelize the execution. And they sort of almost run asynchronous because while you're fetching data, we're doing some other computation and so on. So we sort of quote unquote, like a DAG-like system. Practically speaking, DAG generation and execution is hard. So we sort of flatten the DAG in the sense that it's okay to sort of redo certain things as you go on because if you feel what ends up is if you're keeping all those memory state like now you introduce all of this concurrency and other you know uh patterns that that that slow you down or can induce more bugs so you want to have like so again like just simplifying the the part by rather than a complete dag execution uh if you need to revisit something you could probably revisit it again.
29:14All of your steps are idempotent, so just run it again. Exactly. It's okay that you operated on the data, you made a copy of the data. Think of it as copy on write. Like, okay, I take that data, if I modify it, now I have another copy, etc. If you actually end up keeping or operate on the same copy and modifying now, you got everybody to coordinate. You got so many asynchronous parts wanting to access. It would work in a great single threaded environment, but if you're doing multiprocessing and so many things, just introducing more complexity that you need. We keep the software engineering part simple and leave the complexity of the agent actor orchestration to itself.
30:05I guess taking a step back to this RCA that you're trying to identify, historically, tools like Splunk and those before Splunk even are trying to do correlations, statistical correlation. And like, it's a very different like way of looking at the problem than like an LLM and a reasoning system. Like how do you like connect those ideas? Yeah. I think, you know, statistics, you know, sort of give you hints, but they don't necessarily give you, okay, the sort of causal nature of things. Right. Like, so I think what was important, you know, as we were going and building this was mimic the flow. Like to me, I mean, I've been fortunate enough to sort of build many different ML systems.
31:07And the only way I sort of could think was, OK, how does a human do it? Replicate that. Right. Like so to me, it's like the simplest thing that I could come up with in terms of a framework. and then maybe the machine can do it differently and better. But almost the first approach always is, hey, what do the humans do? So in a way, what we did is, well, if you think about a war room situation where somebody is solving these incidents, you always have a bunch of experts. Everybody's chiming in on, this is what I see. This is what I don't see. And then a core group of people are making determinations on how all of this data is stitched together.
31:47hey, okay, that's unusual that your service is having this issue. I don't think it should. And then suddenly you realize, well, that's a shared component that both of us use, and now that becomes the culprit. So that doesn't necessarily show up with basic statistics unless you sort of have... Basic statistics will show when an incident, everything is broken, so everything must be wrong, which is not true. So now you need to go into the causal connections of like, all right, you are inflicting pain on me, but it's not you. It's somebody else who is this chain of inflicting pain. So that's sort of why the approach that we take and it's turning out to be superior to traditional approaches.
32:40Continuing with that War Room analogy, like the individual teams are seeing anomalies popping up in their systems or events, and they're bringing that up to kind of this core group. And I think of one of the things that that core group is doing is that they've got like this context in their heads of, you know, what's essentially like the runbook, like how the different things connect. And so does your system, like, does it need to have that visibility into that runbook? And if so, like, do I have to define it in some, you know, do you have to be a runbook system? Are we at a point now with like, you know, forgiven domains like an AWS app or a Kubernetes app, like that's all like, you know, software defined architectures that you can infer all that?
33:39Is that stuff you look at? Like, how does that all come together? Yeah, very much the latter where we are automatic in terms of generating that runbook. because I think these patterns are known and we've been able to generalize and with all the training. Of course, if you have a very nuanced system that isn't necessarily a common pattern, then probably we need tuning. But I would say for a vast majority of cases, it's our runbook generation suffices and we're able to do zero shot for majority of our customers. Well, it wasn't the case a year ago. Now it is very much so. Is it zero shot just from observation or is it zero shot including slurping up some configuration that's able to be found somewhere?
Read the full transcript
34:44Yeah, both actually. So, and what helps is a lot of these observability systems have a standard way of defining things. So, you know how to get the service maps or, you know, so you can sort of build on top of existing stuff. Or AWS, you know, there's a thousand ways, but there's a thousand, you know, it's a finite thousand ways of doing things. So you're able to sort of build on top of that. Yeah, and ZeroShot...
35:45gives you the definition so you can use that as a base knowledge while you operate. So it's good that a lot of things are standardized and effort is being put, but it's just extremely complex for people to keep everything in their hand. You kind of alluded to the set of architectures being somewhat standardized nowadays, like how standardized are things? I would expect that there's actually quite a lot of variety in the way people would choose to deploy things. But maybe if you just say, hey, we support Kubernetes or we support AWS, or you have to have, you know, Honeycomb or something, then that simplifies the scope.
36:29The standardization I was referring to more was the way people express their infrastructure and so on, like, you know, your CloudFormation or Terraform or CDK. And, you know, I was referring to that and also like an observability tools. Helm charts and that kind of stuff. Correct, exactly. So that, but then there are variations. And of course, in terms of like you choosing, you know, there are a thousand different queues systems you can choose or NoSQL databases you can choose. So it sort of starts exploding from that point of view. but also there's a certain,
37:14and there's still like a cache system behavior is consistent across whatever caching you're using or a database is fundamentally always a database that needs to do something. So there's certain sort of aspects of things that you can take advantage of. But yeah, purely speaking, if you zoom out from a architectural standard practice uh well that that every that varies for sure right like we have customers who have good old school mainframe sort of uh things uh uh versus uh to all the way to like modern like really proliferated like microservices architecture uh those result in different problems and so you have this kind of collection of agents that you're running you know across many customers like I can only assume that like you it's kind of ground-up built you're not using some off-the-shelf agent framework you know those aren't there yet and you'd want to control all the the pieces like I mean we started building before these frameworks existed so we sort of gotten used maybe an interesting question to ask would be like if you were to take what you had and build a framework based on what you know like what would you be thinking about and how does that differ from what you see out in the framework landscape we're very specialized for what we do but perhaps like i mean one of the maybe the emerging patterns that that sort of i see is um again tying back to software engineering architectures and the two aspects to that.
39:01One is defining clear roles and boundaries and interfaces. So to me, the way I think about when you build really scalable systems is API interfaces are well-defined. You know the request response, you know the faulty behaviors. So I think that part is what's lacking today in most agentic workflows or orchestration systems. and that's what we said is like, no, we got to get this right because it's not going to be like, it can't be like out of 10 runs, we saw one magic. It has to work nine out of 10 times, right? Like, or we want to get 10 to 10 out of 10, but like, that's at least a start. We cannot be one out of 10.
39:47So it was really that effort that we put in into, okay, what are the fundamental pieces? So we need agents have very well-defined input-output structure. They'll always reliably give you the output or saying, well, I failed, and then you know how to retry. So the system sort of ends up becoming just like you're calling a library function, et cetera. So the developer who's using to build these workflows in our system know exactly the behavior to expect. And underneath, we did all the hard work in terms of, and this is why fine-tuning is so important, and a strong proponent of that is now you get to control.
40:32Your failure chances and cases go really low because you fine-tuned the model. So that's like... Just to make sure I understand that, are you saying that... I can imagine you saying several different things. One is that part of the reason why you have so many failures when you're calling generic LLMs is that they don't understand sufficiently or aren't trained to answer the question that you're asking. And sometimes it just takes a weird path. And so if you fine-tune sufficiently, you get an answer more consistently. Another interpretation could be that part of the fine tuning you're talking about is more like controllability, steerability, RLHF style, as opposed to knowledge training.
41:29And you are training for response characteristics more so than knowledge production. It's got to be all of the above, right? When you think about an LLM, you're asking, so first thing is, does it know the answer? Is it even domain? So the domain knowledge needs to exist. So that's the first part. The second is now your prompt, you could ask the same thing a million different ways. Are you going to get it like the right answer? And I think the agentic today in open source or even like proprietary LLM flow that people build, it's like the prompt evolution, right? Like you're trying to ask the same thing in a different way until you actually get the answer you want.
42:15And you take that out of the equation by fine tuning because I can ask the question in exactly the same way. And I know that this is the question that I'm supposed to ask. So you sort of take that dimension out of the equation. And the final part is the response structure. Like now you don't need to say, hey, please always like... Please, please, please give me JSON format. Exactly. Or a cat's going to die. Exactly. So now you're free from, you know... Coercioning tactics. Yeah, yeah. The LLM, because the response structure, the LLM always knows this is what I expect. Or not the LLM, but the particular task, right?
43:01which is why the specialization part is what we do. This doesn't come easy, right? Like this comes with like... You said not the LLM, but the task, like, you know, without the resources to, you know, or, you know, the time to like fine tune an LLM for a given task. Like one of the things that I've done is just like, retry this 10 times and I hope that, you know, like, are you doing that also? or does fine tuning solve all your problems and you're not doing that? We do have, like, it's never perfect. Like you have to have fallbacks, but we don't need to do 10 tries. A couple of fallbacks are sufficient.
43:46Yeah, but I would say the good thing is I think our hit rate is probably 99 % or greater in terms of you getting it right on the first sort of attempt. And yeah, that's the reality, right? Because when you're training hundreds of different tasks and really large LLMs, they're always going to, you know, you can't expect like them to be always correct. And I think it's sort of like you have to take the old distributed, like, you know, say the early 2000s where, you know, there was a shift in thinking about distributed systems and building and saying, no, no, I'm going to assume that everything is going to break all the time.
44:30And I'm going to design for it. And that's how we have the Googles and Amazons of the world working wonderfully. And I think - Yeah, I caught your mention of Chaos Lab or Chaos something or other earlier. Jim, yeah, yeah, yeah. Chaos Jim, speaking to that Chaos Monkey type of learning. correct uh so you mentioned uh hundreds of tasks like is there another implication there that you've kind of learned that um a best practice is is to be very specific and fine-tuning an llm like on a micro task as opposed to you know trying to fine-tune a single lm to do a bunch of different things? Well, we started out with like one LLM multiple tasks.
45:22And then we sort of like, as the mixture of expert things sort of started to gain more prominence. And also different sort of models have different strengths. And we sort of said, hey, you know, we got to use what's best out there and fine-tune on top, rather than relying on one thing. And the good thing is we took that early call where we said, look, models are going to get better. Architectures are going to come and go and change. So from an agent perspective, the entire system, as we swap our mixture of experts to even something else tomorrow, nothing else needs to change of the stack. All of the things are going to work because the interfaces are well-defined.
46:09When you say mixture of experts in your case, are you saying that very specifically like an end-to-end train MOE architecture or colloquially we've got a bunch of agents and we do things like send the tasks to the best agent or a bunch of models? It's the former. No, it's the former. It's the former on an end-to-end MOE. But again, different tasks go to different parts of the MOE. So there's a routing layer that actually goes to the best place. and what I meant was from an agent, let's take the example of an agent summarizing a log that is seen. Now, that sort of agent doesn't need to know the underlying sort of architecture interface.
47:19The interface, the input-output is well-defined. So let's say we take MOE out and Llama 5 comes along and that's like the best thing out there. And we can just swap that architecture. None of the other parts of our system need to change because they have no knowledge of anything specific in the abstraction layer below. Assuming Llama 5 is good at all the things that the MOE does. I mean, the good thing is we have training data, right? So we'll fine tune if the base model is that good, we can always fine tune on that, and then it gets the knowledge. And more importantly, I like to say, we've curated a really good test set to be able to say what model is good or not, should we graduate or not.
48:12And this is sort of a saying that I've had for a good part of my career is like, the training set never matters. I think people obsess over a training set. I'm like, no, no, no. You should obsess over the test set. Because if you know the test set is really good and representative of what you want it to be, then you know it works or not. The training set is just a matter of, I'm not arguing for a training set to be bad. I'm just saying the obsession should be more on the test set. Because then you can sort of quantify and know, am I making a choice on X or Y? And is it informed or not? So you've got these task-specific models kind of baked into this MOE, and you are running this kind of multi-tiered, agentic pattern or architecture.
49:12Like, what are some of the real world realities of like trying to run an agentic system at scale? Like, what do you run into? Yeah, I think we have to learn our lessons a lot on patterns of, you know, how granular do you break the task? can, you know, like if you go too granular, then you end up with too many calls. If you go too broad, then you're asking the LLM to do five things in a single call. And we sort of trial and error in terms of where. And I think the closest analogy I can sort of give to this is from the good old world of RDBMS versus NoSQL systems. So you can think about like the RDA system where you sort of normalize the data and it has all the knowledge that you want.
50:13And so all you need to do is issue this query and then it gives you the answer, right? Now, yes, I don't know. I used to be good at SQL maybe a hundred times in my life and I keep forgetting, right? And it's analogous to a prompt, right? Like where once you get that right, Like magically the LLN answer. Magic incantation. Right, right? Like you get that answer. And then that's sort of like one approach. And then the other extreme is sort of the NoSQL approach where you denormalize the data. You sort of bring the necessary parts and then you control the compute layer to stitch and then finally sort of serve the answer.
50:55And, you know, it's a hard choice. And I feel like the answer is somewhere in between. And you need to know which task is... And again, it comes down to which task is the LLM inherently capable of doing it at a more broader level versus what needs to be... What's a hard problem, it needs to be chunked enough that you now need to sort of make a few calls and then sort of see what it makes sense together. So I think that's sort of the analogy I would say, like, in a practical sense that I think people have to understand is like, hey, what is the LLM really good at? Right? So I can use the RDBMS pattern.
51:43And what is where I'm not going to get the answer. I know I have to break it down. And it needs to be more like a NoSQL pattern. So I think really sort of thinking through these two patterns together, I think helped us because we did end up going the latter way, which is like we went with the whole task breakdown. We ended up being two NoSQL. And what that means is too many calls, more chances of failure, more chances of retry, and the answers being slow. and then we're like, okay, that's not quite the right pattern. We know it's useful. It has its place. We now need to start and sort of aggregating.
52:30So I think maybe 70, 30 sort of split between the two types of patterns that I see in the code base today. Is there a similar trade-off when planning for tool use? Like, you know, it strikes me too, you're defining APIs. They can be, you know, granular or, you know, coarse or broad. Like, do you think about the same things? Yeah, you're bang on there. I mean, this is, again, very similar to software engineering where people are like, oh, what's the right pattern, right? Like some people say, you know, your function should not be larger than 50 lines of code and you got to break down. And then suddenly, like now you're like, as you're looking through the code base, you're like clicking through your IDE through each to know exact.
53:28Then you end up with this too much abstraction and suddenly it's unreadable versus having like a thousand line function that does everything. So I think the answer is always somewhere in between. And I think with tool use and function call, you just have to be like, all right, what is the right sort of way? And it's, again, software engineering practices, it's what's most testable. Is this going to be reused in other parts of the code base and so on? So that sort of defines how fragmented you go versus like how much you pack into a single sort of function call or API call. And you mentioned test set, but it strikes me that it's more than just the test set.
54:24It's like having a really solid end-to-end evaluation system so that you can easily kick those off and evaluate all of these. It sounds like you're evaluating all of these decisions relative to one another. Very much so. Yeah, I use test set in a more canonical way, but in reality, when you have a goal-oriented system, you got to, yes, the individual pieces are good, but it's again, think of it as an integration test, right? You can write all the unit tests you want, but it has to integrate together for your CI, CDC to work. And we take the same approach. And now we use the Chaos Gym in terms of more of an integration test environment where we're bringing up applications in different languages, different architectures, different clouds, different orchestration layers, you know, Kubernetes, ECS, and other, you know, general VM architectures, and then sort of running through, okay, does the system sort of work end to end?
55:33Does it give you what you expect? So that's sort of our integration pipeline by using our chaos gym to sort of certify, well, all right, everything works as expected. And I think that's more important, like if I sort of zoom out in terms of lessons to other agentic sort of frameworks and so on, is like really focusing on both like the unit test and integration test sort of mentality in terms of, yes, you need to check like the individual steps. but it's really important to see how the goal that you're trying to do is sort of affected. And I'll bring back my RL hat again. And the difficulty in a lot of these systems is you don't get sufficient rewards.
56:28There's a whole notion of sparse rewards in reinforcement learning, which leads to not being able to train. is like you don't get like enough rewards through your pipeline to sort of improve. And I think, you know, as an agent is trying to book like, you know, your next flight, how do you know, well, the first search results are sufficiently, you know, enough for you to proceed and so on. So I think bringing that sort of discipline into the system does help you sort of focus and allows you to make an educated call on where in the system is the weakness and does this really work. You talked a little bit about the, I'm trying to remember the three layers.
57:26There's actor, director. What's the lowest level? Agent? That's your agent. Yeah, so you talked about this three-layer system. You talked about the agent. You talked about the actor. I don't think you spent much time on the director. For the kinds of tasks that you're targeting, are LLM reasoning-based systems sufficient? Are they optimal? Do you also incorporate more traditional rules-based heuristics, decision trees, that kind of thing? I think at the planning layer, I would say it's reasoning like LLMs again. You have to take a more practical approach. So I like to say, you know, think about like if you sort of distill LLM or any agentic workflow, think about it as like a really amazing or labyrinth of if-else statements, right?
58:33Like whatever you can decompose into, you can decompose into that. It's humanly not possible. So the way I see that is like. A labyrinth of opaque if-else statements. Correct. and the LLMs are sort of taking these pieces and replacing a lot of these statements so they're making the if-else statement sort of go away and finally there will be a time when the if-else sort of disappears and I think especially with some of these orchestrations and so on some of the playbooks and so on like You can rely on what humans have done before. And these are sort of playbooks and stuff that... So we are able to train on existing data that users have.
59:27So you can sort of mimic that to a good extent. What is hard is... And generally, these are more or less very broad statements. It's like... The director ends up being simplistic in the sense of, go look there, go look here. Well, if you don't have data traces, then do this to sort of get that. So the director ends up being sort of like this canonical or the more common example of booking a flight, right? Like booking a flight traditionally has this well-defined 10 steps, but there are nuances in those steps. And those nuances can be sort of learned by training on data. But if you really distill them out, then they look almost like pretty common, well-known steps.
1:00:25So that's what I mean by the whole if-else labyrinth and LLMs are replacing. because the nuance of this one step of determining, well, should I be looking at Redis or should I be looking at Kafka or something else? That's the hard part. And just saying I need to look at three places is more of an easy thing. So the director ends up not being a very complex piece of software. it's the latter two that actually become like the real sort of workharses in the system it sounds like what you're saying makes the director part work is fine-tuning on a lot of task-specific data and just really solid prompt engineering correct yeah yeah it's patterns of like how would debug a cache issue, debug a database locking issue, and so on.
1:01:29So it's very specific in that sense. So that plan more or less can be generated pretty easily. But it's the other part of like, hey, is Redis broken or Kafka broken? That's the hard reasoning part that is pushed down to the actor layer and just the fine-tuning aspect may or may not solve that. And you talked about unit tests versus system tests. Like, do you unit test the director individually in a sense of like, you know, you just saw this. You know, given what you know, how would you resolve that? Yeah, exactly. like, well, generate me a plan to debug database locking issue, right? Like we just, we know what that sort of should look like.
1:02:28And so we can say like, you generated this or not. And are those plans, like, are those difficult to evaluate? Are you like doing like groundedness testing? like you're looking for n pieces of information to show up in this text that's produced? Yeah, yeah. Those verification ends up being fairly standard ways of doing, yeah. In terms of the model, like we talked about the MOE, Would you say that that is fairly accessible in terms of being well-defined how to train an MOE and computationally accessible? Not so much, actually. I didn't get that sense. Yeah. The fun thing is on the internet, like 99.9 % of the content is very much sort of experimental and tinker, not experimental, it's more educational and tinkering.
1:03:37So you end up finding all PEFT and all other Laura, Q Laura, you know, like people want to tinker, people want to get a feel and learn. It's a lot of material. Like there isn't a lot of material in terms of how do you train sort of things like, you know, like this ground up. It's still very much in a few sort of places. And, you know, we've had to go. It pays off to have good friends in different places and including like PyTorch and being able to sort of talk to them about like, well, we're trying to do this. Like it's not quite what's the best way to, and then discord channels and so on. So those have been helpful in, you know, actually being able to train.
1:04:29Yeah, very much the content in, out there in the internet is very much like, I'm going to tinker something on my laptop or something very basic, but not really like something at scale. And then you mentioned a future version of Llama in the context of talking about the MOE stuff. How do you think about model selection? A lot of people I talk to, they'll start with open AI, the most capable model, and whittle down to something very specific to make it better, faster, cheaper in some way. Is that the way you think about that side of things? Not really, because we sort of have a compute profile in mind in terms of, okay, ultimately, we have to deploy this for our customers or deploying in the environment.
1:05:25We're really mindful, like, you know, you don't have like still very hard to get A100s or, you know, let alone H100s and so on. So you got to be like very mindful of profiles or compute profiles. And also, your customers are not necessarily going to be like, well, I'm going to spend a million dollars to send up this inference cluster. To do my DevOps. It's just different. So the compute profile is very fixed. And that's the reason why Obsess over fine tuning is that we can keep our models small. We can define the compute profile. We know the throughput that we can get. So we're very practical from that sense.
1:06:13So if LAMA 5 ends up being, again, speculating, like it's a one trillion parameter model, probably is a non-starter.
1:06:26So that's sort of how we think in terms of models. But I don't really care about what architectures evolve and so on. Like I'm intellectually, I am, but not as like, say flip, like how I'm going to incorporate that into the product with all the abstractions that we have. Model is just a drop in sort of plug and play part of it. But the efficiency is an important part of it. And so you're, you tend towards smaller as opposed to bigger, like. Totally. Yeah. And also a speed aspect, right? Like a scale, I mean, you're deploying in a large enterprise. You got to have so many debugs happening at the same time.
1:07:14You can't quite have a profile of OpenAI 01 or GPT-4 or even a LAMA 70B. Like those become very hard. And we don't run quantized models.
1:07:35it again introduces more chances of failures and finickiness. So we try to sort of emphasize on the robustness of the model, respecting the whole chain of input-output structures and so on, and at the same time fit the cost profile. Is the MOE, is that like... Is that all of the layers or is that one of the layers? Meaning is that the actor layer only that's the MOE or is it, you know, is it everything? How many models are managed in your system? Yeah, like I would say the MOE has about five distinct parts. but everything is abstracted like in the sense of the agents call like all of the agents are the LLM interface so the agents call the LLM and then so think about that as like the agents have certain prompts or set of prompts to carry out the task so the layers are prompts and the way they think about the task and semantics and the LLM is the LLM, but it's this one LLM for all of the layers.
1:08:57Correct. Correct. Yes. Got it. And that thing is trained, you know, it's just ground up. It's not, there's no part of it that's off the shelf? Yeah, actually, we did start off with ground up. And again, let's say 2022 and so on. And what we found is we needed to incorporate a lot of data for the model to learn English. And then along with log data, because as you generate summaries and reports and understanding the English part was also important. And then as the open source models started becoming better, we sort of said, well, we don't need to train English again to these models. So we sort of take off-the-shelf models now and then do our training, sort of extended pre-training, fine-tuning, and chaos training on top of those components.
1:10:02so what that has done is allowed us to have the model not relearn so you're trying to have this moe where the experts have some correlation to task does that that doesn't necessarily mean that you can take an off-the-shelf moe like a mistral eight by eight or whatever and right like that's going to have any correlation of what you're trying to do like so you are taking individual correct you know three b's or eight b's and like moe-ing them some kind of way yes exactly yes yes uh because uh correct i didn't know you can do that i thought that you that yeah i it was not obvious to me that you can take existing trained standalone models and like end-to-end MOE train them yeah we had to do some interesting stuff there hence the hence the chopping off layers and like wiring stuff together no yeah they have different tokenizers so you have to give the you know reuse the same tokenizers and and and and so on and then you got to pass like the output of one to the output of the other for certain tasks so had to do well also Python is a flexible language so luckily we've been able to sort of get away with that and yeah exactly we had to be creative in terms of what we did because if you recall the time series, the time series is a different model and we had to take the output of that then feeding to the llm to be able to sort of reason through well here is a spike that's happening this is what it is mean this is what it means and so on i didn't even realize that you would be doing that like it within the model as opposed to the time series model is surfacing something and then you're re-injecting that in a prompt I mean, they end up being equivalent in the sense like if you define a single sort of task, then you're just forwarding that information to the next layer.
1:12:24So think of it as like you interpreted something in time series, and then you can send it to a model to spit out the English version of that, right? Like you're taking that representation and now the model is spitting out quote unquote English. Or you can go up. So that means you're doing that, like summarize this log enough that it warrants making it its own kind of flow or module or model. Exactly. And then go up the chain and come down. It's more like do this thing and then it flows through the underlying layer. It also means that you can decouple things easily, right? Because we fundamentally think the models will change.
1:13:12so you cannot be tightly coupled to whatever LLM is that exists underneath. Yeah, that makes sense. Very cool, very cool stuff. Yeah, it's more of a practical approach, right? Like I said, compute profile, there's certain robustness, etc. So we took all of that and, all right, we're going to treat this as more like engineers. rather than, you know, pure scientists who are like thrilled by, all right, I'm going to make this one LLM do everything. Yeah. Awesome. Awesome. Well, Sunil, thanks so much for jumping on and catching us up with what you're doing. It's been great to reconnect. It's been too long.
1:14:04There's a lot of interesting stuff in here. Absolutely. No, this was delightful. and thanks for informed questions. And yeah, I thoroughly enjoyed chatting and reconnecting with you. Thank you.
1:14:38Thank you.
From the publisher
Today we're joined by Sunil Mallya, CTO and co-founder of Flip AI. We discuss Flip’s incident debugging system for DevOps, which was built using a custom mixture of experts (MoE) large language model (LLM) trained on a novel "CoMELT" observability dataset which combines traditional MELT data—metrics, events, logs, and traces—with code to efficiently identify root failure causes in complex software systems. We discuss the challenges of integrating time-series data with LLMs and their multi-decoder architecture designed for this purpose. Sunil describes their system's agent-based design, focusing on clear roles and boundaries to ensure reliability. We examine their "chaos gym," a reinforcement learning environment used for testing and improving the system's robustness. Finally, we discuss the practical considerations of deploying such a system at scale in diverse environments and much more.
The complete show notes for this episode can be found at https://twimlai.com/go/708.




