Redefining Incident Response: Insights from the Chaos Engineer Behind Jeli.io | Nora Jones

25 Apr 2023 · 39 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Dev Interrupted - Episode with Nora Jones

Episode Overview In this episode of *Dev Interrupted*, hosts Andrew Zigler, Ben Lloyd Pearson, and Dan Lines speak with Nora Jones, the founder and CEO of Jeli.io, an incident management platform. The discussion revolves around redefining what constitutes an incident, the importance of incident response, and how organizations can leverage past incidents to foster a learning culture.

Key Themes

  • Redefining Incidents: Organizations often underestimate or misclassify incidents, leading to underreporting. Nora emphasizes that organizations should over-index on labeling events as incidents, as they offer valuable learning opportunities.
  • Learning from Incidents: The discussion highlights the importance of analyzing both major and minor incidents to extract insights that can improve organizational processes and culture.
  • Chaos Engineering: Nora discusses her background in chaos engineering and its relevance to incident management. She argues for a proactive approach to learning from incidents rather than merely responding after they occur.

Key Takeaways

Understanding Incidents

  • Broader Definition: An incident is any unexpected event that disrupts normal operations. This can include minor issues that may not seem significant at first.
  • Data Hidden in Incidents: Organizations should examine the interactions and decisions made during incidents to uncover patterns and improve communication and coordination.

Incident Analysis Approach

  • Post-Incident Reviews: Rather than focusing solely on action items, the goal should be to foster a learning environment. Understanding the context of decisions made during incidents is crucial.
  • Cognitive and Coordinative Work: Interviewing team members involved in incidents can reveal insights into how decisions were made and who was involved, contributing to a culture of knowledge sharing.

Building a Learning Culture

  • Cultural Sensitivity: Adapt the approach to learning from incidents based on the organization’s culture. Some may prefer documentation, while others may benefit from meetings or video recordings.
  • Encouraging Open Communication: Leaders should create an environment where team members feel comfortable sharing their experiences and insights without fear of defensiveness.

The Role of Chaos Engineering

  • Proactive Learning: Nora advocates for using past incidents as a basis for chaos engineering experiments rather than injecting random failures. This helps organizations prioritize their learning efforts.
  • Evolution of Thought: Nora’s perspective on chaos engineering has evolved to emphasize the importance of understanding previous incidents to inform future experiments.

Founder Insights

  • Transitioning to Leadership: Nora reflects on her journey from an individual contributor to a founder, emphasizing the need for rapid iteration and learning in early-stage companies.
  • Importance of Networking: Building relationships with investors and industry peers is crucial for founders. Nora shares her experience with VC Elliot Durbin, highlighting the value of trust and open communication.

Nora's Advice for Engineering Leaders

  • Focus on creating a culture of learning rather than solely on generating action items after incidents.
  • Understand the unique challenges different stages of company growth present and tailor your approach accordingly.
  • Embrace the importance of experimentation and iteration in both product development and team dynamics.

Conclusion Nora Jones emphasizes the need for organizations to view incidents as opportunities for growth and learning rather than just problems to solve. Through proactive analysis and a focus on building a learning culture, teams can enhance their resilience and performance.

For more insights from Nora and other industry leaders, check out Jeli.io or listen to more episodes of *Dev Interrupted*.

Links

  • [Jeli.io](https://www.jeli.io/)
  • [LinearB Free Trial](https://linearb.io/start-free-trial?utm_source=podcast&utm_medium=referral&utm_campaign=devint-shownotes&utm_content=shownotes)

---

This markdown summary encapsulates the significant discussions and takeaways from the episode, providing an accessible and structured overview for readers interested in software engineering leadership and incident management.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I tell orgs, you should actually over-index on calling incidents. A big warning sign is when I talk to a leader and they're like, we don't have incidents. We actually don't have any incidents. And you're like, really? You'd be surprised. I hear this from like very large companies that everyone uses. And that scares. I'm like, I know you have more incidents than anyone else. And incidents are surprises. Incidents are anything that pulls you out of your day that you did not prepare for, right? There are always things to learn from. At DevInrupted, we work to give engineering leaders actionable ways to improve their teams.

0:31That's why we're producing a three-part summer workshop series with Linear B. Each of the three workshops will explore the processes that elite software engineering organizations and executives use to deliver better business outcomes and reduce cycle time by 47 % on average in just 120 days. You'll learn how to assess your current performance and benchmark it against industry averages, streamline processes through automation, and improve business outcomes through resource allocation. Learn from the best and take your team to the next level. Visit our website to learn more and secure your spot today.

1:04We're back on Dev Interrupted, and I'm excited to be joined by Nora Jones. Nora is the founder and CEO of Jelly.io, an incident management platform, and she's also the ex-head of chaos engineering at Slack. Welcome to the show. Thank you so much for having me. I'm excited to be here. I'm really stoked to have you on here because you have a couple of, I would say, almost controversial opinions about how engineering works. One of which is the talk you gave today about leading from incidents and how you think incidents should drive the decision making for a company. Can you expand on that? Yeah, absolutely.

1:39So incidents and unexpected situations are happening all the time. We are constantly getting surprised in our organizations. You know, we have this plan, right? And we set out to do it and we have this feature roadmap. And then all of a sudden a customer is complaining about something. Something doesn't look right. We get this weird spidey sense feeling and we all jump on and we fix it. And then we have to go back to our roadmaps afterwards or back to the things that we had promised to be delivered. And we kind of just leave all that happened there. And there's so much data hiding in that. There's so much investment that we can get out of, you know, this thing we already spent a bunch of money on because it took a lot of time and energy to spend money from the incident.

2:19And there's just there is a lot of value in understanding how people talk to each other during an incident, who they brought in, what they were looking at. Like all that data can be used to understand how your organization works and also help create this this learning culture. So that was kind of what my talk was about today. And then, like you said, I have many controversial engineering. I'm excited. Let's dig in a bit more, though. So I'd love to talk through an example if you have one that comes to mind about how a company can leverage an incident to really learn and grow out of it. Yeah. So I talked about one pretty extensively in my talk today, but I think the real data is hiding in your seemingly innocuous incidents.

3:05You know, the ones where you went, oh, phew, no one noticed that. Like it wasn't a step one. Yeah, we're all good. Right. that's that's actually where you can get the most learnings because emotions aren't as high afterwards when you have a very impacting incident like there was a data dog one recently right and i guarantee you things were emotional internally there afterwards like a hundred percent and if they're not practicing every day you know the seemingly innocuous events the ones that take five minutes the ones that no one noticed like it's going to be hard to glean data from those large scale emotional ones.

3:41And so an example I would take would be find an incident in your organization that didn't take very long compared to your normative standards, but maybe took a lot of people in your org to help fix. And I would ask you, like, did you actually take the chance to learn from that incident, to do a review on it, to interview the people that participated in it? That's an example I would use. And that's one I go over in my talk today. It's like a seemingly innocuous search incident. And then, you know, the team did an incident review afterwards and I'll put it in quotes, but they really just kind of went through a checklist and then got back to work.

4:16And it just it wasn't that valuable to the organization. And so in the talk I give, I kind of take an alternate approach to doing that. Let's talk through that approach a little bit. Do you view the incident report as kind of the key stepping stone to saying, OK, we have to be really extensive here and this is how you start to make these changes? Or how would you approach it? A lot of it is in the cognitive and coordinative work and it's unearthing that. So in every incident you have, you're calling in experts on the situation, right? You can't fix the incident without having some form of expertise.

4:49And so your job after the incident is over is to unearth what made sense for people during those moments, how they figured out what to do, how they learned what to do. And you can do that by asking them certain questions like, oh you know hey Brian how did you get pulled into this incident and Brian can say Joe pulled me in oh how do you tell me how you and Joe know each other oh uh and like it might feel silly asking these questions because you think you know Brian's answers but the thing is Brian wouldn't write this on his own like he wouldn't think to you know and so you're now asking these seemingly silly questions but by documenting this you're able to actually start to learn how your org works and so So Joe can be like, oh, you know, Brian and I worked together on an outage like this a year ago.

5:36Oh, tell me about this outage a year ago. Right. And so then you can start to untangle everything and understand how these patterns are kind of surfacing. And then you can also ask Brian, what were you looking at at this particular time? And he can say, I was looking at this graph I made. And I can say, can you show me that graph? And he might raise an eyebrow at me like, oh, Nora, you know this graph. Like I've seen you use it. But that's not my job in this moment. My job is to like unearth these things and capture them so that other people later on can then learn from this. And like capturing these things, be it in a report, be it in a meeting, be it in Slack.

6:14It doesn't matter. It's like where where does your organization learn? Are you a document culture? Are you a video culture? Asking those questions and sharing those answers in a way that other people can hear and learn, too. And it ends up helping new hires. It ends up helping your leadership teams. It ends up helping your current team. And it just, you know, it ends up growing more experts too. So based on this reporting and logging you're doing and this kind of discovery process to understand where the communication layer is around incidents for your organization, what are the next steps? So I know you mentioned documentation.

6:51How do you enforce folks actually learning from it or encourage that learning from the documentation instead of just having it be buried somewhere? And that's what I was saying before. It really depends on your kind of culture. I've been in cultures that were very meeting averse and they were like, document everything. Everything was in docs. Then I've been in cultures that met all the time and didn't really document anything. And like there's pros and cons to each approach. And I think what you need to do with this is do what feels natural to your culture. Don't try to introduce something totally different, but introduce the way that people are already kind of working, you know?

7:27No prescriptive plan, but instead approach it the way that's going to make sense for your org. I think started a little bit more like started by asking various questions. And on the Jelly website, we have a guide called the Howie that actually like goes through all of this and in a lot of great detail that can kind of give you this this approach. But I would be prescriptive at first. I would start to do it kind of under the radar, like little things like asking questions about having someone pull up their graph, maybe naming names in your incident reviews. Doing stuff like that can help you start to build this kind of culture over time.

8:03But, you know, in a prescriptive plan, I would basically say you should find the responders. You should talk to them one on one. You should ask them various questions about the incident. And your goal when you're talking to them is not to put them on the defensive. It's to make them feel like an expert because they are. And they're an information source for you. Yeah. And a lot of the times people are like excited to tell about their expertise. They are excited to tell you about all the things they experienced in this incident. The trick is on you, on how you talk to them, right? And you're a podcast interviewer.

8:39You're good at this. You're good at making people feel like experts. There's probably all sorts of techniques you use in the way you ask questions, the way you talk to people, the way you research them. It's the same thing with incident interviews. And those go a lot better and you gain a lot more value for them if you take that time to prep. Like, oh, wow, I noticed you were in this incident a couple of weeks ago, too. Like, what was that like for you? And then, you know, they might just be like, it was fine. And you want to start off with a softball question, which is probably what you do in your podcast.

9:07Sometimes, sometimes. And like, but like jumping into those things helps people more comfortable. And when people are more comfortable, they tell you more. And then afterwards, you know, what I would say is amalgamate and aggregate all the things that people told you. And that's that's your report or that's your facilitated meeting or that's your email or your Slack or however your culture works, you know, it's about taking that information and disseminating it afterwards. But that's the prescriptive approach I would take is making sure your goal, I think it goes back to your goal, like your goal with doing this kind of incident review should be to learn, not to report, not to make action items, not to guarantee it will never happen again.

9:53Your goal should just be to learn. And then all those other things kind of fall out of it. So drive the learning culture first and figure out the right way to approach that for your team. Yeah. Instead of trying to force it in iteration immediately and say, oh, well, here's the action item. Let's yeah, let's expand on this. Let's dive in and learn and get people's expertise to your point. And then we should naturally iterate off that. But it's not going to be something you want to force. It's kind of what I'm hearing. Yeah, it's not something you want to force and you actually when when people are just given the space to learn you actually come up with better action items so in my talk today i showed a real post-incident review and like the action items were fix the alerts fix the tests which like that's what happens when you do a shallow incident review and that's what happens when the goal is to report and not to learn it's a bandaid and you could put those action items in any incident that has ever occurred in the history of the software industry.

10:51Right. And so what value is that, you know, and you're actually just wasting cycles and spending more time when you're really focused on action rather than focused on learning and focusing on learning pays dividends in the long run. It helps you build more expertise in your organization. It helps you create an organization that's curious and asks questions rather than keeps information to yourself, it helps you run more efficiently too, right? You don't end up over hiring and over bloating because people aren't aware of how they're all coordinating and you end up focusing on some of the right things.

11:27It also helps you retain your best employees. If people are feeling listened to and have forums to talk about various things in a structured way, it helps you keep them, the right folks that you do want to retain. I totally agree with you about that learning culture, Nora. And it's obviously something that's talked a lot about in a personal sense of like lifetime learning and the value of the importance. But it doesn't always translate to organizations, to your point. Right. And I know, you know, we mentioned some like semi-controversial opinions at the start of this. I know that you have a lot of beliefs about how organizations can take advantage of some opportunities to learn.

12:04And I know one of them is that you believe we need more chaos engineering, which is a great phrase, which is always fun. And let's dive into that. Why more chaos engineering? How does that relate back to, you know, incident manager or anything else? It's a good question. And I think, you know, and I got feedback actually from one of our customers, Jason Copy. He was like, you know, when you originally tried to sell me jelly, I was like, she's the chaos person. Why is she trying to sell me this incident platform? And they're a very happy customer now. But it was like, I think initially a thing that felt different, felt like a departure and incident analysis and incident management is actually an evolution of my thinking around chaos engineering.

12:45You know, I started my career always in risk and always in helping understand incidents and the quality of the situation, depending on the product need at the time. Right. We're always making tradeoffs to hit deadlines. And I think there could be a lot more. I think there could be more beneficial organizations if people were being proactive about learning from incidents and maybe purposefully injecting failure. But my thoughts have evolved over time. Like rather than injecting failure into your systems, you really should be looking at the incidents that you've already had. You know, I saw a lot of people taking that advice I gave like, and I think it was in 2017 to do more chaos engineering, but I saw it in the absence of looking at the incidents they already had.

13:31So they were like injecting chaos to like learn from things and see how things reacted. But the hard part was, it's like, there are a million ways you could inject chaos in your organization. You could do a thousand different things. Are any of those actually going to be useful, right? Is it useful to fail a data center and see what happens? Is it useful to maybe take one of your devs off the on-call rotation that everyone relies on? There's a million different things that could happen, but the question is, is it useful? And how do you know what kind of experiments run. And the way you can learn what kind of experiments run, if it is useful to do that on-call thing, if it is useful to take down a data center, is to look at your previous incidents and to look at, and not just incidents, to look at your surprises, to look at the things that made you go, phew, like those are good situations to maybe inject chaos.

14:21But what I was seeing the industry do was to take that advice and just do any random thing, right, without actually, it was just more based on a gut feeling, which those conversations can be useful, but they're really useful after an investment you already made after something you already spent all this time paying for, which was which is an incident, you know. Was it this evolution of thought around chaos engineering and incidents and what you should be evaluating that led you to found Jelly? Yeah, totally. I got hired at Netflix in 2017 and my title was chaos engineer. And I was building a platform that allowed people to automatically inject failure into production without actually impacting the customer experience.

15:06So we would have like basically two clusters. You know, it was kind of like an A-B test and you would put most of your people into the to the non-experimental cluster. And then you would put, you know, a very small percentage of users into the experiment cluster and you would create a failure. And the hypothesis was always like, oh, if I take this down, it won't impact. It won't impact these users. They'll have the same experience as these users, right? You wouldn't run the test if you thought it was actually going to fail. But obviously, sometimes we learned that, you know, it did fail and they did have a bad experience.

15:40And a Netflix bad experience meant like, oh, I can't press play or I can't, you know, search for my movies or I can't stream something at this particular time. And so we learned a lot about how our users use the system. But that learning ended up being confined to like the team of four people I was on. And none of us were running production systems. We were making code to learn these kinds of things. So dissemination wasn't happening to the rest of the org. Yeah. So it became really important for us to educate people. And then I started thinking, well, the best education is going to be if they do it themselves, we shouldn't be.

16:15I shouldn't be running chaos experiments on search or bookmarks, you know, like it's going to be most useful if they can experience. and see it. And they were all like, I don't even know what to run an experiment on right now. I could, you know, do something related to the code I'm working on right now. And so I started looking at their past incidents and I was like, what was some what was an incident in your head that went badly? And then I started talking to them about it. And my whole goal was to drive them to use my chaos tool more. But then I started realizing that there was so much other data there and how they worked.

16:46And that was sort of like how the evolution came. I went to Slack after I started doing incident analysis at Netflix. I actually spun off from the chaos team and formed a separate team with one of my colleagues, Lauren Hochstein. And then I got hired by Slack around a time when they were IPOing. And when a company has a lot of PR presence, maybe they just raise money, maybe they just launched a feature, maybe they had a big outage, maybe they're hiring a lot, they end up usually having a lot more incidents too. And so I got hired at a time when Slack was having a lot of incidents. And I was running around like, you know, trying to help people learn from incidents, trying to create this dissemination so everyone could learn from it.

17:30And I realized there could be a product in the market for this to really help orgs do this better. So you're a couple of years now in your founding journey with Jelly. And I know you've kind of taken the journey that a lot of software engineers think about, which is, hey, I'm going to go learn at these innovative companies. I'm going to go do big things. And then I'm going to say, OK, here's a problem that I see. Let me jump in and be a founder. Now that I've gained this experience and have this skill set, what have been the learnings for you as a founder? I think one of the biggest learnings is creating this learning organization myself.

18:02Like I've always been creating one from the ground up, you know, as an individual contributor, as a leader of a small team. But now I'm creating it, you know, I'm enabling it from from the top down, too, and like allowing it to be enabled from the bottom up. And so it's a lot of like, I think, putting my money where my mouth is, too, you know, just like making sure I'm doing all the things. And also it's taking into account the phase a company is at. I think that's been a really big learning, too, is like the type of learning from incidents you do is different depending on where your organization is at and your journey to product market fit, too.

18:38How do you think about that differentiation depending on company size or journey? Yeah, I mean, so before we had a product in the market, there were four of us on the team. And there were like certainly surprises and certainly ways we surprised each other, but we had no users, right? And so there was really incidents. And so we had to take some time creating that culture of like asking questions and sharing information and documenting things in a way that people actually wanted to read in absence of user facing issues. So we kind of got to practice it. And then, you know, one thing I usually recommend for folks is to not have the person most involved in the incident run the incident review, because like I was telling earlier, you need to be able to ask them.

19:22It's hard emotionally. Well, and you need to be able to ask them the silly questions like someone that was in the middle of the incident is really bad at telling you what happened, actually. like they can tell you what happened in their own mindset but in order to really get the data out of it that's gonna improve the organization you need someone else asking them questions and so in a perfect world you have someone that was not involved in the incident but when you're a small company hard to do literally everyone's involved in the incident so we had to get really good at like playing these third-party roles taking the person that was like least involved in the incident and also you know really making scheduled time for learning i tell orgs this all the time like it's It shouldn't be easy to schedule over a post-incident review.

20:05Like you should prioritize it because if you're not prioritizing those things, it will come back to bite you. It's not in the next week. It's not in the next two weeks. It's like for years to come, it impacts your culture slowly but surely because you are still taking action on those incidents. But there's been a lot of ways that my mindset shift. I'm not coding anymore. I'm not on the ground anymore. I'm like, I'm leading folks to build this product that I genuinely really believe should be ubiquitous in the industry. Has that transition to not coding on a regular basis been challenging? I started transitioning into that when I was at Slack.

20:44And so I think what is what I miss the most, I do miss getting in the technical weeds. I miss system design. I really liked that. and I miss actually doing long form incident reviews. I actually, I have a lot of fun. I still get that, you know, I get to do it sometimes. I get to do it with some of our customers. Occasionally, they'll let me sit in on an incident review and I'll get to kind of play that role a little bit. Like that can be really fun, but I'm getting really, like it gets really exciting to see those light bulb moments go off for our users now. You know, I'm getting my dopamine hits in my career and like in a different way, you know, I'm not doing it from code or like writing incident reports anymore.

21:27I'm getting to see other folks do it and enabling that to happen, which is it's really it's really cool and different. Are there particular incidents that Jelly has already learned from internally that you could maybe talk about? So, so many. We have them all the time. And I and I tell orgs like you should actually over index on calling incidents. You know, I talk a big warning sign is when I talk to a leader and they're like, we don't have incidents. We actually don't have any incidents. You'd be surprised. I hear this from like very large companies that everyone uses. And I'm like, and that's good.

22:02I'm like, I know you have more incidents than anyone else. And incidents are surprises. Incidents are anything that pulls you out of your day that you did not prepare for. Right. That like those there are always things to learn from. Um, but I say all that, even though we're small and like, you know, we're still newer as an organization, we still have incidents and we, we talk about them all the time. We're doing learning reviews every week right now, actually. And we have one, I think in an hour that I, I'll probably dial into, but, um, we had a unique one this weekend related to, to SVB going down.

22:36That was a, that was a really different type of incident. We're going to get in here. Okay. That wasn't a technical incident, but it was, it was an incident. I had to jump in a situation room. We had to figure out things to do in the moment. It was pretty wild. And then we've had kind of your standard ones as well. We've seen issues in our demo account versus our production account. We've seen issues where core features of our product are not loading. And we take the time to understand how it unfolded. And then everyone on the call, regardless of if they're an engineer, actually starts to understand how the system unfolded.

23:12And who I've really seen it be helpful for is actually customer success because they know when issues are happening, they can actually start to understand how they might have happened and be able to triage a little bit. I like that you've widely defined incidents here. Yeah. You're like, look, our banking system being an issue is an incident. It's not just our own internal technical challenges. It's what are things that are affecting the business? What are things that are affecting our user experience? Can you share a bit about your approach to the situation room and my crisis like that, whether from the SVB incident or something else?

23:47Yeah, I mean, the SVB incident was like very, I think, very different than I would have handled a lot of other ones. But it was also very similar. You know, it was like we were in focus mode. We didn't have time to like pepper things on. And like when you're in the middle of an incident, you have to be very direct. you have to make decisions quickly and you have to make decisions in the absence of the full scope of information and that was what we were having to do this weekend and you know ultimately like with the news on Sunday night you know everyone has had access to their funds but up until Sunday night we didn't know if that was going to be the case yeah and so I was having to make all these contingency plans are we going to make payroll right can we keep our service running?

24:30Do we need to raise a down round suddenly? Yeah. Like how, how, what kind of steps do I want to take in what order? What is up? What is plan A? And then what do you know, what does plan E look like? Right. What's the worst case scenario here and really like plan that out. And I was also like having to communicate with my stakeholders, which in this case was the team, you know, and in usually usually in incidents, it's the team communicating with me or someone from sales. Right. And so it was a really unique and opposite incident. But I was like, hey, here's what we did today. Here's what we know.

25:06Here's what we still don't know. And here's what we're trying to do. I'll update you tomorrow at this time. And so doing things like that, I mean, it was it was nice to like be able to get in that role again and actually just, you know, understand that perspective. And I felt prepared to work in the incident because it was an incident, but it was it was definitely a role reversal for sure. I can imagine that that kind of role reversal in a situation where it's a fairly unique incident. Yeah. Must have been really hectic. Are there any particular learnings you took away from this unique type of incident that you see yourself applying either as a founder or leader moving forward.

Read the full transcript

25:49Definitely. I mean, like I retro to, you know, it was it was me and a few people on my team working through the weekend to try to figure out what we were going to do. And, you know, we're still talking about it. We're still cleaning up some things from it. Right. Just because SVB announced the government announced that it was OK doesn't mean there's like not things to do. But we're not sitting there like what are our action items? Because when you do something like that, you really end up making the wrong decisions. Like I was saying earlier, you know, I think maybe a lot of founders might come out of this with shallow action items, just like, oh, I should have this many bank accounts now or like, but you know, they're not totally understanding the context of their situations, right?

26:31Like when you're a founder, when you don't have a financial background, you do what you can at the beginning of the genesis of your company. It is not a good fiscal responsibility to hire a CFO when you don't have that much money and you're three people like that. Your company will never make it if you're doing that in the early days. And so I think it's it was important for us to understand like how that risk even got there to begin with. But it's also important to understand, OK, what do we know now? Right. Because we learned a lot during the situation, too. Yeah. How large is your team at jelly now?

27:09We're 24 people. Congratulations, first of all. Thank you. Yeah. I see where you're hitting that kind of scale point where it's like, OK, it's starting to get harder to do that synchronous communication, right? It's no longer four people in a room. Yeah. It's still a small enough team that most people know people are doing. But you haven't quite brought in some of those like specialized roles to your point, probably. Right. Where it's like, OK, like, I'm going to guess you still probably don't have a CFO or you have someone maybe a contract in the finance team. And that brings a whole set of challenges.

27:38And the risk reward of that, I mean, I think is on display here. But it's interesting to kind of see that iteration happening for you, where like you clearly, because you've built this into the DNA of your company, have said, okay, we're going to evaluate all incidents. We're going to see how this approach can help us learn and grow and learn from it. So with that in mind, I'd love to know if you were going to go back and re-found Jelly, I wound a time machine, what would you do differently? I mean, I have so much in, I'm a different person now than I was when I started Jelly, you know? Like I have all this expertise now about like being a founder and running a company.

28:18Like I've learned so much. I don't, you know, I don't, I really, like I like where we are today. I like the plans that we have in place, but I think execution is really everything in the beginning. Like seeing what sticks, what doesn't and learning from that quickly. You know, I come from a site reliability background. I have a lot of big network and site reliability engineers. And our whole job is to think of all the ways something could fail, right? Which like, you don't always want to be thinking about that at the beginning of the genesis of a product. And so I had to learn how to take off my SRE hat.

28:55And now I like fully know how to do that. But it, you know, it took a little bit of uncomfort because I had been trained for 10 years of my career on on all the thinking about all the things that could go sideways. Right. But like when you're a founder, you need to focus on on the product and getting it out the door and getting customers. And yeah, I just it's it's a lot of learnings. And I feel I feel really great about where we're at, what we're building, where we're headed. Yeah, I hear this kind of kernel from you that I hear from a lot of the best engineering leaders and leaders in general that I talk to, which is early on, it's so important to execute quickly, iterate and learn from what you're doing.

29:34Yeah. And I've heard this advice before, but it sounds like this is like specific advice for software engineers who are transitioning to that leadership role where it's OK, like you don't need to plan everything out. No, you don't need to have the best technology to get something started. Right. Going in, jumping in, learning, executing, realizing you did something wrong, evaluating it, learning from it and keeping going and doing that quick iteration feedback back loops is really crucial, it seems like, for long-term success. Would that be a coaching summary of your advice there? A hundred percent.

30:06Yeah. And it's like, there's always going to be something to fix. You know, you got to push that first thing out the door, right? And like your team is going to have a million different opinions about things to fix. And like, as a founder, you got to make that decision and you got to own it and you got to, you have to get it out there. And you also like, you have to be willing to kill it if it doesn't work. You know, there's like various features like that. You should try it out, see if it sticks. If it doesn't, what did, you know, you retro, what did we learn from it? It's hard. And it gets harder if you spend more time on it without anyone using it.

30:40And you've spent all these months on it. Like no one wants to kill those things, but it's so important, you know, to be able to and to create that rapid iteration and rapid learning culture at the same time. What about as a leader and kind of people manager? Have there been particular learnings that you've experienced as you've scaled your team from, you know, the four to 25 stage? I know that I talked to a lot of leaders who are kind of later on in their org scaling who are like, oh, you know, I did that a couple of years ago, but now I'm focused on, you know, 100 to 200. So I'd love to kind of dig in to see what nuggets you have around that early stage of like, okay, like execution mode now to some product market fit.

31:23Yeah, I mean, you know, it's it's interesting because I actually feel more organized now than I did when we were with four people. And I think it's because we were pre-product at that point in time. And it's a lot of cooks in the kitchen when you're pre-product, right? And everyone's kind of talking about things. And as a leader, you want to support everyone's thoughts and processes. But I think like figuring out how to get into a mode as small of a team as possible where you're shipping and building the culture quickly has really long term lasting effects. And now at 24 people, we're kind of moving and grooving.

31:58We have a really great shipping culture. We just hired sales and marketing. Two new roles on the team that no one's worked with before. And it's like it's a recalibration. I genuinely think in the early days, every new employee you add, your culture changes. You know, regardless. of who they are or what role they're playing, especially when it's a brand new role. And so now we have more coordination and touch points. We have more things that product and engineering and design impact other than CS, right? And it's also like my job changes all the time. I think in the early days, I felt like my job was changing like every few weeks and we go through these like rapid scale points.

32:38But I was, you know, myself head of sales and head of marketing before. And now I get to take off those hats and focus on something different. And so it's like enabling people to succeed, but also being prescriptive enough as a leader that like they have good boundaries to play in and be creative. Well, the good news is adding these new folks is going to bring up new incidences, winning the team, and it's going to give you opportunities to learn. So that's exciting. And congratulations on these new hires. Thank you. I'm curious about something else that I saw in your bio that you've been involved in angel investing for several years.

33:14How did that inform your approach to starting a company, particularly on the fundraising side? Because I know there are a lot of disparate opinions around what products are right for venture funding versus which ones should be bootstrapped. How did you approach it and think through that process when you were starting out? When I was raising for Jelly? Yeah, it was pretty... I have an entrepreneurial mindset in general, but like this is my passion this is like my life's work I've been thinking about risk in in analysis and incidents and how it can help organizational psychology for for so long it's just like this is my thing this is what I want to do and so I knew I wanted to make this a reality I had several years back when I was an employee at Netflix connected I did a reference call for a VC.

34:06And as I was doing the reference call, we ended up chatting for an hour. And I was so impressed by this VC. They had done so much research on the SRE space. And it was my first encounter directly with a VC, actually. And they asked if I'd be willing to do some diligence calls in the futures for them. And I just, I kept in touch with them. And then they asked, like, would you ever start a company too? And I was like, oh, maybe I have these kernels of an idea. And so as I was like at Slack, I realized there was a big market for this and a big need for it, right, to help companies learn from incidents when they are moving so fast and don't feel like they have the time.

34:45And so that's why Jelly was created. And I, you know, I mentioned it to him and it was just it felt pretty kismet just because I already had that trust built up. Right. I wouldn't advise founders that have not spent time building relationships or anything like that to go that route until they, you know. Spend the time in that space. Yeah. Like, it just, I don't know. Like, I get asked a lot of founders, like, about my experience there and about how they can go raise money. And I'm like, you need to be thinking about it years before you take your entrepreneurial journey, too. because it's a big role, you know, and usually like your first VC takes a board seat, you know, imagine like a board seat with a stranger.

35:35Like that's wild, you know, like you don't have that trust built up yet. And I also, I felt great about adding this first investor because it was someone that I felt like I could ask questions to about building a company. Like I know SRE and I know product really well, but it was my first time being a CEO and I wanted someone that I could, I could ask questions to without fear of looking a certain way, you know? Do you want to shout them out? I don't mind you saying the name. Oh, my VC? Yeah. Oh yeah. It's Elliot Durbin at Bold Start. Yeah. He's great. Yeah. But yeah, I mean, I've done, I've done a couple angel investments, you know, it's, it's folks like the space I know really well or companies I really believe in.

36:17And it's a lot of fun, but yeah, it's definitely, It's definitely a different decision to go the venture route. I would only recommend doing it if you think it's going to scale your growth, which for us, it absolutely has. Well, thank you for sharing that. I know both of our audience find these stories from founders like yourself, from leaders like yourself, really valuable. Yeah. Do you have any final thoughts that you want to share on any of the topics we've discussed today or anything particular you want to shout out? I really enjoyed giving a talk at Lead Dev today. I think, you know, check it out.

36:47Check Jelly out. We're all about helping you move faster, learning from any surprises that happen, but also doing it in a way that enhances the quality of the outcome and in a way that helps your organization improve. I think at this particular time of the industry right now, there's a lot of orgs going through layoffs. And so a lot of expertise has left the door, which like a lot of orgs are still having the same expectations of their employees, even though all this expertise has left the door, which leads to a lot of reliability challenges. So I would say, honestly, this taking the time to do incident reviews is more important now than it ever has been, especially if your org has gone through something like that.

37:30Really well said. It always brings to mind that XKCD classic comic that I'm sure everyone listening may have seen at some point of this giant infrastructure held up like a tiny block that one engineer is like maintaining to get going. And, you know, sometimes that engineer gets laid off and suddenly down the line, that little block breaks. And so does a lot of other things. So well said on continuing to learn and explore incidents. And thank you so much for bringing this mindset and approach and joining us on Dev & Rupp today. I've really enjoyed it. Yeah, thank you for having me. I will quickly bring up another comic that that reminds me of.

38:02There's like, there's all these stick figures pushing a car with square wheels and they're, they're doing it so fast. And this person comes up to them with, you know, a wheel and he's like, Hey, this will help you. And they're like, no, thanks. We're too busy. And I feel like that's what a lot of orgs get into after incidents too. So I just, you know. Great reference. Thank you, Nora. Yeah. Well, thanks again for coming on the show. Yeah. Thanks for having me. If you want more interviews like this with Nora and other founders, let us know on social media. And if you found value in this today, maybe check out Jelly or consider giving the podcast a review.

38:39We always love it when you get those reviews on Spotify or Apple Podcasts. And thanks so much for listening or watching.

From the publisher

If you think your org doesn’t have any incidents, it’s time to change your definition of an incident.

This week we’re joined by Nora Jones, Jeli's founder & CEO, to help us make sense of incident analysis and explain why so many incidents go underreported. Before beginning her journey as a founder, Nora helped pioneer chaos engineering at companies like Netflix and Slack where she developed a passion for understanding the intersection of software and people.
 
A stellar engineer, manager & founder, we caught up with Nora on the heels of her keynote address at the LeadDev conference in New York.

Show Notes:

OFFERS

  • Start Free Trial: Get started with LinearB's AI productivity platform for free.
  • Book a Demo: Learn how you can ship faster, improve DevEx, and lead with confidence in the AI era.

LEARN ABOUT LINEARB

  • AI Code Reviews: Automate reviews to catch bugs, security risks, and performance issues before they hit production.
  • AI & Productivity Insights: Go beyond DORA with AI-powered recommendations and dashboards to measure and improve performance.
  • AI-Powered Workflow Automations: Use AI-generated PR descriptions, smart routing, and other automations to reduce developer toil.
  • MCP Server: Interact with your engineering data using natural language to build custom reports and get answers on the fly.

More from Dev Interrupted

All 208 episodes
Redefining Incident Response: Insights from the Chaos Engineer Behind Jeli.ioDev Interrupted · 39 min
Listen in VO