beeps and on-call for Next.js developers with Joey Parsons

21 Jan 2025 · 48 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: Beeps and On-Call for Next.js Developers with Joey Parsons

Episode Overview Podcast Title: Software Engineering Daily Episode Title: Beeps and On-Call for Next.js Developers with Joey Parsons Description: Joey Parsons, founder and CEO of Beeps, discusses creating an on-call platform for Next.js developers, addressing the limitations of current on-call systems, and building a developer-first experience in a rapidly evolving tech space.

Key Points

Introduction to Beeps

  • Beeps: A startup focused on developing an on-call platform tailored for Next.js developers.
  • Next.js Dominance: Recognizes Next.js as a leading framework for modern development, with many startups launching on it.
  • Developer-first Approach: Aims to enhance the on-call experience specifically for developers using Next.js.

Joey Parsons' Background

  • Experience in Tech: Over 25 years in tech, with specific expertise in infrastructure and reliability.
  • Previous Ventures: Founder of FX, acquired by Figma in 2021, where he focused on developer productivity.
  • Motivation for Beeps: Driven by the challenges of on-call experiences and the gap in existing solutions.

Challenges with Current On-Call Systems

  • Stagnation of On-Call Processes: Existing systems haven’t evolved in line with other aspects of software engineering, leading to inefficiencies.
  • Need for Context: Engineers often lack the necessary context during incidents, requiring time-consuming processes to gather information.
  • Traditional Rotation Issues: Many companies follow a simple rotation for on-call duty, which does not account for modern-day complexities.

The Vision for Beeps

  • Building Context: Beeps focuses on providing immediate context during on-call incidents by integrating with observability tools (e.g., Sentry, Axiom).
  • Streamlined Experience: Aims to reduce the time required for engineers to diagnose issues during incidents using automated context gathering.
  • User-Centric Design: The platform is designed for modern developers who may not have extensive experience with on-call procedures.

Tech Trends and Market Position

  • AI Hype Cycle: Discusses the current emphasis on AI in the tech industry and how Beeps is not solely focused on AI, but rather on solving real problems for users.
  • Unique Market Position: Beeps targets developers on Next.js, a community eager to adopt modern tools and services.
  • Integration Strategy: Simple integrations that require minimal setup, focusing on enhancing user experience rather than complex configurations.

On-Call Process Innovations

  • Incident Management Lifecycle: Acknowledges the need for improvement in the entire lifecycle of on-call management, from preparation to post-incident reviews.
  • Reducing Noise: The goal is to eliminate repetitive alerts that can diminish focus during critical incidents, allowing engineers to concentrate on novel problems.
  • Emphasis on Feedback: Actively seeking user feedback on the utility of the context provided during incidents to continuously improve the service.

Engineering and Development Considerations

  • Building with TypeScript: Beeps’ technology stack primarily consists of TypeScript, ensuring consistency and reliability.
  • Internal Agent Frameworks: Developed in-house to efficiently manage the complex interaction between various services and the context-gathering process.
  • Resilience Requirements: Emphasizes the importance of building a trusted, reliable system, as the brand relies heavily on its performance.

Future Aspirations and Company Culture

  • Support for New Developers: Recognizes the shift in on-call responsibilities to less experienced developers and aims to provide them with the confidence and tools needed to succeed.
  • Cultural Impact of On-Call: Acknowledges the personal toll that on-call duties can take, with a desire to improve the experience for the next generation of developers.
  • Vision for Improvement: Driven by personal experiences and a commitment to helping others avoid the same challenges, Parsons aims to transform the on-call experience in software development.

Conclusion Joey Parsons expresses a heartfelt commitment to improving the on-call experience for developers, fueled by his extensive background in tech and personal experiences. The episode highlights the innovative steps Beeps is taking to create a more effective and efficient on-call system for Next.js developers, aiming to alleviate the burdens associated with traditional on-call practices.

---

Listen to the full episode [here](https://softwareengineeringdaily.com/2025/01/21/beeps-and-on-call-next-js-developers-with-joey-parsons/).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Beeps is a startup focused on building an on-call platform for Next.js. The company is grounded in the key insight that Next.js has become a dominant framework for modern development. A key motivation in leveraging Next.js is to create a developer-first experience for on-call. Joey Parsons is the founder and CEO of Beeps, and he previously founded FX, which was acquired by Figma in 2021. Joey joins the show to talk about the platform, starting a company without an explicit AI focus, the limitations of current on-call systems, building on Next.js, and more. This episode is hosted by Sean Falconer.

0:38Check the show notes for more information on Sean's work and where to find him.

0:54Joey, welcome to the show. Hey, thanks, Sean. Great to be here. Yeah, absolutely. So you're the founder and CEO of Beeps. Tell me about the company. What's the vision? Where are you guys at today? Yeah. So Beeps is sort of like an on-call platform built for modern developers. And when I say modern developers right now, like the area that we're focusing on is Next.js developers, right? What we've seen, you know, if you go to product hunt or hacker news on sort of like a day-to-day basis, they are, you know, probably like 60 to 70 % of like the companies that are being launched today are being actually launched on this platform.

1:29So they're Next.js developers building on top of Rissell. And they actually have very unique challenges when it comes to infrastructure. So Geeks is an on-call platform that's being rethought from the ground up with, for now, the needs of this developer. So it's been pretty fun to build. We're about a year into it right now. Just launched our platform, and it's a pretty exciting time. Yeah. I mean, I guess given that you're seeing that kind of trend with Next.js and Vursell, I guess that's, you know, good job Vursell and the team over there. I mean, that's a good sign that they're doing well. So, and this is not your first company, right?

2:06You've founded companies in the past and you were on actually, you know, talking to Jeff in the past on here about a prior venture. Like, I guess like why put yourself through this pain over and over again? Yeah, yeah. It's an interesting sort of like origin story. So I guess to rewind a little bit, I've been in tech, but largely in sort of like the infrastructure and reliability side of things for a little over 25 years now. Like a lot of people these days might not know Rackspace, but they were one of the early sort of like hosting companies, like one of the precursors to like, you know, Amazon's and like the whole cloud today.

2:37And I was actually in like probably like the first hundred employees there. So like way back in the early aughts and got to sort of see sort of like at that point, how companies were building sort of like globally performant, globally reliable, like platforms at scale. And was lucky enough to be able to sort of like really build a career and an expertise around this. I guess most notably, I ran reliability in a large part of like the infrastructure team at Airbnb for a handful of years and got to see sort of like the evolution of like infrastructure from that perspective. And then, as you mentioned, after that, I left to start a company called FX, which is more of tracking microservices.

3:11Start a company. We were acquired by Figma in 2021. And I spent a year and a half there working on developer productivity. Again, restarting, thinking about reliability. So when I left Figma, I started thinking about the next thing. A lot of people were like, yeah, why would you put yourself through the pain of starting another company and going down this journey? And it's just, on-call has been such a big part of my life. Like reliability has been such a big part of what I've built my career around. And if you look at sort of maybe the last 10, 15 years of like software development, everything that an engineer does on a day-to-day basis has gotten exponentially better, especially in like the last two to three years.

3:48And, you know, whether it's like writing code, testing code, deploying code, like observing code and production, right? All of it's just worlds better than it was, you know, at the advent of like mobile and cloud. But the one thing that literally hasn't changed is on-call, right? There's been a handful of companies that have dominated in the space. And they've built great brands and great trust out there. But nothing has really evolved in the way that actually is looking out for the engineers that are on-call or just kind of making this process better. So, you know, having been through this for such a long time and, you know, been had the weight of deck of corns on my shoulders in the middle of the night and sometimes, you know, being on call for like early stage startups where like I'm literally the only person on call for like months on end.

4:30I'm pretty sure it's taken its toll on me. And I think that I feel kind of like obligated to make this better for like the next generation and at least try to like drastically improve this experience. So, yeah, I think it's a worthwhile endeavor. You know, the fact that you're on call for other people's on call is a little bit crazy. It makes it a little bit different than what a traditional startup would be. But I think, you know, if we can really crack this, then I think that it's going to be worth it in the end. So yeah, I always say during, I founded a company years ago, and I would say I was on call for seven years, basically during that time.

5:01And it's a lot, you know, it's stressful. And I want to get into some of this sort of downside of existing on call today. But before I get there, you know, one other question I had regarding, you know, starting a company, especially right now, like, have you found this time around that? Given that we're sort of in this like AI hype boom, it's kind of hard to start a company that's not focused on AI right now. Like, you know, when I was doing my company back in like 2009, 2010, like everything was about social. And, you know, people would always ask like, you know, what's your viral strategy or what's your social strategy?

5:34And there was all this pressure to do stuff that wasn't necessarily like core to the product, just to have a story around it to tell like investors and stuff like that. So I guess like, and then when these hype cycles are coming on and you're doing something different, it's like hard to sort of bust through the noise when everybody's like paying attention to this one thing. So I'm curious about your experience there. Yeah, I think it is a little bit funny. I think like when I was sharing sort of like the vision for beeps with one of our investors, like, you know, halfway through the conversation, someone interjected and was like, oh, this is the longest we've gone without ever hearing AI in a pitch.

6:05And this was like last year. So I think it's an interesting time. I think that one of the challenges I see these sort of days is folks are so much more fixated on technology than they are actually the problems. And we actually are heavy users of LLMs at Beeps, both for our own development and in Beeps itself. There's a bunch of stuff that's powered by AI, but we actually don't talk about it. Because we're much more focused on the problems and our users care about the applicability of it, much more so than they do the actual technology that drives them. And that's always been the case. Right. But it's not to say that it doesn't get folks excited.

6:40It doesn't get folks ready that they're using to be able to tell people and their peer group or their company that these AI products are making them better. So, yeah, it's an interesting space out there. I think, you know, in some companies that are like relatively in the same space as us, seeing sort of like the seed round that they're raising are pretty astronomical and like the, you know, the tens of millions and things like that. But ultimately, it's going to come down to who builds the most value for their users and regardless of the amount of money that they raised. And we'll see how that shakes out in a long time.

7:12But to me, it's actually like a really exciting time to build a company. It's a really exciting time to be an engineer, like a developer, building products and getting to try a bunch of these tools and seeing how quickly they're evolving and how quickly they're getting better. It's like, you know, every few weeks, my team's asking about some new IDE that they want to try out. and they love it more than the last one. It's just like, okay, when is this going to stop? But at the same time, if it keeps unlocking a lot of potential there, then that's a really great thing. So a pretty long-winded answer there, but I think that's kind of how I've been thinking about it and how it's kind of played out for us.

7:44Yeah, I think we'll kind of know that we're out of the hype cycle when I think companies do go back to somewhere more focusing on sort of the value of what they provide versus the technology that's behind the value, right? It's like when a company can essentially talk about like, hey, we can solve this problem for you. And that's really what the buyer cares about. But maybe there is AI that's powering that plus a lot of other stuff, right? But it becomes less of like, it's in the essentially title of the company's name. Yeah. Yeah. So many people are launching on.ai domains. I don't blame them, but it does kind of, I think it pigeonholes you a little bit in terms of like the solution that you're providing.

8:25But it's a tricky thing. I think it makes sense in this moment. And I guess we'll see how it plays out over time. So back to on-call, what does that on-call process look like for most companies? Yeah, it's kind of growing, right? Or it's kind of evolving, right? Traditionally, most companies have a very simple rotation, whether that is across one team or a bunch of different teams where you have basically your primary on-call. It's like the first person that's getting notified. You have a secondary. And then, you know, sometimes you might have like a manager is like your tertiary on call, you hook up your observability system.

8:59So this sort of like platform. So it ends up sort of like being like a routing layer for who gets notified when something breaks. And for most companies, it's just that, right? And I think that it's a, you know, a pretty standard thing at most companies, once you reach some semblance of product market fit, or you have users that actually matter. and one of the things that's like the companies that have been around for a long time they do this really well like pager duty for example right like their cto tim says all the time right like nobody ever gets fired for choosing pager duty right and i think that it's a powerful thing but they've dominated the market simply based off that brand right and that trust and that reliability and even someone that has run reliability at a company like airbnb or figma I know I can't easily just walk in there and say like, hey, like I helped build this team, right?

9:47Like use my new product. We just don't have that sort of like credibility. And I definitely understand that. So it's like, there's definitely like the technology piece of this, but there's also sort of like the brand building over time. That's going to really matter to be able to kind of like win over the hearts of developers and be that recognizable name that when they begin to feel that fit, when they have those users that matter, they ultimately sort of like have an option to choose there. But it's the most interesting thing about sort of the Next.js space and the Vercel space is that folks are willing to try and actually want to try sort of like modern tools that are tailor-made for that.

10:19Right. If you look at sort of, you know, they're choosing to host on Vercel and Fly.io instead of like running their own infrastructure on top of AWS, instead of rolling their own authentication systems or using sort of like older solutions like an Auth0, a lot of them are choosing to use Clerk or use like SuperBitsSauce instead of like running your own databases. So there's a sort of like passionate community that have all these different new problems, these new challenges, and they're willing to make bets on like earlier companies and kind of ride the wave with them. And they have unique challenges, unique things that we need to think about for how we deliver sort of a platform that helps them while they're on call.

10:55And that's sort of like the path that we're going. And I can get into that a little bit more. Yeah. How did you come to sort of recognize that, you know, the subset of perhaps the Next.js community was sort of the ideal customer profile for you guys? they're open to using these various managed services. They're sort of on the forefront of the technology adoption curve. Like, how did you kind of, you know, I guess like stumble into recognizing that this is probably like a good spot for us to actually go to market with? Yeah, I guess, I mean, like pretty simple answer, just kind of like as a developer myself, when I started hacking on things, I was using like the Next.js tool chain, right?

11:31Like it was just really simple for me to kind of get going, right? And then you look at like the credibility of the companies that have kind of like grown up on those platforms over time. It really reminded me of sort of like the late movement of like the, you know, 2008 or like the late 2000s, where people were choosing like Ruby on Rails, running Django, right? And then you have companies that like Airbnb and GitHub that were built on top of like Ruby on Rails. I think that same sort of movement has been happening over the last few years in the space. And if you look at like tutorials to get started on building a new app, right?

12:00Back then it was like, you know, how do you get started with Rails? It's like how Mongo really kind of got going, by really targeting the Node.js space. But now it's everything out there in terms of how to get going, what people are live streaming on Twitch and what a bunch of YouTube videos are for getting started with building a web app. It's all Next.js. So a lot of it was personal experience. A lot of it was just what you're seeing in the community and the excitement about it. If you go to Twitter and people are talking about it's really popular to have that list of like, oh, what's my tech stack in 2024?

12:33You look at those things And one of the leading things on there is probably Next.js. And it just became really obvious to us. And in terms of the existing on-call systems, what are some of the problems with the current approach? It hasn't evolved, sure. But why is that a problem, I guess, for companies? Yeah. I guess I'll take my time back to some of the companies where we build reliability and things like that. So a lot of the times, these companies do a great job of notifying you. Right. Like whether it's like a push notification or SMS, but it would just kind of drop there. And what engineers actually need are sort of like the rituals around on call, as well as like the things that actually help them during an incident.

13:14So if you break it up like incident management, like there's all the things that, you know, you do before an incident happens to make yourself resilient to them. There's everything that happens during an incident, right? Like getting the context to understand sort of like what's happening. And then after an incident, there's like all the stuff that you do to sort of like follow up with whether it's like incident reviews or postmortems to kind of get going. And every company at some level of scale that has things that matter ends up building a lot of like the tooling around the existing sort of like on-call solution.

13:44And a lot of it's, you know, very similar. Like a lot of people ended up following the Google SRE sort of like handbook, that book that was released, what was this, like 10 years ago, I think at this point. And have, you know, we're really focused on building those processes. And it was kind of strange that some of these incumbents didn't take the ball in front of them and didn't build anything to actually improve the full on-call lifecycle there. So I think that with Beeps, one of the things that we're really focused on is helping engineers build context when something bad happens. So one of the key features in our free product is that we work with a handful of the modern observability tools that folks are using in this space.

14:26So your highlights, your axioms, your sentries. Sentries has been around for a while, but they've kind of like the preemitted error tracking tool in this space. And what happens is most folks, especially at the early stage and as they begin to grow, they just have these alerts sort of piped into Slack. So let's say you have an exception being triggered or a notification coming from Axiom about some trend that's happening in logs. You'll get this nicely printed out sort of message there. and our sort of like our assistant in Slack will actually like listen for these messages and knows how to pull context out of them.

15:01So let's say you get something about like your users aren't being able to log in, right? Like something's triggered in Axiom where you've set up alert that happens. We'll automatically listen for that and immediately begin a thread to help you understand context. So if I rewind a little bit, what most happens, let's say I get, you know, you probably experienced this having been on call for like years on end is that, you know, the first thing you do is you probably open up a handful of tabs in your browser, right? You're looking at, okay, what got deployed recently? What were some of the commits in those deploys?

15:31Like, are there any like services that I use that are down or any servers down, right? What are the most recent logs for my application? And there's kind of like this playbook that you're running through every single time, just to kind of get context to like eliminate potential factors to make you understand where you're going to go do next. And depending on like the experience level of the engineer that's responding to this, you may or may not do these things. you may miss some information that should have just been an obvious thing to look at. And that's really hard when you're woken up at 3 o 'clock in the morning to remember to do all these things and remember which places to look.

16:03And it's a time-consuming process, right? You're navigating all of these different user interfaces that aren't consistent, may change. The runbook's usually kind of outdated when it comes to this sort of stuff. So instead of taking the 10 minutes to do that, we'll actually get all that information for you and print it for you there nicely in Slack in less than 20 seconds. So imagine being woken up at 3 o 'clock in the morning by a notification. Instead of the X amount of minutes that would take you to get to that context, you could, in your bed, understand what you're going to do next when you wake up and get to that context really quickly.

16:36So we hooked into all the different systems in this for-sell Next.js ecosystem really well. I'd have been able to build some really tight integrations to be able to build that context really quickly. So we think that that is going to be a really great lever for people to reduce the amount of time that it takes them to understand the problem, which then in turn helps them understand how to fix it faster. Can you walk me through, you know, what's actually happening in the software behind the scenes when NMRT comes in in order to build that context? Like, how are you actually going and essentially pulling in the right data from all these different places?

17:09How's the integration work and so on? Yeah, yeah. So one of the things that we really tap with beef is that you can get started in literally like minutes, right? So, you know, we just have, you know, simple integrations with Vercel. where most people are hosting their apps, a simple integration with GitHub, and then a simpler integration with Slack. You actually don't need to configure beeps to talk to Century or Axiom or Highlight. We've done all that hard work for you. So through the Vercel API, we're able to grab a lot of really great and rich information about your app, about its deployments.

17:37We have read-only access to your code base to understand the different providers that you knew. So we'll look at things like your package JSON to see the packages that are related to maybe a clerk or a super base and things like that. And then when our assistant is listening for messages in a Slack channel that you direct us to, and we actually will work with any sort of tool that is observability-based, that's setting alerts there. It doesn't have to be Century, Axiom, or Highlight. It'd even work with a data dog where we are looking at the messages and using an LLM to basically decide whether or not this is actually an alert coming in from an observability provider.

18:17And then at that point, we kick off the investigation process across a handful of agents in the background that understand we have a different tool for every integration that basically will go out and fetch this information, send that back up to the primary agent that's communicating through Slack. And then it uses that information to decide what it wants to share, what's the most important context based on the alert that came in to help the user understand what to do next. So again, the background a lot is AI and LLM-based when it comes to the prop moments. It's much more streamlined than that, but it gives us some flexibility to be able to not just give a static list of information back, but actually have context based on the alert that came in.

19:01Yeah. And then you can take advantage of some of the things that these AI tools are really good at, like summarizing information, essentially. Exactly. Exactly. In terms of the context, how do you test that the context that you're providing is actually valuable in the right context? It's a lot. I guess this is probably one of the hardest things to do right now. Obviously, we've written a bunch of tests ourselves that provide inputs. And based on the prompts that we have in the context that we pass those prompts, are we getting information back? And then at the same time, we have a bunch of test apps running wild in production that actually get real traffic and actually have real issues.

19:45And that's probably the best way to test that, okay, was this actually valuable to me? And one of the things that we do at Beeps is that we actually ask you afterwards, was the context that we provided useful to you? And we're using that to power a lot more stuff in the future. I think that there's definitely like a world where as we're able to provide better answers that might give you a hint about what to do next, being able to push that into a direction where eventually a user will have enough trust with what solution we're providing to where maybe they would hand the keys over to us and say, go fix this automatically.

20:22And I think that while I don't think we're there yet, and there's a lot of scar tissue with the whole AI apps promise over the last 10 years, I think there's a world where this could get to the point where there's a large share of issues that are relatively common across a code base that could be solved without human interaction. And hopefully we get there because that's one of the ways you make on-call drastically better. Yeah. I mean, even if you can get to a place where you can provide through an assistive technology, like a pretty accurate depiction of what the solution should be, that's massively more useful than like, hey, here's an alert.

20:59Like there's a problem. Go figure out the problem and the solution. Yeah. And we're very keen on like being suggestive and not sort of like promising the world on this sort of stuff. I think that like, you know, there's other companies in the space that kind of promise sort of like the world when it comes to that sort of stuff. And it's a different game than sort of like your co-pilots of the world that are helping you code, right? Like incident response is very high stakes. And again, like trust really matters here. So we'll see if they're able to have success by doing that. But I think, you know, it's one of those things where if you give someone the wrong answer and you give it to them with a lot of confidence, you can lead someone down a bad path that could have detrimental effects for their users and their company.

21:38And that's a big bet to make at this stage. So yeah, I mean, I think one of the biggest challenges with LLMs is that they're like inherently, in some ways, like the worst form of employee where they're like overconfident and incompetent at the same time. It's like very sure of themselves. Here's the wrong answer. Yeah, yeah. And then obviously that'll improve like over time, right? Like I think we'll get to where that maybe they're not as incompetent as the worst employee. But yeah, we're definitely not there yet, especially when it comes to solving some of these like tricky novel problems. How do we get to a place with on-call where you're not sort of bombarded by the noise of everything and you only have to deal with more, you know, essentially novel issues?

22:19Yeah, I think this is, you know, at least my opinion is that this is sort of like the path to get there. You know, what you don't want to have, you're sort of exactly right, is, you know, like so much of on-call, at least from what I've seen, is just kind of you're getting through the muck of the annoying sort of repetitive things, the things that happen like once a week that aren't actionable, that almost like really ruin the experience for you. And then when something novel does happen, you're too tired or you've been wrecked for like the last few days by these sort of like pestering alerts.

22:50And I think that the process that happens outside of on-call is really important there. And I think the on-call tool actually has a lot of influence on that. And the idea of sort of what happens in larger organizations, even smaller companies, is that you have these handoffs that may lead to action items of things to go look at and things to fix as a result of that. But it's a very human process. It's a ritual that happens maybe not very consistently. And the great thing is that with LLMs and some of these tools, you can power a lot of that. You know the history of alerts that happen. You know that they've happened at some level of frequency.

Read the full transcript

23:28You know that one alert has come in every Wednesday night at midnight UTC and, you know, being able to sort of like track that and understand like what to do about it and what's noisy and what's not. There's an obligation, I think, as an on-call provider to provide that information and make it easy to digest and easy, not have to be like a human process on the outside. And I think that's how you build continuous improvement into this process. Another thing that we kind of have a hot take on is the concept of incident reviews and postmortems. You know, like at a lot of companies, this ends up being a lot of theater, I think.

24:02And, you know, there's probably a whole group of like people that, you know, learn from incidents that are going to come after this. But I think that that's one of the things that we definitely want to improve upon. Because what an engineer really needs is not a document that they're not going to ever read again, that, you know, has a lot of really great information about it. But it's like surrounded by prose that they have to read through in sort of like a thick moment. What they really need is that context when it happens again, right? Like what is the evidence of this incident happening in the past?

24:33You know, what is the historical context here? How was it resolved last time? What are the different things to try? And I think that a lot of that could be actually like auto generated and provided at the moment something critical happens and provide that context that instead of again, this other document that, you know, was discussed and learned from, which is which was obviously great a lot of folks probably aren't going to remember that when the world's crashing around them and they've got to solve this issue so it's a tricky thing and we really want to make that better and i think a lot of these rituals that happen outside of actually being on call are one of the best places for you know helping an engineer have confidence have context and uh you know be able to walk into a very novel situation understand history and make it better but i think part you're right like the first thing is just really getting rid of like the noise, eliminating the non-novel things from having to be resolved by a human and then sort of taking it there.

25:27But we are ways away from that, just to be honest, right? That's something that we want to build into. I'm sure there's other folks building into this as well. But I think that there's a big problem to be solved there, but it's an evolution. So yeah. For where the product is today, do you see the main value proposition for an organization using beeps is around just like saving time, like saving, essentially like engineers, probably one of your highest cost employee resources within a company. This is a way to essentially reduce the cost where they're maybe not building essentially a core product.

26:01Yeah. As a CEO myself and having been a leader of a large engineering team, there's three things that I've always wanted. You want your engineers to have all the tools that they need to be productive, but you want them building products. A lot of times you don't want them to focus on infrastructure and reliability. I'm sure every CEO dreams that they can have an engineering team that instead of dealing with tech debt and incidents and things like that, they wish that they were all building product and building value for users. So the more that you can eliminate there, the better. And then obviously, number two is you want to have a fast, reliable, secure product.

26:37So making sure that you can meet the expectations of your users and And terrible things not happen is vastly important. And then you want to have a really happy, engaged team that cares about its users, that cares about a product. It's usually the top three things that you want as an engineering leader or even a CEO when it comes to thinking about engineering. On-call touches all those things. It really impacts people on the weeks that they're on call if they've had bad experiences. And every minute matters for your users when it comes to building something that's actually meaningful. And we're past the days where, you know, you can have like a, like a nine to five on call because your, your business is predominantly like in the United States or like regional, like everybody's got, you know, global customers that, that care and matter and having a team that being, that's being able to support that is like, is really important.

27:29And I guess kind of to take another step back, one of the other things that we're seeing is, I even saw this at like Airbnb and Figma, is that we're moving past a world where we're like people like me, we're the ones on call, the ones that grew up in like from sysadmin to ops to like SRE, I guess people are calling it like platform now, right? Where you have these highly trained engineers that have a compendium of shit in their brain that they've been through. They know how to deal with these things. They recognize what database instability looks like really quickly. They know how to sift through a bunch of graphs and build context.

28:08They can see a wall of logs and immediately recognize things. Companies are relying a little bit less on those going forward, especially in this Next.js space. people are growing up. There's just not a lot of need for that. And even at companies that are built on top of the cloud and built on top of a bunch of these other providers, you're moving less and less to where you have these highly trained folks that have been in this for a while that recognize these patterns really well as the folks who are on call. And then instead, you have just software engineers on every team, regardless of experience, taking this burden on.

28:42So you have new grads six months out of college, and they're being woken up, in the middle of the night to solve these really hairy issues. And, you know, they're just not like super, super confident about it. So one of the things that we're very prescient about is like, we're building for these developers, right? We want to be able to like, we're focusing specifically on their knees. How do we help build that confidence? How do we help build that context? And building a tool and a platform that does that. So in my experience, like when I was at my time at Google, when you had super junior people on call, you know, it's a learning experience for them, but it ends up actually being like creating more work essentially for everybody else because there's just not that much stuff that they can actually handle on their own.

29:26So they end up sort of almost like a proxy or relay to the more informed people on the team that need to actually jump in there and like actually solve their problem. Yeah. Yeah. So like, imagine if we could just make that better, right? Imagine if those engineers, again, like going back to like opening up the different tabs in your browser, right? They could just not be getting all of that context because they just don't have it ingrained that these are the different places to look and these are the different tools I need to look at. And instead, if we can just give you that context and immediately you know if it's a recent code change and you can just roll back really quickly, that's a huge win for most companies.

30:01And going back to the other big value of Beeps right now and what we're really providing is that one of the things that's very distinct in this space that we're building in is that a lot of times you're not thinking about servers anymore. You're not thinking about servers, but you're thinking about services. You're thinking about these different APIs. So, you know, you look at like, again, going back and using the example of like Clerk and Supabase for off and like Neon, PlanetScale and other places for your database and using tools like Resend for email. And, you know, a lot of these companies that are building on top of like LOMs, they're, you know, they're talking to Anthropic, they're talking to OpenAI, Grok, all these different providers.

30:40And one of the challenges they have is like when sometimes when your app breaks, you don't know if it's you or if it's them, right? Like you don't know if they're down or if they're actually having issues. So one of the big values that we provide is that we keep track of the different providers that you use. So anytime like your package JSON changes or some of the other signals that we look for in your application code base, we update the list of providers that you use. And we have a pretty comprehensive list of all the different tools that folks use in this space. and anytime they go down, we let you know.

31:11A lot of times, within a handful of seconds. So you immediately have that context, you know why they're down and can make decisions appropriately based on that. And it's just that it's a pretty simple thing, but it actually provides an absolute ton of value because again, it's not something you have to go look at. You don't have to keep track of the status of the 16 different providers that you use. And we do that for you and not only tell you when they're down, but when you're down, we'll actually tell you if it's you or it's them. It eliminates a big source of anxiety that engineers in the space have been dealing with for a while.

31:44We did a cheeky thing where we're like, you know, Versailles has been, they had like, are we turbo yet? And some of these things to track sort of like the status of their different like big initiatives. And I was sitting at Next.js Conf, like it was like last year and they were talking about are we turbo yet? And I went out and bought are we down yet? And we have a sort of like portal where you can see sort of like the last time an issue was reported on any of these big providers. It's pretty interesting to see who's down all the time and who's not. And it's kind of like an honest look at this industry and some of the tools that they use.

32:17And hopefully we can help promote better reliability in the space and help this ecosystem grow in a more solid way with a lot of this data. And we're going to be ramping that, revamping that a little bit more to maybe show some comparisons between different tools and things like that. but it ends up being a big value to our users. And one of the things we're really excited about is just knowing that simple information. In terms of the engineering of beeps, what's the stack look like? Yeah, so we're 100 % TypeScript. We want to feel the pain of our users and predominantly our web properties are all in X.js, right?

32:55So again, completely dogfooding the system that we're building for. And then we have a bunch of TypeScript services running on the backend using Node.js and basically running on a handful of platforms that make it really easy to run containers. Like Fly is one example there. And it's pretty simple stuff. But I think in order to really build and understand the challenges that our users have, it was important for us to build in this space as well. Obviously, when we're measuring the health of all these providers, we can't use all of them, right? Because we don't want, like Beeps needs to be more reliable than sort of like even the systems that we're monitoring or the systems that our users are managing.

33:38So it's a pretty fun challenge. But again, just like pure TypeScript at this point. In terms of deployment, then, this is run as essentially like SaaS? It is run as a SaaS, yeah. So we have our web app that is used for like setting up integrations, getting information about like the different apps and sort of like the history. we have a whole status page product where as part of like our free product, you get a status page where you can communicate status of your application to your users. Right. And we have a bunch of different themes that are sort of like fun for the sort of like community. And it's a good way of like expressing your brand through our themes.

34:16And yeah, again, those are all just kind of like, like web properties. Yeah. In terms of what you're doing with some of the LM work, like how do you manage orchestration? I think you mentioned that you have a number of agents that are going off and performing independent work. How does that workflow work? Is that something that you built in-house or are you using some sort of framework for that? Pretty much we've built that all in-house. A lot of the agentic frameworks out there, a lot of them are written in Python and just not matched up to our stack. Whether or not those agentic frameworks have a shared consciousness or each agent has its own consciousness was some of the things.

34:55So we went out and built our own internal package in TypeScript ourself that we're using. I think it might be something that we eventually open source because from a TypeScript perspective, there's just not a lot out there to be able to build these tools. And I think that it's something that I think that we could contribute back. We probably need to do a lot of cleanup and not be so beep specific. But I think there's a world where we hope to be able to share that out with the world and help that community sort of grow. So, yeah. Yeah, I've heard that is a consistent issue with a lot of the, not just even Asian frameworks, but essentially all the libraries that are available.

35:30If they're not Python, solely Python, they're Python heavy where their support for other languages is not as well documented and there's less of the community behind it. So it's kind of hard to invest company resources in a framework that maybe not be there in six months from now. yeah like but we found that like uh we originally built some of our own like our own sdk to interact with the llms to basically do sort of like failover and like racing between uh you know like anthropic or like or like open ai and so that we could be resilient to sort of sort of like any sort of failures there but the brucelle ai sdk is actually like pretty pretty solid when it comes to this stuff and we recently switched over to that to power sort of like our connectivity to like llms But, you know, it's starting to catch up a little bit.

36:18But yeah, from like an agentic framework sort of perspective, it's just not hasn't sort of like reached that level yet. Are you also using DZero from Brazil? I do a little bit of it. Yeah, we haven't used it for any of our like our web interfaces because they're a little complex. But I think like we probably could at some point. But our admin tools, like I'm not like, you know, I was telling you my background about like being infrastructure reliability. I am not a front-end engineer by any world, but I've been able to really build our admin interfaces really easily with vZero. Just very simple prompts, and I can get going and be dangerous enough to build something to be able to help us do better customer service and understand where our users are based on our data.

37:00And going through that process was pretty shocking what I was able to do in a very quick amount of time. So pretty excited about that continued evolution there. And some of the stuff that I've seen demos on where it's using like 3GS to do like, you know, like, you know, spinning world, things like that. It's just, it's kind of mind blowing, right? It's really interesting to see what the power is there. And I'm kind of curious to see how that continues to grow. And as a startup founder, where resources is always, you know, limited, using some of these tools, are you feeling that essentially that you're going to be able to go further with sort of, you know, less people because they're presumably operating a little bit more efficiently?

37:41Oh, 100%. I think that there is a huge sort of like, I don't know if it's like, step level or like exponential, but like, it's a, there's quite a big lift on what we're able to do. I was like, I'll give you an example. so you know we use like kind of off the shelf sort of like background like uh scheduling system through pgboss so like our database is photograph postgraph and instead of like going with like a high level sort of like scheduling system it's like let's use pgboss and kind of go down that path and we were investigating an issue with pgboss and seeing some jobs being backed up and you know they have a very particular schema and it was kind of you know being able to like write a query to like to look what we were looking for was going to take a long time one of one of my engineers came back with a query like really quickly and it had all these different sort of like sql functions that i'd never seen before i forget exactly what they were but it was like wait what is that like what does that thing do and we ran and it worked and it gave us the app what we wanted i like i asked i was like did you write that like did you know all this stuff he's like no i just like i think you popped it into like like cloud three five and it was like came back with like the right answer really quickly and this is probably something that would have taken i don't know if not like an hour or like multiple hours to kind of like really get right it was something that we had in like a matter of seconds that gave us the exact answers that wasn't like our anybody's like particular expertise in terms of like how to extract this information and just those little moments like that even beyond just sort of code generation and like the tab completion stuff, that sort of stuff ends up being, I don't know, things that really, really save you a ton of time and give you sort of like the superpowers that I think AI intends.

39:21And pretty cool to see those specific examples play out very regularly for us. Yeah. I mean, it adds up, right? Especially when you're dealing with maybe technology that you're not necessarily an expert in, or you're not working in a day to day. And yes, you could figure it out, but how many hours or days is it going to take to figure out versus being able to leverage some of these tools where you kind of know what it is that you want enough to vet the output, but you might not know all the like obscure, you know, functions that you've kind of been purged from your memory because it's been a while since you touched that thing.

39:53It's pretty wild. Like one way to think about this from like a, like a beeps perspective is, you know, if you think about it, I definitely see the world where you can build a lot leaner of an engineering team to solve like, you know, really valuable problems for users. Like you're even seeing it today, like some of these early stage startups that have, a ton of traction and build amazing products. And you look at LinkedIn to see how many engineers they have. And I'm like, what? It's like, what seems like a huge company and it's like 10 engineers, right? That's a pretty common thing that you see.

40:23The tricky thing though is you need a decent amount of engineers to have a pretty solid on-call rotation, right? You wouldn't want your engineers to be on-call every other week. So it's a limiting factor in what you're able to do there. And I think that a company like Deep's can help solve that, where instead of having these really rough on-call rotations where you don't want to have them, but every two months, you can tighten the loop a little bit by not making them so difficult, implement strategies that may be different than just the pure week-to-week sort of strategy, or eliminate a whole class of issues and maybe a human's not on-call for some things.

41:00But again, we're not there yet. But I think that like what we were talking about before, there's a lot to be improved upon, but there's a lot of potential there. And I think that it's one of those ways that on-call needs to catch up with the rest of how the software industry has evolved. In terms of the engineering, what has been one of the hardest things to build so far in order to bring this product to market? the easy answer there is like the engineering isn't the hardest part it's like the building the right thing right but that's that's the easiest answer but i think that you know we have unique challenges in that again we're on call for your own call so we have to be incredibly resilient and and have to build that like brand of trust and that can be eliminated in an instant right so having having all those sort of like best practices in terms of resiliency and understanding sort of like our own monitoring of our own system and how we handle those failures I think it's like, has, has easily been the hardest part.

41:54There's not like a specific thing, but it's, it's more so just making sure that like we're redundant across providers, redundant across the systems that we use. And thinking about that way from a, from the start is, is one of the challenges about building a system like this, right? There's, you know, it's easy, it's easy to like launch an MVP of a product that doesn't have those sorts of requirements. But if you're trying to go out for this market in this, in a space where that really matters, it's time consuming, it's challenging but it's a very important thing to do so yeah. Yeah I mean I don't think you can underestimate the engineering effort that goes into building like a really reliable tool and you can't be in the space of on-call and reliability where the product is not reliable itself right?

42:38Yeah exactly it just won't work right it's like it's one of those things like I don't know there's like the everybody thinks it's easier than it actually is right even myself right like every time I dreamt about building a product comparable to a pager duty, I used to always think, that's something I can build in a couple of weeks. I mean, it's a lot harder than that, especially to build that resiliency. Even just working at Airbnb for a long time, everybody was like, oh, Airbnb is the simplest product. What do all those engineers do? It's so much more complex than what you see on the front end.

43:13And so much goes into building a product like that. And then you take the consumer part out of it, and you're building this enterprise product that people need that are relying on. It's, yeah. Look at Twitter. Twitter's like the simplest product in the world. Like the first version of that was probably built in like a matter of days or something like that. But the, you know, if anybody remembers the fail whale era of Twitter, like they were falling over because they basically couldn't meet the scale and reliability challenges. And, you know, luckily they had a talented engineering team enough to like navigate their way out of that and build this, you know, resilient system.

43:46But like the product itself is simple. It's the backend infrastructure to meet the scale needs. That's really, really hard, especially back then when you couldn't just go and stand up a super elastic system on the public cloud. Yeah. I remember like, you know, 10 years ago when people would talk about how simple Twitter was, you kind of asked them like, okay, like, how do you handle the notifications though? And, you know, I guess back then it was like Justin Bieber, right? Like if Justin Bieber tweeted, like, how would you send all the notifications to all the people that like the tens of millions of people that follow him?

44:17Right. Like, and, you know, you could really see somebody's engineering chops by how they would think through that problem. Because that's that was a, you know, a serious thing to think about back then. So, yeah. Absolutely. Well, Joey, thank you so much for being here. This was really great. I appreciate it. Yeah. It's, you know, on call is such a like a personal thing to me that I really want to get right. Because, again, it's a whole different set of developers that are taking on the on-call challenges for these companies. And I would love for them to not have the experience that I had and the pain that I went through.

44:53And I'll share one last story that exemplifies this the most. And it's about 10 years ago now, October 2024. My wife and my then girlfriend and I were in San Francisco. We were at a bar in Pack Heights having a few drinks and decided to walk back to our place in Knopf Hill after having a fun little night. And we're in Alta Plaza in San Francisco, overlooking the Golden Gate Bridge, just kind of admiring the view. And all of a sudden, my phone bus is in my pocket. And instead of like, oh, crap, I was working at Flipboard at the time. I was like, oh, Flipboard, HBase, we're having some issues here.

45:35And I could see the sadness wash over her face because to that point, we'd been dating for a few years. And we'd had family in town. And I had to disappear to go solve issues. We'd been on vacation where I had to run back to the hotel room in a panic. I remember one time we were in LA in Manhattan Beach. And we were at a restaurant. And I literally had to leave her at the restaurant and go sit in the car to help revive a startup that doesn't even exist anymore. and deal with these sorts of issues. So I just kind of see the wave of sadness kind of wash over her face. Instead of reaching into my backpack to pull out my laptop, I pull out an engagement ring and propose to her.

46:19And it's a great origin story for beeps. But if you really think about it, it's pretty sad. I involved on-call in one of my life's biggest moments and used her sadness about all the times that we'd been impacted as a way to kind of like shock her with the proposal. And, you know, it kind of hits home with like how much this has been a big part of my life and how important it is for me to really make sure that we fix this and do this in the right way and make sure that, you know, 10 years from now, I don't have engineers telling me that they replicated that story and had relationships ruined because of on-call.

47:05So if I can just make this one incrementally a little bit better or inspire somebody else to kind of join us and help us achieve this sort of mission of making on-call better for this next wave of modern software developers, then I'll be really happy. Absolutely. That's awesome. And that's an awesome story. Great way to end the recording today. Thanks for being here. Thanks, Sean. Appreciate it. Thank you.

From the publisher

beeps is a startup focused on building an on-call platform for Next.js. The company is grounded in the key insight that Next.js has become a dominant framework for modern development. A key motivation in leveraging Next.js is to create a developer-first experience for on-call. Joey Parsons is the founder and CEO of beeps, and he

The post beeps and on-call for Next.js developers with Joey Parsons appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
beeps and on-call for Next.js developers with Joey ParsonsSoftware Engineering Daily · 48 min
Listen in VO