Complex Workload Deployment with Will Stewart

21 Aug 2025 · 39 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Complex Workload Deployment with Will Stewart

Overview

  • Podcast Title: Software Engineering Daily
  • Episode Title: Complex Workload Deployment with Will Stewart
  • Host: Sean Falconer
  • Guest: Will Stewart, Co-founder and CEO of Northflank
  • Description: The episode discusses the complexities of deploying and managing cloud workloads, touching on infrastructure, scaling, CI/CD, and Kubernetes. Will Stewart shares insights about the challenges faced and solutions offered by Northflank to streamline application deployment.

Key Points

Background on Will Stewart

  • Early Interests:
  • Interest in technology and game server hosting from a young age.
  • Experience with deploying game servers during school and university, leading to the idea of a full-time venture.
  • Startup Journey:
  • Attended an accelerator program in Europe called The Family, which helped kickstart his entrepreneurial journey.
  • Observed infrastructure challenges in a digital agency placement, which inspired the development of solutions at Northflank.

Challenges in Cloud Workload Deployment

  • Complexities:
  • Managing infrastructure, scaling, and database hosting.
  • Configuring and maintaining Kubernetes; ensuring smooth deployments and integrations.
  • Common Issues:
  • Long lead times (weeks) for setting up environments due to organizational, developer experience, and technical problems.
  • Teams are often stuck in outdated mindsets focusing on infrastructure rather than workloads.

Northflank’s Approach

  • Vision:
  • Aiming to provide a self-service developer experience in managing cloud workloads.
  • Building a high-level abstraction for infrastructure and Kubernetes.
  • Development Experience:
  • Northflank allows users to deploy complex workloads without needing deep Kubernetes knowledge, reducing the need to write extensive YAML configurations.
  • Offers features such as CI/CD, logging, metrics, and disaster recovery built directly into the platform.

Key Concepts Discussed

  • Developer Experience vs. Infrastructure Complexity:
  • Many teams focus excessively on tooling rather than product delivery, causing inefficiencies.
  • The phrase "customers don’t pay you to write YAML" highlights the need for a focus on business logic.
  • Graduation Problem:
  • The challenge of moving from PaaS solutions like Heroku to managing workloads independently as companies scale.
  • Northflank provides a solution to avoid this issue by allowing teams to run workloads in their own VPC with more control.
  • Control Plane and Runtime:
  • Control Plane: Abstracts infrastructure specifics and manages workloads.
  • Runtime: Runs user-defined containers and utilizes Kubernetes for orchestration.
  • Northflank separates these to provide flexibility and ease of use.

Future of DevOps and Infrastructure Management

  • Evolving Roles:
  • The role of DevOps is shifting towards providing better developer experiences rather than maintaining complex infrastructure.
  • Organizations are moving away from large DevOps teams toward solutions that enable faster development cycles.
  • Cloud as a Commodity:
  • As infrastructure management evolves, it is becoming increasingly commoditized, with developers focusing on the applications rather than the underlying infrastructure.

Conclusion and Next Steps

  • Future Developments at Northflank:
  • Building a self-deployable control plane for enterprise clients.
  • Enhancing GPU support and other enterprise features.
  • Expanding capabilities to cater to the growing demand for secure, flexible cloud solutions.

Key Takeaways

  • Deploying applications in cloud environments is complex but can be simplified through the right tools and platforms.
  • Northflank aims to bridge gaps in the developer experience by abstracting complexities and providing a self-service model for deploying and managing workloads.
  • The future of cloud infrastructure management is focused on flexibility, developer experience, and the separation of control and runtime to meet the needs of modern enterprises.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Deploying and managing cloud workloads is a complex task that requires developers to handle infrastructure, scaling, and database hosting. Configuring and maintaining Kubernetes, ensuring smooth deployments, and integrating various services efficiently is a common challenge. Will Stewart is the co-founder and CEO of Northlink, which is a platform focused on streamlining application deployment and management. In this episode, he joins the show to talk about the contemporary challenges and solutions around workload deployment. This episode is hosted by Sean Falconer. Check the show notes for more information on Sean's work and where to find him.

0:50will welcome to the show hi sean great to be here yeah thanks for doing this i'm looking forward to it you know i was digging into your background a little bit and it seems like you know outside of some time that you spent in university most of your career has been sort of like founding companies I'm kind of curious, what has been the interest in being a founder? And how did you go from where you grew up to essentially now founding and running a major infrastructure company? That's absolutely true. Most of my professional life have been a founder. Whilst at school, I was always trying to work on side projects, deploying game servers, trying to build a game server hosting platform whilst at university.

1:28And then we realized that myself and my co-founder, that potentially that was an opportunity for us to do full-time. We want to take on that opportunity. And during my university years, we learned how to program and iterating on things were being able to source like Rancher and Mesos. And we were able to really start pushing the boundaries with container deployment whilst at university as a side project. And then we started to think, did we want to get full-time jobs or could we build a company? Could we build a platform? And we were having a look, could we join an accelerator program? and there was an accelerator in Europe called The Family.

2:05And that really put us on a journey of building a startup and meeting the right people. So, transparently for me, I was in a small city in the north of England called Lincoln. There isn't much of a technology hub, especially in technology. It's the home for the Magna Carta and the birthplace of the tank in World War I, but it doesn't have sort of a core software engineering history. So, being able to join this accelerator and learn what building a startup was all about, It was great for myself and my co-founder, Fred, went to university in Switzerland and just always been interested in side projects and building.

2:38And then during university, I studied IT management for business, which gave me optionality in starting with a little bit of business, a little bit of computer science. And during my third year of university, I joined as a placement student at a digital agency in the UK called Clock. That was a great opportunity for me to see how a real business is run and how you deal with customers, how you build multiple projects concurrently, how you hire engineers and have engineers working on projects. So I was able to see how our business was run. And then during that, I also saw how tough infrastructure was, how if you've got lots of different customers, lots of different requirements, lots of staging environments and trying to get releases to production.

3:18I was seeing things break and thought, well, the project myself and my co-founder were building in our spare time for deploying game servers could also be used in parallel to deploy any software. We were thinking about Docker files and ultimately our microservice is very similar to a game server. And that's pretty much what we've had our heads in ever since. So pretty much that's how we got into it, almost by accident. What were some of those challenges you ran into, you know, working in those ID departments that you saw with people trying to do, essentially, run distributed systems and where you were able to identify the opportunity with some of the stuff that you were doing in gaming.

3:55So it also just stems from when we tried to deploy our own backend for this game server hosting platform. I always call that game servers are a gateway drug to cloud infrastructure and Kubernetes because I could spin up a Mesos cluster and deploy containers fairly simply, but it was challenging to the control plane running and then there was no CICD there was no developer experience there was no way to push to version control and have build and deploy straight to prod that was not a thing and Heroku had that Heroku had this great experience where you can build and deploy and at the time it wasn't necessary for complex applications it was more simplistic and in the role as a placement software engineer you're tasked with spinning up staging environments but there was a queue for staging or when you have a new customer you had to spin up brand new infrastructure it would take weeks to spin up and it was like well this could be done in minutes or hours maybe that's where sort of our expertise of trying to build this developer experience here could be quite helpful and interestingly enough the company I was on it my placement with actually leverages our platform now in production so I've managed to come full circle now learning the trade and then now trying to solve some of those problems for them directly but if we step back a little bit a lot of these platforms like heroku and cloud foundry have solved this developer experience and in today's world a lot of those learnings have been forgotten and teams are now having to build internal platforms from scratch and this is what i was seeing firsthand and we set out to try and solve that problem where you could get a great developer experience inside your own cloud account leveraging things like kubernetes and containers and the building foundational technology on top so that teams could self-serve and deploy complex workloads to production.

5:35You mentioned there that it would take potentially weeks to spin up these new environments. Why weeks? Is that an organizational problem, developer experience problem, or technology problem? Maybe a combination of the three. I think there's a mindset where people think in infrastructure and not workloads. I've got a provision EC2. I've got to have all my Terraform. It's got to go to this team. It's got to be approved. It's got to be tested. It's got to be manually provisioned. Maybe there was some bare metal. Maybe there was some public cloud. There's then different workflows for different customers.

6:07Different team uses a different cloud provider. If you think again about this, if you more think about what we're trying to do here, it's deploy a microservice in the database and a cron job. Most infrastructure stacks boil down to three types of workloads. And how do you provide a common primitive so that teams can say I need Postgres, I need a Java Spring Boot application, how do I deploy that consistently for all environments? And that's how we've been thinking about it at Northlank is this high level abstraction of infrastructure and Kubernetes. And now we're six years in and trying to find this right balance of control and ease of use is I think the core problem that internal platform teams and DevOps teams are trying to struggle with.

6:46How do we give the self-service developer experience whilst also allowing the deployment of complex workloads. And I think that fine balance is where a lot of teams are struggling today, whether that be on EKS, GKE, AKS, or even ECS and Cloud Run. There are a lot of cloud infrastructure is, there are so many tools now in the last 10 years, but ironically, it's if not more complex than it was 10 years ago to ship a container. And I think there's a core issue here of teams are focusing too much on building the factory to deliver a product than building the products themselves. And I think we need to change how teams are thinking about their infrastructure, that why are we investing so much on building a factory to essentially host our core product when building custom infrastructure is non-differentiated and we should be focusing more on building business logic.

7:33How do you think of that? What happened that led to us kind of being in this position? In some ways, it kind of reminds me of in the data engineering world with the modern data stack, they make for these pretty pictures. You've got all these logos and different complex pipelines. But the reality is no one wants to have to stitch together 15 different products to just move data from an S3 bucket down to whatever, Snowflake or something like that. So how have we gotten to this place where shipping a container, we just made it progressively more complicated? So I think it's a combination of reasons.

8:07One is deploying infrastructure is inherently complex. You've got to think about performance, security, disaster recovery. You want to give more controls to developers, but you don't want to give them too much controls. It's like a continuous, ever-moving whack-a-mole of requirements. And how teams have solved that is by stitching tools together, because ultimately that's where the market was going with open source technology. Kubernetes winning the sort of enterprise war. Kubernetes is now the default way enterprise deploys software. And it's very easy to start a Kubernetes cluster. It's very easy to run a helm in store, but it's very hard to then integrate all of these very easy things to do into one consistent unified platform that tens of thousands of engineers consume so as soon as a team needs the a to z of requirements of service mesh disaster recovery logging metrics you then start having a hodgepodge of 15 20 different open source tools that then start to break they don't glue well together and i think having a more high level approach where maybe some of these core technologies should be automated at a higher level so that every team isn't building their own data services.

9:13Every team isn't building their own CICD pipelines. Every team isn't building their own, what teams are now calling internal developer portals or platforms. For example, we can't even agree on what IDP means. Some teams think it's a portal. Some think it's a platform. And I think that sometimes teams are focusing too much on the tooling. And so ultimately the company has an objective, build business subject, sell to customer. We have this phrase, customers don't pay you to write YAML. And I think that's more true than ever with, you know, stretch budgets, stress teams, more requirements to deliver more with less.

9:46And a lot of people are spending a lot of time building the exact replication of building an internal platform across teams. Every single organization with more than 50 software engineers is building an internal platform. And what about, you know, you mentioned Heroku, there's other sort of past systems that have come about past later than Heroku. And Heroku, it was groundbreaking when it came out. It was such a transformation for what you could do. It's had its own stumbling blocks along the way in terms of not necessarily keeping up with everything that's happened on the infrastructure side.

10:18But there's also been other newer sort of takes on that idea. But a lot of people who start with those platforms run into the graduation problem where they reach a certain scale and essentially they need to then move the workloads or what they're doing on those platforms over to doing this using the services available in like aws or gcp or whatever and manage that themselves so how do you kind of avoid running into that graduation problem so fundamentally heroku was awesome it still to some degree is awesome still a very large acquisition obviously maybe stifled some growth there but they're still a very successful business a billion dollars in revenue.

10:56That is a wonderful business. But of course, there are restrictions. As soon as an organization tries to scale, they bring this in-house, as you just said. And the graduation problem is something that we've been thinking deeply at. North Flanken, it really stems at VPC. Any company at any serious scale wants to run the workloads in their own VPC, whether that be public or private cloud. And I think there's, at the other end of the spectrum, Cloud Foundry and Pivotal Cloud Foundry sort of found huge success. And then Kubernetes came along and sort of spoiled the story. And at one end, you have Heroku and the other end, you have Cloud Foundry.

11:31Northflank is trying to sort of solve this problem in the middle, which is teams need a self-service developer platform that if you squint, Cloud Foundry and Heroku look fairly similar. Engineering team, self-service deploys a container to production with CICD built-in logs metrics. It's a fairly simple concept, but very hard to implement well. And I think Kubernetes in Northflank's cases has allowed us to have this we leverage kubernetes as an operating system which means that we can run our platform in any cloud whether it be aws gcp azure oracle on-prem open shift on-prem rancher and we're trying to solve the consistency problem where you can take this runtime and run it anywhere and that's where we work with enterprises to run on their own hardware so that they don't have the graduation problem because they can have more flexibility kubernetes has provided us the ability to have more complex, stateful workloads, high availability data services, preview environments, and production release flows.

12:28All of this core technology now works consistently across all clouds. And that means that we've found at least that this negates the issue of the graduation problem because teams feel it's running in my infrastructure. I have more control. I have control over data residency. I can meet compliance requirements and I can fine tune my infrastructure exactly how I want. But I also have a great self-service developer experience. And finding that balance is what we're focused on. Can you walk me through that process? Like, what's it like as someone using NorthLang to actually be building on it? If I have to essentially host this myself, how do I get away from running into the problems, essentially?

13:07Like, I have to host this myself, and now I need to manage it. Exactly. So this easy journey, when I was deploying Mesos as a teenager, as you do, there was a product, T2IQ built Mesosphere, and the installer process was very manual. And that's always stuck with me. how do you install a control plane in any cloud in under 30 minutes and with north flank a user would be able to sign up they can then do a cross account link to their aws gcp azure and then we've separated control plane and runtime so north flank is able to spin up an eks cluster gke cluster fully managed inside your cloud account and then north flank's control plane then allows you to connect your github connect your bitbucket gitlab define your workloads you know you could have 100 of microservices.

13:50We have teams deploying a thousand microservices on NorthLank into their EKS clusters. So you define the workloads, you define the environments, and then a manifest is generated and applied to your EKS cluster. So within 30 minutes, you could have an EKS cluster created inside your AWS account and your workloads deployed to production without ever having to know what a Kubernetes cluster is or write a line of YAML. So the code deployment essentially is going to be based on sort of standard like i commit to github and then it's going to essentially take over from there based on something like github action or however i'm doing cicd exactly so you connect your version control we list the web hooks you push to your spring boot application your dot net your java your javascript no js any application that supports docker or build packs you push north flank does the cicd does the deploy does the release logs metrics disaster recovery and we can even do things like per PR backend previews.

14:47So if a team spins up a new pull request, we can then make a replication of production and run those on-spot instances to reduce costs, but also give every developer their own preview environment, which is one thing we hear time and time again across many engineering teams is they're still stuck with two static dev environments and they want every engineer to have a replication of production for preview and it spins down during the night to save on costs. But that's generally the experience. You push to get a preview environment, merge to staging, have UAT and some acceptance testing. When it passes that, you merge to main, deploys to production.

15:20Could this help me with local development as well, where I can essentially have some sort of environment that I'm running my code against that is closer to mimicking actual production? So we've seen a number of teams leverage NorthLang in different ways. We haven't necessarily designed a local first experience just because running Docker Compose locally is just quite a great experience so we haven't necessarily invested too much time there but we do have things like northlink forward so you can proxy remote containers running in cluster and have them running locally so if you need sort of to access staging securely or some database or some service running remotely you can do that i think one thing we're also building is bring your own kubernetes so in theory you could import a local cluster running on your machine and have northlink running which we do have some people doing but we haven't necessarily designed for a dev first we actually You call Northlank a post-commit platform.

16:11It's when you've pushed a version control, Northlank handles everything. We haven't necessarily designed for a post-commit experience. And then is both the control plane and the runtime are running within my cloud? So currently by default, Northlank has a traditional pass where you sign up, add a card and the control plane and runtime is in Northlank's infrastructure, which we're providing secure multi-tenancy. And then as organizations scale, they need the runtime in their cloud. So it's control plane in our infrastructure and then runtime in their infrastructure. And then we have even customers that are requiring the control plane also in their infrastructure.

16:45So that's something we're working on, which we call self-deployable control plane, where you get a managed experience of running the Northland control plane in your infrastructure. Because, transparently, the largest enterprises in the world, that's exactly what they need. And that's what we've been working on. Your team is generating more code than ever, but you're still stuck with rigid legacy tools, inflexible workflows manual updates and cycled communication are slowing you down just as your engineers juggle more pull requests and context switches than ever that's why there's monday.com's dev platform with fully customizable workflows you can ship faster no admin bottlenecks no clunky add-ons let your developers work in monday dev or right from their ide with ai-powered integrations that keep every task in context get full visibility into progress performance and risk all in real time, fully synced with GitHub and your entire ecosystem.

17:36And with business connectivity built in, Monday Dev keeps engineering priorities aligned with the impact that matters most. No more admin bottlenecks. Visit monday.com slash dev to learn more. That trend of essentially, you got like sort of a SaaS solution, the version where some portion of it's running within someone's own cloud environment via dedicated VPC, and then even a version where I can deploy the whole thing and host it within my cloud seems to be like the trend of where a lot of both data infrastructure and infrastructure as a service needs to go. Like if you look at my day jobs at Confluent, we have all those versions, especially with the acquisition of WarpStream.

18:16It's like, you want to bring your own cloud, you need to essentially have a solution to that because some customers want that. Do you see that as the future that everyone has to, you know, operating in a space has to kind of build against? 100%. This is the most essential thing that most SaaS organizations need to understand right now is enterprise can't buy your software if they can't deploy it themselves. And if you haven't separated control plane and runtime, you better start thinking about it. And what we've started to think about at Northflank is how do we sort of leverage this opportunity?

18:45And we have, we call this BIOC as a service is if every single engineering team has to build sort of a managed self-hosted control plane, that's a huge problem because realistically teams have got to build one product, they don't want to build two. So this is where we're helping some of our customers provide a sort of a one-click managed experience to run sort of Northland Bioc plus their software inside a customer's cloud account. And we're finding, you know, like-minded individuals that understand that enterprise customers want to run their software inside their own cloud. They've got resource commitments.

19:15They've made huge GCP and AWS resource commitments. They've made huge billion dollar investments in on-prem hardware. It's got to be used for something. And it's got to be used for running the SaaS application they really want to run. Are you talking about in this case where like one of your customers who's building some sort of SaaS application can leverage your technology to offer sort of bring your own cloud out of the box without having to think about building all that stuff themselves? Exactly. So it could be I've just built a new SaaS software and I want to have a multi-tenant offering. Then NorthLank has the ability for you to provide secure multi-tenancy just like our PAS is basically NorthLank as a service where we provide the secure multi-tenancy where you can run containers with great security on Kubernetes and inside a customer's cloud account.

19:59And then again, that SaaS provider will may need to run their software securely in their customers account. So essentially for North Flank, it's our customers, customers. And we're doing that through the North Flank API with BIOC as a service. Yeah. Can you break down the control plane? Like what is the control plane actually responsible for? So the control plane is almost, so a lot of teams think about infrastructure as code. We like to think of infrastructure as data. So ultimately, the job of the control plane is to provide essentially this spec DSL abstraction of cloud infrastructure. So instead of thinking of helm charts, NorthLank is thinking about JSON data structure.

20:34What's a workload? What's a database? What's a cron job? Teams are then defining those workloads in NorthLank. That's then being stored in our control plane. And then our control plane is listening and subscribing to all of the Kubernetes clusters. It's observing when things start failing, tries to auto heal, it's listening to Kubernetes events and logs and metrics. And then it's essentially acting as a controller. It's trying to apply and detect drift between the spec that's running in cluster and in our infrastructure. And then we're just applying new manifests as users make changes either via GitOps or through the UI or API.

21:05It's essentially our job to provide and generate the Kubernetes manifests and apply those to the cluster. What about the runtime? What is it doing? So the runtime is essentially just running containers, configuring service mesh, making sure that Cilium and the containers are all running healthy at runtime. So essentially, we don't actually have an agent running on cluster, we're doing everything through the Kubernetes API, which gives us so much freedom to essentially leverage Kubernetes on any cloud provider, because the cloud providers are incentivized to have a consistent experience. This is to become compliant, they have to provide a consistent Kubernetes experience for their customers.

21:40And that means that when we build a feature for AWS, it works on GCP. When we build feature for Rancher, it works on OpenShift. And then our job is through the Kubernetes API is provide that runtime. So the runtime ultimately is just the user workloads. How does someone who's doing infrastructure as code today using Terraform or whatever, play with using Northline? This is something that we're continuously battling. Our sort of stance is you don't need Terraform. Our job is to eliminate your Terraform. Let's think of a high level primitive there's no point thinking in lower level infrastructure primitives anymore because it's now a solved problem if you start thinking about your workloads the infrastructure as code turns from hcl to a json sort of way of thinking and maybe some of our customers would prefer us to provide a terraform plugin and that's something we have been looking at or even driving terraform through the north flank dsl but currently we're finding enough customers that want to just please help me i don't want to have github actions terraform helm charts kubernetes manifests and then there's some plumi people want simplicity and that's where we're finding where the pushback is you don't need your terraform what are the primitives that are defined within dsl so the primitives are are things that are like so bl cluster which defines your structure of your kubernetes cluster either in aws gcp azure things like deployment service which is an abstraction of a Kubernetes deployment, which define where you want your existing image.

23:07Let's say you use GitHub Actions or CircleCI and you use ECR or something like a GitHub container registry. You can then import your existing images, define how many replicas you want, the secrets you want to inject, if you want to enable auto scaling, if you need health checks, sort of defining those configuration options directly on the workload and things like networking, how many ports you want to expose. Do you want to make it publicly available on a network load balancer. All of these settings are being configured directly on a single object, where traditionally in Kubernetes, you'd have to define five to seven different manifests to even achieve that.

23:40A user is just defining a single configuration. For example, in 130 lines of JSON, I can deploy a highly available zonally redundant Mongo, Postgres, and Redis, and three microservices into any Kubernetes cluster in less than 130 lines of JSON. Whereas in Terraform and Helm charts, that would be one to three thousand lines and no one wants to go and write that in terms of like defining you know a database how does that work like if i want to use a specific type of database is that essentially restricted like the number of types of databases is that restricted by what's supported by north flank or does it matter so we have a number of options and we try and think that think about this very carefully because stateful workloads are critical to a business you can't lose the data and if there is data loss how do you recover it so when we think about stateful workloads we have currently we offer six managed stateful offerings which is postgres redis mongodb rabbit mq mysql and minio which are great those are the most popular sort of open source databases and that's what we find most common when we don't have an offering teams want to use atlas or teams want to use rds they can it's essentially a connection string it's a secret so if you want to leverage those offerings you leverage your north flank secret group to inject those securely at runtime.

24:59And how we build our stateful workloads is essentially it's an operator pattern in Kubernetes. And we have an operator that can apply these stateful workloads in cluster. So really, it's up to the customer. Do they want to run stateful databases in Kubernetes, which we think is great for cost efficiencies and simplicity? Some customers don't agree with that vision, and they want to use RDS and Cloud SQL. And if they do, that's okay, we don't care. And then also at the side, teams want to run other stable workloads in Kubernetes, and we're not able to provide those managed services yet. And it's on our roadmap, but then they can run Helm charts.

25:33They can leverage North Bank, what we call bring your own add-on, where they can bring their own sort of tidied up Helm charts that sort of provide data services in cluster. And they're free to go crazy with what data services they want to run. So in that case where you're actually extending this with your own home charts, is this kind of in some ways like to tie this to maybe, you know, programming language, you're providing like the abstract classes, these like primitives that I can operate, build this thing at that level. But if they need to, I can essentially extend those classes and do my own bespoke work to really, really tailor this based on what my needs are.

26:08That's exactly right. And one phrase that we've used repeatedly before is how do we find the right abstractions to Kubernetes? because ultimately it's won the enterprise war. It's going to be here for a long time. I know other startups in our space have not gone the Kubernetes through. And I think that's a mistake. I think for us to deliver the BYOC model and this vision of running in VPC and having this cloud abstraction, Kubernetes is the landing target for today. And in the future, it may not be, you know, there'll be another orchestrator come along in three, six years time. Northlank will be ready to transition that.

26:36The first version of Northlank was running on Mesos. And then we transitioned to Kubernetes when it came along. So our job is to provide the right abstractions at a workload level. And then we will be able to change in and out, you know, the ultimate underlying orchestrator when it comes along. And you're absolutely right that it's our job to enable you to deploy your complex workloads. And that means if we don't do something, someone's blocked from deploying that. So how do we unblock them? Well, through bringing your own add-on, bringing your own Helm charts, and then exposing as many of the features of Kubernetes offers directly in product without ironically realizing it's Kubernetes behind the scenes.

27:09And then if I'm running, you know, certain parts of my code, essentially a different endpoint as microservices. How do I map that code base essentially to the microservice deployment? So in Northlake, we have a couple of primitives called build and deploy services. So essentially, I think of a build service as a repo link. This is my repository. This is going to produce a Docker image, an OCI image. And then you have to build those relationships between how do I get my code from build to deploy? And in Northlake, that looks like a pipeline. And then these deployment targets all have networking configurations, DNS names, and then exposed by the service mesh.

27:46So when we start demoing Northlank to some of our prospects or customers, we show them a small demo of deploying Nginx because that's the most simple container to deploy. How would you get a container running Nginx on Northlank? And we start showing some of the configuration of this is where you enter the Docker Hub address. And then immediately we're doing a scan of the manifesto to take the port. is it publicly available? And then we pre-fill the networking configuration for port 80, and we expose that publicly, and we allow you to configure the resources. So in about 30 seconds, you can go from existing image to networking configuration and deployment running in cluster in under 30 seconds.

Read the full transcript

28:24APIs are the foundation of reliable AI, and reliable APIs start with Postman. Trusted by 98 % of the Fortune 500, Postman is the platform that helps over 40 million developers build and scale the APIs behind their most critical business workflows. With Postman, teams get centralized access to the latest LLMs and APIs, MCP support, and no-code workflows all in one platform. Quickly integrate critical tools and build multi-step agents without writing a single line of code. Start building smarter, more reliable agents today. Visit postman.com slash sed to learn more.

29:00Capital One's tech team isn't just talking about multi-agentic AI. They already deployed one. It's called Chat Concierge and is simplifying car shopping. Using self-reflection and layered reasoning with live API checks, it doesn't just help buyers find a car they love. It helps schedule a test drive, get pre-approved for financing, and estimate trade in value. Advanced, intuitive, and deployed, that's how they stack. That's technology at Capital One. So a lot of your inspiration and journey came from running game servers, but what is the difference between essentially running, deploying game servers and enterprise applications, enterprise software?

29:37So one is obviously revenue focused and deploying game servers for fun was more enjoyment. The difference is if our software goes down, our customers lose money. And that's something we try and ingrain in all of our engineers is that we're trying to build a critical path software. And if a game server goes down, some people can't play an online game. They can't go and compete in the evening. But ultimately, if a database goes down, there are serious ramifications. If a backup doesn't work and a customer loses their data, that's really important accounting information or critical business logic that has gone awry.

30:12And I think it's more about understanding the customers and what are their demands and requirements. Before we were deploying mesos for fun and now it's serious. And we've raised some money, we've got some customers. It's our goal to handle this with care because it is important. How does that change the way that you go about actually building the software? Do you have to essentially spin up new teams that have different responsibilities, have people focused on security, for example, and change the way that you actually test? I'd say that one interesting fact at Northlake is we haven't had to deprecate a single feature in six years.

30:48We've taken sort of maintaining our platform very seriously. It's that if you release a feature into the world in a cloud infrastructure product, you've got to do so with some credibility because a customer needs to know that it's not going to go away. And how we do that is just think very thoughtfully about how we integrate principles, how we integrate primitives into our code let's release a feature but it's got to be well integrated into the overall vision of the product and we have enterprise customers that sign up and go we need this feature to sign this contract and where we don't go sure we'll do it we go why why why and then it's how do we then take that those requirements and then build it consistently into the vision and then how do we build that quickly as an mvp to have this customer sign up and sign a contract then how do we build it for all of our customers so that we could expose it through feature flag to our enterprise customers.

31:39So I think it's quite common across most companies, but that's just how we think about it. So since you're dealing with sort of these enterprise applications, you're building this abstraction layer, you want to be careful about introducing new types of primitives. Does that limit how quick you can sort of adapt to new technology that's coming online? So yes and no. So for example, we haven't immediately, like for example, our deployment target is Kubernetes currently. So we didn't really jump into Wasm or serverless containers because we were trying to achieve a vision of serverless containers without Wasm.

32:14So we didn't necessarily jump into that. We could have integrated Cloudflare workers. We could have integrated Fastly's computer edge, which are great products and they're great developer experience. But we didn't need to go down that route because we're trying to provide this abstraction with this primitive wouldn't necessarily be required. But then on the other side, currently GPUs are hot. some of our customers started requesting hey it's great that i can run all of my cpu workloads but what about all my gpu workloads and we thought well in a cloud platform like north flank it should just be another primitive so then we work quickly to allow our customers to run gpu workloads in their cloud account and it's funny because a lot of companies are building these gpu automation platforms but they have no solution to microservices they have no solution to databases they have no solution to cron jobs and when an organization is thinking about deploying their infrastructure, they've got five types of applications now, with GPU now being an extra one.

33:09And that's where Northland can fit in quite nicely. It's because we have a consistent story for deploying all of your infrastructure in the same VPC. Yeah. And ultimately, even if you're building, you're doing inference, for example, to build AI applications, there's some application part of this that's not going to be running on GPUs. So if you're doing that with a specific, I guess, like GPU platform provider, then you're going to also be doing something different. So it kind of goes back to the same problem that we talked about at the beginning, where ultimately no one wants to have to stitch together like 15 different solutions to deploy their software.

33:41This is exactly right. And we just posted a case study with a company called Waits. They have two engineers working on this full-time. They come from Pinterest, highly skilled engineers. They could have done this themselves. They started deploying their CPU and GPU workloads with Norfolk. And now they're servicing millions of users. and there's just two of them i'm just completely in awe of them that they can scale such a awesome application deploy so much infrastructure there's just two of them and i compare with some teams where they've got 25 devops engineers and they handle about 100 rps i've started to ask some of that when i when i speak to some prospects i'm saying what's your rps and some of them say 26 i say how many engineers do you have and the answer can be can be 50 or 30 but they get like a They need two engineers for each other.

34:30Oh, yes. You know, kind of building on that, like with companies needing, you know, so many DevOps engineers for reaching sort of minimal, you know, scale. How do you see the role of DevOps evolving in the next few years, given if technology like Northflank and, you know, other players in the market essentially take off? So I think that Linux system administrator became DevOps, DevOps became platform engineer. And ultimately, the goal is to provide a consistent and stable way to deploy. Ultimately, their job is to provide developer experience to application engineers. And the concept of you build it, you run it, you know, was hot with Heroku and Cloud Foundry and in recent times has gone away.

35:13But I think with limited resources, organizations just don't have the resources to spin up huge, huge teams to do things that aren't directly moving the needle for the business. And DevOps is going to be seen as a cost center, not a relief. And teams are going to have to adapt to that because ultimately, if you're not driving the critical business logic of the business, there'll be questions there. So I think that platforms like Northlank can reinforce DevOps teams to provide more value to their business faster rather than spending three years building an internal developer platform, which makes no sense.

35:46It's too expensive, takes too long. And then every company that does that then has to re-platform every five years anyway. So all of that effort was almost for nothing. How do you see, I guess, like the future of cloud infrastructure management changing as well? So I think with platforms like Northlank, and there will be others, it sort of makes cloud infrastructure a commodity. So I think clearly, we all know that the three winners in public cloud spending, but there's also a lot of private cloud spending as well. Like there's still many hundreds of billions of dollars being invested in private cloud.

36:17I think that the cost of cloud infrastructure has stopped coming down. It's actually now in theory, you know, probably going up. And we've been used to cloud infrastructure becoming cheaper. And that's not that's not happening anymore. with gpus you know you spend 66 000 for eight h100s on gcp on demand at the moment that's completely insane and with platforms like north flank infrastructure becomes more like a commodity i think that that starts to drive the cost down i think that the cloud provider becomes less relevant because ultimately they're providing you know it's like a water company you you don't really it doesn't really matter who provides you your your water or electricity or gas you you just get some hardware and then it's what you run on top of that that matters and that's where i think the interface that developers consume cloud infrastructure will be more important so what's next for north flank currently we're building sort of self-deployable control plane automation where an enterprise can run north flank on the control plane themselves that's really quite an essential part of our vision going forward self-service gpus we have a great bringer in cloud offering for gpus but many of our customers are demanding you know access to spot gpus in seconds something we're working on very diligently at the moment and then just providing more more enterprise features we we come across organizations using cloud foundry still day in day out you know huge deployments of millions of containers and they're stuck because they then have to go and build from scratch you know the the wonders of cloud foundry or turn to app platform or some OpenShift homegrown solution.

37:50And our opportunity is to try and capture this huge amount of demand and provide an out-of-the-box solution that hopefully you get through some procurement, difficult legal challenges at many enterprises, but that's what we're trying to solve this year. Does the bring-your-own-cloud form of deployment help with some of those procurement challenges? I'd say that for startups, 100%. but then many startups don't have the same legal challenges that some of these enterprises have i think the self-deployable control plane is is the biggest unlock when you go to kubecon or any event almost every developer in the room is talking about air gapped i think we're in startups in in sf and sas startups no one's ever talking about air gapped but when you go boots on the ground and listening to organizations needs to consume all this software air gapped air gapped air gapped air gapped.

38:43So I think there's a disconnect in what people are needing in enterprise and what people are delivering. So I think that's quite important. But the bring your own cloud offering is a huge unlock for North Lank and startups. The ability to self-serve and deploy to production in 30 minutes is huge for us. So investing in our enterprise story and then also investing in how we unlock startups to deploy to production with less resources. Awesome. Well, Will, thanks so much for being here. Thanks so much and have a great day. Yeah, cheers. Bye. Thank you.

From the publisher

Deploying and managing cloud workloads is a complex task that requires developers to handle infrastructure, scaling, CI/CD pipelines, and database hosting. Configuring and maintaining Kubernetes, ensuring smooth deployments, and integrating various services efficiently is a common challenge. Will Stewart is the co-founder and CEO of Northflank, which is a platform focused on streamlining application deployment

The post Complex Workload Deployment with Will Stewart appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
Complex Workload Deployment with Will StewartSoftware Engineering Daily · 39 min
Listen in VO