In short
Software Engineering Daily: Streamlining Cloud Infrastructure Deployments with Jake Cooper
Episode Overview In this episode of Software Engineering Daily, Sean Falconer interviews Jake Cooper, the Founder and CEO of Railway, a platform designed to streamline the deployment and management of applications in the cloud. The discussion covers the challenges of cloud deployment, the unique attributes of the Railway platform, and insights into the future of cloud infrastructure.
Key Points
- Introduction to Railway
- What is Railway?
- A platform for deploying and managing applications in the cloud.
- Automates infrastructure provisioning, scaling, and deployment.
- Known for a developer-friendly interface.
- Background of Jake Cooper
- Journey to the Bay Area
- Moved from the University of Victoria in Canada to the US, driven by ambition and opportunities in tech.
- Experience living in various cities, including Amsterdam and New York, before settling in San Francisco.
- Motivation Behind Railway
- Need for Simplification
- Frustration with the complexity of deployment processes.
- Development experience often joyful, but deployment felt cumbersome.
- Noted that many workflows in deployment are similar across different teams and applications.
- Challenges of Cloud Infrastructure
- Pain Points
- Transitioning from local development to cloud deployment can be painful.
- Issues like managing multiple environments, deployments, and database integrations create a 'trough of sorrow' for developers.
- Cloud Trade-offs
- While public cloud simplifies server provisioning, it adds complexity in managing cloud abstractions.
- Railway's Unique Approach
- Comparative Advantages
- Intuitive UI that minimizes complexity and enables easy resource allocation (e.g., databases, deployments).
- Automated Docker image generation, reducing the need for manual configuration.
- Offers a pay-per-use model that adjusts based on actual resource consumption.
- Understanding Docker Challenges
- Complexities in Docker
- Users often face challenges in building images, including permission issues and cache invalidation.
- Railway aims to abstract these complexities, enabling seamless deployments without deep Docker expertise.
- Avoiding the Graduation Problem
- Sustaining Growth
- Unlike some platforms that necessitate migration upon scaling, Railway is designed to accommodate growth without outgrowing the platform.
- Infrastructure allows for self-hosted services and customizable setups which prevent vendor lock-in.
- Infrastructure as Legos
- Concept Explanation
- The idea of modular infrastructure where components can be easily integrated, akin to building blocks.
- Promotes reusability and ease of composition among various services and applications.
- Security and Access Control
- Zero Trust Approach
- Railway provides a secure environment by ensuring services talk to each other without exposing endpoints publicly.
- Built-in security measures to prevent unauthorized access and leaks.
- User Experience with Railway
- Getting Started
- Users can quickly deploy applications by linking their GitHub repositories.
- Background processes handle resource provisioning and setup automatically.
- Challenges and Constraints
- Current Limitations
- Building trust with potential clients in the competitive cloud landscape.
- Moving beyond initial resistance from companies hesitant to shift from established providers.
- Community and Open Source
- Open Source Strategy
- Significant portions of Railway's stack are open source to foster community trust and contributions.
- Encourages extensive community involvement, aiding in the platform's growth.
- Future Aspirations
- Vision for Railway
- Plans to build on their orchestration engine to improve developer productivity and simplify deployment workflows.
- Commitment to continuous improvement and reliability.
Conclusion Jake Cooper’s insights emphasize the need for simplicity in cloud deployments and the importance of creating trusted tools for developers. Railway seeks to minimize the complexity encountered during deployment while providing users with significant flexibility and control over their applications. The episode underscores the future of cloud infrastructure as a collaborative and modular space, fostering innovation and developer efficiency.
---
For more details, visit the [Software Engineering Daily episode page](https://softwareengineeringdaily.com/2025/07/22/streamlining-cloud-infrastructure-deployments-with-jake-cooper/).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Railway is a software company that provides a popular platform for deploying and managing applications in the cloud. It automates tasks such as infrastructure provisioning, scaling, and deployment, and is particularly known for having a developer-friendly interface. Jake Cooper is the founder and CEO at Railway. He joins the show to talk about the company and its platform. This episode is hosted by Sean Falconer. Check the show notes for more information on Sean's work and where to find him.
0:40Jake, welcome to the show. Great to be here. You know, I'm super excited to chat about a bunch of stuff. I know we've got a couple of things on the docket. So, yeah. Yeah, absolutely. We were chatting before we hit the chord here, but you and I both graduated from the University of Victoria in British Columbia in Canada. So what ended up sort of pulling you south to the Bay Area? Yeah, I think there's almost like this brain drain kind of pull that I think we're both talking about it right before. And I think like half of maybe my graduating class was just kind of like, yeah, I want to generally kind of move down there in general.
1:09So I think I kind of always knew that I wanted to be in the U.S. I think I always knew that I wanted to start a company at some point. So it was a more of a matter of like when, not if, if that makes sense. And so I ended up moving to a few different places before I ended up working out a bit out of like Amsterdam and Italy in between grad and then moved on to New York and then moved to San Francisco. But yeah, I think it's a pretty no brainer in my mind, because you pay the same amount of tax dollars, and you get, you know, twice x the ambition plus twice x the sun, you know, so there's a pretty strong payoff on doing that, you know?
1:38Yeah, as much as I love Canada, I think if you're in tech, there is such a strong pull to the Bay Area, especially when I moved here now like 15 years ago. I think we'll get into it today, but you're running a fully remote company. So I think there's somewhat less constraints on companies and the opportunities people have in tech today all over the world. But that wasn't always the case. But back to what you're doing now around railway, what was sort of the driving factor behind the creation of it? It's a platform focused on streamlining deployment and management of infrastructure dependencies.
2:11Yeah. So, I mean, I grew up like hacking on random stuff. I started actually writing like in my first computer science stuff was writing kind of aim bots and cheats for like video games. And so it was always like a very kind of exploratory, creative kind of thing like that. And so I ended up writing like small programs or anything else like that. And then every single time I kind of like moved to like deploy something, it was just like you're switching from this world of, oh, cool, this joy. You know, it's this like beautiful, happy kind of like thing where you're hacking around. It's nice. And you're like, how do I like move this thing?
2:38Right. And obviously there were like tools like Heroku at the time. And so like finding those tools was like awesome and magical. But there's a whole like class of problems that exist actually outside of like actually getting something deployed that we call like deployment lifecycle. Right. And so it's how do you make changes? How do you like go and get them reviewed? How do you go and add a database in another environment? And then how do you make sure that you're going to actually have that database when you like go in and merge? Right. This whole kind of split universe phenomenon of staging, et cetera, right?
3:03And so when you get into things like that, you end up having to wrangle a lot of stuff, right? And so it goes kind of back to this bit of trough of sorrow, having to go and figure out all of these things, right? And in reality, a lot of the workflows that people end up having are very, very similar, right? And so if you end up building a lot of those workflows, then most people, they want to split into a parallel environment. They want to test their stuff. They want to merge it. And they want it to automatically roll out, right? So they just simply don't want to handle that. So it ends up being that like this class of problems ends up being both interesting from like a systems perspective.
3:31So, you know, we're working some really, really cool like networking, storage, et cetera, all of those other things. We've got our own bare metal servers now, but also just like very, very applicable to the wide swath of people. Right. And I also think that like the compute market is one of those things that will just continue to grow. Right. It's just like we will need more computers. Right. And so anything that we can do to kind of streamline people's like productivity in there, it's like it's one of the highest leverage things that we have in general. Right. And we talk a lot about like leverage, like how do you build leverage?
3:58How do you build efficiency? Right. How do you make it so that a user's action has an outsized return on what they're putting in? Right. Cause that's for us, that's the definition of magic, right? It's like you do a little, you get a lot. Do you think the public cloud has increased that pain in terms of, you know, we have this joy of like building something and you get sort of that aha moment of, you know, building something on maybe on your local machine and running it. And then you got to get to a place where like now I have to deploy it. Does all the sort of like boxes that you have available in the cloud make that even more of a challenge?
4:29I think you're trading pain, if that makes sense. Right. And so like, it's obviously very, very painful to like go and procure servers, get them up, making sure that they don't fall out of the wall and that, you know, somebody doesn't bump the power cable and all that like class of problems of solving those. Right. And then you kind of trade those for like, okay, how do I like manage these machines? Right. And then you kind of have to play with the abstractions that the cloud providers give you. Right. And there's benefits and curses to that, right? In the prior kind of like bare metal world, it's like you want to spin up a server, you want to spin up something.
4:56Well, it's like you're measuring your time on the order of weeks or maybe even months, right? To like get this thing like up and running, right? Versus you go to the cloud and it's like, boop, hit a button and it's kind of just like up and running, right? But you do have to kind of like manage the primitives that the cloud providers give you and kind of like work within those walls in general, right? And those primitives can be, I think, made faster in and of themselves, right? So from making deployment instant, right? moving your builds like immediately beside your compute, moving your storage, keeping all of these things together, right?
5:22That's kind of the whole end goal of like a lot of the stuff that we're building is like all of these things have almost been verticalized, right? If you look at the AWS kind of dashboard, it's like every service, you know, famously Jeff Bezos is kind of like everybody's going to interact with these things over an API. So everything's kind of like vertically sliced, right? So there's no mechanism to kind of like share these things together unless you really want to go and start composing. And then you have to, again, you start bumping into the abstractions there in general, right? So I think it's like you're trading kind of like paint over time.
5:49Okay. And, you know, there's been a number of companies that have tried to simplify this process. You mentioned Heroku. There's other companies in the space like, you know, Render, your Netlify, these various past platforms. Like what is Railway's sort of unique approach to this that distinguishes it from some of the other players? Yeah. So I'd say there's a couple of things. We talk about wrangling complexity a lot in general. So we've built a really intuitive UI that is kind of like layered. It's a canvas. you essentially just go to it and you kind of just like spew out, hey, give me Postgres.
6:18Hey, give me, you know, Redis, give me a deployment of GitHub, right? And I think we've gone above and beyond on a lot of different things. We've built a system for kind of like automatically generating Docker images, right? So essentially, you give us your code, we will go and statically analyze it, we will go and figure it out and stuff like that. So you don't have to actually like write anything to get started in general. Additionally, we've built our own kind of like storage system. So So you can host literally anything on Railway, right? So you can build a ClickHouse database beside your Python instance, beside your whatever, right?
6:48Like for us, it doesn't really matter. And we've like built those primitives in such a way where they feel very, very quick. And they're really, really easy for you to kind of compose together, right? And so I would say that what separates Railway versus like something like Fly or Render or anything else like that on the surface is they may kind of like look very, very similar. But over time, as you kind of like compose these things together, we've tried to almost linearize the complexity of each of these things versus allowing that kind of complexity to spawn out until, oh, you have so many of these microservices and no tools to manage them and stuff like that.
7:16So you mentioned Docker there and being able to abstract away the challenge of putting together a Docker image. What is that challenge that people typically run into? Yeah. So I would say that there's a couple of challenges. Oh, and sorry, another thing to mention is like Railway is only, you only pay for what you're using because we've built our own orchestration engine. So we will go in place workloads. So normally if you're on a cloud provider, you know, you pay for a four gig box. And if you don't use the four gig box, you're billed for that at the end of the month. What we do is we basically allow you to run the code and we pack all of these instances together.
7:44And as instances are scaling, we'll go and move them around in general. Right. So that allows us to get an edge there in general, in terms of like both pricing, as well as like do us doing our own bare metal instances, which is both pricing and performance. But in terms of what makes the Docker process a bit more complicated, it's just another abstraction. You have to wrangle in general, right? You have to like go and figure out where to run place these binaries. What are the permissions? Like which ordering, right? Right. All of these other things. Right. And a lot of people don't know. It's like, you know, Docker is layered.
8:09Right. So it's if you invalidate the cache at any stage upward, it'll invalidate the cache all the way down. Right. And so even constructing the ordering of your like Docker commands. Right. Has an effect on the build times, the image output, all of these other things. Right. So there's kind of like there's a science to it, obviously, which is, I think, on the surface. But there's almost like an art to it where it's like, oh, you have to almost, again, understand the underlying abstraction. And our hope is that we can basically just say like, yeah, you really just you want to make sure that you have node in here.
8:37And then you want to make sure that you also have Python in here. And you basically like you should select almost the packages that you want if you ever use like something like Nineite or something. And then you have access to them, right? There's no messing with permissions. There's no messing with anything else like that. And that's what we built the Nixpacks automated kind of construction engine on. Some of the challenges I think organizations run into sometimes when they invest in a past platform is that if they are successful, they reach this graduation problem where they hit the scale limits of that platform, and then they need to essentially migrate off of it, go to AWS, Google Cloud, or whoever directly, and stand up a bunch of the infrastructure and run it themselves.
9:12How do you avoid that with Railway? Heroku famously had this graduation problem. It's one of the main things that we chat with investors about. And so the interesting thing about Heroku is they built, like the thing that was really, really great for them is also, in my theory, the thing that killed them. Not like killed them, they're doing a billion dollars in revenue and they accepted successfully and all of these other things, right? But like, I think as Heroku goes, like we all know that there is a massive, massive, massive business to be made in there in terms of impact, right? And they were only just like scratching the surface, right?
9:41So anyways, back to the original point about the graduation problem, I think the main thing that happens is you end up having kind of almost this like outsourced state problem. So that marketplace where it's like, oh, I need Postgres and I need Redis and I need to be able to deploy my, you know, call it like Ruby on Rails, like API server, as well as my workers, right? Heroku is really, really great for the servers, right? You just, you spend the state list of things, right? And they'll like scale up or down, excellent. And then you end up going in and integrating with something like an external Postgres provider or they provided one at one point or anything else like that, right?
10:11But it was a very, very bespoke offering, right? So they haven't like solved the generalizable storage problem, right? And I think there's a few more primitives that have come out, namely like eBPF, IOU ring, a bunch of those other things that allows us to kind of solve these things at a more generalizable level on the storage stack, which means you can spin up anything, right? And so instead of bumping into those edges where it's like, oh, I want this specific thing, but Heroku doesn't offer it, right? So I couldn't do like self-hosted click house on Heroku, right? Because there's no way to like do in the Kubernetes world of things like a persistent volume claim, right?
10:42There's no, there's no elastic block storage. There's none of those things, right? And so you end up bumping into these limits of that platform. And we've kind of invested right out of the gate. One of the first things, like Railway didn't even host code at the start. We were just a database provider. We were an underlying database where like you can click and you can get, it was one database at the start and then it was four databases, right? So that's where we've kind of invested in making sure that people can literally just do anything on the platform, right? And then we're making it really, really trivial for them to go and actually go and do that anything.
11:09I mean, can you explain this idea of like infrastructure as Legos? Yeah. So it's interesting. Like, I think if you squint, it kind of already exists in terms of like infrastructure as code. Right. And so if you look at like a Docker compose or a Helm chart or something like that, right, those are essentially like infrastructure as Legos. They have like a variety of environment variables that you have to provide. They have a variety of like inputs and outputs in terms of endpoints that exist. And then they have a variety of like services in there. And they also have versioning that exists over time.
11:38Right. And so when you drop, if you just assume that this thing kind of exists, it's this bucket, and this Docker Compose file has maybe four services, the aforementioned Ruby on Rails service, the worker, the Postgres, etc., that's now kind of like a Lego that you can use and you can piece together. Because that API endpoint actually has an input, and there are environment variables that you can pull from, and there are environment variables that you can provide to. Right. So if you consider that as kind of like a Lego block, then actually you can basically say, hey, I want to go in and import that thing and I want to use it as part of my project.
12:08Right. And so like we allow people to one click deploy things like Strappy or Aki or any of like the analytics toolkits that are open source. Right. Like we're big proponents of open source. We have an open source kickback program where if you build a template, people run the template, you get paid for what people are actually using. So that's kind of the infrastructure is Legos piece of it, where you basically you take that Lego and you drop it inside of your canvas. right and then you can kind of like consume or interact with it right and this ends up solving a like very very interesting class of problems it ends up solving like authentication authorization it ends up solving sharding it ends up solving security because you're not managing this like massive multi-tenant thing it solves like api versioning which is super interesting right so if you go and push changes you can actually go and roll out those changes and we have health checks for your services so if any of those changes were to actually cause those health checks to fail the rollout would fail in general right and so you can actually almost like split these things up over your canvas and consume.
13:01And so that's the Lego aspect of it in general. Can you explain a little bit more detail? Like how does this help solve something like auth? So then you're kind of like not talking with the public internet, right? And so we built this IPv6 wire guard mesh on top of all of our services, right? And so essentially, you're not kind of either exposing your instance publicly, so you can just talk with it externally. And you're also not at risk where somebody says, Oh, you know, potentially you've leaked your keys. And now there's a publicly accessible endpoint, right? Like the best level of security that you could possibly have is you just can't get to it, right?
13:32Without like SSO, right? Like, and so that's like the default level of security that like we're trying to provide here. And I think it's like inspired from the like kind of like zero trust mantra almost, right? Of saying, hey, let's give people the best experience and the best practices right out of the box, right? And we'll make sure that your database is within like single digits, ideally even like, you know, hundreds of, you know, microseconds from your instance. we're going to make sure that it's not accessible we're going to make sure that you have really solid primitives to like go and access these things so we do automated service discovery based on your name so if you have like your analytics service that you've deployed internally right it's just analytics.railway.internal you just make requests to it and then you're the only one that can actually access that within that environment right so that's kind of how it solves like authentication authorization because you don't end up needing it right so you don't end up having to put like a nginx server with basic auth or anything else like that in front of all your things, shoveling it into one pass and then saying, hey, everybody on the company, you know, network, go and do these things.
14:28Right. And then invariably at some point that ends up getting breached and then you having, you know, security posture on those. Right. Yeah. Yeah. So it's kind of like security by default approach. This is a zero cost model, essentially, like let's take the best practices, bake it in. So it's like the guardrails are essentially in place and that's how people will develop against it. Yeah, exactly. So can you walk me through, like, if I'm going to use Railway, like what is that process like? And then can you explain sort of like what's happening behind the scenes? Yeah. So this is funny. It rhymes with what happens when you type something into the browser question.
14:57That's like a very common technical question. Yeah. Yeah. Yeah. So when you go to Railway, so if you go to like dev.new, we will drop you on a page that allows you to basically say, give me a Postgres instance or deploy my GitHub. Right. And so if you hit a Postgres instance, what we do is we like go and we have a fleet of servers that exist all across the world. We will go and make a claim for that volume, create it, and then go and bind that instance there. And then we'll go return to that running instance, which you can access over the either private network, or you can click generate public URL, and it will generate a public URL for you.
15:30So that's the kind of like state full storage one. And if you go and do the GitHub one, what we do is we basically will parse your repository, figure out what applications you might have in there. Maybe you have a Docker file. Maybe you don't have a Docker file. If you have a Docker file, we'll obviously just use it. if you don't have a Docker file, we kind of go down this like tree of decision making where we say, do you have a package JSON? Right? Because if you have a package JSON, it's very, very likely that you have a node application in here. So let's pull in some of that information. Oh, do you have like, you know, or requirements.txt?
15:57Okay, cool. It's obviously Python and stuff like that. So we have this kind of tool that we've built. Again, this is the NixPax engine. It's all open source. It's on our GitHub repository if you want to have a look at it. It's super cool. It's like this Rust engine that we built, but it will go and essentially figure out what is your build command? what is your start command? What are all of these other things? And you can modify them after, right? But the whole goal here is to almost like compress all that knowledge so that the user, when they go to that platform, they just say, here's my GitHub repository.
16:22And we say, excellent, it's already deployed, right? And then you say, wait, what do you mean? Right? Because that's not like the default experience of going to a cloud provider, you have to fill out reams of forms and, you know, select a region, right? In the AWS dashboard, you go, oh, the name, that's pretty good, right? And then you're like, oh, which flavor of Linux do I want? And you're like, oh, okay, I guess I have to pick that now, right? And then you go through this like reams of stuff, right? And so our whole goal is to take all of that config and push it post haste. Anything that you can do later, you should want to go and do later, right?
16:50So even something like setting like a region for like a database, right? We built a system that allows us to like move these volumes around, right? And so if you spin up something in the region that's probably closest to you, which is a good default, right? Like if I'm spinning up from San Francisco, I'm probably in US West. If somebody's spinning up from London, they're probably in the Amsterdam servers, right? So we pick that. And then if you really want to go and move it, you just move it later, right? And so we take all the config and we push it later. And then where's this all running? Is this running in my account?
17:16Or are you sort of like running this within behind the scenes, like railways access to a public cloud? Yeah, so we're running on a few different servers at this point in time. So we have some straddling between Google Cloud and AWS, as well as our own bare metal service that we've spun up over the last year, basically. So it runs in a variety of the, you know, servers that we have. And ultimately, like, the only time you should care is if you're trying to potentially pair it with something externally. So let's say you have, I don't know, like a super base instance somewhere, and you want it as close as possible, because like, ultimately, the only thing that matters is that your computer is beside your storage, just from a latency perspective, because the database calls are so quick.
17:51So then you can potentially go in and select that region and say, like, oh, I actually want to run it on these specific class of instances, right? And, you know, barring any sort of failover, we will go and do that because we built the orchestration engine to go and manage and drop the instances right beside it. So the short answer is it could be running anywhere, and ultimately, you shouldn't really care, except for those latency reasons, at which point you can add that constraint later. How does Railway handle the kind of distributed system dependencies that you can run into? If you have a bunch of services, this can get pretty complex, errors can occur.
18:27How do you manage that aspect? Yeah, so I think in traditional systems, I guess it really depends on the class of error, right? Because there's a variety of different errors that can occur. There's like, I pushed bad code and I flunked the instance and it got past all of the health checks, right? And so we give you a one-click automated rollback, like we'll keep the container around for a little bit. You click that and you say, hey, listen, my health checks didn't catch that. Something is down now, a user's reporting something. You click that, immediately you're back, right? We have automated health checks for going and managing these instances.
18:56So if you define a health check, you just say slash health, and then you go into return, and it's like, oh, okay, I updated my Redis library, and it no longer is able to communicate with Redis for some reason. Okay, cool. Now that fails. And so we'll just actually flunk that deploy and then notify you whether you have emails turned on, whether you're through the in-app inbox. At some point in the future, we'll do an app. So basically get really, really close to notifying you as quickly as possible. And then there's also things that we've done on top of it in terms of solving classes of problems that are kind of interesting and only happen scale of complexity.
19:29So assuming you have a ton of different microservices, let's say that you've modified your gRPC or something like that. I think gRPC is a bad example because it's backwards compatible. But let's say you've modified something and create a breaking change, like you've modified a field on a GraphQL endpoint. You go and roll it out, things start flunking. So what we'll do is we'll actually do a dependent rollout. So if you are saying, hey, my front end communicates with my back end and then my back end starts failing when it rolls out, we're going to flunk the whole kind of like class of that deployment, right?
19:57Because you've made a change to your front end and your back end, right? So we've done a bunch of things in the application that basically say just consume the criteria or config from various different services. And we'll almost construct for you like a dependency graph to like go and automate any of these things. You don't have to do the like Bazel thing where you're defining your dependencies as this and then you forget and then something happens. Or Turbo does this, I think, as well in terms of like build stuff. You don't have to do any of that. You just consume the properties of the services and then we will go and automatically figure that out, including cycle detection, which is super cool.
Read the full transcript
20:27What about like a canary rollout? Yeah, so we can do canary rollout. So you just you can just do it on the command line if you just do railway up. So as an also kind of like a piece of a fill in, we have the obviously the canvas dashboard or anything else like that. But we also have a command line because people like to interact with services in various different ways, right? So the command line is really useful for a bunch of different things, including that. Okay. And then how do I, I think one of the challenges companies typically have with whatever sort of cloud reasons they're using is they'll sometimes have essentially provision too much or potentially too little, which will lead to problems.
21:00Like they're spending too much or maybe they provision too little and then they run into like challenges with like throughput or latency or something like that. How do you solve for that? I think that's a really important thing to solve for. I think it's also a thing that is a core differentiator for us versus other platforms in the sense that the aforementioned orchestration engine that we've built allows us to basically only bill for the what you're using perspective. And so in a production workload, that's pretty cool because assume that you have 2x standard deviation or 10x because you're on the front page of Hacker News or anything else like that.
21:31That's fine. We can go and handle that. We can go and scale up the instances. We can scale them down. We can go and do that for you. Right. The part where it becomes actually like, I think even a little bit more interesting is when you end up with pull request environments that can be served like serverless. Right. And I use the server, the word serverless in like quotes. I know we're not on like a video or we're not going to share the media or whatever, but in a way that basically allows you to send a request to it, the request will spin up the container, the request will be filled and then the container will be finished.
21:59Right. And so we can actually go in and kind of construct that parallel environment of yours with a copy and write database volume. pretty soon, which is super cool. We'll roll that out in like Q1 to 2025, as well as like serverless, serverless spin up of those parallel environments, right? So they instead of spinning up something that's like, Oh, you know, I need 32 gigs for this service, and I need four gigs for all this other service, and you spin them up, and you leave them around over the weekend, and you just incinerate money for like these things that are idle, we will actually only spin them up for the time that you need and only charge you for the kind of usage that you have, right?
22:30And so ultimately, that means that like, you're going to avoid those like random errant, you know,$1 ,500 bills, because somebody forgot to like, you know, actually Terraform apply off of master instead of like their staging environment that they were testing, right? So, and I think that having that posture in place kind of by default means that companies get a lot more cost control. They get a lot more benefit, right? But they still get the ability to move extremely quickly by having the ability to create copies of their environment, right? How does monitoring, logging, observability, these types of things work?
22:58We have templates that people have built that allow you to kind of like exfil logs to Datadog. We have a template for spinning up Grafana or like Victoria Metrics, so Prometheus compatible instances. So you can do all of those things. We've also built from the ground up a observability system inside of Railway. And since we have the edge network, we automatically can kind of add a request ID. So we can kind of give you distributed tracing by default through all of your microservices. We can give you alerting for if things spike or stay high or anything else like that, right? We can give you information that you wouldn't have in other environments without doing, for lack of a better word, a ton of plumbing.
23:38So that's the thing that we've built from the ground up internally at Railway. And I think that obviously Datadog is a massive business and there's tons and tons of stuff in there. So we're straddling more of the 80-20 of let's just give people the baseline amount of things that they want. And over time, we're going to go and ask them, hey, what else can we give to you? But we have people who are just kind of using Railway entirely in terms of build, deploy, observe, scale, all of those pillars and are actually extremely happy on using just that. And I think it's bare bones in terms of where we want to take it right now.
24:09But it's super exciting to say, hey, listen, we have all the building blocks right here and we just need to work with our users to kind of scale to the things that they really, really want. What would you say are some of the constraints or limits today? Constraints or limits? That's an interesting one. I would say that there's like, I mean, it's going to sound weird, but there's not really any constraints or limits right now. And that's kind of like, we've really tried to solve that like Heroku problem of, oh, you know, I'm going to outgrow this thing, right? And so I would say that we're almost limited by trust, if that makes sense, right?
24:40And so, you know, you have AWS, you have GCP, and you have Azure. Those are the big clouds, right? Like, barring like anything, those are the big clouds, right? And you have like Cloudflare that's like trying to do things, right? But Cloudflare is like a$30 billion organization. They've been around for like a decade plus. They've accumulated trust. You know, they're the meme whenever Cloudflare has an issue, it's like software engineering snow day, right? It's kind of like they've continued to build that trust over time and they're still working on it, right? And so for us, I would say like the main thing that's kind of the limiting reagent right now is not like what the platform can do, but it's almost like how much you can trust it, right?
25:15Because when it comes to software infrastructure, it's your livelihood, right? Like there's not much more that you can kind of like trust to people, right? It's your data. It's the fact that if that thing goes down, especially with your data, you're SOL until these people like get back to you, right? And so you're kind of hanging in the limbo, etc. Right. So it kind of pulls more towards the quote of like, you know, nobody got fired for buying AWS, nobody got fired for any of these other things, right? And so I think that's the main thing that we're kind of like consistently working with companies on and saying like, yes, you're going to get all of these benefits, and we're going to give you like a higher order level of reliability and we're going to give you better service sla turnarounds than than the kind of like larger clouds and that ends up being kind of like a very very difficult battle to have with people because they just say that sounds like bs right and you have to just show them over time it's like no like we're going to continue to kind of like work on that and like we will be available for you should anything occur and we've also designed this system such that the there are less and less fault points as you kind of go i mean as a like if you're founder of a business and you're building out a new product, then it can be a lot to sort of bet your product life on another, essentially startup, going back to sort of the trust challenge that you're talking about.
26:23I think actually the startups really, they like are what we call from our like growth master master plan right now is like we're stretching market, right? And so startups like our current ICP, like the normal distribution of that, that kind of like go to market is actually 15 to 50 person teams. Those people like seem to love us, right? They seem to be able to want to move a ton of different stuff over. Maybe it's not literally all of their infrastructure footprint, but it's a large swath of things that are no longer legacy. And they're basically saying, listen, we want to move really, really quickly on these things.
26:53It ends up being those larger organizations that move a little bit slower and want that higher order trust bit really, really flipped. And so when you start going to like organizations, it ends up being most of the sales motion at that point ends up being this kind of like, how do we get past any sort of like trust or compliance or whatever objections, which we've gotten really, really good at and show you the value to kind of like tie it to maybe one of your like top eight, like velocity initiatives where like we just want engineers to be able to ship faster and get get more done, you know? So if you are a larger organization, let's say that you you're able to establish that trust with them, like what is the starting point for them?
27:29Like, you know, obviously, I'm not going to just if I'm not, you know, hybrid cloud or I'm on cloud today, I'm not going to go and like replace everything. Like, how do I get started essentially in a way that doesn't require me to boil the entire ocean? Yeah. So, I mean, the nice thing about Railway is like you can incrementally adopt it. Right. And so if you have services that you want to go in and spin up internally and you want to like pair them with services that already exist, You can do that using like something like tail scale, or we like have a wire guard binary that allows you to like mesh it into your instances over there.
27:56We can also do potentially dedicated instances if you want after chatting with you. So there's a variety of different ways you can get started. But the main point is that people just kind of like incrementally adopt it. They'll start with something they'll basically start with usually kind of it's an EM messing around on the weekend to basically say, all right, how do I go and explore a couple of these things that I know are being pitched as like these faster alternatives to X. but I need to be able to know that like, you know, it's good. Like it's, first of all, it's going to satisfy that like faster initiative to X criteria that I'm looking for.
28:23And two, that, you know, it'll be solid for us to like go in and make a case in the future. Right. And so what we do essentially is we go and we pull telemetry from people as they're, they're signing up and we say like, Hey, when you start getting to points where you want to be activated, we basically just say, Hey, if you want to go and chat with us, you can chat with us over Slack. You can chat with us over email. Right. Like we want to be really, really available without being kind of all up in their face, you know? And then you've open sourced like a significant portion of the railway production stack.
28:48Like what was the motivation behind that? I think the motivation for us is that we want that trust. Right. And so if we can give you this kind of like ability to introspect the service, maybe even like self-hosted in the future or anything else like that, then realistically, there's not much more you can trust us if you can see if you can see the guts and the internals and anything else like that. Right. So that's kind of the main motivator, I would say, in general. it's also obviously excellent to have members of the community be able to like go and contribute we have like people who are like submitting like nix packs prs all the time to like go and add new versions or new providers or new languages or new like you know dependencies or anything else like that and so getting kind of that tailwind of like open source not like tailwind in the like but like the benefit of the open source community to be able to like go in and help us like build this together you know that's i think a big key and i also like kubernetes ends up being open source right and so it's like do you want to potentially you kind of have to go and meet people where they are.
29:41And if they're self-host, if they're kind of self-hosting some of these things, at some point you need to be able to allow them to self-host at least the data plane, right? Does that help also overcome some of the trust issues? I believe so, yeah. Because you can see everything that people are doing, right? You can see what they're doing on the instances. You can see what code they're running. You can see all of the other things, right? So I think ultimately that really helps with the trust issues. Maybe not issues, but like problem there in general, right? But at the end of the day, I think the main thing that you can do to make sure that people trust you is to once the stuff is up, keep it up and keep it running.
30:17Right. And just make sure that it's like a bulletproof experience. Right. So we've been like working, especially over like the last like six months to make sure it's like, we're starting to get, you know, multiple nines on the board of like, okay, cool. Like, this is like a level of reliability, especially with our new bare metal instances. We haven't had any issues so far. So obviously knock on wood there, but like building something that you have higher order control of so you can get that level of reliability so people can say actually i've used it for x i don't really have a ton of problems with it right so and then we've also scaled the like cloud version of it to like you know we're doing tens of billions of requests per month on the edge proxy we have like you know the orchestration engine in the cloud is managing like two plus million microservices on us on like i think four clusters right so that's kind of like been built out so that whenever we go to like have a conversation with larger companies like uber or anything else like that.
31:035 ,000 microservices, that's a lot of microservices, evidently, right? But the cloud version of what we built is like scaled to like 2 million, right? So it's like a few orders of magnitude off. So we can say like, hey, if we were to go and like do a self-hosted version of this for you, we can do that like level of scale. You previously worked at Uber. Did any of that experience sort of like inform or motivate you to start Railway? Yeah, definitely. So there was, you know, kind of this like walled garden of like platform teams where, you know, some people were responsible for like getting code deployed and stuff like that.
31:31And I would go and interact with these things at work and it would be like fine, you know, and then I would go and interact with them at home and there'd be like varying levels of less than fine, you know. And so there was just, it really seemed that, you know, Uber was a high growth company at the time, still is, you know, it's massive. And, you know, they had the ability to retain the best engineers to go in and build this thing. And I still could potentially see ways that they could be like significantly better just in terms of like unlocking developer productivity. So I would say that that is definitely like a driving function for like making this system.
32:05So you've been working on Railway for what, almost five years? Yeah. Yeah. About five years. I think four and a half now. And as like a first time like founder, you're running 100 % remote. Like was that a conscious decision? Yes, it was a conscious decision. So I've worked remotely since 2015, maybe. And I've always like really, really strongly enjoyed working remotely. It's definitely not for everybody. I think it requires people be almost like extrinsically motivated about the problem, right? So people have to be like, they have to really, really like the thing that they're working on. And they have to kind of like see the vision, they have to see all of this other thing, like the where it goes, and everything else like that.
32:41So you can't kind of like get a lot of the benefits out of, you know, being in an office and, you know, kind of being able to breathe down people's spine and about saying like, Oh, we got to like do all of these things, right? So I was chatting with somebody, one of my friends the other day, and I was like, I think maybe my one of my most controversial opinions is like, I think despite running a remote company, I'm like, I think it's a terrible idea for probably about 90 % of people, right? Because you have to be extremely deliberate and you have to like hire people who are really, really excited about the problem space.
33:07We're going to like self-manage, we're going to go and do all of these things. Right. And so that's what we've done so far. We've got like 25 people. We're like remote spanning all the way, you know, like we have some people in like the Western hemisphere of Canada and I'm from like Vancouver Island, as we like previously mentioned, right. I'm in San Francisco now. And then we expand all the way, We've got Thailand, we've got Dubai, we've got Japan. So there's a lot of different time zones that we're spanning in general. And so we've had to hire people who are autonomous and who can push these things.
33:35So that's been excellent from leverage. But it does also mean that there's a specific class of individual who does really, really well at remote companies. And there's a specific class of individual who does well in in-person companies. It's kind of more of a chocolate or vanilla. But yeah, it was definitely a conscious decision because I do enjoy the benefits that remote has not in like a, you know, you can sit around and kind of twill your thumbs. But like you can almost like meet with people at almost various different times in the day. And, you know, their morning can be your evening. And so you can almost have this like almost like time compression handoff of saying like, oh, we've really got to go and do this.
34:08And then you go to bed and then you wake up and it's done. And you're like, damn, I love working with excellent coworkers, you know. And so it means that you almost get twice as much time in the week if you can get these handoffs like 100%. In terms of hiring people that are extremely motivated by the problem and also are able to self-manage, I feel like that's kind of a requirement for any relatively small stage startup. Because you just can't afford to have to micromanage people in order to do their job. You have to hire people who are motivated to be there, believe in the mission, and also can essentially just get things done and understand what needs to be done.
34:41I totally agree by the 100%. And the corollary of that is that like most people don't become like good managers at the start, especially like first time founders. You'll hear this from like, you know, Emmett from from Twitch has talked about like how he was not a great founder originally. I think Brian or Joe have also talked about it from Airbnb. Right. And so it almost like forces you to develop those like skills while you're also trying to assemble the airplane. Right. And so you don't get a lot of that kind of like benefit that you can kind of like smear this over time and learn those as you go through all of these stages.
35:10you really have to kind of like learn and compress a lot of those learnings and you have to do it in there so like that's the only kind of like thing that i i generally see is yes you all have you'll have to do all things eventually but we also talk a lot about you know internally about focus right like how do we stay focused on just doing the things that we can be like top number one in the world at right because that's where we're going to have all of the compounded returns right so those are kind of like some interesting benefits and drawbacks to remote that i think that maybe people don't necessarily consider, you know, they consider the, yeah, sure.
35:39I can go out and, you know, go for a bike ride in the middle of the day and work slightly later and stuff like that. And that's kind of like copacetic with my, you know, operating cadence or, you know, we have, we have people who have families, you know, and they spend the afternoon with their families and then they'll come back and, you know, polish up the stuff once the kids are to bed. Right. And I, I personally, I'm a big fan of treating people. It's going to sound a little bit condescending, but like treating people like adults because everybody is adults, right. It's like, they're going to manage all their time and they're going to like go and do all these things.
36:06Right. And you shouldn't kind of be around and basically say like, oh, I'm going to approve all your PTO requests. I'm going to like go and do this. Right. Because again, those kinds of people that you mentioned, like they're not going to be successful at startups in general. So we've almost kind of erred on the side of like, let's just pretty rapidly remove any of the guardrails in terms of onboarding and say like, hey, this is how we kind of operate. If you really, really like it, like here we are. And if we don't, then like, you know, we can, we can hopefully, hopefully find a really, really awesome spot for you to like go and land in the future.
36:32How do you test for that? Like, how do you find these people that are going to be a fit? Yeah, so there's a couple of problems that I like to ask people. There's a couple open-ended kind of like technical problems in terms of like design. I don't think you test for this using like lead code, you know? There's a lot of problems with using things like lead code. Oh, yeah. And so like we prefer to go for kind of like maybe the Montessori school of management of like open-ended kind of like problems of saying like, hey, listen, like how would you like solve this class of like real world problems, right?
36:59So I think that's like one way to go and do it. I think also sitting down and chatting with people about what drives them, you know, are they passionate about X, you know, stuff like that. And it doesn't need to necessarily be dev tools, right? Like it could literally just be like, I'm so passionate about networking, right? Like it's just networking. I got, I got a wrist of switches in my basement and all these other things, right? And stuff like that, right? And I think if people, they have that thing that they're really, really passionate about, and you can almost like see it and stuff like that, it's going to sound like super realistic.
37:26But like, you can almost tell when like people really, really care about these things, because the moment you poke them, they almost say it kind of like expands. And you're like, Oh, my God, like, that's so much stuff, you know, like, where did that all come from? Right. And where it came from is this kind of deep passion, this life experience, all these other things, right. And so that's like one way to go and test it in the interview. And then as part of like onboarding, what we do is we rapidly kind of will just remove the guardrails as you go, right. So like, we have six weeks of onboarding, which like may seem long, but the goal is almost pushing out the funnel on what you can do, right.
37:56Because by the end of the six weeks, we consider onboarding to be like, how do we get you from good to doing great work and then being able to be fully autonomous? So that's the whole goal is we do two weeks of tasks. So it's like five tickets. You just have them in linear. They'll be really straightforward. It's like, hey, this thing is actually broken. It's like, you go and you make a few lines of code changes. And then that's your first two weeks. And then you move on to the problems, which is, hey, our current experience sucks. Right. Very different class of problems that we've just given you.
38:23Right. It's like, well, what sucks about it? Like, what are the problems? All those things. Right. And then people have to like go in and, you know, maybe they'll ask their coworkers, maybe they'll go and ask the users, maybe they'll go and put together a document that says, I think these are the things. And then people will be like, no, I think, what about like, we could probably drop this requirement. Right. And then they'll go and solve those problems. And then the third one is like opportunities, which is you've been here a month. What do you think the company needs? That's the only prompt.
38:45Right. And so you get to that open ended kind of like line of thinking by the end of it. And, you know, that's that kind of pushes people more towards like, oh, I can actually do pretty much anything here. Right. And so it's just a matter of what. Yeah, yeah. I think there's a certain like, you know, back to what you're talking about, where you're trying to look for, you know, what is the person like really passionate about or interested in? There's a certain obsession behavior that I think people need to be successful in startups. And it doesn't even need to be historically an obsession that relates necessarily to what the company is doing.
39:17But you have to kind of be manically obsessed about solving something in your life. Because a lot of it is really about knocking down doors and solving problems. And the level of abstraction also hits you much faster earlier in your career, I think, at a startup because there's just less people around. To your point, by week six, you're asking questions about what could the company be doing better. If you're at a really large organization and you've worked for some, I've worked for some, Like you could really be in that first stage of like, here's a really small problem that you need to solve.
39:49And that could be like the first two years of your life in existence there. Yeah. Right. So we aim for trying to find those people because we think that I think even when it comes down to like solving problems excellently, right? Like that last kind of like 5 % of the problem is kind of where most of the progress is. Right. And if you're not like focused and you know, you're consistently oscillating and bumping between all of these different things, right? You're probably not going to like get actually to like the meat and potatoes of like what that problem actually is, right? I'm a very, very big believer of sit there and run a ton of different revisions, right?
40:23And then figure out like why these things were either better or worse than your previous revisions, right? And make sure you have a clear goal that you're kind of like entering towards, right? And I think that doing that and doing that well, you know, that focus and that passion is like, it's almost like a necessary precondition. It's like, it unlocks like pretty much all of those other things, right? And I don't see how people can potentially do it without it. I've seen it once in a blue moon, but it's also very, very rare. So, you know, it's we aim for kind of saying like, OK, well, like how do we generalize this class of problem of like finding these extrinsically motivated people?
40:54Right. And saying like, this is kind of like in general, the archetype of individuals that we've seen be quite successful here. Yeah. I mean, there's a quote about it. I don't remember the exact quote, but it's like, you know, we've completed 90%. So now we get to start on the other 90 % of the project, essentially. because exactly it's really that like last 10 % where all the hard work the meat and potatoes has to exist in order to actually whether it's a product or whatever it is that you're building that to do that that's where the polish happens that's where you know you run into your scale problems and so forth it's also important like not just from like an individual perspective but like from working with other individuals perspective because that last 10 % it's the hardest like that last extra mile of like going and doing the thing it's like yeah well like this is like good enough right like it'll work right and if you work with people like who have you know outsized talent or outside standards or anything else like that, they'll basically say like, no, we can do better, right?
41:44Like we can push this thing a little bit farther, right? And then they'll have some like tools in their toolkit, they'll have, you know, some skill or some emotional acumen or some way to kind of like judo your thinking and basically saying like, what have we thought of it like this? Like, I think this is really, really close. And this is where the reason why it's really, really good. Like, can we try this? Right? Versus, you know, I think the conventional thinking at larger companies is like, okay, cool, it's done. Like, let's just move on to the other thing. Awesome. Jake, thanks so much for being here.
42:07Cool. Awesome. Yeah. Thanks so much for having me. This was great. All right. Cheers. Cheers.
From the publisher
Railway is a software company that provides a popular platform for deploying and managing applications in the cloud. It automates tasks such as infrastructure provisioning, scaling, and deployment and is particularly known for having a developer-friendly interface. Jake Cooper is the Founder and CEO at Railway. He joins the show to talk about the company
The post Streamlining Cloud Infrastructure Deployments with Jake Cooper appeared first on Software Engineering Daily.
