Bringing Vitess to Postgres (Interview)

23 Jul 2025 · 1 h 15 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The Changelog Podcast Episode Notes

Episode Title

Bringing Vitess to Postgres (Interview with Sugu Sougoumarane)

Overview In this episode, Sugu Sougoumarane, the creator of Vitess, discusses his return from a three-year sabbatical to bring Vitess to Postgres through a new initiative called Multigress. The conversation covers the motivations behind this project, the technical challenges involved, and the current state of Postgres at scale.

Key Participants

  • Sugu Sougoumarane: Creator of Vitess and head of Multigress at Superbase.
  • Podcast Hosts: Not extensively mentioned but include discussions with Kyle Galbraith and others.

---

Episode Breakdown

Introduction

  • Welcome back to The Changelog.
  • Introduction of Sugu Sougoumarane and the focus of the episode.

Sugu's Sabbatical

  • Duration: Three years.
  • Reason: Burnout from long-term work on Vitess (12 years) and seeking personal growth through gardening and volunteering.
  • Return: A renewed interest in technology and realizing the ongoing need for database solutions.

Transitioning from Vitess to Multigress

  • Vitess Origin: Began as a solution for YouTube to handle database scaling issues; evolved into a full-fledged project.
  • Motivation for Multigress: Address the growing need for scalable solutions in the Postgres ecosystem.
  • Technical Challenges: Translating Vitess concepts to Postgres, understanding the significant differences between MySQL and Postgres.

Technical Insights

  • Vitess Features:
  • Intelligent connection pooling.
  • Sharding solutions and query routing.
  • Evolution of Vitess from a simple connection pooler to a fully distributed database system.
  • Multigress Goals:
  • Adapt Vitess concepts for Postgres while ensuring a Postgres-native experience.
  • Improve high availability and scalability for Postgres users.
  • Challenges in Migration:
  • Addressing differences in implementation between MySQL and Postgres.
  • Need for a new approach to sharding and query processing in a Postgres context.

The Future of Postgres and Multigress

  • Postgres Community: Noted growth and excitement surrounding Postgres over the last decade.
  • Absence of a Similar Solution: Sugu discusses the lack of a scalability solution for Postgres comparable to Vitess.

Development and Implementation Plans

  • Team Composition: High-caliber team including previous contributors to Vitess.
  • Timeline: An MVP (Minimum Viable Product) is expected in 3-6 months, with a focus on approachability and user-friendliness rather than excessive flexibility.

Thoughts on Open Source and Collaboration

  • Open Source Philosophy: Emphasizes the importance of community and adoption for the success of open-source projects.
  • Potential for Collaboration: Multigress could benefit other platforms, including Neon, demonstrating a willingness to foster collaboration within the database community.

---

Conclusion

  • Final Thoughts: Excitement about the potential of Multigress to fill a significant gap in the Postgres ecosystem and improve scalability options for users.
  • Future Announcements: Sugu plans to share more insights and updates through blogs and community communication.

---

Key Takeaways

  • Sugu's Journey: From burnout to renewal and innovation in database technology.
  • Vitess to Multigress: A strategic shift to address the needs of the Postgres community.
  • Technical Challenges Ahead: Bridging concepts between MySQL and Postgres while maintaining a native feel.
  • Community Focus: An emphasis on open-source values and collaboration to drive success.

Resources Mentioned

  • Vitess: [Vitess GitHub](https://github.com/vitessio/vitess)
  • Superbase: [Superbase website](https://supabase.io)
  • Postgres: [Postgres Official Site](https://www.postgresql.org)

---

This episode provides a deep dive into the motivations, challenges, and future directions for bringing Vitess-like scalability to the Postgres ecosystem through Multigress. The conversation highlights the ongoing evolution of database technology and the importance of community-driven solutions.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05What's up friends, welcome back. This is your favorite podcast, the changelog. Yes, we feature the hackers, the leaders, and those voting Vitesse for Postgres. Multigress. Today we're joined by Sugu Sugimaran, the creator of Vitesse. And you know we're Postgres for life around here. We talked to Sugu about his sabbatical he took, what Vitesse is, taking it from side project at YouTube to full on company. Now at Superbase, bringing Vitesse to Postgres with Multigress. We cover what's required to take Vitesse to the Postgres world, voting the team to do it, implementing multigress multigress at superbase and so much more a massive thank you to our friends and our partners at fly.io that is the home of change all.com learn more at fly.io okay let's multigress

1:05well friends it's all about faster builds teams with faster builds ship faster and win over the competition it's just science and i'm here with kyle galbraith co-founder and ceo of depot okay so kyle based on the premise that most teams want faster builds that's probably a truth if If they're using CI providers with their stock configuration or GitHub actions, are they wrong? Are they not getting the fastest builds possible? I would take it a step further and say if you're using any CI provider with just the basic things that they give you, which is if you think about a CI provider, it is in essence a lowest common denominator generic VM.

1:45And then you're left to your own devices to essentially configure that VM and configure your build pipeline, effectively pushing down to you, the developer, the responsibility of optimizing and making those builds fast. Making them fast, making them secure, making them cost effective, like all pushed down to you. The problem with modern day CI providers is there's still a set of features and a set of capabilities that a CI provider could give a developer that makes their builds more performant out of the box, makes their builds more cost effective out of the box and more secure out of the box.

2:21I think a lot of folks adopt GitHub Actions for its ease of implementation and being close to where their source code already lives inside of GitHub. And they do care about build performance and they do put in the work to optimize those builds. But fundamentally, CI providers today don't prioritize performance. Performance is not a top level entity inside of generic CI providers. Yes. Okay, friends, save your time, get faster builds with Depot, Docker builds, faster GitHub action runners, and distributed remote caching for Bazel, Go, Gradle, Turbo repo, and more. Depot is on a mission to give you back your dev time and help you get faster build times with a one-line code change.

3:02Learn more at depot.dev. Get started with a seven-day free trial. No credit card required. Again, depot.dev.

3:39Today we're joined by Sugu Sugumaran, co-creator of Vitesse, co-founder of PlanetScale, and now the head of Multigrass at Superbay Sugu. We are excited to have you here. I'm very excited to be here too. I'm excited to see you back at work. You've been off work. You just told me you took a three-year sabbatical. Is that accurate? Yeah, yeah. I even wondered if I was going to come back until, you know, I started feeling the itch. And yeah, that's what brought me back. Gardening, you know, ripe tomatoes. What kind of vegetable did you start growing?

4:20I actually went on self-reinvention of some sort. I kind of, you know, did a full debug of myself of some sort. Do you climb a tall mountain or how do you do that? I don't know. Just sit at home, think, walks, and a lot of thinking. and I even got into doing some voluntary work because, you know, I realized, you know, as a human, I realized basically I thought I should be more human and that led me into a bunch of voluntary work which I actually started to do. And I even thought of going full-time. So then I realized it's not something you can do full-time And that's when I realized that I'm starting to miss the fun of doing technical work.

5:19So I started doing some advisory work and stuff and slowly I got pulled back in and realized, oh, my God, we got to do this project, which is how I am where I am today. Did it take how long did it take before you started getting that itch? Because three years is a pretty long one. I think by that point I might be detached enough to not get the itch again. It was strange. Like, I think last year sometime, late last year. So for about six to eight months, I've been feeling the itch. I think it started because people, because I'm still in touch with friends and stuff, they used to ask me for questions and help about things.

5:58And then as you answer, you kind of enjoy that you did it. And you say, OK, maybe I can. And then I talked to some people and they said, oh, you know, someone like you is still useful. And as you realize, the last three years, everything changed in the industry. Right. AI actually, the first GPT came out just when I went on sabbatical. Around mid-2022, I think. And when I thought of coming back, I thought, you know, maybe I'll be obsolete or something. because I don't know anything about AI. But then it turns out that infra is still way behind. And they say, oh, what you know people desperately need.

6:50Databases still matter, turns out. Databases matter, infrastructure matters. Nothing much has changed in that area. So I came back and slowly got, you know, roped back into things. What pushed you away in the first place? Why did you go on that sabbatical? Had you accomplished all your goals? Were you burning out? What was the situation that made you? Both of those actually. Both. Okay. So this was 2022. I had been working on Vitesse for 12 years straight. We started Vitesse in 2010 at YouTube.

7:27And Vitesse is an enormous project. And for a while, I was the sole maintainer of this. And it's a pretty big load on your brain. And after I founded PlanetScale, hired a team and trained them, and finally the team was outproducing what I could do. And that's when I realized, okay, you know, it's now time to recover from this huge load that I've been putting myself through. And that's what led to my taking the sabbatical. What year was that again? Was it around the time that you, I'm not familiar with the story behind the scenes really with Sam Lambert, but I'm, I guess, tangential friends by having him on the show and hearing his backstory and PlanetScale's early story and mainly taking over the CEO role and what that did for PlanetScale at that time?

8:26What time was that you stepped away and was that the time that you brought on Sam as a CEO? So I think my memory may be hazy here. I think he came on board around 2020. Sam has been a fan of PlanetScale from the beginning. So he even did a Sequoia Scout investment when we launched. And so he has been following us. He's been actually advising us also in the early couple of years. And then in 2020, he came on board as, I think, CPO, Chief Product Officer. And he changed the whole image of PlanetScale. Until then, we were, you know, very, what I would call a back-end, serious back-end database, but in some respect, you know, boring of some sort.

9:23We were pretty serious. We would do things, but now we were like, he brought the excitement in when he came on board. And the first feature he launched was the schema branching and merging, which the dev community completely loved. And that was around, after about like around 22, Jiten at the time was CEO and Sam talked. and they came to an agreement that if Sam became the CEO, that would be really effective because then he could control the whole story. He became the CEO probably much earlier, probably 2021 would be my guess. I'll have to look it up. It was probably announcements. So he had been the CEO for a while and then in 22 is when I left.

10:24uh also by that time the team was all wrapped up like vites had a tech lead uh so everything was looking great so i almost felt like i was a warm body that's a good feeling i mean you've done a lot of the things you set out to do for those of us playing catch up can you describe the tests what it is and how it helps scale databases Yeah, yeah, definitely. With us, actually, when we started with us, we thought it was a six-month project. So YouTube was actually in deep trouble scalability-wise. There were more than a handful of outages every day, like sometimes 10, 12 outages per day, and they were all caused by the database.

11:16I wouldn't say outages, 10, 12 pages per day. Many of them would lead to outages, and they were all caused by the database. That's the time when my co-creator, Mike Solomon, and me kind of decided to solve this problem once and for all. So we decided to take ourselves out of this firefighting mode and think of a solution that will leap us ahead of where we are today. Because it felt like what we were doing now was a losing battle. It was like whatever we were doing, things were getting worse by the day. And that's how Vites was born. And we wrote code for about a year, actually. It took us about a year to launch the first version.

12:06And the first version that we launched was just a connection pooler, which was actually the biggest problem we were having with MySQL at that time. and that immediately gave some breathing room for YouTube. And the way we solved it was we didn't solve it like how you would solve a problem, which is let's solve connection pooling. Because we sat and brainstormed about all future problems that we're going to have, we actually built an intelligent proxy that would not just connection pool, but would also understand your queries. So we actually built a parser in that connection pooler. And once this got launched, people started noticing that this so-called connection pooler is intelligent and it can do other things beyond just taking care of connections, like, for example, fielding bad queries or watching queries and killing them if they run for too long.

13:07And so we slowly started adding those features and it became an intelligent connection pooler that would protect the database completely. And that eventually evolved into, then we added sharding solutions to it, like how do you reshard? And then we added one more layer, which is the routing layer. And then we realized, oh my God, this is now almost a fully distributed database. Why don't we make it a fully distributed database? which was the final step where VTES became a fully sharded solution. By the time we started in 2010, I think what we call as the V3 was around 2014 or so. About four years later, VTES could be adopted by anyone outside of YouTube.

14:00And then the rest is history. The rest is history, right? Open source, mass adoption, planet scale, build a business around it. Yeah. It's very successful business to this day. Really cool stuff. Kind of an open source success story. I was gonna say prototypical or stereotypical, but like not necessarily, but just a good one. You know, just a good open source success story with, you know, money involved in a business that could be built and kind of what everybody wants is like build some cool tech, open source it, help companies scale their businesses, is like provide massive values once and for all like you said and then also build a business around and be able to be successful and then take a three-year sabbatical when you're all said and done you know like all right i'm done for now yeah i think to extend from what you're saying there jerry i think it's wild you know suga one thing you mentioned was like the nine month or six month you thought the project would take yeah uh in terms of like it's done this but then it turned into a company and then obviously PlanetScale has been able to and I think others too have been able to hire and you know maintain full-time open source maintainers on the the test project not just to support PlanetScale but to support the project I think that's just so wild how you think like it's just this little thing or it's a certain amount of time and solving these skill problems the next thing you know it's something much bigger than that.

15:29Yeah, yeah. Like the first adopter was actually a company called Flipkart. I believe they still use the older version of Vitesse. Then HubSpot came on board. They still use Vitesse and they have their own operator. they were the first ones to launch among the first ones to launch with us in kubernetes and then later slack came on board and github came on board so now it's like uh they are pretty they are still pretty uh substantial contributors to the project at this point well you said once and for all and i said once and for all but the technical facts about vitess is it's not for all right?

16:14It's for all who are running specific database systems. And one of those alls who aren't part of the all is the Postgres ecosystem, which has been awesome for many years, but has been exploding, I would say, in the last five or 10 years in usage, in features, in capital support. I mean, there's just lots of excitement and growth in Postgres, none of which Vitesse could provide for right it just wasn't it was just my sequel specific right yeah yeah so that was uh uh i think the postgres community took note of vitess a long time ago as far back as 2018 was the first conversation i had with someone from the postgres side and they said hey we need something like vitess for postgres uh and uh there have been a few false starts on this project i think the one where we got pretty serious about it was actually in 2020.

17:15Someone called, I don't know, he's pretty well known in the Postgres community, Nikolay. So he and I talked and we felt that, no, you should, we should do this thing for, you know, Port with Test for Postgres. We even like formed a sub team and started talking about it. But, you know, 2020, PlanetScale was on the rise. It was just too busy and unfortunately I had to make a call that, you know, like I have this company to take care of, I have this company to build. I can't afford this distraction at this time and I actually kind of had to back out saying that I don't think I can do this at this time.

17:57I still feel bad about having done this.

18:02And so Postgres has been in the back of my mind all along. Even after I went on sabbatical, I reached out to some of the older contributors and told them, you know, like, maybe you can start this project. But no one at that time felt like they all had their jobs and this would all be something on the side. And until this year, when I said, you know, it's now or never. And that's when I said, OK, I'm going to take this plunge. I think mimicry or ports or whatever you want to call it is kind of a tried and true part of the open source culture. You know, take this good idea from this ecosystem, which I appreciate, but I'm not a user of and bring it to my piece of the world and reimplement it over here or steal the ideas.

18:56and it's kind of, Vitesse has been so successful and so helpful, especially to large companies who are well-funded at helping them scale. I'm surprised that all these years there is not already a Vitesse for Postgres community built or coming out of a startup or something. Like, no, Sugu has to do it, right? Like it's back to the Vitesse people are going to do it. I don't get this. Yeah, I don't know why I had to come and start this. Is it that hard? Because usually you look at something and you think, oh, it can't be that hard. I'll do it over here. Is it that hard of a technical problem? Or is MySQL and Postgres implementation-wise so dramatically different that even the concepts don't necessarily come across?

19:46Like what's that chasm look like? I think it is more of the first problem. The fact is like sharding, And if you go and ask in a university, sharding, they'll say sharding is one of the oldest concepts in databases. It's over 30 years old. But the issue with sharding is that nobody ever worked on it. Nobody ever developed it into something practical that will actually work in production at scale with performance. which is actually strenuously hard. Like for me to make the leap from an application sharded system to a system where the application does not care about sharding and generically implement it in a database, it took me about like I literally sat and thought for three months straight, nothing other than just pure thinking.

20:51Understanding SQL, understanding query complexity, understanding relational algebra, going back all the way. Because, okay, how do I shard it? You go hit the books, you go look at databases. Nobody has ever done this before. And so then I realized, you know, this needs to be basically reinvented from the ground up. So that was one of the hardest, one of the most strainful thinking things that I have ever done in my life. There were times when I used to get headaches thinking about it. It was that hard. And finally, one day I had the aha moment. It looks like I have an answer to all the questions to make this thing work as a sharded environment, as a sharded system.

21:44and it's possibly that is that like a typical, when you get a query, the typical way you break it down is different from how you would break it down in a sharded system, because you get a select statement, you pass it, and then you identify each of those nodes and you just assign operations to it in a traditional database. And then you run the optimizer that rearranges these things. But in a sharded system, you have to identify groups of things that actually don't need to be changed at all because there is an underlying database that can execute that entire query with no changes. Because what's underneath is also a database.

22:33So being able to analyze queries and identify these subparts and be able to outsource it to a database and then do the rest of it on your database, on your proxy that's actually sending the queries, required a different kind of thinking. I think that is one reason that I could think of why sharding is hard. Can you describe your aha moment, like what your realization was, or did you just try to describe it in other words and I missed it? Because it sounds like you're describing the problem and the system with the proxy in there. But like, what was it that cracked the nut that said, aha, we can do it this way?

23:14Yeah, so how did I do it? The aha moment was, actually there was an early aha moment, early aha question, which is until the day when I said that, so before sharding, before this generic sharding was there, the way the application worked was it would actually send a query and along with the query, it'll tell you where to find the rows for this query. We called it the keyspace ID, which basically helped you. It says there's a lookup where you look up the keyspace ID. It tells you, oh, this keyspace ID can be found in this shard. The application had to guarantee that all the rows for that query lived in that shard.

24:08so that made it extremely simple and keyspace id was also a value of a column in every table that we had which made it safe and simple which means that like you can't go wrong and the question i asked myself is this keyspace id is a crutch no can we move without it Right. And it was scary because the application knows the keyspace ID, but how would this system know the keyspace? We test does not know the keyspace ID. So if you just sent me a query, what do I do with it? I don't know where to find this. And so I went to the application and figured out the application is figuring out what this keyspace ID is.

24:58How is it figuring it out? and that's when I realized oh, it's figuring it out based on the user ID so there is a source and it always turns out that that source is also in the where class of every query that it sent and that's when I realized oh, well, if the application is doing it what if I did it for you you say where user ID equal to X and I will do this computation, which is a hash function, and figure out where it is and send it. So then the question is, oh, then in that case, how about the other queries? There were other complex ways by which application was computing the keyspace ID. And that's like, what if we internalized each and every one of those things?

25:51Then the application does not need to ever compute the keyspace ID. We will do it. So that was the first aha moment. And then the second aha moment was, the second question is, but in that case, only these queries. So that limited the number of queries the application could send. The next question was, but then if we took this over, what if there's a query that came where a keyspace ID could not be computed? What would we do? That's when I realized, okay, then I have to understand the query. What if it's a join? What if it's a correlated subquery? What if it is a non-correlated? So that's when I had to go into relational algebra and find a way to map any random query to a query plan where I would say, oh, this query, I need to send it to all shards.

26:47Or this query, I can send it to just this one shard. or this query is in multiple shards, but I cannot send it to all shards because I need to take data from here and then data from somewhere else. So the aha moment is when I identified all these primitives and I could prove to myself that using these primitives, I call them the route primitives or the sharded primitives. If I used these, I can satisfy any query. the full SQL language can be supported eventually. We just need to do it. So the first implementation only supported some constructs, and then we slowly started adding more and more. But we knew that theoretically, all SQL can be supported in the sharded system.

27:41So to the key space ID point, so I understand the relational algebra and the entire SQL dialect, like supporting all the keywords over time. That makes sense to me. The Keyspace ID, as a non-vitess user, I'm not sure how it works. Do you have to go, does the infra person or the team that sets this up, do they have to go in and somehow configure it to know how you should shard against this particular application so that it can do the Keyspace ID thing? Or is it just, because different applications, different databases have different, as you know. Correct, yeah. So there's a cool story there. The first keyspace ID was actually a random assignment.

28:23Okay. So we would take a user and say, I don't know if it was a salt, what kind of, I actually don't remember how we initially computed. We used to randomly assign them to shards. Actually, oh, now I remember. A user was assigned a shard and we would put the shard number in a lookup table. It's random. So when a user gets created, they are assigned a shard, and the shard number is put in a random table. And then came the time when we had to reshard it. And we realized, oh, my God, how are you going to reshard this? Because now there is this table. That's when we transitioned to keyspace ID, where we assigned a unique ID to a user.

29:12and then we actually assigned that keyspace ID. We changed actually this lookup table to replace the shard number by the keyspace ID. And at that time, we populated every table with the keyspace ID. And then we realized that we are wasting this lookup for no reason at all. Then we did a full resharding where we did ranges based on keyspace IDs. in which case you just hash the user, you have the keyspace ID, and then that goes to that shard. And so that's how we evolved at YouTube from where we started to... But at that time, the keyspace ID was a column. And the reason the keyspace ID was important there was because we were using MySQL replication to do the resharding.

30:08and we had to use that value to decide where the row should go when we resharded. And that was actually another hurdle I had to cross, which is, can you reshard if you get rid of that row? In which case, then we'd have to make MySQL do that computation to do the replication, which means that it had to be a function that MySQL supported. The interesting story here is at that time, there were a large amount of arguments about what is the best sharding scheme? Is range-based sharding good or is a hash-based sharding good? We spent hours fighting on it. There was no order lookup-based sharding, right?

30:56There's like so many sharding schemes. And that's when I realized that there is no correct answer. And anything you choose, somebody is going to hate it. Right. And that's when I came up with this idea of no first class citizens in sharding, which means that Vitesse will actually, every sharding scheme is a plugin, is an extension. Gotcha. So you pick the one you want to use. You pick the one you want to use. And you don't have any opinions on that as Vitesse. Exactly. Exactly. That actually goes back to, goes back to actually a story from, I don't know if you've heard of Illustra. So, this is 1990s.

31:55Michael Stonebraker had, I think, just launched Postgres. And from that, he founded a company called Illustra, which is a commercialization of Postgres. And the claim to fame for Illustra is that indexes need not be defined by the database. You can define your own index and do it as a plugin. And at that time I was at Informix and Illustra was actually winning deals against Informix as a startup.

32:34And Informix actually ended up acquiring Illustra. And I love that concept so much that I actually transferred myself into the Illustra team to develop pluggable indexes. And that inspiration remained with me all these years. And when I said, you know, we have to develop indexing for sharding schemes, I said, well, you know what? I'm going to use this idea where all indexes in the sharding scheme are externally defined. Yeah. It makes sense because, I mean, oftentimes opinionated software is better. Like we do have a golden path or we do have a way that you should go. you can swap it out if you want, but this is what the framework or the tool thinks you should do.

Read the full transcript

33:22But when it comes to this, like I said, and like, you know, like the shape of people's data is so different and custom that there isn't necessarily a best way of sharding as you guys found. So it's what's best for your circumstance. And so why have a first class citizen when you can just say, no, you're going to, you're going to figure this one out yourself because we can't actually figured out for you. I think that makes a lot of sense. Yeah. The, uh, the thing that I'm now starting to realize, uh, I'm now learning Postgres, you know, like, uh, a little bit more deeply. I find the same ideas still there in Postgres, uh, where indexes are pluggable.

34:04You can plug in data types, you can define operators. So I'm kind of feeling like, you know, oh my God, there's a lot of similarity between you go full circle yeah yeah so it's pretty exciting to see this well friends i'm here with a new friend of mine harjot gill co -founder and ceo of code rabbit where they're cutting co-review time in half with their ai code review platform so harjot in this new world of AI-generated code, we are at the perils of code review, getting good code into our code bases, reviewed, and getting it into production. Help me understand the state of code review in this new AI era.

34:55The success of AI in code generation has been just mind-blowing, like how fast some of the companies like Cursor and GitHub Copilot itself have grown. The developers are picking up these tools and running with it pretty much. I mean, there's a lot more code being written. And in that world, the bottleneck shapes the code review becomes like even more important than it was in the past. Even in the past, like companies cared about code quality, had all this pull request model for code reviews and a lot of checks. But post-gen AI, now we're looking at, first of all, a lot more code being written.

35:26And interestingly, a lot of this code being written is not perfect, right? So the bottleneck and the importance of code review is even more so than it was in the past. You have to really understand this code in order to ship it. You can't just wipe code and ship. You have to first understand what the AI did. That's where CodeRabbit comes in. It's kind of like, think of it as a second order effect where the first order effect has been Gen AI and code generation. Rapid success there now as a second order effect. There's a massive need in the market for tools like CodeRabbit to exist and solve that bottleneck.

35:57And a lot of the companies we know have been struggling to run, especially the newer AI agents. If you look at the code generation AI, the first generation of the tools were just tab completion, which you can review in real time. And if you don't like it, don't accept it. If you like it, just press tab. But those systems have now evolved into more agentic workflows where now you're starting with a prompt and you get changes performed on like multiple files and multiple equations in the code. And that's where the bottleneck has now become code review bottleneck. Every developer is now evolving into a code reviewer, a lot of the code being written by AI.

36:29that's where the need for code rabbit started and that's being seen in the market like code rabbit has been non-linearly growing i would say it's a relatively young company but it's being trusted by 100 000 plus developers around the world okay friends well good next step is to go to coderabbit.ai that's c-o-d-e-r-a-b-b-i-t dot a-i use the most advanced ai platform for code reviews to cut code review time in half, bugs in half, all that stuff instantly. You got a 14 day free trial, too easy, no credit card required, and they are free for open source. Learn more at coderebbits.ai.

37:11So what's different then? It seems like the concepts, the breakthroughs you had with Vitesse, the proxy with the parser and the smart way of handling these key spaces and the bring your own sharding like all this stuff seems like it would apply at least conceptually across to postgres um but i guess at the nitty-gritty it's just wildly different or so the the the part that i'm still uh i'm still trying to figure out is so in other words every idea that exists in with this will just smoothly port over to Postgres. But as you said, there may be details where it may not port as well. So there's a few.

38:02At this point, we have made a few policy decisions. Let's put it this way. The biggest policy decision is this has to be Postgres native. So that part we do not want to compromise, which means that the end product should feel like it should be built for Postgres. So that is one policy decision that we have made. I don't know how far we are going to get there, but that's our North Star of some sort. and the part which means that from Vitesse there are some parts we will definitely leave behind for example anything that's MySQL specific is not coming over what else will we leave behind we will leave behind anything that is legacy in Vitesse anything that we built in Vitesse but continue to support because there are users using it but if you built it from scratch, we wouldn't do it that way.

39:09So that we want to leave behind. And also anything that's not well implemented. If something is, it works, so we left it alone, but it was not well implemented. Given a choice, I would do it differently. So that we are not going to bring over. And everything else we want to bring over for sure. So there are actually, and there's a fourth category, which is if there are things we can improve, we should. More specifically, I do want to improve the high availability story in Postgres because that is currently a big challenge that users are facing. There is no good high availability solution. So that is something that definitely want to improve on the Postgres side of things.

39:59So we'll bring some components over from Vitesse because Vitesse already has some HA components. So we'll bring some of those components, but we'll also build something that is very specific to Postgres. And the place where I'm still confused about how to solve the problem, but we will figure it out, is how do we solve this sharding and query processing? Because Vitesse can do query routing. Vitesse has the concept of pluggability. The part that I'm thinking about is the Postgres pluggability is more at the binary level where you actually build extensions and link them into Postgres. Right. Whereas Vitesse is built in Go which if we brought that in, how do we make these plugins work?

40:56Do we have to translate them or do we figure out a way to make them still work if we ported the Go code over is the part that I'm working through right now. Actually, we now have some people hired, so I'm working with some of those people too. Does that compromise your desire to be Postgres native considering what you learn with Vitas being written in Go, et cetera? How does that compromise your maybe Postgres for life? I'll maybe put words in your mouth, but this native Postgres way. So if you look at the Vitesse design, it was designed to be for the longest time, we didn't want to implement anything specific to MySQL.

41:41We thought this should be its own database system. And we actually restricted ourselves to just SQL 92. Because SQL 92 is generic, all databases must support it. Therefore, if we did SQL 92, then it'll work for anything. So we could even swap in Postgres underneath. But what we found out is that people were hesitant to migrate over. They said, oh, I want my application to run as is when I move over, which is what forced us to add MySQL extensions. So, in other words, we could port the SQL 92 part of the test directly over into Postgres and it'll work. And the problem is it's a subset. So if you do anything Postgres specific, it's not going to work.

42:33And so I'm beginning to think that then porting that means that, yes, we can port it, but then we have a lot to build to catch up to everything in Postgres. And looking at 10 years from now or five years from now, when, let's say, we manage to get full compatibility of Postgres, we will still face the issue of these plugins, these extensions. So that's the problem that I want to think about right now to make sure that when we get there, there is a solution for these extensions. How important is it to reuse a lot of what you learn with Vitesse in this new world? Because I can imagine like all this work you've put in and countless hours of open sourceness, et cetera.

43:21Like there's a lot in there. How hard is it? How important is it to bring a lot of it over? The learning is totally valuable. and without the learning, we wouldn't be able to build it. Rebuild, like if you say rebuild with tests, to what it is, it is a lot of code, but it won't take the 15 years that it took us to build. It may take us, even if we were to rebuild from scratch, it may take us a year maybe. Especially with AI, right? I mean, that might be a big help there. I mean, Go is really useful in in Cogen tooling and stuff like that. So I can imagine that that might even be a leg up for you to speed up the process.

44:08Yeah, we have been we've been using Claude to, you know, to to analyze Postgres. Yeah. So we you're probably right. We can probably do this much faster than even what I'm talking about. We had to rebuild this whole thing. So there's this common movie trope, you know, the retired badass. I'm not sure what this is called, you know. Like the guy who. Exactly. For one last. John Wick is like, comes to mind here. I'm sorry. John Wick, Bruce Willis often plays this role. It's true. Yeah. Or the Tom Cruise. Or Tom Cruise. Yeah. I mean, pick. I'm not sure who would play you, too. But you're coming out of retirement for this because, like, we got to get.

44:52What is it? Armageddon. And we're like, it has to be these people, you know. Right. which was funny because those were like oil drillers weren't they yeah interesting setup but no they come out of retirement for that one last mission you know i'm getting too old for this it was a common thing they say um and here you are but you're coming out with super base i mean it's a different story like super base came from where they're obviously their postgres maxis were fans of super base and uh no paul koppelstone had on the show a few times and so we know about them. We know what they're up to. Obviously, PlanetScale was on the MySQL side, but why Superbase?

45:29Why come out and hook up with this ragtag group of ProScrash Maxis to build this thing? So I actually engaged with Koppel on a whim, you know, saying, like, so when I said, this problem needs to be solved, needs to be solved now. My first thought, I should start a company. Let me do a startup. I actually started talking to some looking for co-founders. I was even talking to my to Jita and my PlanetSkill co-founder. And I was kind of starting to engage this. then I thought you know like what am I trying to achieve this time like I'm not trying to maximize profitability what I want to achieve is give this project the best chance for it to succeed and I was thinking if I do it as a startup I'm bringing so much risk and so much distraction right because then I will have to make sure that I raise money I have to you know keep people happy and and instead like what if I found a place where I don't have to worry about these things where I don't have to worry about raising money where I don't have to worry about and even like if you're a you know seed stage startup it's also hard to hire people you know there are people who are really, really happy, you know, where they are, they are being very highly paid.

47:15How do you convince them to come and, you know, like toil here and take the risk to, you know, we had this struggle at PlanetScale, by the way, when we left Google and we thought, oh, we know so many engineers, so hiring is not going to be a problem. And then we start the company, nobody wants to join us. So I thought of that. I said, you know, like maybe there is value in not going all the way to a big company, but at least a company that can pay people well, still give them a good upside. And so I thought of, you know, I'll start talking to some people who are, you know, in the startup stage and see see where they are.

48:00And that's how I engaged with Koppel. And I was like pleasantly surprised to find out he's like almost a perfect fit for what I was looking for. And the biggest one is the open source part. Yeah.

48:21Even in YouTube, we kind of had to push hard to open source with tests. for me, the success of an open source project is adoption. Open source and adoption is true success. And that's exactly what Paul believed in, Koppel. So it was so perfect. And talking to him about all the other values, he was like, at that time, it was almost like a slam dunk obvious that this is a place where I should be building this. Nice. So it's not exactly like the movies because in the movies, he's, you know, he would have sent out his henchman to find you and you would have been like, Oh, I don't want to come back.

49:02Yeah. You would have been like fixing up some old boat or, you know, practicing Tai Chi and like, they would have found you and you wouldn't have wanted to, but you had to, to save the planet, you know, but in this case, you actually reached out. Like you were, you were, you were ready to solve this problem once and for all Postgres users. And so that, that does make sense. I I think that jives, Adam, with my impression of Paul and Supabase. Does that make sense to you? Yeah, I like the fact that it's, you know, Postgres native, open source. I mean, that's the roots that Paul shared here. He's like, I think he said, the board of directors will have to fire me or pry this role from my cold dead hand.

49:41One of those, something elaborate like that to express his love for open source and his love for Postgres, Postgres native. So how'd it go from there? So he hired you. He's given you a team now and he says, have this on my desk in six months or how's it work? I'm giving myself about three to six months to come up with an MVP. It may vary. The reason why it may vary is because of my, I don't want to compromise on the long-term plan. So if for the sake of delivery, I feel like I'm deviating too much from where I want to go, I would rather delay the delivery. But I don't think it will take us longer than three to six months to come up with an MVP.

50:31We have a very well-defined set of objectives to reach for the MVP. And I think we'll hit it all in about three to six months. and then iterate from there. The idea is that from there, it should only be iterations, no major changes. That's the reason why I want to make the right decisions today so that it becomes iterative from then on. So that is the oversimplified plan, but I'm pretty sure we'll have, there are about five big features that we want to launch in this MVP. So is Go still on the menu then? or does that not quite get you Postgres native enough? I'm not sure if you can write. Can you write Postgres extensions in Go or do you have to use the C Go or how does that work?

51:21So that's the part that we are debating about. I don't want to start a brainstorm here, but I can tell you what we are talking about. Sometimes it happens. Yeah, just let it happen. Yeah, so the one option we are looking at is C Go. Whether we can, like if you have an extension, like can we link to call you through CGO is one thought that we are having the other thought we are having is well AI is so good nowadays can we like translate this for you and run it as an extension within Vitesse is the other thing that we are debating so there's the other crazy idea which is translate leave go and do it in C directly so all cards are on the table gotcha the last one is the least liked one at this point but uh it doesn't matter no whatever whatever gets whatever is the right solution whatever gets the job done so that's fun it's gotta be fun to be back at square one to a certain extent with a technical problem in front of you now like a big gnarly technical problem that's gonna take three to six months for an mvp and much longer if it's successful just to be back at where you started, but in a different context, you had a different part of your life.

52:39Yeah. And the best part is I get to fix all the mistakes I made in the previous project. Oh, yeah. But what about the mistakes you're going to make this time? That's the problem. You know this, V2 is perfect always, right? That's right. You're not going to make any bad decisions this time around. Yeah, that's got to feel good because there's so many landmines that you just can avoid this time around. Yeah, yeah. that's awesome what's your team looking like who are you looking for you looking for hardcore postgres people or did you hire from the test there were some uh so i cannot name names yet because uh some of them are still in the pipeline but it's a pretty exciting pretty high caliber team some some are older uh with test contributors that uh yeah there's a diaspora you know of with as contributors.

53:33Some are from there. And some are actually people that were that have handled this type of scaling problem before. So for whom we test is an obvious thing. So it's a pretty high caliber team. I think we'll be announcing the team very soon. They're going to be kind of the founding team of this multigress project so we'll probably be announcing in maybe about a month or so but it's hiring is looking very good at this point i'm pretty optimistic what is implementation like in this new world is that still to be found out as you make these choices that have long-term permutations into maintainability etc like how I know how, for the most part, how Superbase works as a service.

54:29But if I wanted to use what you're calling multigress, Vitesse for Postgres, essentially, if I wanted to use that on my own outside of a Superbase context, what do you know about the future thus far to kind of showcase implementation details for end users? So in this, I should talk about something that we did in Vitesse which we are not going to do in Multigress. Vitesse actually is what you would categorize as one of the most flexible enterprise solutions you can ever think of. If you want to ask a question to Vitesse, I want to change this very specific behavior of Vitesse, there is probably a command line option for it.

55:20so that made Vitesse extremely powerful and flexible that's the reason why Slack could adopt it and completely deploy it using their rules GitHub could use a completely different approach and I believe Etsy has an even crazier approach and Vitesse works for all of them but the problem it created by solving these things is approachability if I came in and looked at Vitesse I wouldn't know like as a layperson who was just oh my MySQL is getting too big I want to use Vitesse it will take you a few months of learning how Vitesse works before you can start to use it confidently so I realized that's a problem and so what I want to do in the multigress is make it a lot more approachable.

56:16Actually, at the expense of even reducing those flexibilities. Because I think adaptability is more important, approachability is more important than flexibility. And if we add flexibilities, we want to hide them away from the user that they shouldn't have to find it until they need it. so that is that is kind of the approach we are going to use and what that means is that it should be easy for a user to just take a multigress and just deploy it in their environment and it should just work for them do you know how that will manifest inside a super base or is that more on the super base product team to decide like will that be invisible to super base customers Absolutely.

57:07Yeah. So the... Makes sense. The idea is what we are going to be building is a Kubernetes operator. So if you want to deploy it yourself without Kubernetes, you may have to do a few things for yourself. But if you just took Multigress and use the operator, you specify, you know, this is what I want. and then the entire cluster comes alive. And we plan to do that within Superbase also. So this operator will be deployed within Superbase. And all Superbase has to do is, you know, hey, bring up a cluster and we'll just spin it up for you. And they have to be operating at significant enough scale that they can act as a test ground for you, I would think.

57:58Or maybe they can't do that with their customers. But I'm just thinking like you had YouTube to build the test with And that was like awesome, right? Like you had a real scale database that had consistent users, et cetera. And you could build the system. I guess in that case, you built it incrementally over time as you figured it out. But building this in a vacuum, you got to have some sort of real scaled things to test against. Maybe these are some of your, your core team that you'll be announcing later that I'm sure they're going to come from scaled up companies. Yes, yes. So there is a philosophy in, so we had that luxury and that's the reason, I think that's the main reason why Vitesse succeeded, where others have failed, is because nothing that would not scale could be even written in Vitesse because the next day YouTube would go down.

58:52Right. Right. And we would so we have created a few outages due to those things. So if something was not going to scale, we would know right away. So we had that luxury. So nothing that went in with us, everything that went into with us had to work at scale. So that was that was one thing. And the advantage we have is we have those learnings. So we we know what not to do. Sure. But don't you want to test it as you go and make sure that you're doing it right? On the testing side, what we over time developed, that is something we had the luxury of not having to do in the early days at YouTube, is we had to develop really, really, really, really thorough tests.

59:36The reason was because at some point of time, we would be making changes in the Vitae source code. And we had to be confident that people are going to take this. Slack is going to take this and run it. GitHub is going to take this and run it. It's going to come back into planet scale. And there is no waiting until it goes live to find out what is going wrong. So we actually, in Vitesse, we had an extremely rigorous testing policy. Extremely strict with 100 % code coverage, performance tests, backward compatibility tests. so that's actually one of the it also makes it hard for an engineer to contribute because you change five lines of code, now you have to write all these tests to make sure that you won't break anybody but the confidence is people now have developed the trust that if a test releases code, I can just take it and run and it will not fail so I'm hoping to bring that same philosophy here which is to test thoroughly to the extent that we are confident that this will work.

1:00:53And we'll obviously still ease that into production. Could you actually bring the test suite over or at least parts of it from the test? I mean, because that's a pretty robust test suite. If you could port that somehow relatively easily. Anything we can copy from the test, we are going to. Yeah. Definitely, yeah. But there are some parts which will just copy. There's no reason to change anything there. So those parts will just copy. Yeah. Have you looked at deterministic simulation testing, similar to what Terso is doing? And what was the other database, Chair, that you had called out me? Tiger Beal.

1:01:30Tiger Beal, yeah. Have you looked into that? And is that translator? Is the test space? I need to look this up. Okay. But in VITES, we probably have the opposite, which is called the fuzzer. Is that the same thing? Uh, similar. Yeah. Similar. So we do, we do run the fuzzer and the fuzzer does not miss. As far as I know, uh, it, it misses nothing. From, from, I, I wasn't there when they implemented it at planet scale. It was, uh, someone called Vicent who did that, uh, test suite. He said that, um, uh, the fuzzer was so good that it found bugs in my SQL. That's good. So, so I have, pretty high confidence in that, but yeah, whatever it takes.

1:02:16Well, we have a couple past shows we can recommend for you as well that might give you some insights to just some of the behind the scenes, not exactly that and how to implement it, but more fodder for your requests. Yeah, yeah, definitely. Well, it's all very exciting. I assume the licensing is going to be similar to Vitesse or similar to what Superbase normally does. Is there any sort of gotchas in there in terms of the open source side of it? No, Apache is what we are planning to do. And yeah, the test was actually BST once upon a time until we joined the CNCF and we changed it to Apache. But yeah, it's going to be Apache.

1:02:55Does Multigrass make sense in the CNCF as well or no? It's an option. It's definitely an option. we even thought of as either a fork or could be a sub-project of Vitas but then I felt that that will restrict its freedom right now like starting it as an independent project these are very difficult decisions and I think once we make these decisions and bring it to a good shape, we could look at actually moving it back to CNCF as its own project. Yeah, I think you want the space to reinvent and rethink without feeling at all burdened by your past decisions with Vitesse. And I think obviously copy and paste the stuff that makes sense, but don't make it a fork where you're basically starting at this foundational, which you may end up wildly different in implementation or you might be the same, but at least give yourself room to come there organically, you know?

1:04:03Yeah. I can't like, uh, uh, that is something that was pretty obvious to me. Uh, although the fundamentals are the same. Uh, so you can, even if you don't copy code, you can copy ideas pretty easily. So those, those don't change. But, uh, when you get into the details. There's so many small things. And the thing is, if you, uh, uh, if you don't do them the Postgres way, uh, that will keep nagging you, you know, you say there is, there is this thing where, you know, I have written this, there's this if statement hanging there, which is there only to differentiate between Postgres and MySQL. Right.

1:04:41And so you don't want any of that. None of that. Yeah. That kills the joy part, right? It takes away the joy that you're trying to enjoy what you guys are building here. Well, it's very exciting. Obviously you're just getting started. So there's still a lot of question marks for us, Postgres fans. It's just a lot of waiting, seeing, you know, watching maybe the GitHub project as you put it out there and the announcements as the core team comes together, you know, you're get the, the post, the Vitesse Avengers, you know, get the whole team together. But what else, what else is on your mind or things that are important for folks to know before we let you go?

1:05:20The thing I feel is, I feel a bit anxious that there's not much activity in the open source project. But, you know, we may see cool, but we are paddling hard underneath, really, really hard. So I will be actually publishing a few blogs about some of the thoughts. There's one area which we didn't cover much of, which is the high availability and the consensus part. So actually, my initial blogs would be covering those. I feel like this is something, the database world kind of went a little bit astray with mounted storage, disaggregated. I feel like that needs to be brought back in. I feel that storage should be with the database because database are IOPS hungry and having your data local to your database is very important.

1:06:24And that story can be complete only if you have a good consensus story where your data will not be lost if you lose your node. So that is something that we want to solve really, really well in Postgres. so that people, because I have seen the good days of the database, where database goes screaming fast. Like Slack and GitHub, they run hundreds of terabytes of database size. They run millions of QPS, and their latency is like millisecond, one millisecond, two milliseconds. That's the kind of latency they get. And they can't tolerate any higher latency. And I want to bring those days back into Postgres.

1:07:11So that is kind of what I'm shooting for. So are those things that Vitesse tackled or did not tackle? High availability. Those are, yeah, those are like pretty, very core to Vitesse. Yeah, I thought so. So those parts we are going to bring into Postgres. Interesting. Adam, anything on your mind? Any other questions? I'm thinking about Neon, honestly. I'm thinking about, you know, just this kind of bigger picture. I'm curious if this shake up, this change in the database world, like for a while there, we had obviously PlanetScale leading, you know, massively scale MySQL, which was great for MySQL, but not so great for the Postgres folks like us.

1:07:52And then we've had Superbase, obviously. We've had Neon. Neon was acquired by Databricks. I don't know how that's going to shake out for a product that's usable by the public or if it's a Databricks thing. like this seems like a good time for Superbase to do this because there's change amongst the Postgres world. And I feel like this is the clincher really for Superbase to finally have this kind of feature set that wasn't available unless they built it. And why build it in a non-Postgres native way, in a non-open source way, given Paul Copplestone's way of thinking as a leader there. I'm just curious your thoughts on that Sugu is just the thought of like Neon, Postgres, Superbase this shakeup, you coming back out of retirement basically to avenge this scenario and shard for life with Multigress help us unpack this shakeup I suppose are you excited?

1:08:54Obviously you came back but how does Neon's change change the landscape for Postgres specifically to open the the floodgate for Superbase's mission with Multigress. Yeah, going back to why I started this project, right? The reason why I started this project was because I believed, I still believe actually, that there is no scalability solution for Postgres. Not Neon, not Aurora.

1:09:26Like I've been, I'm an advisor to Metronome. which is a Postgres shop, they have one of the largest Aurora instances and they are struggling. They have no way out beyond where they are looking at moving data out of their Postgres instance because they've hit the limit. And now they're saying, oh, we need to do something about this scalability thing. And that's the reason why I wanted to start this project. And one of the things I thought is with Superbase, if not just like, so this is a problem that is meant to solve it for the industry overall. In other words, Neon could use multigress if they wanted.

1:10:14If they hit the limit of a machine, you could deploy multigress on Neon and run it unless there was incompatibilities, for which you have to make changes. But even in those cases, the changes would be minimal. So that's what I came for. And for Superbase specifically, is that if Superbase is going in its trajectory, there are going to be users that are going to max out a single instance very, very quickly. As a matter of fact, I won't be surprised if some are already hitting those limits. and their way out is you could go to Aurora maybe and that will extend your runway by maybe at best 2x. I'd be surprised even if you got like a 2x runway.

1:11:09As soon as your QPS doubles, you pretty much are at your limit. So the only way out of this is a sharded solution and that is why i started this project i i said that you know like in my sequel if you hit your limits you could go to vtess but there is nothing like that in postgres and vtess for postgres at even mediocre level of functionality will allow these people to stay in postgres otherwise you have to find a way out of postgres once you hit that limit well after you accomplish the test for Postgres, maybe you can build the test for SQLite. No, I'm just kidding.

1:11:57You know, for your second after your next sabbatical, you know, the only thing that gets your back is Dr. Richard Hipp approaches you and says, Sugu, we need you. We need you to SQLite. That reminds me of I think airplane, was it? where there's a background scene with a movie poster which says Rocky 83 and a really, really old guy is standing there. Yeah. Airplane. You're going back in the day into the 80s to mention movies. I haven't seen Airplane for years. I mean, like, don't call me Shirley, okay? Exactly. Good stuff. Yeah, good stuff. Suga, thanks for coming back from retirement for this.

1:12:45I know that, I mean, considering what you just said, that there's no way out, that sharding is the only way. I mean, you're coming back at the right time. I think you're coming back to the right team with the right motives at the right time. And, you know, somewhat full disclosure, Jared and I are both angel investors in Superbase. So we kind of have, you know, this desire for our own upside, but like technologically, this is going to be amazing for Superbase to essentially be one of the only leading at scale Postgres flavored, Postgres for life, open source for life flavors out there to choose from.

1:13:21And so I'm super excited about that for you and for them and for Postgres because we need it. It's amazing. I am excited too. Get building, man. Get building. Get building. All right. Thanks, Sugu. Bye, Sugu. Thank you.

1:13:37well friends it was awesome talking to sugu diving deep into the test diving deep into what that means for the test for postgres aka multigress all the work is going into it and what's to come for us postgres maxis postgres for life but yes we are also in denver this weekend if you're in denver around denver want to hang out with us on friday or saturday for the live show learn more at changelog.com slash live. Big stuff is happening. Don't miss it. Come hang out. And it's also free to attend for Plus Plus members. It's better. You know what? It's better. That's why it's better. changelog.com slash Plus Plus.

1:14:17Big thank you to our friends at Code Rabbit. Check them out at coderabbit.ai. Our friends at Depot, depot.dev. And of course, Fly. Check them out, fly.io. That's it. The show's done. We'll see you in Denver.

1:14:41Thank you.

1:15:10Game on.

From the publisher

Sugu Sougoumarane, creator of Vitess, comes off sabbatical to bring Vitess to Postgres. We discuss what motivated Sugu to come off sabbatical, why now is the time, the technical challenges of doing so, the implementation details of Multigres (Vitess for Postgres). We also discuss the state of Postgres at scale.

More from The Changelog: Software Development, Open Source

All 232 episodes
Bringing Vitess to Postgres (Interview)The Changelog: Software Development, Open Source · 1 h 15 min
Listen in VO