Distributed databases with Peter Mattis

30 Sep 2026 · 1 h 42 min · 44 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Distributed databases and storage systems, centered on why B-trees matter, plus Peter Mattis’s career history (GIMP, Gmail storage/indexing, Google build systems, GFS/Colossus) and his views on AI-assisted coding and testing.

Guest

Peter Mattis, co-founder and CTO of Cockroach Labs. Background includes building the original storage system behind Gmail (message threading, indexing, and storage), contributing to Google’s distributed file system evolution from GFS to Colossus (metadata scaling via Bigtable; Reed-Solomon erasure coding), and earlier work on GIMP and related graphics libraries (GTK). He also worked on build tooling at Google (Google 3/buildfiles leading toward Blaze/Bazel/Buck concepts) and later contributed data-structure work influencing Go’s runtime (Swiss-table-inspired hash map improvements).

Key claims

  1. B-trees recur across real systems because they efficiently support ordered data, indexing, and metadata needs.
  2. AI can produce high-quality production code, but AI agents are “lazy” about testing; the fix is easier than it looks.
  3. Storage/database performance is constrained by hardware realities (HDD vs SSD/NVMe, network latency, cache effects), so “speed of light” napkin math matters.

Notable examples

GIMP (built in college), Gmail launch (April 1, 2004) with message threading and B-tree-backed unread/thread counts, Colossus metadata stored in Bigtable to remove GFS master bottlenecks, and Go map/hash-table performance work inspired by Swiss tables.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Peter's Journey into Technology

1:58 to 3:20

Peter discusses his early interest in technology and switch to computer science.

“I wanted to get into how did you get into tech originally?”

The Birth of GIMP

3:20 to 5:24

Peter elaborates on creating the GIMP image editor during college and its impact.

“Well, I mean, the big thing that I did, along with my roommate in college, we had this course, was it a compilers course?”

Lessons from the GIMP Experience

5:24 to 6:30

Peter shares insights from his GIMP experience regarding competition and innovation.

“And I just took a little lesson from that.”

Transitioning to Google and Gmail

6:30 to 8:43

Peter recounts his journey to Google and contributions to Gmail's development.

“Somehow one of them, you know, looked me up, asked me to come interview with Google.”

Building Gmail: Challenges and Innovations

8:43 to 14:00

Discussion on the technical challenges and innovations behind Gmail's launch.

Building the Google Build System

14:00 to 15:05

Learn about the evolution of Google's build systems and their design challenges.

“But it also kind of drummed up this excitement.”

Introducing Colossus: Google's Distributed File System

15:06 to 18:19

Discover the design principles and capabilities of Google's Colossus file system.

“And what gconfig spit out at the end was a monolithic makefile, but one you didn't have to write.”

Erasure Coding and Its Benefits

18:20 to 21:56

Understand the concept of erasure coding and its advantages in data storage.

“So in GFS, and there's varying costs as far as like, well, you're not storing two copies or one copy of data or two copies, you're storing three, triplication, three times as much storage you're having to use.”

Metadata Management in Distributed Systems

21:57 to 23:18

Learn about the role and challenges of managing metadata in distributed file systems.

“There's the Colossus using the big table, and then there's the normal big table sitting on top of Colossus.”

Latency vs. Throughput in Distributed Systems

23:19 to 26:24

Explore the balance between latency and throughput in distributed file systems.

“And to me, it's always a bit conflicting of like, well, it's pretty easy to, I guess, build a distributed file system with my limited knowledge.”
Show all 44 chapters

The Speed of Light and Data Transmission

26:25 to 27:25

Learn about the fastest ways to transmit data and how physical laws impact technology.

“or if you're playing a game, maybe you need to have like, you know, frame rates of like every four milliseconds.”

Optimization Techniques in Modern Systems

27:26 to 28:00

Discover techniques to improve performance in multi-threaded and distributed systems.

The Complexity of Performance Tuning

28:00 to 29:26

Explore the intricacies of performance optimization in programming.

“The high frequency trading, I mean, they do this all day long.”

Innovations in Data Structures

29:26 to 33:39

Learn about data structures like B-trees and their applications.

“You made some contributions to the standard library, right?”

Contributions to Go and Hash Tables

33:39 to 38:13

Understand the evolution of hash tables and Peter's contributions to Go.

“Hash tables are like the, the, one of the earliest things you've learned about in college and data structures.”

Peter's Journey After Google

41:08 to 42:05

Hear about Peter's transition from Google to startups and his motivations.

“and why he left Google after building Colossus.”

Founding Viewfinder and the Acquisition by Square

42:05 to 43:30

Learn about the founding of Viewfinder and the circumstances of its acquisition.

Technical Foundations: Colossus and Spanner

43:30 to 45:12

Discover the differences between distributed storage systems and databases, focusing on Colossus and Spanner.

“And how is Spanner different to Colossus?”

The Birth of Cockroach Labs

45:12 to 47:18

Explore the journey of founding Cockroach Labs after experiences at Square.

“I mentioned earlier, he was working on the GIMP.”

Naming the Database: The Cockroach Inspiration

47:18 to 48:36

Learn the story behind the unique name CockroachDB and its intended resilience.

Customer Experiences and Use Cases for CockroachDB

48:36 to 50:34

Hear about how companies have adopted CockroachDB and the incidents that led them to seek it.

“The campaign was performance under adversity.”

Understanding Sharding: Manual vs Automatic

50:34 to 53:10

Dive into the concepts of sharding in databases, comparing manual and automatic approaches.

“We felt the burden for that belongs on the database developer.”

The Ubiquity of B-Trees in Databases

53:10 to 54:48

Discuss the role of B-trees in database structures and their importance in indexing.

“you can imagine all your keys in a system.”

Strong Consistency in CockroachDB

54:48 to 56:00

Explore the concept of strong consistency in databases and its significance in transaction management.

“So CockroachDB offers strong consistency.”

Understanding Isolation Levels in Databases

56:00 to 58:12

Learn about different isolation levels in databases, focusing on serializability and eventual consistency.

“And the benefit of this, the serializability is, again, it's a very simple model for the application program.”

The Evolution of Database Performance

58:12 to 59:26

Discover how database technology has evolved to improve performance while maintaining application simplicity.

“meaning as soon as you're writing it, when you're reading from the database, you already get the written value back.”

Consensus Protocols and Raft

59:26 to 1:01:43

Explore consensus protocols in distributed databases, with a focus on the Raft consensus algorithm.

“Yeah I mean consensus protocols the original consensus protocol is called Paxos and it was famously hard to implement.”

Mission Critical Applications Explained

1:01:43 to 1:05:17

Understand what constitutes mission-critical applications and their importance in various industries.

“And one of the things that is just a general truism of distributed databases is you don't want to have a lot of back and forth.”

The Transition from Coding to Leadership

1:05:17 to 1:07:18

Learn about the journey of moving from hands-on coding to a leadership role in a tech company.

AI's Impact on Coding Practices

1:07:18 to 1:10:02

Discover how AI tools like GitHub Copilot are changing coding practices and improving developer productivity.

Innovations in Software Engineering

1:10:02 to 1:12:03

Explore the evolution of coding and how AI is changing software development.

“And then I just remember having this feeling of like, well, the code is just like materializing before my eyes.”

Harnessing AI for Innovation

1:12:04 to 1:14:48

Learn how AI tools are facilitating innovation in database management.

“Well, it started with this tool that was like this kind of benchmarking tool that tested this matrix.”

The Future of Software Development

1:14:49 to 1:17:38

Discuss the future trajectory of software development influenced by AI.

“Like, I think this is a thing where like a tech lead who has no management responsibilities, but they have a group of interns, but they don't need to do with their performance, with their anything.”

Current Technologies and Frameworks

1:17:39 to 1:19:59

Discover the tools and frameworks being utilized in modern software development.

“Earlier this year, we kind of rolled out this internal platform where non-engineers could write kind of mini applications.”

Quality and Security in the Era of AI

1:20:00 to 1:23:33

Understand the challenges of maintaining quality and security with increased output.

“I can't remember what the codex thing is called, but they have various ways of spinning up graphs of agents and prosecuting them.”

Addressing Industry Challenges

1:23:34 to 1:24:00

Examine how to overcome common challenges faced by the industry due to AI integration.

Challenges in Testing and Security

1:24:00 to 1:25:14

Discussion on the difficulties of maintaining rigorous testing and security practices in software development.

Evolving UX Practices

1:25:14 to 1:26:01

The conversation shifts to how UX design processes are becoming more integrated with coding practices.

“It's like, oh, that UX element is wrong.”

The Future of Code Review

1:26:01 to 1:28:26

Exploration of the changing landscape of code reviews in the age of AI and automated tools.

“I think it's pretty, it's starting to get a bit controversial.”

Cognitive Shifts in Coding with AI

1:28:26 to 1:31:24

Discussion on how the presence of AI is changing the coding experience and cognitive processes involved.

“I think the model kind of likes it because it can work within that guardrail.”

AI's Role in Amplifying Domain Expertise

1:31:24 to 1:34:06

Analysis of how AI tools can significantly enhance an engineer's domain knowledge and capabilities.

“I mean, I sometimes think it's like, you know, everyone's in this kind of spectrum of capability of software engineering.”

Learning and Growth as Engineers

1:34:06 to 1:38:01

Advice for engineers on leveraging curiosity and AI tools to enhance their skills and knowledge.

“but I feel like the analogy that comes to mind to me though is the domain experts, they're a little bit like sorcerers in their particular domain.”

Harnessing AI for Productivity in Software Engineering

1:38:01 to 1:39:24

Learn how AI tools can amplify coding productivity and quality.

“You said, like, you wouldn't believe, but my coding output is insane and it's high quality and it's database quality level.”

Peter's Journey with AI and Engineering

1:39:25 to 1:41:19

Explore Peter's transition back to coding and the impact of AI on his work.

“You know, the stuff you might have had to compromise on in the past, and you can take away some of those compromises.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00He built a Gimp image editor while at college, designed the original storage system behind Gmail, and built many more large, complex, and widely used systems. This is Peter Mattis, co-founder and CTO of Cockroach Labs. Before we sat down to talk, Peter told me, For the last 30 years, I've always been a prolific coder, but my current output is a bit insane. And this isn't Vibe Coded Junk, but database-worthy, high-quality, high-performance code, thanks to working with strong coding models. Today we cover why B-trees are so important when building databases, and why Peter kept reaching for the data structure over and over throughout his career.

0:31How he wrote 100 ,000 lines of production code per year pre-AI, stopped coding between 2022 and 2024, and why he is back now. Why he thinks AI agents are lazy about testing, and how this is an easier fix than it looks. And many more. If you're interested in distributed databases, distributed storage systems, or knowing how Peter manages to use AI to produce unusually high-quality and production-ready code, this episode is for you. This episode is presented by TurboBuffer, a ridiculously scalable fast and cheap hybrid search engine built on top of object storage by an engineering team I've really gone to like after spending time with them.

1:05The TurboBuffer engineering team is doing something really really cool. They're completely redesigning their storage architecture from first principles to make storage faster, cheaper, and more reliable at scale. If you follow TurboBuffer, you know that their storage architecture was a massive part of their early success. Redesigning a winning architecture is a big deal. It's one thing for your query plans to pass unit tests. It's another to bring each query plan to performance parity or better while maintaining correctness and reliability in production. Here's the cool part. TurboPuffer is documenting the whole thing.

1:36Their new storage architecture, which they're calling Teepuff V3, demands some hardcore systems engineering and they're building it in public, sharing design decisions and benchmark results as they ship. They're keeping a work log of their journey and the first post just dropped today. follow along at turbopuffer.com slash v3. That is turbopuffer.com slash v3. Peter, welcome to the podcast.

1:59Peter Mattis:Oh, I'm happy to be here. This is awesome. I wanted to get into how did you get into tech originally? When did you figure out computers are interesting? I figured that out in kind of elementary school, high school. Gaming was a little bit of a gateway drug for me as a software engineer as many other people. I remember like early on, my mom did programming at IBM at some point. I'm never even quite sure what she did, but we had computers always around our house like apple 2 plus apple 2gs i'm from that era on up and you know go to the live uh the bookstore i'd find a book on basic or magazine type in programs no idea what it was doing but you know just kind of like i was addicted to like you can produce put stuff into these computers you get interesting stuff out and then i got to college and i was like i didn't think there was any money in computers i didn't know anything about it i started as a mechanical engineer following the footsteps of my dad you started mechanical engineering as your specialization.

2:48Yeah, as my major. Yeah.

2:49Peter Mattis:And I got in there and I was like doing, like I'd done some computer stuff before. It was a real foolish move to do this. But first semester of doing these homework assignments in mechanical engineering, they were god awful, like six pages for a single problem. And I happened to take a CS course at the same time. And it was so easy. And then everybody almost was failing it. I'm like, I'm in the wrong, I'm in the wrong field. Let me switch. And then you switched. Then I switched. Yeah. What was the first software that you built either at college? it must have been a college, that you were like, all right, this is a piece of software that I'm kind of proud of.

3:20That's a complete piece of software.

3:21Peter Mattis:Well, I mean, the big thing that I did, along with my roommate in college, we had this course, was it a compilers course? I'm not quite sure anymore. This is like 30 years ago. And we were kind of bored with it. So we wanted to do something like kind of fun on the side. And I'd done journalism in high school in my junior and senior year. I knew stuff about kind of computer graphics and wanted to do something like Adobe be photoshop so started just kicking the tires and we built up this program that a lot of people know of called the gimp and along with the gimp i did a lot of the the graphics library gtk this has since evolved just massively since then it's kind of interesting because after college i kind of stepped away from it didn't really stay involved much past my first year out of college but definitely people still know me and it uh it led to other some interesting uh events in my career and and it was like starting gimp it was literally just you saying all right i wanted to do something like photoshop how hard could it be and then you just this wasn't through what you learned in college right this was like you figuring out how to build you know like a graphical i guess engine rendering drawing data structures all of that stuff all that stuff figured it all out i remember trying to look at some papers back then my roommate was looking at papers we're just fingering it all out and it's one of these things you almost kind of need to be a little bit naive to start anything like that because if you know how hard it's going to be you would never start it so i think there's it's like you have to have this level of like if you're really truly wise about the effort you like never get into it but then when you get going it just keeps on snowballing it gets far further and further along and then at some point you're like wow this is really awesome but there's an interesting little tidbit associated with our work on the gimp which is kind of a little bit known um i think we've talked about this before but it's just kind of fascinating where we were uh getting to the point where we're like we should release this to the public and back then there was uh usenet news groups where people would post stuff and there's one on graphics and remember like just a couple weeks before we were going to release our first version of the GIMP, someone else came on there like, I've been working on this graphics program and it did everything the GIMP did, every last thing and then some.

5:18And we're just like, oh, that sucks.

5:20Peter Mattis:I guess we'll just keep on working on it. It's been fun. And then we release the GIMP. Never heard from this other guy again. And I just took a little lesson from that. There's always going to be someone else working on your idea. You can't get dissuaded if they pre-announce it. Nothing ever comes of it. And a lot of the marketing behind some of that stuff that you might hear is like, I mean, I don't know if we like, we stole his thunder. I'm not even sure what happened. Never traced down what happened, but there's a real lesson there. So there's a possible future where you read this announcement where someone's saying, I'm going to build all of this thing.

5:51And you go like, ah, you know, someone did it. And you kind of go back and just do something else. And the game never happens. That's right. Wow.

5:57Peter Mattis:Yeah. I guess, especially with today with startups, you know, it sounds like just do your thing, put it out there at the very least right the advice i give people is like you might have a unique idea but most likely there's like a dozen people out in the world who've had the same idea and there might be a couple of them working on it but a lot of people don't even work on it there is like whatever they can't get going on their idea so you know i wouldn't be concerned at all if like you hear someone else working on your same idea that's probably the case you know we're working on some cool stuff now my current company i guarantee you there's other competitors out there working on the same thing and just know it's like it's a competition you got to enjoy that aspect of it not be afraid of it and what were some fun things that the gimp led to well you know um i got kind of tired with graphics and that's why i kind of moved away from it after college i got into storage systems and i bounced around uh you know i worked at one of the early search engines inked me and then i i got over to another startup and um i met through that time period in that first startup i was doing um met larry and sergey from google because it turns out that larry i think it was maybe sergey i'm not quite sure the very first version of the google logo was done in the GIMP.

7:01Peter Mattis:So they knew of us. Somehow one of them, you know, looked me up, asked me to come interview with Google. I did back in 2001. And, uh, this is when Google was three years old. Yeah. Three years old. And, uh, I said, no, you did not. I said, no. Yeah. No, here's the calculation in my mind. I was like, Google's down in Mountain View. I was living up in San Francisco and I didn't want to do the commute. So I, uh, went and worked at another startup for a year. And, uh, that one, after a year, I like, I saw it wasn't going anywhere. and they called me back up and they said hey do you want to interview again i was like sure like we're not gonna be able to offer you the same stock options you got before i'm like okay i can't remember what it was and i still can't remember what they offered me the first time um but it probably would have made a lot more money if i had taken that first but i did all right i'm not like complaining but it's one of these things like yeah now i think back about that i'm like oh a lot of good stuff came out of it some point along the line red hat was going public they actually offered friends and family stock during their ipo to a lot of people we got offered i made a little bit of money off that not a lot but it was like just after college it was actually quite significant at the time you know so i mean there's a number of good things that kind of came out of that that and i also used the gimp as a you know i was looking for like oh i can't report photoshop and here's the gimp and it did so many things so like i'm sure there's so many people like you had a positive impact of being able to use this thing and the nice thing that i what i really liked about it it is it was free but i wasn't like stealing any software from anyone you see what i mean this was at a time where free software was not as common as today free and open source was not as mainstream as it is today yeah so the other cool thing i've heard from a lot of people because people still like i mentioned this like oh i've used it and then some people you know software engineers like i learned how to program by looking at your code and i'm just like wow that was the code i wrote 30 years ago and i wasn't nearly as good as software engineer as i am now and then you got into google second time around you you you said yes what did you start to work on yeah well i got in there and they said like hey you know this was early day 2002 was april 1st auspicious date um your start date yeah april 1st 2002 and i got in there and they're like well you know we're gonna actually build email google email and it wasn't called gmail at the time there's a code name internally it's called caribou caribou yeah um and got in there i started working on it and i was kind of tasked with working on the the back end threading and message storage and indexing system and worked on that you know pretty hardcore for the first year and a half actually i was on to i think it was three years um on it through the launch which actually happened to be april 1st 2004 do you remember this there was a uh it was like considering april's full joke yeah i remember so correct me if i'm wrong but the the launch said gmail with one gigabyte of storage or something or infinite storage and like star one gigabyte i'm not sure what it was but back then most email providers would give you about 10 megabyte of storage the free email providers and then you could pay for maybe 50 megabytes or 100 megabytes but that was really expensive this launch it looked like an april fool's joke because who could possibly give you a free email service with 20 to 50 times or 100 times more storage we were talking about this internally it was like kind of a shock and awe campaign for the industry i think it was actually four megabytes for like hotmail or yaki mail and then you had this and not only that we had uh so much more storage but it was indexed really fast so you know you could do a search you know come back almost instantaneously but can you tell me can you tell me internally like when the project started okay we're gonna do email for google how did you and the team arrive to the point of like okay we will offer all this storage and we'll do it fast because this was at a time if you can take us back but what i remember is hard drives were still expensive they were relatively slow we're talking hdd we're not talking necessarily about ssd if i remember but if you can take us can you take us back of what what it was like what the constraints were like and then how you innovated like to actually like do something that's never been done before yeah yeah so at the time google internally had this large distributed file system called gfs google file system yeah and uh we were like basically looking at numbers and being like yeah we think we can build on top of this they They had a lot of knowledge about how to do search and retrieval.

11:08Peter Mattis:It ended up being that like some of the existing systems they had for search and retrieval, they started building the prototype on there and then that was completely rewritten. That was actually what I got involved in because I got in there and there's already a prototype and then it was completely rewritten. So I was on that, you know, threading side. We decided to have message threading right from the get go, which is also kind of innovative. That wasn't common in email systems. And threading to have the message thread, which I assume needed data structures on the back that it needed like you know storage and and figuring out the read intensity of those kind of things yeah you know there's uh some b trees involved in that i think that might be the second time i implement b trees and i think i've implemented them like a dozen times now how do b trees relate to email threads uh it's just like you have to you know have some storage there where you have a thread id a message comes in you have to look up it was using the search index to match actually the the subject of the message id into the thread there are some other things taking place in the b-tree as well because we were keeping track of the unread counts of of threads and whatnot so you know i can't even remember all the details now this is ancient history this is 2004 what are we in 2026 22 years ago so it's like left my memory but there was definitely b-trees there was also this you know inverted index taking place there all this code has since been completely rearing what were the economics like in the team you must have done the economics of being able to offer free email which of course i'm sure you did some maths of like how it could be subsidized but to like not make a terrible terrible loss yeah yeah yeah no i mean we were kind of uh kicking around ideas and um you know one of the ideas that had come up at some point was hey maybe we can put ads on this and it was one of these crazy things so i didn't wasn't involved in this this is paul bukite went on to do some cool stuff at uh friend feed and facebook and ended up i think a partner white combinator at some point but he just like one night he's like no i think i can just take some of our existing ad functionality and incorporate it in there and it just went like gangbusters and there's a huge business that you know kind of grew up from that which is kind of incredible so i mean i think there's a lot of things that it's just like you kind of have a sense of what can be learned like oh we think this is going to cost you know per user a couple dollars per year how do we monetize that and you know we didn't want to charge for it we eventually they did charge for it because you have the whole google workspace stuff but at the time it was more like ah can we do this for free oh it looks like it can be economical and uh it took a little bit of a leap of faith and do you remember the launch uh the reason i'm asking because i remember that it was an invite-based system like not everyone could get in and i assume that must have been to control the expected demand because again you were you're offering something for free that was paid before and it was kind of pretty obvious that there would be massive demand how did you think about it kind of monitor demand decide how many people to onboard Honestly, this is a team effort.

13:50Peter Mattis:And I wasn't involved in that. I mean, I do remember being worried about like the load that would happen. And then someone came up the idea of like, hey, we should do an invite based system. And this also had like it played dual roles. So it kind of constrained the growth. But it also kind of drummed up this excitement. Oh, you can get me a Gmail invite. I remember people asking me at the time was like, yeah, I can get you as many Gmail invites as you want. And then after you built the system, where did you move on to? What was your next project? Was it the build system? Well, there was the build system.

14:18Peter Mattis:I mean, that was kind of all muddled in my mind because I was kind of doing that part-time when I was doing Gmail as well. At some point, you know, Google actually had this large monorepo. I think they actually still have the monorepo. They still have the monorepo. And they started out with the Google One repo. That was before my time. Then they moved on to Google Two. That was what was there when I got to Google. And at some point, we saw the strains of Google Two. And Google Two was just one single monolithic makefile. I think it actually had some submakefiles, but it was this really large unwieldy makefile.

14:45Peter Mattis:Someone came to me and was like, I think we could do something. and I'd had some interest in the build systems. And I, you know, kind of put the foundation in for Google 3. There was a bunch of people involved. I was kind of doing like kind of the initial work on Google 3. And Google 3, the initial insight was like, hey, makefiles kind of sucked to write. We introduced this thing called buildfiles. And I decided to do it as just this stripped down kind of Python language. But it was still Python at that point. And what gconfig spit out at the end was a monolithic makefile, but one you didn't have to write.

15:13Peter Mattis:Over time, this evolved. It became Blaze internally. There's some other systems that are associated with now. I don't even know how complicated it's got. I haven't seen it for a long, long time. Then that became Bazel externally. It became Buck. There's some other, you know, people went to Facebook and they're like, that was awesome at Google. What was the reason that MakeFiles or Make didn't really work? Was it, are we talking about build performance? Are we talking about maintainability or readability? Yeah. So like, I think make itself is like, it's okay in declaring dependencies, but it's a little bit like kind of assembly language.

15:48Peter Mattis:And you didn't want to write, you know, all your dependencies in assembly language. And if you didn't know what you were doing, it was easy to make a mistake and miss dependencies and whatnot. And so the kind of my, the thought behind the build files is like, hey, you need to express these dependencies, but at a higher level and just like kind of cleaner semantics associated with it. And from that, you can compile down into the assembly language. And then eventually folks were like, oh, you don't have to compile down to that. we can just kind of implement you know the dependency kind of update engine directly and then that's where the performance improvements were able to come in so like because i google scale like when you have a large repo in general like my understanding is that the reason basil and buck are so popular for large code bases is it can help you improve your build performance it gives you a lot more lever levers to play around with from caching from being smart about cache generation to obviously just raw performance yeah that's exactly it yeah so you kind of just dabble and like okay i'll make this build file did that people took that over but what was your next main focus oh well so i mentioned gfs earlier google file system and at some point we realized that there were some limitations in gfs scalability bottlenecks because i've been working on the storage system for gmail i actually dabbled in another storage system you know kind of a research thing they never went anywhere um but because we're working on that i'm going to participate in like the founding team of colossus which was the successor to gfs and as far as i know colossus still exists it's gone through multiple iterations at google but it's like the second generation you know distributed file system so what is colossus yeah so when i say distributed file system externally you might think of something like s3 kind of blob storage it had a flat name space kind of like s3 you give names and there's like a minor hierarchy there but it's like very limited but it's not like a posix file system so you don't have the full directory hierarchy you didn't even have the full like kind of permission system a lot of that stuff got it added later but files are not stored on your local machine there is a fleet you know kind of a service out there that has all the files they're writing it down to their hard drives and now ssds and your client can access that and it's all replicated so if there's any crashes and whatnot you're not losing your data i don't know when um at some point s3 got erasure coding we did erasure coding in classes that was kind of a big breakthrough which coding we used re-solomon um so for the audience who's not familiar with erasure coding you might think of like i want to have replicas um and there's a this thing in hard disks called raid where it's like i don't actually have to have full replicas i can actually you know if you take a plus a and b you can x form together and then you can have this kind of third kind of version and there's more and more complicated versions that reed solomon is kind of you know i think there's actually better codes now but it's like one of the known ways to do the known ways yeah yeah and we we had to you know kind of pioneer internally like oh how are we gonna actually make this work in distributed file system where you focused on latency, on being able to store data more efficiently.

18:33Peter Mattis:It's storing it more efficiently. So in GFS, and there's varying costs as far as like, well, you're not storing two copies or one copy of data or two copies, you're storing three, triplication, three times as much storage you're having to use. And with Reed Solomon, you can get that down quite a bit lower. I can't remember offhand exactly what we used for Reed Solomon, but I think it was like essentially 2x. But you get that with also the redundancy too. So it's like, it's smaller and the redundancy is higher. And so, because I guess the naive thing, if you're saying, all right, I want my data to be replicated at three places, you take three nodes, three machines, physical machines, and you say like, copy one, copy one, copy one, I'll have it three places.

19:10Great. If one explodes, I still have two. Wonderful. And then you're saying that the algorithm here is you could take not three X the data, but two X the data, split it smartly across machines, or maybe you could take it and take it lower, and you still have the thing where like, oh, one of them explodes. I still have all my data because it's split in a...

19:28Peter Mattis:That's exactly right. And like kind of the mental model, you know, if you just want to understand at a high level, which is essentially you might want to say like, I want to have eight replicas of this data. I think I can't remember offhand. I think S3 might use nine. They've actually talked about this publicly. You have kind of nine chunks of data, but any five of those chunks can be used to reconstruct it. And what this means is you can lose any four copies and you can still reconstruct your data. And oftentimes it's more like, you know, the first five chunks are exact replicas and the other four are kind of parody ones.

19:58Peter Mattis:I've kind of forgot some of the details that escaped my mind, but it's along this line. But when you come up with an algorithm, you can then prove that this algorithm will work, right? Like this is a little bit like, I know in software engineering, like maths and algorithms is a bit out of fashion. But in this case, this is really important because once you can prove it, that this algorithm works, it will work. It will work. And, you know, it's like the math behind here is like Gawah fields over like GF2, something like that. I don't even know that. I never actually understood the full math behind it.

20:27Peter Mattis:I always regretted not doing more math in college, but you didn't have to. You like Reed and Solomon proved how this work. I think it's like back in the 1970s, something associated with like communication network. So you just take that and, you know, kind of use that expertise, but leverage that and have to do all the engineering behind it to make it work in a storage system. And then when building a distributed storage system like Colossus, what were other things beyond, okay, you want to store data in a resilient but efficient way? What were other things that problems that you needed to solve?

20:56I'm thinking things potentially like sharding and re-sharding or metadata being important, those kind of things.

21:04Peter Mattis:Particularly for the scale that Google wanted to operate at, GFS kind of had this scalability limit. I believe it was like, you know, you could have a thousand machines in a GFS cluster and it kind of... Which I guess sounds big until like today, it's kind of relatively small, right? It sounds big for most people. And then Google was like, no, we need to have this scale up to 10 ,000 machines. And there's some bottlenecks. there was this GFS master as a single node. It was a bit of a bottleneck. We're like, ah, we need to have a distributed master to store the metadata for all the objects. And Google at the time happened to have this system called Bigtable.

21:35Peter Mattis:And so Colossus stored its metadata inside Bigtable. And one of the things that I'm kind of proud and kind of like also, you know, a little bit embarrassed by, one of the design choices I went down, but actually worked out, is we wanted to use Bigtable for the metadata for Colossus. And the metadata, just for those of us not as into this reason what is the metadata and a distributed file system yeah it's like the uh the names of the files and for each of the files the files are broken into chunks what were they 64 megabyte chunks and then you have to have the list of chunks for each file yeah and then you you have to um periodically you know the master has to be scanning over this and doing repair um repair work and you know but there's more metadata than that but that's it in a nutshell so we wanted colossus to have this you know kind of scalable you know service big table to store metadata the biggest user of gfs is bigtable so we want bigtable to work on top of classes so so you have the circ look defendants yeah well uh there's a bootstrapping thing right bootstrapping yeah oh yes yeah which one starts up like you need to mock something somewhere right yeah no i mean the way it actually worked at the time and they've since replaced this you know because this is like the way to get started and leverage what you have and eventually you kind of get rid of it but there is the uh kind of a foundational bigtable that bigtable didn't use Colossus.

22:50Peter Mattis:There's the Colossus using the big table, and then there's the normal big table sitting on top of Colossus. And it all worked for years, so I don't even know when they got rid of that. They got rid of it at some point, but that was... But I guess it sounds like you can make hacks that go really long knowing that they're hacks and they get you off the ground, right? They get you off the ground, right? Because if we had to implement that kind of big table layer from the get-go, it would have just delayed how long it took to get Colossus built. One of the things that strike me about a system like Colossus is it promises, or this was internal to Google, but even distributed file systems that are external, they will promise high throughput, high availability, and low latency.

23:33And to me, it's always a bit conflicting of like, well, it's pretty easy to, I guess, build a distributed file system with my limited knowledge. I could probably do something where I have either high throughput, but high latency, because whenever I write something, I write it out to all the replicas. You already mentioned one technique of doing it, but how did you kind of reconcile? How did you get low latency while you have high throughput, while you also have a replication going on on the file system?

23:58Peter Mattis:Well, I mean, these distributed file systems, I mean, the latency isn't super low. in particular when they're running on hard disks, which Colossus was doing at the time, which S3 does, you actually notice the latency. So like S3 is a high performance system. Google GCS, the competitor from Google, which is built on top of Colossus. It's high performance, has incredible throughput, but the latencies are like 20 to 30 milliseconds for first read. And that is bounded by your hard disk latency. If you put it on SSD, it gets down to closer to SSD latencies, but not actually kind of the state of art SSD latencies, which is kind of this crazy thing that's been happening in our industry, it's just like how much faster the hardware has been getting.

24:34Peter Mattis:Reading from a hard drive, maybe five to 10 milliseconds nowadays. I never touch hard drives. Reading from an SSD over NVMe, 30 microseconds, 50 microseconds. So that's a microseconds. There's a thousand microseconds in one millisecond. So we're talking like a huge, huge difference. Well, there are now startups or infrastructure companies that are starting to take advantage of the fact that they can, they can have an NVMe layer and they pull things up either predictively or not, But as you say, like when the physical reality changes, you can build systems on top of it that that should take advantage of it.

25:06Peter Mattis:Yeah, absolutely. You know, the disks are so much faster with SSDs. The networks are so much faster. I mean, just kind of crazy how fast like the intrazone leaks these are to Google or Amazon Center. And it wasn't like that. But when we were building Colossus, you know, I can't even remember what the numbers were, but milliseconds to do a network roundtrip. And now it's down in, you know, 100 microseconds within a zone. I mean, I just look at these things and I'm like, holy crap, you know, the hardware guys have really done a good job. Yeah, sometimes I feel that our software should feel way more snappy.

25:37And there are some snappy software, but sometimes I almost wonder if we're getting too complacent with all the abstractions or not even doing this like napkin maths. Simon Erickson at TurboPuffer talks about this napkin math word like you. It was like, all right, here's the theoretical limitation of the hardware. Reach from SSD might be, I don't know, 30 microseconds. and then like how can i build a system that is as close to this as possible as opposed to the other way around saying okay like you know a human will notice like 20 milliseconds like 100 milliseconds let's build around that yeah yeah and sometimes like when you're architecting something

26:10Peter Mattis:you need to think about the human kind of perceptible latencies but oftentimes when you're dealing with the storage system layer you know i know simon working on turbo puffer they're doing great stuff over there you have to think about the machine scale and the machine speed which is a lot faster than human perception. Like a human can tolerate maybe 100 milliseconds of delay, or if you're playing a game, maybe you need to have like, you know, frame rates of like every four milliseconds. But the machine wants much, much faster than that. Someone says, napkin math, I call this speed of light numbers.

26:37Peter Mattis:And like sometimes it's literally the speed of light bottlenecking you. The cross zone latencies between zones, cross region latencies is speed of light and fiber. You want to hear a kind of crazy fact? I love hearing crazy facts. The fastest way to send um pack of data across the globe is to send into space is it because the speed of light is faster in vacuum quite a bit faster or well it's not fuel vacuum no way yeah so so you cover a higher distance yeah well you actually uh i believe the the way to do this you send it straight up and then with starlink you send it straight up you bounce around between starlink you send it down to the other side so you want to get it out of the atmosphere as quickly as possible possible but now if we're ever talking with speed of light this will also go there's like a digital transformation happening so you would need to calculate how long it takes for that system to to process and do it and and but you're saying that even if you do this like really well it will be faster than beaming it through an optical kit and an optical cable slows down the speed of light right yeah yeah yeah the speed of light is only the speed of light and vacuum in every other medium it's slower i was talking i did a deep dive on the hedge fund industry and they didn't tell me exactly they said that they do use satellites and microwaves and some of these things that they will not tell you because you know this is their thing but i had a suspicion that they might have found a faster way and i think this is like somewhat well known but the the details they're not going to get into because again that but like yeah so yeah they're probably balancing stuff in space yeah yeah so i mean one of the things that they do in the high frequency train is like between new york and chicago it's actually not far enough distance wise to make it worthwhile to send it in space so they were doing microwave beams there yeah but i mean if you really want to get faster you need to built like a vacuum tube between them and it's like send a maybe they're doing that yeah but i mean the thing that i think is like just kind of awesome about performance nowadays is i mean there's so many layers you have to be paying attention to in terms of performance that the rabbit hole just goes so deep there you're almost certainly running on multi-threaded systems right and you're like oh i need to have a multi-threaded program i need to have synchronization in there well you get the best performance if you just kind of avoid the synchronization and part of this is lock-free programming but part of it is arranging so you don't need locks at all you need to carry about your processor caches and there's this whole like kind of setup of caches above the cpu like you can think of registers as a cache and you have l1 l2 l3 even your memory and onto disk and there's just you have to pay attention to all those levels and if you do your performance gets way better and if you ignore it and you're like oh i'm not worrying about kind of kind of you know kind of the cache accesses i'm just accessing data all over your program will just be way way slower And some folks pay a ton of attention to this.

29:11Peter Mattis:The high frequency trading, I mean, they do this all day long. And in so many other places, like no one pays attention to it. And you get this kind of gradual degradation of the performance of the hardware or the software. I did want to talk a bit more about low-level stuff, but not about the speed of light, but low-level data structures and programming language features. You made some contributions to the standard library, right? Well, I've done a couple. Not quite standard libraries. So, I mean, I just like have been always fascinated by data structures. It's kind of awesome. I mean, I think just the algorithms in general, you're just like sorting algorithms.

29:44Peter Mattis:They're just kind of awesome. You could probably explain, you know, insertion sort and, you know, kind of everything. Could you explain quick sort as well? Quick sort's a little bit harder. Right, that's the thing. Yeah. No, no, but it's a smart one. Yeah. And then you get to these levels of like, you know, it's like, oh, wait, someone really smart came up with this. so one of the things i worked on just as a little bit of a side at google at some point was a colleague came to me and he was like you know what we're using the stl map structure all over the place and the stl map structure is a balanced binary tree i can't remember if it was red black trees or one of these other balancing algorithms that many cs college students implement and he came to me he's like i think we could do better because you know there's actually a cache problem here every time you every node you're traversing down you're going to a different cache line he was thinking about using something else a skip list to do this um which um that's another awesome awesome data structure everybody should kind of look at but at some point i was like actually this feels more like a b tree so i implemented b trees a couple times before and figured out like how to implement a b tree that implemented almost all the semantics of the stl map it couldn't quite do it perfectly and the reason is uh when you insert into a b tree node you have to shift stuff around so you don't get pointer stability this is just kind of fundamental but if you can you know you don't need that for your use case you actually can pack more data in so the thing about a btree like the the real easy way to describe this is like you just have a small list of items like eight items the best way to store that if you want to you know kind of fast access in sorted order is just to sort the items right literally not to have a tree at all yeah for like eight items just have it in the very simple list very very just an array, sort the array, and then you can either do a linear scan over it, you can do a binary search, and oftentimes linear scan is faster.

31:27Peter Mattis:And then you think about that, I'm like, well, if I want to store one eight items, I can just have one node that has eight items, and then I have another node, and then you have a parent node that connects them together. And that is essentially like the, you know, you build it, you think about building it bottom up, you start with just one node of eight items. Oh, I need to insert the ninth thing, I'll split into two chunks. And the two chunks, one will have four, the other will have five, and then you have a parent node that points to them, and then you just kind of recurse on that that's the b tree algorithm in a nutshell everybody go implement it actually nobody should implement this anymore because nowadays we we have something else that will implement this in all the optimizations because there's a crap ton of optimizations that you can do on the b tree but but i just want to go back to this like there was already an existing implementation for maps and then so your your colleague looked at the code and said we i think we can do better what i want to figure out is like in my mind again someone sitting outside of not involved in how some of these libraries or data structures are built.

32:25I always thought, and again, this might be naive, but really smart people sit down, they kind of look at the state of the art, they implement it, and there's no way it can be faster. In fact, I've had arguments in the past saying, like, oh, let's write a faster sorting thing. Surely it is the fastest. But if you could bring us a little bit of how you've been inside how it happens and how other people like yourself and and your colleague can say like oh what what if we what if we try

Read the full transcript

32:52Peter Mattis:something else yeah yeah so i mean my recollection here is he was working on this kind of the big um internal system i think it was called gaia that actually had the mapping from you know you log in you have your user id and you have to look this up and they were storing you know all like this map from user id and email to whatnot to the metadata about the user in stl maps you just know it's like well there's a lot of memory usage here and it shows up on profiles and then we're like well what can we do to do better and that was kind of the genesis of it and you know he happened to be working on it and he happened to be working with me like we just started kind of noodling on this problem like oh can we do something better and it's not one of these things like i think now with google they have a whole team working on their kind of internal libraries at the time it was more of like you know everybody working on their own systems and contributing to a shared base but but i guess it still goes back to what you were just saying of like just go down the layers try to understand and if something just doesn't add up like suddenly like oh there's this big explosion memory usage like you know just ask the questions what why is this and if you're able to or you happen to be like oh can we do something about it right and one of the things that you know he observed earlier on i think part of one of the things was it was like a map from integer id to something else and like if you look at red black tree every node you have your kind of value that you're storing the map and then you have two pointers you might have an integer id that's like four or eight pipes and then two yeah and you look at that and you're like oh that seems like a lot of overhead and you're just you could use a bit well the bean tree actually has a lot better um it has better spatial locality and that's what made it faster um but it was actually smaller as well at the same time because you had less pointers involved you also contributed to go right yeah well that came later yeah yeah it came a lot later but can we talk about that yeah i just one of these other things you know i pay attention to like you know when there's research papers coming out about new data structures and like hash tables.

34:41Peter Mattis:Hash tables are like the, the, one of the earliest things you've learned about in college and data structures. Like how do I map keys to values where the ordering is unimportant? That's when the hash tables come in. And there is like, you know, the, the very earliest ways to do this. I've implemented hash tables multiple times. It's like you take your key and it might be a string and you, you put through a function and it spits out an integer, and then you map that into an array of buckets. And if multiple things map to the same bucket, you have to have a link. You have a linkless. Yeah, this is a naive implementation, right?

35:09Peter Mattis:Naive implementation used quite frequently. And over time, people discovered a lot better ways to do hash tables. That's called chaining of your hash. There's another technique called open addressing, where instead of actually having a linked list, you just kind of hash it again and move on or kind of walk down to subsequent buckets or subsequent slots to find out where you should be. And I remember reading about this new technique. It came out of some folks at Google. I believe it came out of their Swiss office. It was called Swiss tables. I believe that's where the naming came from. I'm not 100 % sure about that, but I remember reading about it.

35:45Peter Mattis:And then I was working on Go for a long period of time. And Go has this built-in map structure. And it's a hash table. It's a very highly optimized hash table because the Go team is very competent. The Go runtime team. And various folks have taken an attempt at putting together a Swiss table implementation for Go. and i tested some of them and i was like this is kind of fascinating what swiss tables do and i'll explain how it works in just a second but i looked at it's like well it's really hard to beat the performance of the runtime the runtime was really good and i kept on a new lift on this for a little while and uh eventually i ended up having to take this business trip to uh to india to bangalore and so i was on a long flight no yeah it always starts like this and i just like i'm just gonna to try to pull on this.

36:29Peter Mattis:I pulled on it sufficiently that I could get some of the benchmarks to be faster. Wow. And then I'm like, you know, that like is kind of like catnip for an engineer, like kind of make it all faster. Can I figure all the rest of it? You know, got some help from the runtime folks. There's a, an issue on the Go issue tracker that, you know, where other people have been attempting this because people propose like, hey, let's use Swiss tables and like, like, well, you know, we don't quite know all the details. You're going to have to navigate this and that. And there were some ideas there that combined them together and got to the point where it had kind of a complete implementation that was faster on most benchmarks, not quite all of them, but most of them.

37:02Peter Mattis:And then the Go folks eventually picked this up and pushed it over the finish line. And then so you came with the idea, you got to the point where you were able to show an implementation that showed how some of the benchmarks were faster. And then you started to work with some folks on the Go team too. Well, it wasn't quite that. I came up with an implementation that we ended up using at my company. It was good for our use case, but actually putting it into the runtime whole other you know kind of level but but then you just showed like here is this implementation and then they they kind of took the inspiration and the ideas yeah they're like well this is great we want to make all they always are looking for ways to make it run time faster and there's like you know oh wait and then so so you did this in this contribution this was a lot of years after you left google right yeah this was from from the outside this is from the outside awesome yeah and other people contribute stuff to the outside as well you know we had another colleague um he contributed one of the CRC implementations, you know, adapting some stuff.

37:55Peter Mattis:CRC? Cyclic redundancy checksum. You know, Intel published some papers about here's how to do the CRC and assembly very fast and contributed one of the implementations. You see a number of those things where, you know, people just like are contributing externally. It's not a lot, honestly. I mean, I actually don't know the full details, but, you know, people are regularly contributing to these things. Peter just described how the Swiss table work came together using an issue tracker in the the Go team pushing all of this over the finish line. This is where I need to mention our season sponsor Linear, which is a place to coordinate work between humans as well as agents.

38:30One thing I've noticed about how most of us work with agents is how it's a pretty single player thing. You open a terminal UI, go back and forth with an agent, and it usually produces a PR. But the rest of your team has no idea what happened in that chat unless you tell them or copy the whole history. And when everyone in the team works like this, a lot of work happens that's invisible to the rest of the team. Linear's take is that agent work should be teamwork. Even today, teams already use Linear to define the work to be done. Now, Linear can already delegate an issue to a coding agent. This agent could be an AI agent that Linear integrates with like Codex or Cursor, or Linear's own agent, or a custom agent.

39:05Either way, the engineer in delegating stays responsible for the outcome. What I really like about how Linear works is how the work stays visible. Your teammates can follow the session of the agent, check out the PR it produces, and join the review. We've gone from single-player work to multiplayer engineering work with agents. Oh, and one more thing I like about Linear, a focus on costs. Linear agents' autorouting chooses a model that is the best suited for the task. Teams can also inspect usage and set limits, track usage, so you can use capable agents without having costs balloon out of control.

39:35Hop on board at linear.app.pragmatic. Peter also previously mentioned Gaia, Google's internal system that mapped logins to users. It's not surprising that Google custom built all systems, including this one, but most of us won't build our internal Gaia. This brings us to our seasoned sponsor, WorkOS. You can think of WorkOS as something like Gaia for the rest of us, identity infrastructure you've otherwise spent quarters building yourself. WorkOS includes single sign-on, skim directories, sync, audit logs, role-based access control, basically everything a big customer security team asks for, delivered as a handful of clean APIs.

40:08It's how companies go from we have a login to we can sell to a fortune 500 without standing up their own internal identity platform and work was is already building for the next version of the agentic authorization problem their newest product is airlock the authorization layer for ai agents think about what happens when you hand an agent a task like clean up the sale opportunities in our pipeline the last thing you want is for this thing to have standing permissions to delete whatever it likes airlock sits between your agent and the tools they call. It evaluates every request against the agent's intent and your rules and then it allows it, denies it, or routes it to human for approval.

40:46You write down the policies in plain language, the agent never sees your credentials, and every call and verdict gets logged. It works with coding agents like Cloud Code and Codex and with MCP gateways. So if you're working out how to let agents do real work in production without over-permissioning them, take a look at WorkOS Airlock at workos.com slash airlock. And with this, let's get back to Peter and why he left Google after building Colossus. You're at Google, you're building Colossus distributed file systems. You're at this point probably working on probably the largest system on the planet, honestly.

41:20Why did you even consider leaving?

41:22Peter Mattis:Yeah, yeah. Well, after Colossus, I kind of dabbled in this other project called Google Goggles for a little while. Remember the glass holes? Yeah. It keeps coming back, the idea, by the way. so yeah yeah now it's still here and present it seems like google's early yeah yeah and you know i think that was a technology before its time i don't think it was ready to do at that point and looks like the uh actually doing the glasses is quite a bit harder you know the android phones we were trying to power it on were you know not powerful enough and then you know i just kind of got wanderlust you know like you know am i just kind of stagnating here google which is a strange thing to say but you know some other people feel it as well and decided to go off and try my hand another startup that didn't work out we got aqua hired by square so yeah i want to pause for a second so this company was the company name uh the company that we founded is called viewfinder viewfinder yeah it was in the mobile photo sharing space which should sound familiar this is like instagram this is like snapchat this was in 2012 yeah right as as instagram and all the more we're taking off yeah yeah no we were right there in the play and uh we just didn't have the the right go to market kind of like how to attract the users how to get viral growth kind of thing because from the outside like what i read when i when i checked you know the story just like oh you know like you you you co-founded a startup it got acquired by square hooray like it sounds like you had bigger ambitions and this was a decent outcome but not the the dream right yeah no it wasn't dream at all so uh the term i used was aqua hired so sometimes a company will get acquired get bought for you know their ip for their product for their business and other times they get bought just for the talent the people the people and we got bought just for the talent they acquired the ip but i don't think square ever did anything with it it wasn't kind of like where they were working but we built up a kind of strong technical team and that's what we were hired for and you know like i can't remember the details we'd raised a small amount of money we were able to pay our investors back make them whole maybe they got a little bit of a haircut but maybe they actually got a little bit of a but it was it was not they were okay essentially they got their money back which is like you like you know as a founder you uh you know investors are big boys they're used to losing their money but you kind of feel bad if you lose a lot of money for them so you know getting them paid back kind of makes you feel a little bit better yeah yeah and then you you spent a little time at the company that acquired you were just square yeah and then you started itching a little bit again yeah yeah yeah because you know we've been working on these you know distributed file systems and storage systems you know colossus uh one of the kind of sister projects to colossus is Spanner.

43:48And how is Spanner different to Colossus?

43:51Peter Mattis:Well, Spanner is essentially a distributed database. Colossus is a distributed storage system. And the way I think about the difference between a distributed storage system and a distributed database, you might think, oh, they're both storing data. I was about to ask. Yeah, yeah. Because a distributed database will at some point be a storage system. Right, right. So what's the difference? Yeah. So for Colossus, it was targeting large files, large append-only files. You can't update them. Append-only, yeah. Yeah, large append-only files. So, you know, 64 megabytes, maybe up to gigabytes in size.

44:23Peter Mattis:But if you're a database, you want to be storing like kind of small, like, you know, kind of, you know, if you're using SQL or like the relational data, you might have a table with, you know, billions or billions of, you know, rows. Those rows are broken up into columns. The columns are typed. That just has a very different nature to the engineering challenge for a database than it does for a distributed storage system. And usually distributed databases are implemented on top of some kind of distributed storage system. And that was the relationship. So Spanner was implemented on top of Colossus.

44:52Peter Mattis:And some of the design decisions in Spanner kind of directly fell out of the append-only nature of the files in Colossus. You can't update a file in place, so you have to make the files immutable in your database. And this is where log-structured merge trees come into play. and they weren't invented at google um but google really popularized them with level db which emerged out of the work on big table and spanner um that got popularized into rocks db i subsequently re-implemented one of these things and this is what we use at cockroach labs it's called pebble so i'm very familiar with the internals of that but it's kind of all based on this idea that the the data is kind of stored in immutable files so how did you decide to found cockroach labs while we were working at square um my co-founder and i it was actually three of us we were all at Square and one of them is Spencer.

45:41Peter Mattis:I mentioned earlier, he was working on the GIMP. He's my college roommate. And he was also at Google. He was also at Google. This other guy, Ben Darnell, who was also at Google, he joined us at Viewfinder and ended up at Square. We, you know, like we're kind of just noodling on a project to do. And we'd actually had this design back in Viewfinder. We were like, ah, we didn't, we looked around for a data base to be using in order to build Viewfinder on top of. We didn't really like the things that were out there. The technologies that said Google looked better we had the big table we had spanner and whatnot we were looking around you know like each base existed but i wasn't quite happy and there's some other systems like react and others and you know at one point we're just kind of like you know came up with the design for um cockroach to be an initial design and i was like no no guys we're doing a mobile photo sharing site we shouldn't build a distributed you know database so we put it on the back burner which i think was absolutely the right thing maybe or maybe we should just pivoted away from doing the mobile photo sharing site given the way things worked out and then we got to square and we saw some of the same problems that they were experiencing with data storage systems and spencer is very convincing managers convinced some of the management that like hey you can just work on this part-time you know to see if it had life behind the the design and then kind of conscripted uh ben and i into it and eventually started getting you know attention externally and we're like hey can we go and spin this off into a company and that's what ended up happening and so you started a company but i understand you didn't raise vc funding initially right that was the case at viewfinder at uh we did differently at viewfinder we kind of eschewed the vc money and you know in hindsight i wouldn't recommend that um but you wouldn't recommend yeah i wouldn't recommend taking the vc money because my experience the vcs are very very intelligent they can help you navigate a lot of challenges you know i think sometimes there's this perception you know the vcs will push you into various areas and maybe there's some bad ones out there that do the vcs i've had experience with just like some of the sharpest people you know i've ever met and so you kind of have like an extra like like person helping you on the team pretty much mentoring you giving you guidance telling you what they're seeing they give you advice seeing what they're seeing in the market where things are going that is very hard for sometimes for a founder especially the technical founder right that you're focused on the engineering part yeah yeah no we actually took money right away for cockroach labs um it was almost like as soon as we left we got did a little road show you know kind of in the the bay area and got some interest and you know got an investor right away i have to ask about the name though yeah how did the name of cockroach yeah i mean uh so you know we named the gimp yes that was uh that was mine pulp fiction come out in college and like which we named this thing oh the new image manipulation program i think we're thinking image manipulation program initially and i'm like oh gimp it's obvious and it just stuck at some point you know we're newly on the this new database and you kind of want to give things a name um you can't just say oh we're working on this distributed database you kind of need something and this master was like oh cockroach db like cockroaches are unkillable i want these things this database to be unkillable you know cockroach is going to survive the nuclear apocalypse apocalypse so that was where the genesis was and just stuck yeah we're we're whenever a bunch of our nodes goes down this thing will still be up yeah yeah and you know this is where we're at today with cockroach db it's like one of the things that i point out i'm like holy crap this is awesome you kill a node we did this whole campaign last year which was really just to prove out something that It had already been present.

48:56Peter Mattis:The campaign was performance under adversity. But just like you can run a workload against it, you can kill a node, you can sometimes kill a whole region and the system keeps on going. And it's like stories like that, you know, like what we did on the marketing side there, but also we hear this from our customers too. They've had fires in data centers and all the other data systems go down and CockroachDB keeps on going. I'm like, that's awesome. When you started out, who were companies, startups that who wanted to use CockroachDB and how has it changed since? because, you know, like just making the case, like, okay, I'm starting a startup.

49:28It's a small startup. Like I will need a database and I'll, I don't know, I'll typically choose a Postgres, right? Because it's free. Everyone's using it. I'm running it on Node. At what point did you see that typically tech companies are like, okay, like this is not enough for me that it's running on a Node either because it can go down or because I'm outgrowing it. What was it they are growing? Like I'm trying to get a sense of like, at what point did companies say like, tell themselves like we need something distributed yeah in a database yeah i mean oftentimes we have companies calling us up after they've had a disaster so so it's like a node went up or a

50:06Peter Mattis:hard drive failed that kind of stuff you know we it's not quite like we're ambulance chasers but if you see an outage like a big outage from some you know company you know like we'll sometimes be trying to knock them up but also they they uh will call us you know i mean like there's a very big bank who's now a customer um don't think i can name them but you can go read they had a very serious outage due to a weather event and after that which probably knocked down i'm assuming a region or a database or an networking cable or tree fell on something it knocked down a whole region you know it was a region-wide power out you knocked down the region and there's a mandate from the ceo it's like no we just have to you know be able to survive these things that gets pushed down all the way and you see this in other places where you know one of our early customers they're running on AWS and they just got to the maximum size you can run in Aurora instance on and then what would typically happen at that point is then you have to shard your database you know this is very standard practice you you take your single node database you create 10 or 20 or 100 shards and this is what Google done for some period of time and that's a heavy burden on the application developer and the way we always phrase this is like I mean the application developer is becoming a database developer at that point and they're doing it poorly you know they're trying to implement distributed transactions or indexes and whatnot.

51:18Peter Mattis:We felt the burden for that belongs on the database developer. Can we talk about automatic sharding? I think it's safe to some, most of us will know what sharding is when you're, but actually let's start from like manual sharding and then how you can implement automatic sharding. And if you can tell us like, you know, tactics that a database like CockroachDB can do to actually just take that load off of you. Yeah. Yeah. So I think the very basic form of sharding is a little bit like the hash table. let's say, you know, you have a fixed number of shards. Like, let's say it's just 100 shards. Your data model is a user with a lot of data associated with the user.

51:51Peter Mattis:You just take the user and you say like, oh, they map into one of the shards. And, you know, you kind of just rely on the hash function to get like fairly even distribution. The problem with this is at some point, you know, one of your shards will get full and you have to kind of re-shard. And that's a very, very onerous process. And re-sharding, I guess, simple way to do is like, if it's just a hard drive, I don't know, per node, where you write the user data, gets it gets full and you're like okay well i now need to split it somehow i need to move it i need to remap it i need to rejave my metadata which knows where this data lives yeah yeah and depends on exactly how you're doing that mapping from like you know the user id or whatever your shard key is to the shard you might have to remap them all right this is very typical so exactly so i mean this happens in hash tables where you know oftentimes in order to grow the hash table you just have to essentially create a new hash table double the size and copy all the data over now That's kind of like the very basic straightforward way.

52:42Peter Mattis:And there's various levels of complexity you add on it. One of them is called consistent hashing. And there's various techniques to do this. It's kind of fascinating, like how they all work. But in consistent hashing, you can add an additional node, and then it only moves a fraction of the data from each shard over there. There's various systems to do that. And I believe this is like 100 lies Cassandra. The way CockroachDB does it is more into Bigtable, more into Spanner, more into HBase, where instead of actually hashing, we actually take, you know, you can imagine all your keys in a system. And this is always true in any system.

53:14Peter Mattis:You can imagine them just in one big contiguous key space. And then you kind of partition contiguous spans of that. And then you have to build up an index on top of those contiguous spans. And what I just described there actually sounds a lot like a B-tree. So there's this index on top that is like that maps you from, you know, like I need to have this range, which node is it on? And this is a little bit like a B-tree. You know, he kind of squint, you know, it's like, I think you squint and everything's either B-tree or it's a hash table. And, you know, that index structure. But it's now starting to make sense because when I remember when I read, it might have been the Wikipedia article on B-trees.

53:47It said this is data structure that is frequently used in databases. It's now coming back to me because I didn't think too much of it. I'm not someone who builds databases. But now that we're talking, we just organically keep touching on B-trees again and again.

54:00Peter Mattis:And the other place it comes up in databases. So this is where it kind of comes up in distributed databases. and no one ever really calls it a B-tree. I just kind of squint sometimes and I see like actually kind of a B-tree. But the other place it comes up in databases is for your indexes. So if you have a table and you have like an index on your email address and you want to be able to scan over those email addresses in order, that's a B-tree under the hood in a database. And, you know, like any kind of index you have usually provides sorted order. There are hash indexes, but oftentimes it's the B-tree index and they're ubiquitous in databases.

54:32Peter Mattis:There's actually a paper called the ubiquitous B-tree and basically just identify that. I think that paper was written back in the 80s and they're still ubiquitous today. They're the foundations of single node databases and pretty much every data system I've worked on has had B-trees at some point in place in them. I want to ask about strong consistency. So CockroachDB offers strong consistency. Now for people who are a bit more newcomers to distributed systems, can we talk about the consistency models and then why strong consistency is important and why it's hard to implement it? So, I mean, there's multiple ways to kind of approach this, but I mean, if you've used a database, you probably heard of transactions.

55:09Peter Mattis:And transactions are a way to perform a whole bunch of mutations atomically. So databases talk about atomicity, consistency, isolation, durability. The durability is really easy. It's like when I write data to the database, it has to be durably written. So if anything crashes, it comes back. The atomicity is just referring to the fact I want to do a whole bunch of changes. I want them all committed or all aborted at the same time. I don't want to have like some kind of partial operation. And why is this important? Why is the atomicity important? Well, the atomicity is what gets you to the point where it's like, as an application, I can do a bunch of operations.

55:40Peter Mattis:And if there's an error and an error occurs, it all kind of gets rolled back. And it's a much simpler development model to work within. And then there's the consistency and isolation, which kind of get, you know, kind of muddled. The isolation is referring to isolation between transactions. I don't just want to run one transaction at a time. That's easy to do, right? I want to run a lot in parallel. Oh yeah. and make it so that when they're running in parallel they are running as concurrently as possible but you want to have the appearance when they're running as concurrently as possible that there is kind of a serial order to them so this is like kind of the whole trick the kind of the gold standard for isolation is called linearizability don't worry about that there's a step down from that is called serializability and that literally refers to having a serial order of your transactions but you're having to construct that in a way that you're doing everything as concurrently as possible.

56:29Peter Mattis:And the benefit of this, the serializability is, again, it's a very simple model for the application program. They don't have to worry about weird kind of defects occurring in their program. And some of the ones that can occur, it's like the classic description is of a bank, right? Where I might want to read, you know, like have an operation that reads and says, like, do I have$100 in my bank account to transfer somewhere else? And you could like arrange for lesser isolation levels that you might be able to subtract that$100 twice. And that's bad, right? You know, we want to keep accurate track of your bank account or, you know, what's in your shopping cart or, you know, it's kind of like the use cases are endless there.

57:07Peter Mattis:And you do this wrong and you have very egregious bugs. But now going back to weak consistency and strong consistency. Anything less than linearizability or serializability might be considered kind of weak consistency, but there's also like, you know, kind of strong consistency and eventual consistency. so the eventual consistency is like sometimes i can do an operation and it might not be immediately i might not be able to see all the updates the read results will not necessarily give me the current update but they'll come back you know at some point you know i've written part of it and the rest of it will show up at some point and oftentimes the easiest thing is your credit card balance right yeah yeah yeah and uh it's faster to do it that way faster in terms of just what the performance you can get out of the system but again it's a little bit harder for the application to deal with.

57:49Peter Mattis:One of the places this often comes up in distributed databases or databases that have any sort of replication is that I could write to the primary replica and then I read from the secondary and it's not there yet. That's eventual consistency. Yeah, that's eventual consistency. And you can often work around this, but just like puts bigger burden on the application developer because you have to pay attention to that. And then with CockroachDB, you have strong consistency, meaning as soon as you're writing it, when you're reading from the database, you already get the written value back. You get the written value back.

58:19Peter Mattis:You know, it's like you read whatever you wrote. It doesn't matter if you're reading from the same, you know, node you wrote it to. If you read it from another node, you still actually get the data you just wrote. Is the trade-off logically not that you would have higher latency? Because clearly to implement strong consistency, you would somehow need to, in the naive approach, you would need to write all replicas, right? Yeah. What are you doing inside Cockroach TV? Well, we are writing to all the replicas. Well, yeah. Yeah. You're just doing it fast. You're just doing it fast. You're making that efficient.

58:48Peter Mattis:I think this is one of the things that kind of also fascinates me about the software industry is we keep on finding ways to be more and more sophisticated in order to provide, you know, like do things that make it easier to write the applications, but do it in a very high performance way. And we've gotten, you know, very, very good at this over, over the years. And this is the area I know about the databases. This is happening everywhere. Like I'm just fascinated by how fast graphics have gotten, where when I entered the industry, like you were literally like writing out each individual pixel.

59:14Peter Mattis:And now you have these GPUs that are doing like billions of triangles per second, whatever the current numbers are and just like kind of astounding like how much sophistication I've got has gotten into every area of computer science wherever you look at it. One more thing on CockroachDB I want to ask about is the raft consensus. What is the raft consensus? Yeah I mean consensus protocols the original consensus protocol is called Paxos and it was famously hard to implement. Raft was you might think of it as a variant of Paxos it was kind of like an alternative to Paxos but in my mind today. And another consensus algorithm being that you have like a number of no's like three to a lot more and then how do you get them to agree on what do you typically agree on yeah yeah so you agree that the write occurred so like the way to think about consensus so you might think i want to replicate data and i i write it to a primary and write it to a secondary and you you can't actually have consensus when you only have two replicas and the reason you can't have consensus is if there's a crash and i come up like if i'm on the secondary how do i know if something was written to the primary if i'm on the primary how do i know it was written on the secondary right and you're always going to be in this kind of confusing place where it's like you either have to roll back a little bit or like you know you lose some data and consensus requires at least three but you can have consensus across more than three replicas and the idea with consensus is i'm going to write to three places and normally you don't actually when you're doing a read you don't read from a multiple of them but if there's a crash i have to do recovery then i'm reading oh i can read from uh two of the uh any two of the three and i know I can kind of determine what had happened previously.

1:00:44Peter Mattis:And it's usually just on the recovery time that you're actually doing that consensus read. So reads are typically just happening from one replica. You have to write to all three. And on a crash during that kind of failure is when the consensus read occurs. And inside CockroachDB, how many replicas do you choose for the consensus? It's typically three. For some system tables, it can be five. And then customers also have control of this. At the database level, you can write to five, seven. And five is like, you know, if you're really concerned, you know, about the data durability, you might use five.

1:01:14Peter Mattis:But there's a slowdown, you know, the more you're running to, it's like the more storage space. And obviously, like we're talking like slowdown with nodes. But of course, if they're like between regions, there's now you have a lot more resilience for, let's say, an earthquake or power hours or whatever. But now you will have additional latency, just speed of light basics, right? Speed of light latency, right. And it's, you know, tens, you know, or up to hundreds of milliseconds or even higher if you're going across the globe. So, you know, you have to be very careful with that in terms of how you architect your queries.

1:01:43Peter Mattis:And one of the things that is just a general truism of distributed databases is you don't want to have a lot of back and forth. You want to kind of like do all your reads in one kind of parallel read set, get them back, then do your writes. Right. But if you're doing like kind of serial operations where I read a row, I write a row, I read a row, I write a row. I mean, the latencies just add up. We talked about founding CockroachDB, but how has the company grown and where are you today? Yeah, yeah. I mean, we're being used in like, we're powering mission critical applications. This is our bread and butter.

1:02:10But by the way, can you elaborate on mission critical? Because it's like, if you're not, you are in the industry where you know what this means, but from the outside, it can feel hard to put a thumb on what is mission critical? Is my SaaS that is like showing as mission critical? Probably not.

1:02:25Peter Mattis:Yeah. So mission critical, in my mind, is like these kind of, the other term of art is tier zero applications the ones that are like just the the core crown jewels of what's running a company you know like a trading system you know your trading system can't go down if the trading system goes down this is kind of a critical problem or the firm that is running the trading system you know banking systems um but also you know like we work with doordash you know like some people might think that delivering your burrito is a kind of a mission critical system it certainly is for doordash right you know if that goes down is problematic we power shopping carts um you know other stuff like this where it's like well if the shopping cart goes down you know that company is losing you know hundreds of thousands millions of dollars per hour so that's kind of the criticality you think about yeah i guess of course that they're losing but but this is like when yeah their customers are also like they're used to this just working like running water and then when it's not the same thing as when your utility breaks right your water electricity is out you'll survive but it's not what you expected Yeah, yeah, yeah.

1:03:24Peter Mattis:And everybody's like, what? What age are we living in? The electricity goes out. We were kind of had this ingrained into our heads at Google. It's like Gmail cannot go down. People are relying on it. Search cannot go down. You know, if it goes down too long, people are going to move to other systems. And it's not like, you know, like in some ways search isn't as mission critical, except, oh, wait, every single search is ad dollars behind it. And you can actually notice the blip in the revenue. And it's not just the blip in the revenue. It's the blip in reputation as well. I mean, I think that's the one that really poisons companies is like, if your bank is down for a serious amount of time, the reputational damage there will be horrifically bad.

1:03:59Peter Mattis:And, you know, like we often talk about like, oh, and then grandma won't be able to pay her rent and she'll get evicted. You have to take this like responsibility really, really seriously. like no but also like just gmail being mission critical just on the way here we only exchange numbers later but we were communicating over email like oh i'm i was telling you that that you were telling me that you're here i emailed and i never for a second thought that it could go down and i think we were like responding within 30 seconds right and it's just i just know it's there like i i didn't like bother setting up a secondary communication line so yeah yeah it's like when people just like when you have that kind of level of trust with your users you got to maintain it and invest in it but then it leads to this kind of freedom for the user as well where you just don't think about like i don't have to worry about it it's just going to work and then in terms of the the company like how many engineers do you have roughly we have some hundreds i don't actually know the precise engineering number 110 but there might be 150 in r &d overall um there's other folks besides just engineers i'm clear as engineering managers those are engineers as well and then you know like we've been growing steadily it takes quite a while to build a distributed database not for the faint of heart yeah so it took a couple years a couple times yeah uh yeah well i did distribute storage system i did distribute database it's not for the faint of heart right so there's a lot of work getting it to a level of stability then a level of kind of quality beyond that level of stability um getting all the bugs out um and then continue to innovate and put more performance into the system adding functionality to integrate better within within enterprises a revenue has been you know kind of steadily growing over the years and you know it's at this place now where we see a path to future success as well and i want to ask about your coding habits so when you co-founded the company how much code did you write for the first few years i wrote a lot so i've always been a very prolific coder there were a lot of code early days and early days i mean i was uh kind of we were all technical co-founders ben spencer and i um and we were already a lot of code and i was no exception but you know if like i look back at my github output you know it's like kind of peak years maybe a hundred thousand lines of code in a year which is a lot yeah yeah no so uh we're talking pre-ai pre-ai right this is when you're back doing this manually right you know at some point you know we started out using a system called roxdb which is an lsm at some point i think it's back in 2019 you ran into limitations with it i decided um we wanted to rewrite it and did a big push to you know right that might have been like 40 50 000 lines of code and then a bunch of other people come up and helped and like you kind of look at that output i'm just like oh my goodness that was a lot to keep in your head it's a lot just to type you know 100 000 lines of code the average kind of like uh that the industry talks about is 3 000 lines of code from an engineer in a month and so if you multiply that out maybe 36 000 in a year that's good right so i was doing a lot you know i kind of look at that it's like there's kind of a max that you can hold in your head at a time um the tools have gotten a lot better since i first entered the industry we gotten better debugging techniques better testing techniques but still um quite significant and you were cto from from from the beginning co-founder of cto but there was a time sometime around like 2022 when you decided to kind of be a bit more hands-off right yeah yeah can you tell me about that i mean the general rule of thumb for uh engineering leaders is well you got to have your team you had to manage your team and we had a vp of engineering um but i was kind of getting to the point of like okay is my are my coding days done you know like i just direct from a higher level and you know i got this advice for a long period of time and i pushed back on it but you know i kind of acquiesced at some point and i think it was the right advice i'm not saying that the advice was wrong at the time um but there was time period from about 2022 to 2024 whereas like my my output declined i think i actually did the swiss tables thing in that time period but i wasn't doing much on the the core the business yeah the the core business you know i would get in there and do some work but like it's really hard that if you're in meetings all day to also do coding i think this is the fundamental tension if you like so you kind of took on the kind of the meeting burden the coordination burn and the yeah the the stuff that was if i'm reading correctly before you spent a lot of your head in the code and now you're spending a lot of your head like above the code the business the engineering the or the whatever customers that kind of stuff the customers and just being an executive as well you know so very hard to wear all those hats simultaneously and then i got back into it because ai started to emerge so how what when did you get when did you start using and when you start to find it useful or in terms of coding yeah it was interesting because you know those initial versions of like kind of glorified autocomplete came out and we're talking about the github copilot the cursor yeah the early one pilot that was the one i i had the first exposure to you we dabbled with cursor at the time but they're all like kind of glorified autocomplete and it was of crazy that you could just start typing something you know like build the rest of the function you look at you're like wait you kind of got that right this is crazy right and um you know we're trying to encourage our engineers to use this and at some point you know i can't remember if this is my idea or my co-founders or someone basically like you know like in order to like really guide people about how to use it you have to be a user yourself you know i think this is true of like engineering management in general like you want to like manage engineers you have to know how to be an engineer like if you if you don't know how to be a good engineer it's like really hard to manage other engineers i feel you'll have a hard time like just relating to them at the very least exactly exactly so kind of took it on me like no i mean this is like it was clear very early on like this is probably going to go somewhere but it wasn't quite clear how far how fast it would go and you start dabbling this and i was like oh okay well it's not quite good enough it's not quite good enough but you know let me start getting back into the coding very rapidly you know you started seeing the signs of life like you know kind of the opus models coming out it was sonnet first then the opus and you're like you're looking at these and you're like oh wow okay well they seem to be able to do quite a lot but you know the code is still not great but then it was just like the just this cadence of continuing improvements and you know i was starting to do a lot with sonnet and then starting to use opus and you know i had that same moment everybody else did and this was last year last uh november november december winter break right yeah it was thanksgiving i distinctly remember it because you know you were not chilling you were coding weren't you you were agenting i was uh agenting just like everyone else i mean i had this thing i'd wanted to do for a long period of time um on cockroach db which is like i wanted to like test like all these configurations of cockroach db across like you know different vertical scaling like how many cpus you have on a node how many stores you uh discs you have on a node how many nodes you have in the system and test across all this huge matrix of it this is one of these things they couldn't ever quite get prioritized appropriately because it never seemed like kind of critical but i've like i always had this intuition there was something there and then over this like four day span the code just like materialized as i was using um i think that was opus four seven or four five whatever the number was like the recollection i had is you know like a pretty fast typer.

1:11:08Peter Mattis:And then I just remember having this feeling of like, well, the code is just like materializing before my eyes. You know, you would ask for these things. I got out of the habit of actually typing and you just kind of like ask for something materializes. If you've ever seen like some of those, you know, a movie where they had this kind of archetype of a software engineer gets in front of the keyboard and they start typing, show the screen. It's just like, it's going way beyond human speed. And then it was gone that, right? Like it wasn't. I think as software engineers, we used to laugh at these.

1:11:37I still remember swordfish, the thing when they're visualizing the things are going, or the code appearing. You're seeing that the person's typing, and then it's a big line. And as software engineers, we're laughing because it's not how it is. But it's crazy, that effect, right? Yeah, yeah.

1:11:54Peter Mattis:But now it is. It's even better than that effect, right? Because it actually works this time. And it's that slow in comparison. It materializes faster than that. like you can literally go like we've been talking about b trees a whole bunch i implemented another b tree in in the past month you did yeah and it took about 30 minutes to implement probably 10 000 lines of highly optimized rust i mean it's just like it boggles the mind um i mean i we should look up later on 10 000 lines it's like a crazy amount you physically cannot type that fast you you start getting back to coding like are we talking about kind of vibe coding prototyping or are we actually talking you start to contribute like proper production ready code that is up the level of what you're doing at HomebookDB?

1:12:35Peter Mattis:Well, it started with this tool that was like this kind of benchmarking tool that tested this matrix. I am the CTO. We have an office of the CTO. The office of the CTO's mandate is to innovate. And I was looking for places like we can have innovation. And one of the things that we want to innovate in was, you know, better auto-scaling of a CockroachDB cluster. And kind of in the January timeframe, I came out with like what I would think is like kind of a research breakthrough, you might say. And it came about because I was dabbling in this area and just working with models, trying to understand it.

1:13:09Peter Mattis:And they're not just good at coding. They're also good at like helping you explore design ideas. And I think this is kind of the fascinating thing where it's you have to get out of the mindset of like, I know exactly what I'm going to build. But more like, hey, we have this problem. Talk through it. Be a partner with me. And be a sparring partner. Be a sparring partner. you know like there's this advice you might have heard that like you know if you're stuck on a problem you should go yellow duck it go rubber duck rubber duck it uh i think i've heard his yellow duck you know go just go talk to something you don't even need it to respond now you can talk to this system uh this intelligence and it will give you back stuff and it's not always right i mean this is the thing even today these models are fantastic fables fantastic astros fantastic and they're not always right but they they kind of like they know so much it's just encyclopedic and you can explore ideas super super fast and then be pointing out well that doesn't sound right to me and it'll be like oh yeah you're absolutely right you know like i hate that sick fancy it's like it kills me or your right to push on back on your right to push the back on that that's a new episode right yeah and yet like just able to make such fast progress um and i i i find it absolutely incredible so it quickly moved from just doing kind of side things of building up some tools and whatnot to building you know essentially over the last eight months it's january um we've been building towards a you know a new launch and you know like a lot of that has been produced a gently powered engineers and like the entire company's gone on board with this now where i think it's like probably 100 of engineers are using it to a greater or lesser degree and a lot of code is being produced and i think it's very high quality code these models will not test things adequately they don't look quite intently enough about performance but if you're like you can guide them in the right way and i think there's actually a huge advantage anybody's done management before it's a little bit like being a manager of people where you're a manager of a large group you're not looking at every line of code but you're definitely kind of helping architect the system i think there's a very strong analogy there i sometimes push back on this analogy the reason being i was an engineering manager my take is that working with these agents it's not like management because management has so much of the human stuff like as a manager when i think of all this a bunch of stuff i dealt with which was the people side of things the conflicts between people the performance reviews the meeting etc and you have none of that you do have the orchestration like again this is like almost like i guess a really like naive way of management where like you know they don't push back they they start to do it sometimes they're like unreliable but you know i almost like to use like orchestration a bit more because i feel management is is so much more involved.

1:15:45Like, I think this is a thing where like a tech lead who has no management responsibilities, but they have a group of interns, but they don't need to do with their performance, with their anything. Like, it's a lot closer to that, if you know what I mean.

1:15:55Peter Mattis:Yeah, I do agree. I use the engineering manager shorthand, but it's really about being a tech lead for like a 30 or 40 person organization, you know, or being an architect for one. I think that the term architect gives me a little bit of a distaste, but it's been an architect. He's also on the ground. A hands-on architect. A hands-on architect. And you don't have any of the management and stuff which is a blessing and curse but it's kind of remarkable that you can spin these things up if they make a mistake you can keep on correcting them until they get it right we've always been able to do that on the human side and i can do it faster i think the ultimate result for me is you just have to be more ambitious about everything you do you can produce more higher performance higher quality more secure so our ambitions have to raise up you also mentioned with AI, like you're ever using and you're building stuff, but also with CockroachDB, you're now building something that is also related to AI.

1:16:47Can you talk about that?

1:16:49Peter Mattis:Yeah, yeah. We're building multiple things. I mean, AI is the future. It's here to say. I think that's for sure. Yeah. I mean, like, you know, one of our theses, which is not crazy. Every application in the future is going to be written by AI. You know, I think there will be some kind of bespoke software, you know, handcrafted software, we'll see that, continue to exist just as people write assembly still, you know, but it's going to be diminishing in size. So, you know, asymptotically approaching 100 % of software will be written by AI. Generated, right? Yeah. It's hard to say if like when it's going to end where the humans are the agent, you know, kind of like supplying the agency and the vision behind it.

1:17:25Peter Mattis:I think that might exist for many, many years, but like, I think the code will like fundamentally be written by the AI. And I think we're going to just see this explosion of applications. And we're seeing that inside Cockroach Labs. We're seeing it elsewhere. Earlier this year, we kind of rolled out this internal platform where non-engineers could write kind of mini applications. I mean, this is like, you're hearing this at other companies. We did the same thing. And over the course of just a couple months, you know, 500 applications, a thousand applications. By non-engineers. By non-engineers, primarily by non-engineers.

1:17:58Peter Mattis:And I thought it was awesome. Like our HR team is building like these little applications. Like this is the thing I've always dreamed about doing. and like they were never serviced like your cfo's all over the place and ours is no exception producing dashboards that they can never create it before and i think it's very empowering i'm married i have a wife she needs software she cannot produce that software on her own and i have never actually helped her produce the software which is my own failing but um i think there's a world in the future where she gets like everyone's getting custom software built for them and you're just going to also see greater and greater applications and higher quality systems being produced by every company as well.

1:18:33Today, what is your stack? What do you work with in terms of harness, model, how you run agents? Is it one agent? Is it multiple? What kind of terminal do you use?

1:18:43Peter Mattis:Yeah, yeah. It's evolved over time. So, you know, like when I first started dabbling the AI stuff again, it was, you know, GitHub Copilot. I was an Emacs user for like 20 something years. I got convinced to move to VS Code. But that's all gone now. At some point, I moved to using Cloud Code. That was the thing that whatever reason I just got started using it was cloud code in the terminal of late I use a mixture so I'll just grab my current setup I use the cloud desktop app cloud code via the cloud desktop app. It's fantastic. Good job anthropic I also sometimes use the codex desktop app just so I have an alternative model to turn to For some very critical things we're working on I will get you know one of those models sometimes like the best model oftentimes I'm using the best model like Fable.

1:19:26Peter Mattis:You know, I've recently started using that sometimes it's Astra, sometimes it's the other ones, but you like you have one producer design, you have the other one being like, Hey, my colleague produced this. Can you kind of tear it up? You know, adversarial review it. And it's not always perfect. Right. You know, but I think there is utility, especially for something that's super, super critical. We're pushing towards the launch really soon and kind of rushing towards the finish line. I'm telling folks on the team. It's like, you know, especially our very senior folks, like you should use the best model right now yeah um it's worthwhile to do that i'm generally just using the best model and part of the reason is i don't actually know that i'm getting a lot more intelligence from it but i don't want to have the cognitive overhead of deciding on a case-by-case basis yeah should i use sonnet should i use opus should i use fable should i use soul or astra and i think you know maybe if you like i really need the speed i would make that decision but oftentimes i'm like i'm doing things in parallel so you're asking how many agents i'm spinning up well it depends like oftentimes there's like a certain number of sessions you might be using i find like kind of cognitive overhead is about five to ten sessions concurrently but sometimes those sessions will have many sub-agents doing things yeah and they'll be long running yeah it'll depend on what i'm doing so as i was coming here i was on the train i spun up something to do a little bit of research and that was just probably using three sub-agents.

1:20:43Peter Mattis:Other times, it might be 20 sub-agents. Anthropic has dynamic workflows. I can't remember what the codex thing is called, but they have various ways of spinning up graphs of agents and prosecuting them. I think sometimes I'm probably using 100 sub-agents. Others, it's only five. It depends on what's happening. Sometimes I'm not doing any of that agentic stuff because I'm often thinking about something else. Since you were an Emacs user for 20 years using the terminal, right? Well, Emacs, but how come you moved to graphical interface with agents a lot of people still use terminal uis yeah yeah well i so i moved from emacs to vs code i was just like a colleague was like hey these ids are really good you need to try them out i was like emacs is an id and then i i kind of moved and like you know some of the stuff was just more seamless and like emacs keeps on catching up it always feels a little bit you know six months to a year behind and like every time they have a new version my setup would break and i just got tired of that so i moved to vs code but then i was like i'd have many terminals open in vs code my setup would always be like terminals in vs code tabs the terminals and tabs you know have like eight of them open and then they started taking over used to be like just cloud code in the terminal yeah yeah and then you know just some point recently is like six weeks ago colleague was like well have you tried out the cloud desktop app recently it's really good i was like yeah no i'm very very happy and i made the switch and it was a little bit awkward at first and then suddenly i'm like holy crap this is awesome you can like manage the agents a bit better easier i mean it's like you have the sessions there the sessions i'm down the side you know and it's like they're kind of like tabs in some regard and shit but the whole integration that anthropic's been doing in open ai is doing the same thing now by the way like i'm not dissing cursor and factory and cognition they're all pushing the same thing it's incredible how fast these systems are innovating right now and evolving What do you think good software engineering looks today compared to like, you know, four years ago before we had AI?

1:22:37Has it changed? Has good changed or that's not really changed?

1:22:40Peter Mattis:Yeah, I think the ambition has to increase, you know. I think quality has to be higher. Security has to be higher. I'm glad you mentioned quality. Can we talk about that? Because I'm seeing across the industry, just quality declines, which you cannot fully put your finger on AI, but oftentimes it is people pushing out more and more and just not paying attention to just small regressions here and there. Again, like you're building a database. Like, have you noticed any or have you gotten feedback of any quality regressions? Or if not, like how come? Because like when, like this is just basic, you know, like we talk about the law of physics.

1:23:19This is an observed thing that when you start to have more output, you increase your deployment frequency. you often not always but you often have uh like more regressions and more bugs if you produce a

1:23:32Peter Mattis:certain number of lines of code you're probably going to have a certain number of defects per line of code and now you can produce more lines of code so you probably would have more defects right but the thing that pushes against that is that you can be telling these agents you have to give them quite a firm hand at this this is a thing that i i hope the the model providers are listening to but you need to give them quite a firm hand they kind of get a little bit lazy on the testing side and you have to make sure the tests are comprehensive but also they're using all the testing techniques and there's a lot of testing techniques out there we have a lot of knowledge about how to do testing well and you know what the agents are lazy humans are a little bit lazy so getting the humans to actually be very disciplined about their testing is also challenging i think it's actually easier with agents you know you can kind of instill that you can get them set up you have to give them the guidance use property-based testing use metamorphic testing use like these advanced testing techniques deterministic simulation testing you know there's like technique after technique you can use and humans have always like i'm lazy myself like i'm pretty good about being disciplined about testing at some point you're just like okay that was enough we gotta ship and like now you can get a little bit like you know a little bit stricter and firmer i think the same thing applies over on the security side the performance side i mean we've always had at cockroach in the industry we've known security coding practices yep you can literally have every single line of code on every commit reviewed by a security expert doing like an adverse or adversarial uh security review by yeah an agent or multiple agents or multiple agents right and you know we're going through a tough time in the industry right now with like kind of um hacks and security leaks and whatnot there's only a limited number of bugs that can be in software i think we'll get it you know out of it on the security side and on the quality side and then also just on the pulp when i say quality such as bugs but it's also like you And it's just a little thing.

1:25:18Peter Mattis:It's like, oh, that UX element is wrong. And there's no excuse for that right now. Like fixing it is so, so easy. And what we're also starting to see evolve is like, you know, designers using Figma. Well, I think that should be a thing of the past right now. Our designer just deals with HTML and CSS and JavaScript directly. And sometimes is even just producing pull requests, you know, PRs, which she loves. We all love. Everybody's happy with this. There's none of this waterfall. So she's producing pull requests for the production code base. Yeah. Yeah. And you know, this is on the UX side, not on the core database.

1:25:51Of course. But this is the area that she owns, right?

1:25:54Peter Mattis:Yeah. It's just, I mean, she loves it. Everybody loves it. I mean, there's no downside. This is not new. Yeah. For sure. Yeah. What about code review? What's your take on code review? I think it's pretty, it's starting to get a bit controversial. Like, is it going to stay or not? Cause there's been a practice that's been around like, think about, I mean, you, you worked at Google. Google has been so big on code review. As I understand, there are two code reviews, right? The domain experts reviews it, and then there's a language expert reviews it. So you went through that. How did you think of it, and how are you thinking of it now?

1:26:26Peter Mattis:I used to be part of the... I was one of the early members of the C++ coding review team. I'm asking the right... Okay, you will have strong opinions. Tell me. Yeah, yeah. I mean, I see that code review, I think, was fairly essential before the age of AI. but now we're getting to the point where i'm reviewing and looking at code i look at most of the code that the agent's producing still and yet it's like you're not giving it the same level of scrutiny right and this has always been the case if you get a pull request from a junior engineer you have to give it more scrutiny than if you get from your most senior engineer and you ask any you know tech lead any engineer manager any software they'll say yeah yeah so the junior engineer or someone new to the code base you know you just have to give more scrutiny and what i'm finding is the agents are getting better you're having to give less and less scrutiny and yet you do still have to give scrutiny to some things it's like i said it's like the testing will be incomplete and maybe they didn't follow the security stuff maybe you know like there was a performance regression and i think it's like you know assuring there's kind of constraints on the system so it's hard for them to do the wrong thing i suspect that you know i don't know if it's going to be this year or next that we're materially going to stop looking at the code in the same way that we don't look at assembly anymore yeah like you will we trust the compiler to produce pretty good assembly we trust the compiler to produce good assembly unless you're one of those people where it's really important you're and you're maybe a game developer or or someone and you look at it and but there's fewer and fewer of those folks yeah and even then i think what will be migrating to is why why does the human have to look at the assembly the ai should look at the assembly like i'm doing something right now where i want to have a zero overhead abstraction, you know, for doing something that is done test time.

1:28:07Peter Mattis:And what it's in compression is compiled out. They will set up for me a little system where it actually looks at the decompiled code to verify like that. There's only a few extra instructions put in place. I never would have put that in place, but now it has this guard rail that every time it changes this, it can verify that there's no regression. And I think you're going to see more and more of this where you kind of put guardrails in place. I think the model kind of likes it because it can work within that guardrail. now for what like 20 plus years you were writing so much code like i'm sure you were in the zone you remember being in the zone and just churning out you wrote a lot of like production ready code now that you're coding with ai like do you get into the zone yeah absolutely and how how is that zone is it the same is it different feels it definitely feels a little bit different it's Maybe a little bit less intense, but you're managing more things cognitively.

1:28:59Can you describe me, like, what is this right now when you're in the zone?

1:29:03Peter Mattis:Yeah. Well, I'm thinking of ideas that normally would have taken me a week to experiment with. And I think of multiple of these experiments, and then I fire them off all simultaneously. And then I'm kind of, like, reviewing, like, what else should I do while that's being complete? Sometimes I'm kind of reviewing, like, they'll be giving me progress updates of, like, oh, hey, this is coming in. We're seeing this stuff. and i'm being like well that doesn't sound right hey what about this or did we do this you know correctly you know maybe i had design and it's not implementing the design quite perfectly but i have this little feeling that you know i haven't been a college professor but maybe i'm a i was a you know if i was a college professor i had a whole swarm of research assistants and they're all got off doing things and it's coming back but it's coming back just really rapidly rapidly you're not waiting months or weeks i'm not waiting months or weeks and then i'm iterating like oh that one failed that's fine you know you just gotta let go and this this is the nature of the the software i build i think there's other pieces where it's like you just whip out a website you can whip out something that doesn't have this level of kind of scrutiny you know very quickly but i i talk about this in a i have a whole bunch of analogies about like the let's talk about what analogies do you have about about using ai i mean the one i was uh you know advocating for uh i had kind of two that i was advocating for like late last year and this year which is you know ai is coming for us it's here right and it's like uh you know you're producing software like walking down the road and sometimes you know someone will pass you they're running they got an efficient gate and whatnot but everyone's under their own locomotion and the these agentic coding agents came and it was like a car pulled up next to you you get into the car you don't know how to drive you don't understand the controls but you got to get in you start you might figure out the gas pedal and it takes off and crashes into a tree but you have to learn how to drive you know i think using all these coding tools it just isn't doesn't just happen naturally it's learning how to drive i think we might be in the era of f1 driver right now which is like the really good people can drive these systems a lot harder a lot faster than the people who just picking them up if you've never used an authentic coding tool there's a vast difference between someone who's like really expert in them knows where they break can pay attention that versus someone who's just picking up for the first time one analogy i've heard is we used to talk about 10x engineer you remember like this used to be a debate for a very long time is it or is it not but now what i'm hearing is the 100x engineer yeah and so you're saying that you are seeing some folks who maybe let's not use the term 100x engineer but like this like f1 driver who is just really really good at it like how would you describe a person who you've seen this is is it just like rock solid engineering basics and they picked up uh they lean into using these tools or what are they like yeah yeah i mean There is quite a bit of a, it feels like, you know, it's directional that they're good at software engineering before.

1:31:48Peter Mattis:I mean, I sometimes think it's like, you know, everyone's in this kind of spectrum of capability of software engineering. And this is just like, you know, taking that line and spread it out. And it's not quite true. You know, I think it's helped some people more than others, but, you know, it feels like it's just stretched it out. So, you know, your ability before is now amplified. I'm always interested to learn that, for example, Boris Churny, Thibaut at OpenAI. they both have been really really good software engineers boris wrote one of the first typescript books the first typescript book for o'reilly he built some massive system same same with tbo who built it and now you're kind of seeing oh these people are building all these tools and innovating like yeah they worked really good before yeah well i mean this is a i mean i have two other analogies to give you about like the what feels like an ai you you've undoubtedly heard about the term pair programming and programming yeah you know and the idea behind pair programming is it's good to just like you know have one keyboard one monitor and two engineers at it one of the keyboard and the other one sitting beside them kind of like looking over the shoulder and giving guidance and i think there's an aspect of that feeling where i actually got chachapiti to do a little image of this where you know it's like the android is at the computer typing and you're just there giving instructions but that was maybe the way it felt like a year ago i think it feels a little bit different now um the one i i've just started recently saying is that i feel like the domain experts the people who were really strong before are now massively amplified have you seen all this mathematical stuff coming out like the crazy proofs well i don't understand it but i i don't understand them either yeah yeah so i told you i wasn't very good at math you know it's not complete i just was like i stopped in the freshman year of college right yeah but you know i kind of like watch along with these uh advancements and you know terence tau he's like probably the most famous living mathematician you know super genius he actually posted this session there's this recent breakthrough it's called the jacobian conjecture i don't even know what it meant but he posted this chat gpt session where he's interacting i think it was chat gpt and you can see him interacting with this intelligence and it was crazy because he's talking to it as a peer colleague it's responding and literally it honestly looks like i'd encourage everyone to go look this up it looks like you know almost a foreign language it's like his domain expertise is getting amplified by the system he's interacting with you can see how he's like kind of learning and exploring ideas just really rapidly.

1:34:05Peter Mattis:Now there's all this controversy about AI mathematics, but I feel like the analogy that comes to mind to me though is the domain experts, they're a little bit like sorcerers in their particular domain. You know, you have the earth sorcerers, the earth wizards, the water ones, whatnot. And if you know the magic incantations, the right words to say the right order, you actually get something kind of magical. And if you don't, you just get sparkable stuff that doesn't have anything there behind it. You know, you would probably admit, you don't know much about distributed databases. If you're going to ask Fable or Astra, build me a distributed database like Cocker TV, you will get something out, but it'll kind of be ultimately hollow inside.

1:34:41Peter Mattis:But if you're an expert in databases and you ask to build a distributed database and you can point out all the various things you have to know about a distributed database, here's what you have to worry about, the storage layer, the networking layer, here's the various data structures, runtime inside, you can actually get something quite magical very, very rapidly out of it. So what would your advice be for engineers with like mid-level to senior level who have been figuring out how to do coding to become strong engineers in this, I guess, age of AI? Well, the first off is you have the most amazing tutor kind of readily at hand.

1:35:16Peter Mattis:And I mean, one of the things that I would always do throughout my career, and now I've kind of stopped doing it, but it's like the reason why is going to become obvious, which is like, I would always look at other people's code. So I was at Google early on. You probably heard of Jeff Dean. jeff dean was amazing coder his colleague sanjay gamelot also just an incredible coder and you know i would be looking at their pull requests i'd be looking at their changes it wasn't called pull requests at google there's a different name for it but i'd be looking at my cl right yeah cl um this is uh p4 it's a different version control system but i'd be looking at their changes be like how'd they do what they did right be looking at their code oh my goodness sanjay's code is really always very elegant you know jeff's high performance how's he doing that you know like how's he going about it it's almost just like you acquire via osmosis but now what you can do is not just acquire via osmosis but you can literally like i mean i would be encouraging if you're a junior engineer you know there's a senior engineer nearby like you could ask him how they're doing what they're doing but you can just ask the ai to dissect what they've done and explain it to you and explain to you at various different levels like how does this code work what is it doing give me a diagram explain it to me like i'm five explain it to me like i'm 10 explain it to me in French, whatever you want.

1:36:30Peter Mattis:I mean, fundamentally, to some degree, AI is a translation tool. They translate it from whatever kind of language or understanding it's in and keep on interrogating it until it increases your understanding. You were saying how domain experts are very much amplified. I guess one strategy as a software engineer is like, obviously become a great software engineer and use it as a tool. You can get a lot faster.

1:36:57You But I use AI to explain a few things for me to understand upfront, which was very helpful. And it would have taken me a lot longer time beforehand. So I can use this. But I wonder if there's another part of like, as a software engineer, you can use it to become more of a domain expert wherever you're working. If it's a payments company, I mean, use it to learn about payments as well. So you can help the business, you can help your team. And honestly, you'll just learn more, right? Yeah, yeah.

1:37:21Peter Mattis:I think, you know, I would encourage everyone. You have to have a little curiosity, right? don't don't be bound in by you know the area you're working on explore outside of it you know i was working on gmail but i was fascinated about how like the the main google search engine worked i was fascinated by like how the internals of big table worked even though i wasn't directly working on a big table just explore in look at those things and now it's so much easier because you have this super advanced patient intelligence there to explain to it like why do you think it was done this way and then like i mean once you become an expert you can be integrating well i see us done that way why don't we change this would this be helpful and you know that's where you go from just kind of learning to actually contributing back i think everyone has to have the personal agency to do this you know if you're just sitting there waiting for someone to educate you on how to do this it's going to be really hard right now because anyone who's coming in explaining how to use ai or explaining how to be a better software engineer they're going to be out of date right you just got to get in there be using these tools all the time yourself and using it like easy to learn i feel like i've learned more in the past probably even year than the previous five years combined which is weird given given your trajectory and given the environment you were working right i mean like anybody's been in the industry for a while like i'm definitely a better coder i was a better coder 10 years ago than when i first got in the industry it's like i can look back every decade and realize like i got a lot better and i feel like i just got a lot better over this past year and and this was also one of the reasons i was really excited to talk to you Because when we started to just exchange messages, the first thing you wrote to me and I asked you, like, hey, you know, how are things going?

1:38:54You said, like, you wouldn't believe, but my coding output is insane and it's high quality and it's database quality level. And those were the things I don't really usually see. I usually see, okay, I'm not producing more code, but it's slop. But again, like, to me, this is a bit of an inspiration. Like, look, like, you can use these tools to just, like, amplify yourself as a software engineer. Like, you are one example, right? Hopefully one of many.

1:39:15Peter Mattis:Yeah, yeah. Yeah, no, and I'm not the only one in Cockroach Labs. We have other people doing this as well. I find it very exciting. You know, it's a little bit exhausting right now, but it's very exciting. Like I got into software engineering because I like building stuff. I can build stuff faster. You know, the stuff you might have had to compromise on in the past, and you can take away some of those compromises. I mean, you see this in the UX of software coming out. I think the UX is a lot higher. You see all the fancy like web animations and whatnot. That's only just like the surface level.

1:39:43Peter Mattis:It just extends way, way below that. Peter, this was awesome. Thanks for coming on the podcast. Yeah, this is wonderful. Thanks for having me. One reason I was excited to talk to Peter is because he's been a very high-profile and productive engineer, pre-AI, building some of the most resilient distributed systems in production. CockroachDB is known for...

1:40:20I found both stories a good reminder that you can improve the existing library or even the language, especially if you measure which parts feel. slow. Another part of his conversation that I liked was how Peter came a bit of a full circle. He used to write 100 ,000 lines of code per year, being a very productive engineer and CTO. He didn't stop writing code aiming to coach engineers between 2022 and 2024. And then he started to code again because with AI tools, he wanted to coach his engineers better, but it's hard to do if you don't use the tools yourself. And now he finds himself being extremely productive.

1:40:55And this time, the team around him is productive as well. And we're not talking about vibe-coded software, but database-worthy, high-quality code-generated and committed to production. Peter is convinced that AI amplifies existing expertise, and this is one reason why he probably learned more this last year, building with AI, than the previous five years combined. And I find it a valuable reminder that learning and building deep expertise in software engineering, this is very valuable. And as closing, I appreciated that Peter said that not only is he excited, but he's also exhausted. There's a lot to learn, but it's tiring and neither him nor anyone I know is immune to this.

1:41:30So if you're also exhausted with all of the things going on with AI, know that you're not alone. Check the show notes for more of the pragmatic engineering deep dives on Google's engineering culture and on distributed systems. If you liked this episode, please make sure you're subscribing to your podcast player and a special thank you if you leave a rating. Thanks and I'll see you in the next one.

From the publisher

Brought to You By:

• turbopuffer – a vector and full-text search engine built on object storage. It’s fast, cheap, and extremely scalable

• Linear – the product development system for teams and agents

• WorkOS – everything you need to make your app enterprise ready.

—

How is it that a software veteran who regularly shipped ~100K of database-grade code to production each year, pre-AI, feels like he’s even more productive today, with no drop in quality? Peter Mattis is co-founder and CTO of Cockroach Labs, and an original creator of GIMP. He also worked on Gmail and distributed storage at Google.

In this episode, Peter reflects on his journey from open source to Google to founding a database company, and we explore how to keep systems fast, reliable, and correct at scale, from Gmail’s early storage challenges to the tradeoffs in building distributed databases.

Peter tells us how AI has brought him back to writing code after his work shifted toward management, and why he believes AI can improve quality and multiply the impact of domain experts. We also consider the future of code review, and Peter has some advice about how to level up our engineering skills.

Timestamps

00:00 Intro

02:42 Peter’s path into tech

04:00 Building GIMP

09:30 Working on Gmail at Google

14:51 Google’s infra: google3, build files, Bazel, and Colossus

21:30 Distributed storage bottlenecks

23:59 Latency, throughput, and availability

30:04 Contributing to libraries

41:52 Google Spanner

46:10 CockroachDB

52:00 Manual vs. automatic sharding

55:28 Consistency models and strong consistency

1:00:03 Raft consensus

1:06:15 How AI brought Peter back to coding

1:19:12 Peter’s tools and agentic workflows

1:23:08 How AI can improve quality

1:26:39 Code reviews: are they done?

1:29:17 100x engineers

1:35:33 Peter’s advice for leveling up your engineering skills

—

The Pragmatic Engineer deepdives relevant for this episode:

• Inside Google’s Engineering Culture

• Resiliency in distributed systems

• How to debug large, distributed systems: Antithesis

• Pushing software engineering limits with “napkin math”

• Designing Data-intensive Applications with Martin Kleppmann

• Formal methods with Hillel Wayne

—

Production and marketing by ⁠⁠⁠⁠⁠⁠⁠⁠https://penname.co/⁠⁠⁠⁠⁠⁠⁠⁠. For inquiries about sponsoring the podcast, email podcast@pragmaticengineer.com.



Get full access to The Pragmatic Engineer at newsletter.pragmaticengineer.com/subscribe

More from The Pragmatic Engineer

All 45 episodes
Distributed databases with Peter MattisThe Pragmatic Engineer · 1 h 42 min
Listen in VO