In short
James Cowling (ex-Dropbox most senior engineer; now CTO at Convex) discusses building “great systems” by deeply understanding why they work, what to do when they fail, and how to design for simplicity, validation, and long-term maintainability. He also covers distributed transactions, durability/replication at Dropbox scale, and technical leadership/career incentives in large companies.
Guest background
Cowling led major Dropbox infrastructure work, including the multi-year migration away from AWS S3 and the design of Dropbox’s storage system (Magic Pocket). Earlier, he did PhD research at MIT on distributed systems: distributed transaction coordination (Granola), Byzantine fault tolerance/consensus, and Viewstamped Replication Revisited (influential to companies like TigerBeetle). He later worked on large-scale transactional systems and storage reliability.
Key claims
- Best engineering comes from understanding “why,” not just making systems work.
- Simple systems are harder to design than complex ones; simplicity is scalable over years.
- Performance bottlenecks come from coordination points; eliminate coordination to improve throughput.
- Teams should be oriented around the problem they solve, not defending the system that survives.
- Avoid “artificial complexity” driven by promotions/OKRs; durability is non-negotiable.
Notable examples
- Granola “independent transactions” to serialize atomically without consensus voting; contrasts with two-phase commit/locking.
- Dropbox multi-homing: primary/secondary for continuity vs active-active (often too costly due to latency).
- Durability via erasure coding (e.g., ~27 fragments) and fast reconstruction by reading enough fragments.
- S3 migration: dark launch/double-writing for six months; thick validation; validators in production; “reset the launch clock” after a production bug.
- Storage mapping validation: using a single large MySQL table instead of complex distributed hash structures.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Challenge of Designing Simple Systems
0:45 to 2:34
Discussion on the complexity of creating simple systems and the mindset required.
“Someone will argue back, no, what you can do is get really good at using Claude.”
PhD Thesis and Distributed Transactions
2:34 to 4:52
James discusses his PhD thesis on distributed transaction coordination and its significance.
“I'm not actually sure if that was a standard term at the time or whether I made it up.”
Independent Transactions in Granola
4:52 to 7:58
Exploration of independent transactions and their impact on performance in distributed systems.
“serial throughput across large numbers of transactions.”
Challenges of Testing in Early Systems Research
7:58 to 9:10
James reflects on the difficulties of testing systems in early academic research environments.
“Now, this was right at the cusp when I think in some respects Google ruined systems research.”
The Shift from Academic to Industrial Research
9:10 to 12:39
Discussion on the transition from academic research to practical industrial applications.
“obscure intellectual ideas right because i'm all for like papers about systems but i think there's also value in a paper about an idea.”
Multi-Homing: A Key Engineering Concept
12:39 to 14:01
James explains multi-homing and its importance in data redundancy and availability.
“I mean, in academia, your goal is to advance knowledge.”
Replication Strategies in Data Storage
14:01 to 18:02
Learn about different data replication strategies and their trade-offs.
“because the data would be replicated in other regions, but with a window of vulnerability, with a window of time where there may be some data lost.”
Erasure Coding and Data Durability
18:03 to 21:48
Understand erasure coding and how it ensures high data durability.
“And we actually had our own custom encoding matrix we had developed.”
Cost Efficiency in Data Systems
21:49 to 25:06
Explore the importance of cost efficiency and workload understanding in data systems.
“experts at this and most people should not be experts at this you know most people should focus on their applications.”
Choosing the Right Technology for Data Systems
25:07 to 28:00
Discover how technology choices, like programming languages, influence system performance.
“whereas S3 has to design the system for everybody.”
Show all 47 chapters
Migration Challenges: From Python to Rust
28:00 to 29:00
Learn about the migration challenges faced when moving a storage system from Python to Rust.
“So initially the prototype was in Python, if you can believe that.”
Cascading Failures and Congestion Collapse
29:00 to 30:50
Discover how cascading failures and congestion collapse impacted system reliability.
“At a certain point though, we would have, you know, let's just pick a number, let's say a million.”
Building Buffer Against Unknowns
30:50 to 33:30
Understand the strategies used to create buffers against unexpected capacity challenges.
“Congestion collapse is when workload to a system crosses a threshold where it all kind of collapses.”
Trampoline: An Innovative Solution
33:30 to 35:50
Explore how Trampoline helped manage storage capacity during peak loads.
“it's the equivalent of your disk filling up on your laptop, except it's a million disks.”
Double Writing and Migration Strategy
35:50 to 38:30
Learn about the double writing strategy for ensuring data integrity during migration.
“Eventually, we built this system called Trampoline.”
Simplicity in System Design
38:30 to 42:00
Discover the importance of simplicity in system design through concrete examples.
“And as a result, it's going to launch later and it's going to cost us some amount of money.”
The Importance of Simplicity in System Design
42:00 to 44:30
Learn why simplicity is crucial in system design and how it affects long-term maintainability.
“all the data is where it's meant to be, I just walk over the table and check.”
Incentives and Complexity in Engineering
44:30 to 47:10
Explore how the industry’s incentive structures can push engineers toward unnecessary complexity.
“I think it's the long-term beneficial thing to do.”
The Role of Purpose in Engineering Work
47:10 to 49:50
Understand the importance of aligning engineering work with meaningful missions rather than systems.
“A lot of folks there, though, I'm not that interested in hiring because, you know, if I'll do a deep dive with them and they'll say they build a system and I'll say, well, why did you build it?”
Managing Projects and Team Identity
49:50 to 52:50
Discover how to effectively manage projects by focusing on team identity and problem-solving.
“riding their outdated system to the grave, like the captain going down on the Titanic.”
Navigating Inertia in Large Organizations
52:50 to 56:00
Learn strategies for overcoming inertia in large organizations and driving meaningful changes.
“And so I think it seems like such silly management philosophy, but I think it's really, really important to orient a team and an identity around solving a problem and not owning and defending a system.”
Accountability in Team Dynamics
56:00 to 56:40
Learn about the importance of accountability within a team and personal ownership of projects.
“I'm like, well, we're going okay without it.”
Framing Priorities as a Tech Leader
56:40 to 57:40
Understand how to prioritize projects effectively rather than just answering yes or no.
“I do not think someone will have good luck going around complaining about stuff and just saying this is a dumb idea and being negative.”
Career Progression in Engineering
57:40 to 58:20
Discover how the focus shifts from programming skills to broader company impact as you advance in your career.
“It's not like, should we redesign the database?”
Choosing Integrity Over Promotion
1:00:16 to 1:02:50
Explore the importance of doing the right thing over seeking immediate rewards in your career.
“So if you really want to maximize long-term income, I don't think you should try to do that.”
The Value of Long-Term Commitment
1:02:50 to 1:05:26
Learn why staying in a job long enough to see the impact of your decisions is beneficial for growth.
“There's no shame in just wanting to have a regular job and just be chilling and just be getting paid.”
Finding Fulfillment in Your Work
1:05:26 to 1:08:06
Understand the importance of enjoying your work and the impact of career satisfaction on life quality.
“So I do think there's an actual structural incentive towards staying in a job.”
Navigating Management Roles in Tech
1:08:06 to 1:10:04
Insights on when and why to move into management, and the value of technical expertise first.
“I like having weight on my shoulders, I suppose.”
The Case Against Early Management
1:10:04 to 1:11:39
Discover why early career professionals should avoid management roles.
“circumstances in the first three years of your career.”
Balancing Leadership and Credibility
1:11:40 to 1:14:08
Learn how to maintain credibility as a tech lead without authoritarian methods.
“And you actually say it's bad to lead by example, but there's a quote in there.”
Empowering Teams Through Communication
1:14:09 to 1:16:48
Understand the importance of why communication in tech leadership is crucial.
“And then largely, I just trust the team to do the right thing.”
Navigating the Ownership-Accountability Slider
1:16:49 to 1:21:18
Explore the balance between oversight and accountability in tech leadership.
“the right thing, not get promoted, right?”
Mentoring Senior Engineers
1:21:19 to 1:23:50
Learn approaches to effectively mentor high-level engineers.
“I'm not going to leave you high and dry.”
Finding Your Unique Engineering Style
1:24:01 to 1:27:39
Discover how engineers can identify their strengths and personal styles to excel in their careers.
“You've got to be your own brand of engineer.”
The Evolving Landscape of Engineering Careers
1:27:40 to 1:31:40
Learn about the shifts in engineering career advice and the challenges faced by junior engineers today.
“Something that you used to say five years ago, you don't say anymore or vice versa?”
Wisdom in Engineering: The Importance of Learning
1:31:41 to 1:34:00
Understand why synthesizing knowledge and experiencing challenges is essential for engineering growth.
“I go home and I'm tired and I watch a YouTube video.”
The Human Role in an AI-Driven Future
1:34:01 to 1:36:29
Explore the potential future of engineering and the ongoing relevance of human ingenuity amidst rising automation.
“It's like, you know, if you were doing a math degree, right, a big part of doing advanced mathematics is proving theorems.”
Building Convex: A New Engineering Platform
1:36:30 to 1:38:00
Learn about the founding of Convex and the challenges of creating robust systems for application development.
Designing Convex: A New Abstraction Layer
1:38:00 to 1:40:18
Learn about the design principles behind Convex and its advantages over existing systems.
“So what the client sees is a consistent view of what's on the server.”
Innovating in System Design: Challenges and Opportunities
1:40:18 to 1:43:59
Discover the challenges faced when innovating in system design and the excitement it brings.
“And we worked a lot together and a lot on abstraction and, you know, the value in clean designs that minimize complexity.”
Scalability and Background Processing in Convex
1:43:59 to 1:46:04
Explore how Convex handles scalability and background processing through innovative practices.
“And I would encourage engineers to try to find this kind of stuff to work on, where it's like you're on that edge of like, I'm really liking this, but also it's a bit tricky.”
The Most Stimulating Technical Challenges in My Career
1:46:04 to 1:48:08
Reflect on the most intellectually stimulating challenges faced throughout the speaker's career.
“When you reflect on your career, and it sounds like you've done a lot of gnarly technical work across your PhD, Dropbox seemed like pretty intense systems work and Convex is also doing a lot of cool stuff.”
Consulting for Silicon Valley: Behind the Scenes
1:48:08 to 1:52:01
Get insights into the speaker's experience consulting for the show Silicon Valley and its accuracy.
“He started his career as a software engineer at, I think, Lockheed or something.”
Reflections on Career Sacrifices
1:52:01 to 1:53:27
James discusses the sacrifices made throughout his career and the impact on his personal life.
“because it made my visa more complicated.”
The Fallacy of Hustle Culture
1:53:28 to 1:55:46
James critiques the hustle culture and shares insights on work-life balance.
“But yeah, I think everyone has to decide how much they want to really drive their career.”
The Importance of Learning by Doing
1:55:47 to 1:57:41
James emphasizes learning through practical experience over theoretical knowledge.
“Do you have a best technical book recommendation for people?”
Navigating the Noise of Tech Trends
1:57:42 to 2:00:37
James addresses the distraction of tech trends and encourages focusing on real projects.
“I've been an engineer for several decades.”
Transcript
Automatic transcript. May contain errors.0:00James Cowling:The best engineering comes from a deep understanding of why. This is James Cowling, formerly the most senior engineer at Dropbox and now CTO at Convex, and we dived into all the technical work of his career. It's not about getting a system to work. It's what do you do when it doesn't work. Simple systems are way harder to design than complex systems. He also had some interesting takes on technical leadership. Really, a team should be oriented around what problem do they solve. They should not care about the system that survives. I've had many friends whose promotions were rejected because their work wasn't complex enough.
0:37James Cowling:Yeah, I mean, it almost angers me. I just like it so much. Is there any career advice that majorly changed in the last five years? Someone will argue back, no, what you can do is get really good at using Claude. But guess what? It's not very hard to use Claude. Here's the full episode.
0:59Well, first off, your PhD thesis was huge. It's 156 pages. I didn't know thesis. I thought papers, maybe 10 pages or something like that. It's its own book.
1:12James Cowling:It's very similar to Spanner, ultimately. And what's funny is I did that work. It came out a little bit before Spanner, so I got a citation in the Spanner paper. But then Spanner came out and everyone forgot about my paper. you know about the research like if you could kind of describe the the problem that granola was solving and you know how it solved it those types of things we talk about that yeah absolutely so i mean my my interest in my career has been um two parallel threads um one has been just abstractions in general the idea about how to build simple models for complex problems and And I find that a very, very difficult and very intellectually stimulating design exercise, how to design APIs, basically.
1:59James Cowling:And the other has been large-scale transactional systems. I'm a big fan of transactions. A transaction is do a bunch of things at once. And again, I think transactions are just one of the most incredible abstractions we've invented because it allows us to manage probably the most difficult problem in computer science, which is concurrency. And so Granola is a, you know, that was a long time ago in my life now, but it was an algorithm for how to do distributed transaction coordination. And in particular, transactions that I described at the time as one-shot transactions. I'm not actually sure if that was a standard term at the time or whether I made it up.
2:42James Cowling:but a one-shot transaction meaning you kind of send some code to the server and say run this function all these reads all these writes on two three whatever nodes at once and how to make sure this commits atomically across multiple shards in a distributed system but yeah that was my my phd thesis and my master's thesis was on a on a um on byzantine fault tolerance uh so a consensus protocol in the presence of malicious nodes, so how to achieve agreement in state across multiple parties when there's malicious entities. And I guess the most well-known work I did as part of grad school was work on this paper called View Stamp Replication Revisited.
3:27And that was a paper on redefining a protocol called View Stamp Replication, which was a
3:35James Cowling:predated Paxos a little bit. Very similar algorithm. Paxos, Raft, VR, they're all virtual synchrony, all basically the same thing. And that was a paper I wrote at the time, which ended up being, I guess, influential to a few really great companies like Tiger Beetle. You mentioned transactions, and I saw on the paper, there's this idea of an independent transaction. And if I'm understanding correctly, a lot of the efficiency in a distributed system is lost by needing to get consensus and to vote for consensus. And I saw that in your research, you found out a way to avoid needing to vote to reach consensus.
4:18James Cowling:And how did you do that? Yes. I mean, a lot of people, I think, think about performance maybe from the wrong angle, because I think of performance as a factor of just of its raw horsepower. How fast are your disks? How fast is your network, et cetera? And no one has particularly faster disks or memory than anybody else. What really matters to performance in a large-scale system is eliminating points of coordination. So it's how to allow systems to progress without having contention between parties, where you're basically reducing parallel throughput to serial throughput across large numbers of transactions.
4:59James Cowling:And so in Granola, we had this idea called an independent transaction. Again, that was the terminology at the time, where there were two entities that are basically processing pure functions. So they have independent state, and they will both come to the same conclusion as a result. And all they need to do then is serialize those transactions. So they have to decide if it was to happen atomically across multiple nodes, what timestamps should it get? And in Granola, the paper was mostly about how to exchange these timestamps very efficiently. So how to have multiple parties each propose a timestamp and then choose the maximum of these, basically, so you could safely serialize the transaction.
5:42James Cowling:And there's a lot more complexity that goes into this. But really, the focus there was about how to maximize throughput in a large distributed system without resulting in the alternative, which is two-phase commit and two-phase locking. And two-phase commit with two-phase locking is basically the standard approach for when you want to have multiple nodes agreeing on the same thing, where they both agree to lock their state, not process any other data, and then commit a transaction. Two-phase commit can be quite, one, it can be low performance because you're blocking basically the systems for the duration of the transaction.
6:18James Cowling:And it can be high risk too, because you are taking a dependency on another node. You basically blocked waiting for another node to return. And so that was what Granola was. Now, it's funny you asked about Granola because I forgot about it. That's so far in my history, I kind of forgot about that work. And I even forgot the phrase independent transactions, to be honest. So it's great to hear you bring it up again. But I guess all this stuff does feed into all the work you do later on in interesting ways. When I think about distributed systems, I think about you're at a big company and you got a bunch of machines.
6:54But when you were at MIT building this thing, how did you build and test it? Did you, you know, were there spare machines that the college had or was this in the cloud?
7:05James Cowling:Yeah, this was a long time ago now. I'm showing my age. This was pretty early in the days of AWS. And so we had a rack of servers in our office that we could use. And I was fortunate enough to be at MIT where we could afford a rack. And there was also a service called Planet Lab. And Planet Lab was like a big communal set of nodes that academics could use to run tests on. But Planet Lab was a communal system. And so it was continually having problems. I mean, It's just a free for all. And so for the longest time, I would just sleep next to my desk and wake up whenever Planet Lab was free because, you know, you need to run some benchmarks.
7:47James Cowling:These days, you just spin something up on AWS. But then I would, yeah, I would sleep next to my desk and I would wake up in the middle of the night or random times and check Planet Lab status and then kick off a benchmarking job. Now, this was right at the cusp when I think in some respects Google ruined systems research. and I say that with affection and respect to Google because before then most of the research was coming out of academia and a lot of distributed systems research was done on smaller scales and a lot of the value prop was the ideas. So like hey I wrote a paper here's an interesting new idea and yeah it's all in theory and so the value that came out of it was the idea.
8:31James Cowling:around about this time you started seeing papers out of google amazon and these various companies where it wasn't necessarily a paper about an idea it was a paper about a system hey i built this giant system it has a whole bunch of features some of which are interesting some of which are not and by the way it powers gmail or whatever and that that um kicked off a pretty interesting transition because at that point then you know um program committees reviewing papers started to expect to see realistic benchmarks that frankly grad students are not able to at least at the time produce i mean we i wasn't running gmail on my system um and i think it did in some respects obscure intellectual ideas right because i'm all for like papers about systems but i think there's also value in a paper about an idea.
9:23James Cowling:It's not we built this big thing and it works and up to you to figure out what's interesting about it. But instead, here's a new thing. There's a paper, an old paper just called Leases, right? When someone invented the idea of a time-based lock and it's just a paper called Leases. And that was a cool era of systems research. I think a lot of it now has shifted towards industrial research, which people are building stuff for practical purposes, whereas I think it's hard for academia to compete on pragmatism. And academia is really a great place to do impractical work. When you look back on you getting the PhD in academia, being enabled to completely explore an idea versus going into industry, but maybe you worked on Spanner at Google or something like that.
10:14When you look back, which path do you think would be better and why?
10:17James Cowling:I think that's a bit of a misconception a lot of folks have about PhD programs. I think a lot of students are like high achieving students in college and they think that the PhD program is like college plus plus. It's like, hey, I like learning things. And so I'm going to go do a PhD, but it's not. A PhD is training to be a researcher. And for most people, they shouldn't do that. If someone wants to be a professional software engineer, they probably shouldn't spend their time trained to be a researcher. But with a really important caveat, I think there's a really interesting and challenging developmental experience you go through in a PhD program at a top university, at least, which is you get to a certain point in time when you have a problem that you're facing that no one else in the world knows the answer to.
11:04James Cowling:You have a problem that you are the world expert on. And you can chat about it with folks, but you can't ask your advisor because your advisor doesn't know either, right? So you're faced with these difficult challenges that you can't read a book, you can't ask ChatGPT, right? And so you're forced to go through a quite difficult and frankly, emotionally challenging experience of being unsure, being uncertain and learning to think for yourself. I think that is really valuable for all engineers. I think that was really, really valuable in my career. I do see a lot of folks early in their career think that all knowledge comes from reading or, you know, absorbing it from someone else or that there's a right way to do things.
11:51James Cowling:But I think being in a PhD program or being faced with really demanding open questions does train your mind to be comfortable with that discomfort. it. I mean, look, I feel a little bit lucky that I went through grad school without LLMs existing, because I had to deal with this uncertainty. There was no crutch to help me out, which I think was very valuable. So, I mean, a lot of, you know, if you want to be a researcher, go do a PhD. If you want to build stuff, go build stuff. And since leaving academia, I've done a lot of work that I guess one could call innovative, like advance the industry in certain ways and did novel stuff.
12:43James Cowling:But our goal was never research. Our goal was just to solve problems. And I think that's the difference. I mean, in academia, your goal is to advance knowledge. In industry, your goal is to solve problems. and I gravitate more towards just solving problems and I find that a more comfortable environment within which to work. Going to your time in industry, you know, at Dropbox, you became the most senior engineer at the company and looking through all the projects that you did, I had a series of just technical curiosities. I saw this idea early in your career about multi-homing and I wasn't familiar with that concept.
13:24What is multi-homing? What's the problem it solves?
13:28James Cowling:Yeah, I mean, multi-homing is the ability to have data in two locations, two homes. And multi-homing can be valuable for a variety of reasons. One is called primary, secondary, or lazily replicated multi-homing, whereby all rights commit authoritatively in one region and get replicated to a secondary region. And so this is normally used for business continuity. One thing we did at Dropbox, we made sure that if the entire West Coast blew up, which hopefully wouldn't happen, Dropbox would keep running because the data would be replicated in other regions, but with a window of vulnerability, with a window of time where there may be some data lost.
14:13James Cowling:So that's kind of primary secondary replication or multihoming. There's something else called active-active multi-homing, where there is truly an authoritative copy of the data in multiple locations. So basically, when you, for example, write to a system, you don't externalize that write as having succeeded until it has landed in all the regions. And so, for example, in the storage system at Dropbox, the block storage system, we truly had a multi-region replicator system where we could take down an entire region, say take down a region in Ashburn, Virginia, where everyone's data centers are, and there'd be zero downtime for the company, and the data was still safe in multiple regions.
14:54James Cowling:I think multi-homing is something a lot of engineers aspire to work on, because it seems like the right thing to do. But I think the reality is, for most companies, it is not. For most companies, not, because there's very high latency costs for this. very high because the speed of light is just not getting faster. The speed of light is fixed. And if you have to, you know, synchronously write data across multiple regions in the United States, you're going to have, say, 60 milliseconds in the commit path of your protocol, which for most applications is not tenable. So there is a lot of desire. You know, a lot of engineering teams reach out for kind of advice on how to adopt active, active multi-homing for the company.
15:40James Cowling:I would normally say, don't do it. I would normally say, frankly, if US East is down, if Amazon is down that day, that's okay. If Amazon's down that day, your company will be down. That's a shame, right? But by avoiding that complexity, you're going to be able to move much faster and build a much better product. And so in my mind, systems is all about trade-offs and making the right ones. I would generally recommend most people do not make the trade-off to have partition tolerance or multi-region availability. Even though for a company like Dropbox, yes, that does make sense. When your job is storing data and you have several exabytes of it and hundreds of millions of customers, I think that's when it starts to make sense.
16:26So it's just another term for, I guess, data replication and having it available in other regions. Okay. I imagine also, I mean, the cost of storage is a concern. how many replicas would you keep for something? Like, let's just say I stored something in Dropbox. It's my document. Is that on the order of one or two, or is there multiple? No, it's on the order of many, many.
16:49James Cowling:And so if you were to store a file in Dropbox, I modeled the storage to be, as we advertise, at least 12 nines of durability. Internally, the models look around 24 nines of durability. And so that means the data is secure with 99.99999 % where there's 24 nines, right? Which means at least according to the model, you know, the universe will be extinct before any data is lost. And the way that is done is by a combination of what's called erasure coding. So erasure coding is how you take several blocks of data and combine them together with an encoding scheme and spread them around in different locations.
17:32James Cowling:At that scale, at Dropbox scale, you're taking things into consideration like putting data in different racks, different rows of a data center because they're on different power feeds. You're taking into consideration different eras of hard drives and different manufacturers of drives because they can have correlated failure patterns. And then replication across regions. So, you know, you could be looking at 27 fragments, for example. That doesn't mean you're storing 27 times the data. It's kind of encoded in many regions. But it's really, the replication schemes get really quite sophisticated at that point.
18:06And we actually had our own custom encoding matrix we had developed.
18:11James Cowling:It's called a Van der Mond matrix, where you kind of take a bunch of data and you combine it together to produce outputs. And we would plug in some variables like how much do disks cost? How much does network bandwidth cost? Because there's a trade-off. You can either store, if you don't want to lose your data, you can store more copies on more disks. or you can store fewer copies and anytime a disk fails, you re-replicate it really fast. And re-replicating really fast costs network bandwidth. So these kind of variables go into this equation. It ends up being an extremely complex field of endeavor, but it's kind of abstracted away into a part of the system that doesn't leak into anywhere else.
18:53So with Erasure Encoding, my document in Dropbox is fragmented fragmented into a bunch of different chunks of data and loaded potentially from many different machines? Yes, absolutely. I mean, the first thought I have is now there's, I might be waiting there and one of the 27 machines is slow and I can't look at the whole doc.
19:16James Cowling:So how do you prevent against that? It's actually faster than not replicated. Because if you imagine, I'll pick a a simplified example. Imagine this is not the encoding scheme Dropbox uses, but imagine to reconstruct a file, you have to read six out of nine fragments. So if you read six out of nine fragments, you can just ask all nine and reconstruct and return the data as soon as you've heard from the first six, it's actually faster and you can construct these encoding matrices so it's actually faster to have a ratio-coded data than none. I see. Okay. So you can like oversubscribe, over request, and then you complete on a portion of them being received.
20:02James Cowling:Yes. Now in practice, it was a bit more complex than that. We'd often have a copy in a single disk to provide fast access. We'd often try to make sure that you could serve your data out of a region close to your home region. So we'd make sure that your data was mostly served with low latency. But if that region had failed, you could reconstruct it from the remaining regions. So there's a lot of people, there's a lot of talk about building your own infrastructure and you can save money by moving off the cloud. Almost definitely you can't unless you, either you have very small requirements or very fixed requirements or very, very, very heavy investment.
20:43James Cowling:Because if you want to compete with Amazon, if you want to build a more efficient storage system than Amazon, you have to have a supply chain team that's working with Western Digital and Seagate constantly negotiating on prices of disks and buying shipments at certain times and capacity teams and data center teams. There's a lot of work that goes into optimizing this because ultimately our desire was to use the disks to 90, 95 % of the disk size to maximize storage efficiency. To put it this way, I mean, that was, you know, I guess it's a ballpark figure, a billion dollar project. Right. And was, you know, I think at the time, as far as I know, was the largest ever data migration in history, I think, at the time.
21:28James Cowling:And so extremely large engineering project with very high technical investment. So, yes, if you have that scale and you have the engineering team to do it and you're willing to keep innovating, if you're willing to keep optimizing and investing effort, then, yes, you can do it. but I think the reality I mean the cloud has been an incredible innovation right most people are not experts at this and most people should not be experts at this you know most people should focus on their applications. When you say that most people shouldn't my immediate thought was but one of the big projects you worked on at Dropbox was migrating away from S3.
22:06Yes yes.
22:07James Cowling:So you know why did Dropbox migrate away from S3? Yeah that was a desire at the company for a long time you know from even before I was there so I I started Dropbox in 2012, and I spoke to Drew, the Dropbox founder, I think in 2010, about this project. And so I think there was a desire to control the destiny of the company from a strategic perspective. I mean, at the time, this is before Dropbox kind of reshaped itself as being more about collaboration. At the time, it was a file sync and share category. That was the market sector. And owning the file system was really valuable to the company.
22:46James Cowling:ultimately we saved a huge amount of money. I mean, and this is before the company went public. We really drove massive cost efficiencies through the project, but it was hard, in a way that I think it'd be very difficult to emulate without a huge investment. And I do think that there is a benefit to an organization from having hard problems to solve. Because if you have a company with extremely hard technical challenges, you can attract engineers who like working on those hard technical problems. And when they've solved those problems, they cycle off and work on different parts of the system. So, you know, after we all worked, we had such a great team.
Read the full transcript
23:29James Cowling:It was a very, very small engineering team. And after we kind of shipped the story system reliably, we all went off and, you know, Jamie went and redesigned the sync protocol, the desktop client. And I worked on the file system and the distributed databases. And so there's value to a business to have that level of technical investment. But it's like having a baby and then you've got to raise the baby. You can't just build a system like this and be, that's it, we're done. You own it and you have to keep investing in it. Did S3 do any counter negotiation before you set out to leave them? They say like, oh, you know, we'll cut you a deal.
24:12James Cowling:I guess I'm allowed to talk about this now. It was a long time ago. Yeah, for the longest time there, I don't think they were particularly aware that this was happening. But, you know, the data center folks talk and certainly it was noticed that Dropbox was buying up a lot of data center space. So, yeah, we had obviously at the scales that we were at, I mean, we were negotiating very good rates with Amazon. You know, we weren't paying sticker price. were paying very, very, very good disc counter rates. But yeah, at a certain point, they weren't able to meet our cost efficiency because when we launched the system, it really was more efficient than S3.
24:51James Cowling:And that's for a variety of reasons. One was that we were using kind of new experimental disks called shingled magnetic recording. We were the first ones, I think, to use these disks at scale. And two, we had a very tight understanding of our workloads. So we were able to design the system specifically optimized for our workloads, whereas S3 has to design the system for everybody. So it got to the point where Amazon would not have been able to offer us a more competitive deal because we had a more efficient system. I wouldn't recommend another company do this right now. But I think at the time, it certainly made sense for us as a company.
25:28Can you give an example of a tight understanding of your workload leading to something you could do that S3 couldn't?
25:36James Cowling:Yeah, absolutely. So, for example, I know that, well, I'll try not to leak any confidential data, right? But when you upload a file to Dropbox, there is a pattern of access, right? So typically, people access the file very quickly, shortly afterwards, because you're sharing it with someone, or maybe Dropbox is processing that file to generate an image preview. And then it decays at a certain rate. And so we understand in general the average block size and we understand also the access pattern. So we could do things like at a certain point, we had these two clusters. One was designed for contemporary storage that was kind of storage inefficient, but access efficient.
26:24James Cowling:So it was very cheap to read and write to, but it was inefficient to store. And data would get written to there first. And in the background, it would get moved in bulk to this coldest storage system. And this coldest storage system was far more static. And so it was able to have kind of more efficient algorithms and be written to in bulk. And if it went down for writes, that was no problems because it wasn't in the live path. And so we're able to trade off. Again, like systems is all about tradeoffs. So we're able to trade off the live data write path from the long-term read path. That was one of many examples where knowing the size of your data, where it's accessed from, how frequently it gets accessed, how long it takes to delete that data, you can really tune a system to your workload.
27:13Stuff like looking at even things down to knowing how much power to put in a rack.
27:20James Cowling:You have a rack of hardware. There's a power distribution unit, a PDU, at the top of that rack. It has a circuit breaker, which can handle a certain number of amps. we would have to figure out, you know, how many amps are required for that rack based on access patterns. And there were times where we got it slightly wrong. There was times where we got a, maybe a bad batch of hardware and disks were failing too frequently. And as a result, we were re-replicating the data more regularly. And I was getting, you know, messages from the data center team saying, Hey, we're running the racks really hot right now.
27:52James Cowling:So that's the level of optimizations you can make when you get to that scale. But again, that's multi-exabyte scale, you know, million hard drive scale. Yeah. When I was reading about this migration, I saw somewhere in the migration, you initially started with Go and then the racks or something that you're running on the hardware itself was ooming too much or is requesting too much memory. And then you migrated to Rust. So initially the prototype was in Python, if you can believe that. And to be fair, Python is actually pretty efficient for I.O. I think people give Python a bad rap for I.O.-bound workloads.
28:31James Cowling:It's pretty good at I.O., but obviously not great for concurrency, not great for memory management, and very hard to refactor. And it was critical the system was correct. And so we migrated everything to Go, and we built most of the storage system in Go. This is before Go was in general availability. I think this is before Go was GA. and we built the system in Go. Go is a great language for concurrency, a great language for proxies. You know, it's really well-designed for like servers that moved out of one place to another place. At a certain point though, we would have, you know, let's just pick a number, let's say a million.
29:08James Cowling:Let's say we have a million nodes in the system and every node has some amount of memory, some amount of disk, some amount of sheet metal in the chassis, right? And we would itemize all these things. you'd have a pie chart of how much money is spent on all the things. And it would have stuff like, yeah, sheet metal and screws and stuff. And so you're trying to optimize the storage. And a big problem for us was the amount of memory these nodes were using. Not just the memory they were using, but the unpredictability of it. With Go having a runtime and, you know, in a storage system, an out-of-memory error is pretty bad.
29:49James Cowling:because if a node runs out of memory and restarts, that looks like a disk failure. So that looks a lot like a disk has failed and has to be re-replicated. And so a batch of nodes zooming can lead to cascading failures throughout the system because maybe I remember a time, it was a band, it might've been De La Soul, I can't remember, there was a band released an album on Dropbox. So it was a big spike in load to a few files. So it was very, very high bandwidth, and it caused the machines to oom. So those machines oomed, those disks oomed. And so as a result, no problems. The system went to try to recover that data from a whole bunch of other replicas, right?
30:34James Cowling:Now, all of a sudden, you've taken one amount, one fire hose worth of load coming in, and you've turned this into seven fire hoses worth of load coming in because now you have to do a more expensive reconstruction operation, right? So now you've 7x the load. And these are the kind of cyclical kind of behaviors that can lead to something called congestion collapse. Congestion collapse is when workload to a system crosses a threshold where it all kind of collapses. So designing against congestion collapse is really the hardest part or one of the hardest parts of Magic Pocket, which is the name of the storage system.
31:08James Cowling:So ultimately, we switched to Rust for the storage nodes themselves. And this, again, this is before Rust was in GA. So that was a bit of a risky move. but we rewrote the switch to rust coincided with getting rid of the file system entirely on the disks and directly addressing the the disk heads so there was a there's an instruction set i think it's called zbc zone based block control something like that there's a there's a there's an instruction set for accessing disks that we were using the disk manufacturers gave us the draft specs of these new disks and we were operating off the draft specs and directly controlling the disks And so all that product was tied up together into a product called Discotech.
31:52James Cowling:It was the disk technology project. And ultimately, if you look at that pie chart, one congestion collapse stopped and reliability improved. But if you looked at the pie chart of where all the money was going, it really shifted to be almost all disks. And what we wanted to do is get that pie chart to be almost all the money spent on disks. and as little money spent on RAM, on compute, on network, on power, on sheet metal, et cetera. The cascading failure you mentioned, was there, so that happened and then there was a postmortem. I don't think we needed a postmortem, but I think we knew as it was happening.
32:28James Cowling:I think when I was getting paged in the middle of the night on these things, you'd be pretty evident. And so, look, again, this is an argument in favor of the cloud. and imagine you've spent hundreds of millions of dollars on a storage system and it's out in production and it's on physical hardware you can't just go and put new memory chips in every one of them i mean you can but you have to pay people to come in and swap them out you know and so that's a tricky place and so as a and this is what this is what i love about industry you know Because some of the more challenging moments on Magic Pocket were stuff like in one week, just a weird coincidence, two trucks crashed that were delivering service.
33:13James Cowling:So there's two trucks showing up to deliver racks. And they both crashed. I don't know. The drivers were okay. So we lost capacity for two weeks. So we lost capacity for more than two weeks, for probably six weeks. What do you do? What happens? it's the equivalent of your disk filling up on your laptop, except it's a million disks. And you can't tell the customers to go away. You can't delete their files. So a lot of tricky capacity work. And ultimately, what that led to was trying to build in all these protections against the unknowns. Making sure we had the right amount of buffer planned out for anything bad that could happen.
33:59James Cowling:making sure we design the system so there couldn't be congestion collapse, so there wouldn't be memory spikes. And, you know, at one point, there's a process called, I think it's called FMEA, which is like a threat modeling process where you had a big spreadsheet, and you kind of write down every bad thing that could possibly happen, and then, you know, how bad it would be if it happened. Existential risk, does someone die? You know, there's a fire in the data center. All these kind of things get put into the spreadsheet. and you kind of do a bit of a almost pre-mortem kind of work to figure out all the all the potential failure modes and then design around them and i mean i that's i love that work i i really um i don't know i do like the firefight i like the i don't i can't say i like getting paged because i've spent my whole life on call but i do like that rubber hits the road stuff i like that wow there's congestion collapse and and there's no one that can help you and so you've got to think through this problem when i worked on infrastructure instagram we had this concept of defcon knobs which are basically these configs that you could flip that would gracefully degrade your system where you can still operate it but you know maybe in the case of Dropbox, you store less replicas.
35:18So you take on a temporary increase of risk in losing data because you need to. Did you have something like that? And did you flip it in that type of case?
35:29James Cowling:Never when it came to durability. So, I mean, we just had an absolute zero non-negotiable, like there was no room for negotiation on users' durability. And so we had those knobs for background processes for CPU and memory, for example. So if there was like a spike in load, you could turn off background processes and turn off the test load, for example, on the system. Eventually, we built this system called Trampoline. And what Trampoline did was, which actually was that it saved us a ton of money, right? When if we ever got too close to the threshold, we would just start writing data to S3, because this is right there, right?
36:11James Cowling:So So you can run your capacity way closer to the edge if you're willing under worst case scenarios just to dump 30 petabytes on S3 and then move it back when it's done. And now that didn't happen very often. We would do it to test it. We would do that just to make sure the system worked. But yeah, being able to have an escape hatch for worst case scenario was really nice. I see. So S3 is kind of, it's like elastic storage. Exactly. When I looked at this project, it's such a massive migration. And my first thought is, how do you coordinate this whole project without breaking the system as it's running?
36:53Yeah. What are your thoughts on doing such a large migration without breaking things?
36:58James Cowling:Yeah. Now, I mean, in terms of the engineers that built the initial version of the system, it was a handful, three, four, five, six engineers. It wasn't a team of a thousand. It was a very, very small team. And so how do you build such a large system with a small number of people? And then how do you do a high-risk migration?
37:21James Cowling:One thing you do is you keep things simple. You try so, so, so hard to build very cleanly abstracted, simple systems that their failure modes are very understandable. And they're decoupled so they don't have congestion collapse, et cetera. So focusing on simplicity, it gets tricky, by the way, with people doing agent development. They're not the best at building simple systems. Simplicity is still the domain of human beings for now. But a big focus on simplicity. The other was this very thick layer of validation checks during this migration. And in fact, when we did the migration off of S3, we had something called the dark launch, where we would be moving data off of S3, but we were keeping it in both locations.
38:11James Cowling:And we had this kind of contract with the Dropbox founders. We'd have to kind of keep this system running with no incidents, no downtime, no data loss, whatever, for six months before we would delete any of the data from S3. So we have double-rided. and um and there's one point in time halfway through this process where there was a bug got through to production um and it didn't nothing bad happened with the bug but it was a it was like a it was it was like a bugs slipped through our multiple layers of you know release process etc and um so i went to the vp and i said hey a bug made it through to production and we're going to reset the launch clock.
38:56James Cowling:And as a result, it's going to launch later and it's going to cost us some amount of money. Let's say double digit millions, you know? And they were like, great. Thank you. That's good. I trust you. And that was a cool thing. Dropbox had a lot of incredible cultural values, but that was like, okay, cool. If you're prioritizing user safety, that's the right thing to do. And no one was mad about that. It was almost like they were proud almost. You know, it was just like, it was like, yes, you're operating in accordance to the principles of this company. Yeah. Okay, so it was double writing both systems for a while.
39:35And then -
39:35James Cowling:Some subset of data, yeah. Okay, and then you switched over reads. Once we were sure the system was durable and we had all these validators running in production 24-7, they were just migrating as fast as we could. I think at some point we got up to 700 gigabits per second, 764 gigabits per second of peering bandwidth between Amazon servers and ours. Certainly someone on the network team over there noticed that there was that much data moving out. At one point, I got a slightly nasty email from someone saying it was super weird. I think the phrase super weird, it's like it's super weird that you're doing so many reads and not that many writes.
40:14James Cowling:And I didn't respond to that email. But to be fair, Amazon were a great partner for Dropbox. Dropbox still uses AWS. They always, there was no concern that they would do the wrong thing by us as a company. I mean, we only had excellent experience with AWS. But yeah, it was a moment for us. It probably strained the relationship somewhat. You mentioned simplicity and intuition-wise, it makes sense. Do you have a concrete example, though? Yeah, I've got a concrete example for you. And maybe it shows the difference between academia and industry. So the storage system is a giant distributed system with a file stored in various locations.
41:02James Cowling:And so you need a mapping from the file to where it lives on these disks. And so all we did was had a cluster of a thousand MySQL nodes. all right big giant database and it was indexed by the block id and it said this block is on these disks and that's a pretty like simple it's not sophisticated you know and it and every time we'd hire someone out of academia or maybe from other companies they say oh this is not very sophisticated because like you could use a patricia try or you could use a distributed hash table and that would map a block to a set of locations. And I think that's optimizing for the wrong thing.
41:51James Cowling:Because the really nice thing about dumping a list of files and the locations in a giant database, it is written in one location. If I want to validate what happened, if I want to check all the data is where it's meant to be, I just walk over the table and check. And we did. We had services constantly walking over the table and checking. Whereas if it was a distributed hash table or some giant complex data structure, it's very hard to validate. So, I mean, designing for validation is very important. Designing for understanding is very important. It's not about getting a system to work. It's what do you do when it doesn't work, right?
42:25James Cowling:And so having a very simple boundary, that's a very basic example. There's more sophisticated examples that take more time to explain. But something like that is a lot of engineers will want to do interesting work, will want to advance in their career. They want to be seen as an intellectual problem solver. And so the tendency can be to design complex systems. And my argument is always that simple systems are way harder to design than complex systems. Like simplicity is so hard. And I think to like maybe the untrained eye, a simple system can seem like obvious. And the best compliment you could ever get about anything you design is people say like, oh, isn't that the obvious way of doing it?
43:16James Cowling:It's like the same as convex. You know, people say, oh, isn't that, what's that? That's just like the obvious way of structure. I'm like, great. Because it wasn't obvious when we did it, no one else was doing it, right? Everyone thought we were idiots. If after the fact, people think it's obvious, then you really nailed it. I think, but I think that's a, it requires an understanding that simplicity is the hardest thing in systems. And because simplicity is scalable. And I don't, yes, simplicity is scalable in terms of numbers of queries per second, right? But what I really mean about scalability is you can take a simple system and have it run for five years and have people work on it for five years and have all sorts of features added to it and have requirements changed because the company realized the product didn't work the way it wanted to work and it wants to change things and it still stands the test of time.
44:08James Cowling:Whereas a complex over-optimized system will not. And I think that's the tough thing about distributed systems design, especially LLM, augmented distributed systems design, is just because something works doesn't mean it's maintainable over a long period of time. It doesn't mean it's understandable. It doesn't mean it's cleanly architected and abstracted. That stuff's really very hard. Absolutely. And I agree with you. I think it's the long-term beneficial thing to do. One unusual thing, though, in the industry that I've seen is the incentive system for engineers is actually, I mean, you mentioned the desire for an engineer to want to be seen that they can do something difficult.
44:50there's that but there's also the incentive system of promotions and I've had many friends whose promotions were rejected because their work wasn't complex enough and so that kind of forces it's a forces complexity which is kind of unusual I wanted to know what you thought about that yeah
45:08James Cowling:I mean it almost angers me I just like it so much partly why I started my own company you know I think the ideal for anyone is to be doing work where you're being appreciated for solving the problem. Like, you know, and this is, if we get philosophical, this is what it was like going back to the farming days, right? There was no incentive to make it really complicated to milk a cow because the goal is to like milk the cow and then the reward is you got milk, right? And I think it sounds so silly, but at a startup, that's the same thing. The startup, the goal is to build the system, have it work, have the users like it, have it grow.
45:51James Cowling:And everyone gets rewarded and celebrated for solving the problem. It gets hard to scale that. So at large companies, you end up with so many layers of organization that people end up building alternative incentive structures. Right. It's like I'm so far away from whatever the hell we're trying to do over here that my goal now is to get all green checkmarks on my OKR plan. But who cares about your OKR plan unless it solves the problem? And so, you know, the thing that really drives me, really drives me insane is when people try to chase artificial goals. right and and i understand that if you're in a company with um like this that you may have no choice in the matter but what i want to tell people there is a better way and you know and that that better way may not be available to you may not have job opportunities near where you are for example but if you do have the ability to go work at a company where like you are being appreciative of problem solving, that will make you so much better as an engineer.
47:03James Cowling:Like, and I see this when I interview people, you know, and if I, you know, do a, I mean, everyone knows Google has tremendous engineering and tremendous engineers. A lot of folks there, though, I'm not that interested in hiring because, you know, if I'll do a deep dive with them and they'll say they build a system and I'll say, well, why did you build it? And they're like, I don't know. The VP told me to. and like, oh, how's the system used? And they're like, I don't really, I think Ads uses it, I'm not sure. And this is a caricature, but I think it's very, very hard to do good engineering in that environment.
47:38James Cowling:You can do competent engineering, but the best engineering comes from a deep understanding of why. And this is something we just drill into, you know, the team here at Convex, or, you know, the team embodies so strongly at Convex is like, everything exists for the why. Like, don't build a fancy load balancer unless it's not needed. It turns out we do need a fancy load balancer. We're building it right now. But you should always start with why are we doing this? What's the point? And I feel for people stuck in environments that are not like this. But you know what? I don't know. Try to fight the system a little bit.
48:16James Cowling:I think I do see a lot of maybe nihilism, a lot of defeatedness sometimes amongst junior engineers, a lot of this cynicism. Like, what does it matter? like who cares it's just a big organization and nothing matters and but i think it does matter like i've just if i think of like the happiest times in my life it's been like just dedicating myself to a cause and and and trying really hard and and trying to do the right thing and and i felt good when i went home and and not trying to get promoted just trying to do the right thing and then and then assuming i'm gonna get promoted if not go somewhere else um i know it does sound quaint when I'm saying this, but I think it's possible to do this.
48:57James Cowling:And especially possible if you surround yourself with people like this. And if someone is in a big company and they're feeling frustrated by politics, look around and see if there's a team of folks who just seem to want to do the right thing, just seem to want to do good stuff. I don't think that's selling out. I think that's being true to yourself. That's what real engineering is. Not trying to make a complicated, fancy thing to get promoted. Just build the coolest thing that solve the problem. This really reminds me of something you had written. I thought it was really good writing. And in the writing, there was this idea of system bias.
49:31And you have this quote you're writing. It says, here's some examples. It says, the team is spending six months to improve performance by 10 % when it was completely fine to begin with. Or the team is trying desperately to force their tooling on clients who don't need it. Or the team is riding their outdated system to the grave, like the captain going down on the Titanic. I've definitely seen examples of all those types of things in industry. And so, yeah, I think it was in the context of your writing about what you should orient your team around, not systems, but actually missions. And maybe that's a way to fight system bias.
50:19James Cowling:Yeah, I mean, one of my jobs at Dropbox wasn't the most fun job, but it might have been one of the most impactful jobs is shutting down projects, you know, looking around and being like, huh, that thing over there that has had 60 people working on it for two years doesn't seem to make a lot of sense to me. And then it wasn't like a hostile thing, but I'd go and chat with a team and I'd say, hey, what are you all doing? Do you believe in what you're doing? Does this make sense? And the team in private would say, oh, I don't really know. I don't really. But inertia is so strong. This whole desire to not get in trouble, to just keep doing what you were previously doing is so strong.
50:59James Cowling:And talented people can end up doing things that don't make a lot of sense. One of the things I said in that article is it's kind of a cheesy story. But when we started building the storage system at Dropbox, the team was called the Magic Pocket team because that was a silly code name for the system we built. And so the team was oriented around building that system. But as soon as we shipped it, I renamed the team to the storage team. And that actually took a bit of work because you had to, like, rename all the email addresses and the channels and the repos and the things. And so it seems like a waste of time.
51:35James Cowling:And my argument to the team was that the responsibility of the storage team is not to advocate for magic pocket the storage system. It's to solve the needs of storage for the organization. Because who else in the company knows more about storage than the storage team? And if there was a point in time where S3 was a better idea, it would make sense to move back. or maybe there's a different kind of storage system that we were meant to use, right? It's the job of the storage team to advocate for moving back, right? And so I've seen this before. You have like a team called the puppet team, you know, for people don't really use puppet that much anymore, but like for, you know, the puppet, the job manager or puppet versus chef.
52:21James Cowling:And then they'd be kind of advocating for their team's thing when really a team should be oriented around what problem do they solve? They should not care about the system that survives. Because if you are on the magic pocket team and someone says we should move back to S3, that's pretty threatening to your identity and to your career, right? But if you're on the storage team and it turns out it makes sense to move back, I don't think that's the case, but then that's an exciting new product for you to own. And so I think it seems like such silly management philosophy, but I think it's really, really important to orient a team and an identity around solving a problem and not owning and defending a system.
53:05James Cowling:Because you just see this in big companies. It's just inertia is so strong. Yeah. And you see people doing things they don't believe in because that's just what they do. If inertia is so strong, how did you fight it and close down all those projects? I think I got lucky insofar as I was there pretty early on. And I worked hard enough. I was putting in probably 16-hour days at the start. I'm not advocating for that, but I was dedicating my life to the company. And I think it became pretty obvious to people that I cared. Here's a guy over here that really wants to do the right thing and cares about the company.
53:41James Cowling:At that point, you build up enough confidence, capital, you know, that you feel comfortable saying things. You know, I wasn't at that point. I wasn't afraid for my career. I just I was afraid for the wrong decisions getting made. So I so so I felt psychologically comfortable making observations. And then at a certain point, you know, there's another, you know. a lot of engineers so at one point i would mentor um you know all the a lot of the staff plus engineers at the company and and people would sometimes get grumpy you know because they were like oh we're not doing the right thing over here or this is inefficient and every every engineer listening to this has a story like this they're annoyed about some inefficiency at the company and my response was generally like do you think we should solve this problem right now because if we should, let me know the team to take some engineers off and the product to shut down and I can redirect resources and we can solve this problem right now.
54:43And they'd be like, oh, well, we shouldn't shut down any other stuff.
54:46James Cowling:I'm like, cool, we just have this many engineers right now. And so if there's anything lower priority, let's stop doing the lower priority thing and do the higher priority thing, this thing. And oftentimes the answer was, oh, no, nothing else is lower priority. and then the answer is we just have to accept just have to accept right there's no point in being angry or upset that we're not doing the right thing all the time so i think there's a dimension to you break it you bought it i don't think like when i said part of my job was shutting products down it wasn't going around just causing problems right it's ideally like solving problems like oh this product is not going the right direction let's redirect it and do this alternative thing And so I think the thing that helped me, I guess, was having a sense of ownership that like instead of complaining, I just wanted to go fix problems.
55:39James Cowling:And that was part of the culture. I mean, there was when I started at Dropbox and I remember being at Dropbox and Infra was, I don't know, seven, eight, nine people. And, you know, we'd have someone join from Google, for example, and they'd say, well, someone should go build. I can't do anything without this logging framework. Couldn't possibly do anything without this logging framework. I'm like, well, we're going okay without it. And like someone needs to build this thing. And everyone would be like, well, who is someone, right? Because it's just us. It's just us. We build it or we don't build it.
56:11James Cowling:And they pretty quickly come to understand, oh, wait, it's just us. There's no other idiots out there. We're the idiots, right? So that was, I don't know, I loved that time because it was just a time of accountability. Life gets easier and harder when you realize that everyone else is not an idiot. When you realize everyone else is just dealing with their own stuff, right? And so I think, I do not think someone will have good luck going around complaining about stuff and just saying this is a dumb idea and being negative. I think people will have, everyone wants problem solved though so if you're someone in an organization who is who's willing to put their head up and say you know what i think this thing over here is a bad idea but here's a different idea and i'm willing to own it and put the effort behind it i think that's a that's a recipe for success from that article you had a great quote or i guess a great question to think through this it said you went to everyone or a bunch of people and you would say if we could be spending these resources working on any project at the company right now, would this still be the best use of time?
57:27And I feel like it frames exactly what you just described really cleanly.
57:32James Cowling:Yeah. I mean, I guess this is like maybe a trick for being a tech leader or a manager is almost nothing's a yes or no question. It's a prioritization question. It's not like, should we redesign the database? I don't know. Maybe. I guess. Is it the most important thing going to do right now? No. Cool. Let's not do it. I think it's much easier to have those conversations than to think about it as a yes or no. I often hear, oh, my VP won't let me do blah. Okay. Well, it could be that your VP is dumb. That is probably not. It could be your VP has a different set of priorities. Maybe their VP knows that you need to ship these features.
58:09James Cowling:And if you don't, then the company is going to be struggling. I don't know. But to really frame it around prioritization. And the way I think about the career ladder for engineers is, I think some people, maybe it's not common, but think that like becoming a senior engineer means you get better at programming. But like, I don't know, I think my programmabilities went down from like level four onwards. You know, I think once I got to like level four, that was peak programmer for me. And then I probably went downhill from there. And I got wiser, whatever. But I think the real thing that happened is the scope that I cared about increased.
58:46James Cowling:So at a certain point, you're just thinking about, cool, what matters most at the company for the next five years? And I don't think a junior engineer on their first year of the job should try to do this because you probably don't yet have the wisdom, insight, knowledge to be able to make a good assessment. And I think it probably would be a mistake to try to come up with redirecting company strategy. But as you grow, I think the real key part about growing is the IC678 engineers are not necessarily the best programmers. They're just getting better at having broad perspective in decisions within a company.
59:29OpenAI, Anthropic, Cursor, and Vercel all use this product to make their lives better. And the problem it solves is when you're building SaaS or an AI product and you want to sell to other companies, there's all these requirements you need to meet. There's SSO, there's SCIM, there's RBAC, there's audit logs. These are all things that take time to integrate but aren't the main focus of your app. WorkOS is an API layer that lets you meet all of these requirements in just a few lines of code. So let's say you have a new SaaS product and you want to sell to other companies. WorkOS will solve all of these critical feature gaps for you.
1:00:06You can check them out at workos.com to learn more and get started. And I appreciate them for supporting my work and sponsoring this podcast. You mentioned this idea of do the right thing and then the byproduct, you also get promoted. And I think that's the dream. Do the right thing, get promoted. but in reality oftentimes people would have to make the trade-off imagine a two by two matrix of doing the right thing doing the wrong thing getting promoted not getting promoted obviously that you know do the right thing get promoted great do the wrong thing don't get promoted
1:00:43James Cowling:obviously bad but I'm curious about the other two quadrants which one would you have picked when you were earlier in your career let's say I came to you I said hey you could do the wrong thing but you're going to get promoted or you can do the right thing and i guarantee you're not going to get promoted absolutely the second one so i'm going to say a really tacky thing right um there are a set of engineers um who at a certain point just make infinity money like the amount of money they can make you know going and whatever going to work at anthropic is is a huge amount of money right and so there's a certain point where it just doesn't matter anymore like the money is whatever, right?
1:01:25James Cowling:So if you really want to maximize long-term income, I don't think you should try to do that. But if you wanted to maximize long-term income, if you get to a certain level of experience and skill and seniority, you've made it. That's it. You're done. Money's fine, right? And so I think there's this desire amongst junior engineers early in their career to be kind of over-optimizing for promotion and salary, et cetera, as opposed to investing in themselves. And I guess they probably don't want to hear me say this because it's easier for me as an old guy to say this, right? But I think, you know, for the longest time at Dropbox, there was a time at Dropbox where like, we just didn't have our level system right.
1:02:10James Cowling:I was tech leading the team and I was making the least amount of money of the team because I was at some point a tech lead manager, I saw everyone's salaries. I was making the least money, but whatever, right. It all worked out in the end. Right. And that was a happy story. But I do really think I benefited so much from working with the best people. Like if you have two choices, like making 20 % more money now or working with like the best people in the world, right. Just be around the best people, because like you will be, that's going to set you on the ship to success. And now, not everyone wants to do that.
1:02:49James Cowling:Not everyone wants to go on that. There's no shame in just wanting to have a regular job and just be chilling and just be getting paid. And fine, there's nothing wrong with that. But if you do want to maximize, if you want to maximize your career growth, Then the way to maximize that is to maximize your skills. And hopefully you're not so cynical to think that there is no correlation between talent and compensation. If you think that's the case, okay, I don't know what to say. But there is a correlation between talent and compensation and growth. And so my advice to people early in their career is land at the best company with the best people doing the most important problems.
1:03:37James Cowling:And your life will be great because we are so lucky as engineers. Like where else can you work solving problems for a job, like solving puzzles for a job and getting paid so well? Like engineers get paid so well compared to most jobs. The privilege, the real privilege that we get is to work on something cool. and that's what I advocate for. I love the idea you mentioned earlier that's kind of counter to the common opinion I often hear. Like the common opinion I hear someone working in a big company, they're in this machine, it's kind of defeatist, they're going, ah, I shipped this thing, but I hate this or I don't believe in this.
1:04:21And I really liked your perspective on doing the right thing and like basically giving a damn that someone outside of you is doing the right thing too and expanding your sense of ownership. And I want to know, what is your motivation for that? Because there's so many other people who are faced with the same inputs and they get to a very different conclusion. So what motivates you to actually do the right thing when the machine doesn't necessarily incentivize you to do so?
1:04:51James Cowling:Yeah. Well, I think the machine does incentivize you to do so, but I don't think it's visible immediately. I think people, for example, who do less job hopping really do grow the most because I would say very strongly, if you're not in a job for three years, you're not going to see whether your decisions were good. You can get more money as a junior engineer, but you cannot become a very talented senior engineer without being around for long enough to own the consequences of your decisions. It's like playing a basketball game and leaving before the game's over. Like you're just not learning, right?
1:05:29James Cowling:So I do think there's an actual structural incentive towards staying in a job. Now, if you're in a bad job, leave the bad job, right? But there's a structural incentive to be able to stay long enough to have an impact. But also, I don't know, like, again, I guess it's easy for me to say. And also I joined the tech industry when it wasn't like a very lucrative field. Like, you know, it wasn't like it is now, you know? But I just like it. The reason I left academia was because I wasn't confident I was making the right. So what would happen how I was in academia? I'm not saying everyone was like this.
1:06:06James Cowling:So firstly, you don't know what to do. So you have to make up a problem to solve. And you make up a problem. And then you make up a solution. Hopefully a good one. And then you write a paper where you try to convince everyone that it was a really good idea. and I hated it because I didn't want to convince people. I just wanted to build it and see if it was good. Like I just wanted to build it and ship it and see. I didn't want to play pretends and argue about it was good. That's what makes me feel good. I don't know. And like obviously if you're an engineer who's living paycheck to paycheck, you should maybe ignore what I'm saying.
1:06:50James Cowling:Like, I don't want to come across as unempathetic to anyone who's really struggling financially. But if you are not struggling financially, the best thing you can do for your quality of life is, you know, be enjoying what you do every day. You know, like, I don't know, you could have a fancier car or you could enjoy what you do every day. And I love cars. Like, I'm a motorhead. but I promise that enjoying what you do every day is going to have a much bigger impact on your life and um personally I don't enjoy like going to work and and working on pros I don't believe in like I can only operate and say Jamie my co-founder's the same way like we we really get along really well in this respect we can only um we can only put up with doing things we actually care about and believe in.
1:07:43And I think that's the luxury.
1:07:46James Cowling:That's more luxury than taking a first class flight. That's more luxury than going to a three-million-star restaurant. The luxury is you don't get up and have to do a terrible job. You get up and go to a place that you like your coworkers and you solve cool problems and you go home and feel proud of yourself. If you can if you can construct your career that way and it's not easy you have to it requires very active effort then you're going to have a good life i saw in your career journey it said that you're occasional manager and you seem like someone who really enjoys the the technical aspects of things so how'd you decide to occasionally dip into management i feel like most people should not want to be managers and i sometimes think it's a bit of a red flag if someone wants to be a manager too much right because um being a manager is a hard job and uh so i think most people should go into management because they need to because it's necessary and so what happened every time i went into management i'm in a management role now as well because you have to right um but um someone needed to manage the team and and so i do genuinely enjoy accountability i like responsibility.
1:09:00James Cowling:I like having weight on my shoulders, I suppose. And so I went into management several times. But as soon as I had the opportunity to get out, as soon as someone else would come in and manage the team, I would kind of bounce out of that role back into engineering. Now, going into management is tremendously educational. Being a manager really lets you see the world in a different way. You realize companies are more complicated than you thought. you realize the engineers are more complicated than you thought you realize everyone on the team is going through something and you're like oh wow now i understand why that thing happened so being a manager is very educational one thing i would say i would strongly caution people against is going into management too early in their career and i see this happen um i see this happen a lot with well-intentioned people who want to push people into management as a means of career advancement, right?
1:09:55James Cowling:I think it's doing people a disservice. I think you should let folks take their time in a career journey. I don't think you should be going into management under most circumstances in the first three years of your career. I think you should get to, ideally, get the staff engineer before you do that. That's not going to happen for everybody. But ideally, you take the time because if you go into management before you're an excellent technician, it will limit your ability to influence strategy later in your career. How does that play out? It plays out with people who are stuck or have roles as people managers where they see their job as maybe making a team happy or maybe coordinating a team or dealing with all the day-to-day challenges of people on a team, but they're not organizational leaders.
1:10:49James Cowling:They're not like pushing the team towards excellence. They're not like, hey, how can we reframe what this team is about? How can we help influence technical strategy? And again, people management is a fine job, but I think the best companies are ones where all the managers are very technical. And so they're able to make sure the company's doing like a certain point. If you're not a particularly technical manager, you're not going to be able to evaluate the work of your team. You're not going to know whether your team is even doing well. And I guess maybe you might contribute towards some of the stuff you were saying about, you know, cynical organizational attitudes where like my manager doesn't understand me.
1:11:29James Cowling:Maybe your manager doesn't understand you if they haven't spent enough time developing technical skills and struggling. Yeah, there's another piece that you wrote I thought was really good about, you know, leading by example. And you actually say it's bad to lead by example, but there's a quote in there. It says, modern tech workers can be an anti-authoritarian bunch at the best of times. Let's say you're a tech lead or tech lead manager or manager. How do you strike that balance so you don't lose credibility as a tech lead? Yeah, there are command and control companies that the manager just tells people to do things and they just do them.
1:12:12James Cowling:Those typically are not excellent companies. And not every company needs to be excellent. But if you want to have a company where your team is really innovating, or the engineers feel very personally responsible for making really high-quality decisions, you have to have them believe in what you're doing. And so at Convex, I'm the founder and CTO. I can't really go to someone's desk and say, do this thing. Now, they might do it just because they like me. There's a good chance they would do it, but not because I'm the boss. They'd be like, well, why? And that's because that's our culture. I mean, we don't just do things because we're told, like, if I want someone to do something, I'll spend time with the team talking about why we're doing it, where the market's trending, where's the gap in our product, what, you know.
1:12:57And then once we've really well articulated why something matters, people are just going to do it anyway.
1:13:04James Cowling:Now, sometimes you have to have hard conversations, but I think the way I think about it, certainly in terms of conflict in an organization, there's this kind of hierarchy of the values you have, and then why, what, and how. And engineers are very often debating the how. They're very often debating what algorithm should we use for this? And should we use this container service or that container service? And these are kind of like the implementation details. but most times when I see organizational conflict it's because well-intentioned people are debating as best they can about how to do something but they don't agree on why we're doing it.
1:13:43James Cowling:So if one team thinks the most important thing we can do right now is get more features out to expand our customer base and the other team thinks the most important thing we can do is increase reliability because there's a risk to the business. They're both very valid perspectives but it's going to lead to them doing very different things. And within an organization, I'm a strong believer that everyone needs to have 100 % why alignment. So the stuff we argue about or we debate or we talk about ad nauseum is why. And then largely, I just trust the team to do the right thing. And I think that's the case for any tech lead.
1:14:18James Cowling:If you want to have credibility on a team, if you think that you can just tell someone to do something and they'll listen to me because I'm the senior guy, and no, they won't. They won't. If they listen to me, it's because I've come in and explained it in a way that resonates with them. And again, one of my jobs at Dropbox was to resolve the situations. Someone would say, oh, that team's an idiot. They won't do blah. And then I would go talk to that team. That team was not an idiot. And we'd talk through it, and they would decide to do the project. And it wasn't because they were scared of me.
1:14:51James Cowling:I hope they weren't, at least. It was because I'd take the time to figure out what their motivations are. And that's something you really have to learn. If you're a tech lead, you don't really have that much authority over people. And really, you have to kind of encourage them and get them to believe in what you're doing. And if you are leading through authority, you're not going to have a culture where good ideas arise from within the org. People are just going to do what they're told. Influence without authority is huge. I mean, even big companies like Meta, they technically don't have titles.
1:15:27James Cowling:everyone's just a software engineer um is that also how convex is run i mean we're we're a very flat organization there are people in tech lead roles i don't completely buy the everyone's a software engineer thing it's like you know it is you know anthropic everyone's a member of technical staff because it's kind of like a wink wink thing you kind of know it's like no one's mentioned the title but you kind of know if like the most senior person in the company comes to your desk, you're probably going to notice, you know? And so I think there's no point in, you know, playing pretends. People know, you know, but I will say that you should have an organizational culture where you don't do something just because a senior person says something.
1:16:12James Cowling:Now, I do think that reputation matters. Like, I certainly think that, you know, I don't know, if a very experienced person at a company comes and tries to explain something to you, you probably should listen to them because they probably have some wisdom. There might be something to be learned. They're usually open to it. But ultimately, yeah, even the most senior person I don't think should be leading through authority. They should use the benefit of their experience to be able to win the hearts and minds. And that's how you, engineering culture is so important. And if you want a culture of ownership and innovation and drive and enthusiasm and people trying to do the right thing, not get promoted, right?
1:16:55James Cowling:You have to have a culture where everyone believes in what they're doing. And no one's going to believe in what they're doing if the senior principal engineer comes to the desk and says, I'm not going to tell you why, but you have to delete this database and do this other thing. That's not an empowering statement. What the empowering statement is, hey, let's spend some time together to talk about where this product's trending and how this is probably not going to work out and how there might be a different way we can solve this problem. On the topic of tech leadership, I mean, in that article, I mentioned that, you know, the title was don't lead by example.
1:17:27And I think that kind of might be confusing for people. Can you explain why you think you shouldn't lead by example?
1:17:33James Cowling:Yeah, leadership by example is a very passive thing to do. And engineers are passive people at the best of times. You know, I think that's stereotypically a little bit part of our personality. So, I mean, the very concrete example is, you know, when I first started becoming an engineering leader, I was trying to lead by example. So I wanted to kind of demonstrate the behaviors I want everyone else to have. And very specifically, like with regards to on-call and people getting paged, I wanted people to have high ownership. I want people to jump on issues as soon as it happened. So I would do it.
1:18:08James Cowling:I'd be always the first one to respond to a page. I would always be writing up the reports. I would always be jumping on all the bugs and stuff and really falling over myself to kind of show how I want people to be. But from their perspective, all they see is that the lead's just doing all these jobs and they don't know. They're like, oh, maybe that's James' job. Or maybe James knows how to do it and I don't know how to do it. I didn't know. I was just kind of figuring it out. or maybe he likes doing those things. And I think at a certain point, being a leader is about understanding human psychology.
1:18:48James Cowling:And I don't think you can just act a certain way in front of people and wait for them to copy you, right? I think you should, obviously, you should act with integrity and values and you should own the values of the team. But you have, sometimes you have to explain stuff. Sometimes you have to tell people, hey, very specifically, there's a trajectory almost every leader goes through, every high achieving leader, where they become a tech lead and they care so much that they become a micromanager. Like they review every line of code, they're involved in every decision. And at a certain point, they are the bottleneck for the team.
1:19:26James Cowling:So this person is so overwhelmed. You might've gone through this yourself, you know? They're so overwhelmed and they're like, wait, my team doesn't even seem busy right now. And I'm so busy. I'm reviewing all this code. I'm doing the strategy. What's going on? And then a manager will come along and say, hey, like you're micromanaging. You got to let your team have more ownership. And so the tech lead says, okay, sure, whatever. I'll let them own stuff. And they'll just take their hands off the wheel. And then the team falls apart because you can't just stop doing the things you're doing, right?
1:20:00James Cowling:You have to go and have a conversation with people. And so I see it. I see leadership as this kind of slider between oversight and accountability. So when someone's new to the team, when they're very junior, they're in a mode of oversight, like you're checking their work, right? But at a certain point, you have to dial down the oversight and very importantly, dial up the accountability. So instead of saying, I'm not going to look at what you're doing anymore, you say, okay, cool. You've got this project. Let me know when it's going to get done. Next Thursday. Cool. All right. What's the plan? Is it going to happen?
1:20:38James Cowling:Okay. How are you going to know it's correct? Great, great, great. It's on you. I expect you to do that. Great. Let me know if there's any issues. And then the ownership relationship is explicitly on them. Because what you want to do is basically encourage ownership within teams. You can't go from owning something yourself to not owning something yourself and expect that to develop. You have to to go have those conversations. You have to go and give people accountability. Now, what I have found, even though that can feel like an awkward conversation, most people genuinely like accountability. Most people like to own their work.
1:21:15James Cowling:Most people like to say, hey, this is on you. We're all going down with the ship. I'm not going to leave you high and dry. I'm the tech lead. I still take responsibility to. You're accountable to this project and go for it. And that's how you can develop people within your team. You know, a tech lead shouldn't be about you as the boss and the team as like the people who, you know, do the work or something, right? It's about this kind of flow where you're the more experienced person typically, maybe the more organized person, the more strategic person, and you're working on developing your team members so they can take your job.
1:21:53James Cowling:And then you can do something else. That slider you mentioned, what if you give someone accountability and they blow it? Do they go back down the slider to where you start micro-managing? They have to recognize it, right? So this is growth. I mean, this is growth, right? It doesn't always work. And so I think I had another article about paper cuts or something. So we'd call these paper cuts. So some decisions don't matter that much. You get them wrong and maybe it sets you back a week. You get them wrong and maybe something's a bit suboptimal. And these are great candidates to give people accountability for because they do it.
1:22:34James Cowling:And if it doesn't work, they get to experience it and it will feel bad. And I guess that's good. It's good to feel bad. You know, you shouldn't be demoralized, but it's good to try something and it doesn't work. And you're like, oh, wow, that didn't work. And then you feed that back into your little LLM in your brain and you get better next time. Right. That's a growth experience. I think what you don't want to let people do is lose their arm. Paper cuts are fine, but I would not let someone, a junior person, especially design the replication system at Convex. You have to have safeguards. And one of the arts as a CTO, a leader, a tech lead, is figuring out the right level of altitude for how to know whether you can trust someone to take on ownership for something.
1:23:23You mentioned, because you're the most senior engineer at Dropbox at some point, that you had to mentor other very senior engineers. And, you know, when I think about mentoring a junior engineer, it's relatively straightforward patterns. But when I think about, let's say I need to mentor a senior staff engineer, maybe even a principal engineer, how do you mentor someone like that that's already so polished? They already know how to take ownership.
1:23:49James Cowling:um yeah how do you mentor someone so senior yeah i mean there's in two ways what one is there is still just a whole bunch of commonality everyone goes through the same problems they they don't know how to deliver harsh feedback they they don't know you know there is a bunch of just standard stuff people working on but i do think that um that at a certain point everyone needs to become the best version of themselves sounds so cheesy right but there is not an archetype for what, according to me, there's not an archetype for a senior principal engineer. You've got to be your own brand of engineer.
1:24:26James Cowling:So you might be, for me, I'm like the strategic kind of collaboration, simplicity, abstraction engineer, and my coding just fell off a cliff. Some engineers are like the deep science fiction, the super hardcore science problem-solving engineer. So everyone's going to find their own brand. So oftentimes what I'll be working on them is a combination of two things. One, finding their strengths and helping them be more spiky, helping them to like really excel in the area that makes them special. At the same time, almost always working on the personal side of the engineering, like the organizational, personal, understanding people, understanding the why.
1:25:10James Cowling:It's just, I don't know, we're so mathy, you know, as an industry. We think that somehow like even silly things like like, for example, I have to tell so many senior engineers that you can't change someone's mind in a meeting. You just can't do it. Like you can you can you can disagree with someone and they're going to be mad or whatever. Right. But like but to change someone's mind involves going through a complex series of like of neurological processes. Right. That like don't happen live in front of 10 people in a meeting. And so if you force someone to agree with you about them being wrong, they're just going to go along with it and their ego is going to get bruised and whatever.
1:25:52James Cowling:So one thing, firstly, meetings generally aren't for decision making. Most of the time when you identify a disagreement, I would point out the disagreement and why I don't agree or where the problems are, provide enough information. and then let it sit and let them go and reflect on it and like come back and have a conversation a week later after they've gone back and thought about the why, right? And so this is kind of very psychological, but it makes no sense to be like, well, they should be able to change their mind. I mean, well, they won't, right? That team should do this, but they didn't, right?
1:26:28James Cowling:You know, people should like my API. Well, they don't, right? And I think it's like that whole like, it's almost um it's almost like a capital a capitalist attitude towards like interpersonal behaviors right it's like doesn't like you know capitalism rewards success you know or or or impact it doesn't matter how well intentioned you were if you didn't manage to convince them you could be the smartest person in the world but if you can't change someone's mind then you were kind of useless, right? And so I think a lot of senior unions still struggle with this. It doesn't matter how smart you are.
1:27:05James Cowling:It doesn't matter how right you are. Are you effective? And a lot of being effective at that level is about uncertainty and project management and simplicity, blah, blah, blah, blah, blah, but also kind of interpersonal dynamics because most hard problems happen with a team. Not always, but generally when you're at that level, you're going to have to have 10, 20 people working with you to get something done. And that's a whole different ballgame. On the topic of career advice, the industry's changed a lot in the last five years because of all these agentic tools. And I wanted to know if you thought, is there any career advice that majorly changed in the last five years?
1:27:46Something that you used to say five years ago, you don't say anymore or vice versa?
1:27:51James Cowling:No, I don't know if I've changed my perspective much, but I think the industry has changed this perspective dramatically. I think, I mean, let's just be honest about it. There is incredible demand in Silicon Valley for senior engineers, and it is getting harder for junior engineers to succeed and grow for a variety of reasons. Why is there demand for senior engineers? Well, because LLMs can't do everything. Well, you know, the architecture and simplicity and design, they're still the domain of human beings, despite what you might hear on Twitter, right? And every company, including the labs, are desperately hiring senior engineers, right?
1:28:29James Cowling:But junior tasks are getting a little bit commoditized, and that worries me. Because I do think that that learning, it's very, you know, wisdom is kind of facts put into practice and then synthesized, And so a lot of people would argue that, oh, well, it's easy to learn now because of chat GPT because you can just go ask it a question about how does two-phase commit work or what's the difference between snapshot isolation and serializability. And it will give you probably a pretty good answer. But I think growth as an engineer does require wisdom. And wisdom only really happens when you synthesize it, in my opinion.
1:29:13And so I think I'm still very bullish on young people.
1:29:20James Cowling:We're hiring junior engineers at Convex. I'm very excited about, I mean, I love working with junior engineers who are really hungry to grow. But what I would say is train your mind. Like, do not listen to anyone who tells you that there is an advantage to having less knowledge. I'm not sure if you've seen people say these ludicrous things that like, oh, maybe in the future, like not knowing engineering will be an advantage because you won't have biases and you'll just use Claude. I think these are like ludicrous statements. Like software engineering is an intellectual discipline that helps you think.
1:29:56James Cowling:Like the software engineering, the best software engineering is not about knowing syntax. And it's not about knowing an algorithm. them is being really good at conceptualizing problems and be able to break them down to building blocks and coming up with clean solutions to them. And that requires experience that requires like it's like doing doing weights with your mind. And just like if you went to the gym, and you just like, it never hurt, like if you went to the gym and just picked up really light weights, you're not growing. And also, if you went to the gym and picked up a heavyweight and like, okay cool i think i can do it and you let the robot pick up the weight for you the right you're also not really growing right you have to do the reps and so i would say and it's tough but find a way to stress your brain every day find a way to to avoid now obviously agentic coding is here right obviously that's you know i could never tell someone to like never use a coding agent because that'd be silly, you know, but I can say that it is easy to fall into a passivity trap, like where you're just being passive about your learning.
1:31:06James Cowling:And I do think you need to spend some time in the intellectual wilderness of not being able to solve a problem and struggling. I would say, if you're running into a new problem, try to think of a solution yourself and then go check with an LLM, it's going to be hard for you. It's almost like, for example, I'm not very good at reading anymore. I've got to be honest. I would find it hard. I mean, I can read, obviously, but I find it hard to sit down and read a book because my brain has been fried by the stimulation economy, right? I go home and I'm tired and I watch a YouTube video. I don't tend to go home and read a novel.
1:31:51James Cowling:I should be better at that. But it takes discipline to do that. I would say similarly, it's getting harder to solve a difficult problem without reaching for the help. It's getting harder and harder every day to be faced with a very difficult intellectual problem and not being like, well, I'll just do a Google search or I'll just ask Claude, you know, whatever. I'm still doing the work. I don't know. I would really encourage people to practice using your brain every day. If I was just devil's advocate or just thinking from the junior engineer perspective, I might think, well, the proof that Claude can do this work today means that I don't need to know it today or tomorrow or in the future.
1:32:33So why even build that skill in the first place? And also these agentic tools, they have positive trajectories
1:32:42James Cowling:too so yes so firstly um the agents are not particularly good at a lot of parts of engineering currently um and so probably everyone should agree there's you know claude is not good at designing distributed systems protocols right now for example you know or managing a three million line code base you know um and so there are parts of engineering where engineers are valuable and like i said you know this because the labs are all hiring engineers no matter what they say they're still hiring engineers, right? Desperately hiring engineers. Really, really aggressively hiring engineers. So engineers still have value.
1:33:21James Cowling:And maybe they won't in the future, but I doubt it. I really do think there's a role for human ingenuity in engineering. And so here's the trade-off I would pose to people. I mean, there's kind of three paths. One, get out. Get out of engineering. Go mow lawns and do whatever if you want. that's the that's the most nihilistic attitude i i don't believe in that i believe i really love engineering i believe that there's promising roles for human beings in engineering so the second is to realize that um there's value right now for humans in engineering and maybe one day agi will be here let's say in the year's time agi will be here and there's two people one person just gave up on problem solving right now and they're just like feeding the machine and they're running 17 coding agents in parallel and they're just accepting everything Claude says and they've given up right and one person said you know what I still want to have an active role in my learning I want to understand what it's doing I want to think hard about problem solving right play it out one year AGI arrives who is going to be better place for the future right the person who's been training their mind right like there was there's you don't engineering is not a means to an end it's a mechanism for for improving your mental processes it's like like i still do whiteboard encoding interviews and by the way anthropic still does whiteboard coding interviews just in case you were wondering they don't say like just use claude right so i still do white whiteboard coding interviews with candidas candidas are getting worse at coding that's absolutely true but i don't know any better vehicle for evaluating someone's intellectual capacity for problem solving than seeing them solve an engineering problem.
1:35:07James Cowling:It's a great mechanism for doing so. It's like, you know, if you were doing a math degree, right, a big part of doing advanced mathematics is proving theorems. And by the way, the theorems are already all proven, right? Part of the exams is proving a theorem that has already been proven. And so you might say, well, what's the point of proving that theorem? It's already been done. Well, the point is that the act of proving that theorem improves your mind. And then you can go and do more innovative stuff. So I would say, sure, you may not. I still do think that there is a very, very promising role for human beings in engineering.
1:35:48James Cowling:Maybe not in coding, but coding and engineering are very different things. right um but even if you don't believe me you can either give up now but or you can just keep trying to improve your brain and you'll be better off anyway imagine if you're wrong imagine if you take the defeat of you imagine if you say well ag is coming next week why bother doing anything and imagine you're wrong oh my god you just got off the you just got off the off the ride you got off the ride there's so much cool stuff i mean this is a cool time for engineering right this is a really cool time there's so much cool stuff happening and i don't and i i hate this talk about like ai means humans don't have a role anymore ai means no one's gonna have a job ai means we're gonna do nothing else new and there's no one's gonna i what i like is hey look at this all this new cool stuff we can build like look at how there's ways we make people's lives better and to be honest i really wish my peers and and my cohort would stop it with the the real doomer like um human elimination kind of narrative and because i think there is a little bit of a like a you know a psychological anchoring you know i and as part of convex want to make the you know we didn't start convex to eliminate jobs we started commerce we want to make it easy for people to build cool stuff and and i think the more we as an industry get behind let's do everything we can to make it possible for people to do more cool things i think it's a really exciting future we have ahead of us i wanted to talk about what you're working on now or convex and why did you quit dropbox to build convex hey dropbox is a great place at a certain point i i was there eight years uh whether i i don't know if i say i outgrew the company but i'd been the most senior year for a while there um you know at a certain point i wanted to grow and do my own stuff you know and so so i left the company very amicably um and um started convex now why convex convex started pre-agentic era because my observation was that that um the real differentiator in the success of a project especially a large project is the quality of the abstractions the quality of the design the quality of the architecture right systems that thrive over time and are extensible of systems that are architected well and in my mind the most difficult challenge in engineering according to me is distributed state management how do we store state reliably and reason about it you know modifying it concurrently with other users and so so we designed a platform for application building based on our experiences building large-scale systems so convex is a is a transactional database where the transactions are written in typescript They run a stored procedures, TypeScript stored procedures.
1:38:41James Cowling:They're serializable. There's automatic reactivity. So what the client sees is a consistent view of what's on the server. And I would say Convex is a very designed platform because Convex is designed to be very composable and fit together well. And so we designed Convex for developers to use. In particular, we wanted to make it so that application developers were able to build complex full stack applications. That was the goal of Convex. Now, all of a sudden, a janky development came along. And that's been really interesting for us because it turns out that what humans find hard is also what agents find hard.
1:39:19James Cowling:You know, coding agents are not particularly good with large code bases. They're not particularly good with reasoning about action at a distance, you know, race conditions across services. They're not particularly good at simple architectures over time. And these are the things that Convex gives you as a developer. So the idea now, and almost everyone using Convex is using Convex because they have the coding agent doing their front end, but they need a backend abstraction, a higher level abstraction than something like AWS or something like Postgres, which makes their problems go away. And that's the company.
1:39:55That was one thing I wanted to ask, because immediately when I think, oh, I just need a backend or something like that, just go to AWS or just host something like that. So this is a layer of abstraction on kind of on top of those types of primitives. Yes. That makes it easier for an application developer.
1:40:11James Cowling:It ties into a lot of the stuff I said earlier about making problems go away. And frankly, I mean, I watched the interview you did with Barbara Liskopf, and Barbara was my advisor in grad school. And we worked a lot together and a lot on abstraction and, you know, the value in clean designs that minimize complexity. And so AWS is a fine tool. Postgres is a fine tool, although none of the mainstream databases are that great, frankly. They're fine tools, but they don't make problems go away. And so the idea of convex is a higher level set of abstractions that you can use and not reason about state management, not reason about concurrency, not reason about scheduling, not reason about transactions, not reason about polling and data sync and type safety and all those things.
1:40:57James Cowling:So convex is a, if you think about the history of engineering, over time, the abstraction floor raises. You know, when Barbara was first starting, she was using punch cards, you know. I don't know if she mentioned it to you, but when she started as a programmer, she never heard the word programmer before, you know. That was the first time she heard the word, right? And then, you know, you went from punch cards to, like, you know, having proper operating systems and, you know, then, you know, languages like C and then higher level languages. and then you had cloud computing. And over time, the abstraction floor raises and you largely forget about what's going on beneath the surfaces.
1:41:37James Cowling:Most people don't think about how S3 is implemented. I do, but that's what I used to work on. But like most people just use it and it just stores your data and it gives it back and that's great. That's a successful abstraction. But I do strongly believe that the world is and has been overdue for a new abstraction. One level up the stack. And especially now that people are doing agent development, largely they don't want to own a Postgres instance. Largely they don't want to think about Kafka versus RabbitMQ. They don't want to think about what set of tools to use. They want it just to work so they can focus on building their application.
1:42:11When you talk about the abstraction, there's obviously a lot of stuff going on behind the scenes in Convex and the technical side. And what is it that Convex is building behind the scenes that you're most excited about and why?
1:42:27James Cowling:Basically, Convex is a new operating system in some respects. So we have the primitives, queries, mutations, actions, subscriptions. What I think is kind of cool is how we built this. We have our own database that we built. We have our own distributed database that tracks read ranges and write ranges and does very efficient subscriptions over web sockets, et cetera. So that's the current operating system set of primitives. But Convex is getting much larger workloads now and much more interesting workloads and more high-performance workloads. And so we're in the process of developing a slightly lower-level API for doing very efficient kind of background processes, singletons, APIs like fork, like operating system primitives.
1:43:17James Cowling:And I'm pretty excited about launching these and how much faster it's going to make various Convex components like the workflow system. And to be honest, the thing I find exciting every day, challenging every day, I still find Convex very hard. Like, to be honest, like, I struggle every day. I don't find my job easy. I mean, I feel confident at my job, but it's not easy. Like, designing the new API for this is super hard. I can't just go ask ChatGPT. It's not going to give a good answer, right? And because it's innovation. It's new ideas. I really enjoy it. I find it stressful sometimes. I find it challenging and tiring, but I also find it exciting.
1:44:05James Cowling:And I would encourage engineers to try to find this kind of stuff to work on, where it's like you're on that edge of like, I'm really liking this, but also it's a bit tricky. You know, it's a bit tough. You mentioned fork, and in operating systems, I'm familiar. You just take the existing process and kind of split it. What's the idea of fork in a distributed system? So Convex almost never has scale issues with regards to live traffic. Live traffic is typically bound by user-facing interactions, people clicking on stuff, running a website, acting on a website. Every now and then, someone will come to Convex and want to kick off a million background jobs to do something, background processing.
1:44:46James Cowling:It's a big workload. Programmatically, you can trigger huge workloads. And so one of the things we have to scale is kind of these background workloads, and a lot of them involve things like scheduling. And there are a lot of workloads in Convex that would be very efficient if you had a background singleton process to perform things like aggregates. I'll give a very silly example. Let's say you're building an election on Convex, a voting system, and every vote is a new row in the table. and you want to show a tally of the votes. One way of doing this is having a bunch of background processes or crons adding these things up.
1:45:26James Cowling:One way is doing a table scan, which is the obvious way to use Postgres, which doesn't scale. The other is to have a background job, which if there's new votes, it adds them all up, keeps a tally. If there's no new votes, it goes to sleep and waits on like a condition variable to wake up again when there's a new job to perform. And so these are the kind of primitives that we're working on right now. Most people won't even know they exist, but allow us to build this very high performance primitive for scheduling, aggregates, background aggregations, et cetera. And I'm pretty excited about the next generation of workloads we can support as a result.
1:46:06When you reflect on your career, and it sounds like you've done a lot of gnarly technical work across your PhD, Dropbox seemed like pretty intense systems work and Convex is also doing a lot of cool stuff. When you look back on your career, what was the most technically stimulating work you've ever done? And why was it hard? And what'd you learn from it?
1:46:29James Cowling:There were certainly times in grad school where we were formally modeling consensus protocols and stuff. And I'd be on the phone with Barbara on weekends and talking through, trying to reason about this in our heads. That was pretty intellectually stimulating and fun. but I think the stuff I found most stimulating was stuff like working on very large storage system with a team where you know things are going wrong you know where where the rubber hits the road that's where I find and this is every day at convex you know the rubber hits they're like you know hey we have a compaction process that runs in the background but it's running into issues we might have to redesign it using partitioning etc I feel most intellectually stimulated where where there's a really clear constraint in front of me.
1:47:16James Cowling:And that to me is engineering. Like if I don't actually know what the definition of engineering is, but I'm just going to make it up in my mind. Engineering is science with constraints. It's like how do you solve problems in the presence of resource constraints? I'm not particularly interested in constraint-free environments. That's art. I like craft and engineering. and the more visceral and difficult the constraints, the more fun that is for me. And I've been lucky enough to, whether it's luck or intention, I don't know. But I've always placed myself in those environments. You know, like let's go get on the hardest team and own the hardest problem and then put the effort in to survive.
1:48:09this question might be a little bit off topic but you know i know you were a consultant for the show the tv show silicon valley i love that show and i gotta hear how'd you get involved with that
1:48:21James Cowling:yeah that was a lot of fun so um a lot of folks might not know this um i had nothing to do with season one so a lot of tv shows they don't know whether they're gonna survive as a tv show so So Mike Judge, who wrote Silicon Valley, also was of Beavis, butthead, and Office Space fame. He started his career as a software engineer at, I think, Lockheed or something. So he actually was a software engineer that a lot of people don't realize. And so Silicon Valley was like a throwback to the kind of work it did. And if anyone's seen the movie Office Space, you would get this. That's really a dystopian cubicle era tech industry film.
1:49:01James Cowling:um so they did season one of of silicon valley and then it was very popular and they got picked up and so they had to figure out what to do for season two but they didn't know what to do because they designed they've written a storyline that gets to the point where there's a compression algorithm and and then what happens and so they needed to find um an expert on compression and i guess ostensibly that was me and i don't know whether i was an expert on compression i guess I was an expert on storage at least. And so, and so they came to the office and, and, and we just chatted and it was so much fun, you know?
1:49:36James Cowling:And so I was involved in, yeah, pretty heavily involved on the show. A lot of it was, you know, storyline. And so first of all, yeah, sure. What would you do with the compression algorithm? What would, could you design a storage system? And so coming up with story ideas that are technically accurate, they were really, you'd be surprised to know how much they care about accuracy a lot of um people i know can't watch that show because it's just so creepily accurate they find it so cringy and partly why silicon valley can this tv show can be so cringy is because it's it's real like those stories are almost almost maybe not everyone so many of the stories looking about are just real stories they just they went to a bunch of companies just farmed everyone for like stories of crazy things that happened in the tech industry and they wove them into the series and all the characters are based on real people and real archetypes um but they also cared very much about technical accuracy so i would also do technical consulting and then they'd say stuff like oh we're building a data center in our house what should the rack look like and you know what should the diagram on the wall look like and and partly you know i wanted to be like oh well it doesn't really matter no one's gonna care and they're like no no no it matters that they really cared to get it right um so yeah i love the show but it it can be hard to watch just because of oh my god how how real it can feel yeah are there any easter eggs where you look at and go that's unusually accurate or or you know that system diagram actually is like very spot i can't i don't think i can even say them because um there were stories of like early dropbox there were stories of having ideas ripped off by other companies and being tricked into having meetings with folks only to you know it may be the the competitor's team there to steal the information a lot a lot of those stories are real and so there's people who watch silicon valley and like oh wow that was that was something i went through and probably it was because it was about that situation.
1:51:48That's such a cool experience. Did you get paid for that? Yeah, that's a complicated question.
1:51:56James Cowling:I got paid because I had to get paid because it was like a Hollywood union thing. I didn't want to get paid because it made my visa more complicated. So I've got a green card now. I'm all good. But yeah, but there was something where they had to pay me$400. Anyway, I made a grand sum of$400 off that show. looking back on your career is there any regret that comes to mind that maybe other people can learn from uh yeah i mean oh i think i under invested in my personal life to be honest people are probably gonna say that much you know because i think you can be all about like growth and you're sure i could have grown more i could have dropped out of grad school say three years in or four years in i probably would learn just as much i could have taken a job at dropbox back two, three years earlier and made a lot more money.
1:52:48James Cowling:Everyone who's been in the industry long enough has been offered to co-found several billion dollar companies. Everyone has a story about the times they could have been a billionaire several times over. I don't really regret those. I think, yes, it's been a lot of sacrifice to be blunt. I've been on call my whole career. I've carried a laptop almost every day. There's many dinners and parties and events I've had to skip. And there's people in my personal life who have suffered as a result. And I really appreciate those people. And they love and care about me. And they know that I have a passion for this.
1:53:27James Cowling:And so they accept me for who I am. I think this is old person talk. But yeah, I think everyone has to decide how much they want to really drive their career. because there's a trade-off like absolutely like i made a tremendous amount of sacrifices in my career and i've really prioritized building as probably the number one i mean values first and then building but if yeah went back in time yeah i would have had i had a good life but i probably would have done more vacations and you know i just had a bit more of a balanced life i i really do think i mean i see that with all the 996 stuff and and this kind of performative like you know photos of being in a bar and a laptop and stuff and i'm like that's not that's not real like that's not i mean sure i was working more than 996 back then but like i probably still do work more than 996 but that's i but i don't do it like um as a checkbox you know i do it because i just really want to be doing stuff and and so i think just um i would i would caution people against that hustle culture.
1:54:37James Cowling:Firstly, you're only young once. You should have fun. I've got quite a few grays in here. But also, that's acting. I mean, focus on solving problems, work hard, be passionate. Yeah. When I studied your career, I mean, there's mentions of you working 16 hours a week in various places but um you know you you like on call you like fires um you like ownership like all those things are recipe for working obscene hours yeah and you know i go home and i'm tired and i wind down by building stuff now like i i'm lucky enough to have a little workshop at home and so i go home and i you know make things with my hands uh that's just as you know that's just you become an infra person um you become an you become an engineer like in all aspects of your life um but yeah i would just i would just say like the cool thing is doing cool stuff that humans use the cool thing's not working long hours the cool thing's not like showing off about about that you were running a coding agent all night who cares the cool stuff's the cool things enjoying doing important things.
1:55:54Do you have a best technical book recommendation for people?
1:56:00James Cowling:I mean, I have to be honest, I haven't read almost any technical book. I mean, obviously, I was in academia for a long time. So I read a lot of papers, read a lot of papers. Learning is awesome. Reading is great. But balance it, right? read something and then go put into practice and most of my career I have I mean sure I did a PhD so I guess that that is like the academic side but after that most of my learning has been by doing because there's nothing like being faced with a real problem in your face to like to really to really develop as an engineer yeah I think if I answered the question I'd probably say the same I think learning by doing is that's that's really where where it matters most um and then last question for you is if you could go back to the beginning of your career and give yourself some advice or what'd you say i'd say you know what it'll it'll be okay like don't sweat the small stuff as much it's hard because like my My whole engineering brand is about caring about details.
1:57:03James Cowling:And I love Dieter Rams and design. And so being obsessive is a little bit part of my DNA. But I think I would go back and say careers are long. I mean, there's just any story anyone reads about 22-year-old billionaire, blah, blah, blah. Just ignore that story. That's not real. That's not repeatable. That's not normal. And it's not that healthy. And it's not that good for the people either, right? You probably won't, ideally won't max out your growth for 20 plus years as an engineer. I'm still learning all the time. I've been an engineer for several decades. So my advice would be, don't sweat it.
1:57:52James Cowling:There's time to grow. And I do think, again, it's a little bit of a modern phenomenon, but there is a feeling right now, oh my God, AGI has come and better max out my growth in the next three months. Well, guess what? You ain't going to do it. It's not going to happen, right? You can't do it. You can't max your growth out in the next three months. It won't happen. You don't have to be running 17 agents at the same time. all you gotta do is orient your career around learning every day and getting better at what you're doing and trying to solve things in the most simple ways yeah i mean on twitter i i see this take all the time of this permanent underclass idea where you know if you if you don't make it in time for agi then you know you're going to be part of this permanent underclass yeah i mean look it's a bit challenging economic times for a lot of folks.
1:58:43James Cowling:And I don't want to be, I don't want to be unsympathetic to people who are, you know, having financial difficulties. At the same time, it's just not an instructive attitude. There's not much you can do with that information other than feel bad about yourself. And, and, and I, and I don't, and I, and someone will tell, someone will argue back, no, what you can do is get really good at using Claude. Well, guess what? It's not very hard to use Claude. I just want to like grab some, sometimes people tell me like, oh my God, I'm finding it hard to keep up with all the new models, the new model drops.
1:59:13James Cowling:I don't even know what the new models are. Like I use them, but I forget the latest model because it doesn't matter, right? Like there was this thing called Ralph, right? I guess it's still Ralph. It's like a loop thing. I don't really know what Ralph is. And I haven't heard anyone mention Ralph in the past few weeks, but it was the biggest thing in Twitter for like a month. and it just doesn't, this is noise. Somehow, sometimes I feel like it's like tech tabloids. It's like people think that they're learning by somehow knowing, like listening to what Jensen said today or like, oh my God, Boris said that Claude writes herself.
1:59:53James Cowling:I don't think people, that's like, it's just like reading about Beyonce, but you're a nerd. And so you're reading about Jensen, right? But it doesn't matter. You're not growing. You don't have to know. It's okay. The new coding agent could come out and you could miss it. And then next year, if it turns out it's the big one, you'll just use it. It's not hard. I haven't seen any skill so far. I guess there's some skill, but it's not a hard skill. If you're good at engineering, you can figure out how to use code code or open code. So just watch out for tech tabloidism. It doesn't matter. Just be building stuff.
2:00:34James Cowling:Just do real crap. I love that mindset. And yeah, well, thank you for your time. I'm on Twitter too. I'm part of the thing. But just ignore me too.
2:00:51Oh, God. All right. Well, thank you so much for your time, James. I really appreciate it. This was a lot of fun. Ryan, it was great. Thank you.
2:01:02with a comment or a like. Also, if you have any recommendations for people you want me to bring on, please drop a comment. Guests like Barbara Liskov, Mike Stonebreaker, Mark Brooker, these were all people that I brought on because someone left a comment. On another note, aside from the podcast, I'm working on building the ergonomic keyboard that I wish existed. Here's a glance at the prototype. It's a split keyboard, so there's two sides. This is in the case but yeah we launched on kickstarter and we hit our goal within eight hours of launching i really appreciate it if you were one of the people who grabbed one of the early units we're now working on the long journey of building the tooling now and so if you still want to pick one up i've left the late pledges open on kickstarter so you can grab one there i'll put a link in the description thank you again for watching the podcast and i'll see you in the next episode
From the publisher
James Cowling is the CTO at Convex and was previously the most senior engineer at Dropbox. We discussed technical details of his past projects, simplicity vs complexity, and career advice given where AI is today.
• My ergonomic keyboard project I mentioned, you can follow along here: https://read.compose.llc/
𝗣𝗼𝗱𝗰𝗮𝘀𝘁 𝗹𝗶𝗻𝗸𝘀:
• YouTube: https://youtu.be/3XkmNSuHFmY
• Apple: https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835
• Transcript: https://www.developing.dev/p/dropboxs-former-most-senior-eng-building
𝗧𝗵𝗮𝗻𝗸 𝘆𝗼𝘂 𝘁𝗼 𝘁𝗵𝗶𝘀 𝗲𝗽𝗶𝘀𝗼𝗱𝗲'𝘀 𝘀𝗽𝗼𝗻𝘀𝗼𝗿 𝗳𝗼𝗿 𝘀𝘂𝗽𝗽𝗼𝗿𝘁𝗶𝗻𝗴 𝗺𝘆 𝘄𝗼𝗿𝗸:
• WorkOS: makes your app Enterprise Ready with easy to use APIs to add SSO, SCIM, RBAC, and more in just a few lines of code, check them out at https://workos.com/
𝗧𝗶𝗺𝗲𝘀𝘁𝗮𝗺𝗽𝘀:
0:00 - Intro
0:53 - Systems work during his PhD
13:05 - Dropbox technical deep dive
21:57 - Why Dropbox migrated from AWS
36:40 - How to do massive migrations
44:31 - Simplicity vs complexity in promos
49:23 - What technical teams should be focused on
1:00:25 - Doing the right thing vs promo hypothetical
1:08:13 - Why he dipped into management sometimes
1:11:36 - Why you shouldn't lead by example
1:23:23 - How to mentor Senior Staff+ engineers
1:27:30 - Career advice for the AI era
1:37:21 - Why he started his own company
1:46:05 - The most technically challenging work of his career
1:48:10 - How he got involved in Silicon Valley
1:52:16 - Career regrets
1:55:54 - Top technical book recommendation
1:56:36 - Younger self & permanent underclass advice
𝗪𝗵𝗲𝗿𝗲 𝘁𝗼 𝗳𝗶𝗻𝗱 𝗝𝗮𝗺𝗲𝘀:
• LinkedIn: https://www.linkedin.com/in/jcowling/
• Twitter/X: https://x.com/jamesacowling
• His company: https://www.convex.dev/
𝗪𝗵𝗲𝗿𝗲 𝘁𝗼 𝗳𝗶𝗻𝗱 𝗥𝘆𝗮𝗻:
• Newsletter: https://www.developing.dev/
• X/Twitter: https://x.com/ryanlpeterman
• LinkedIn: https://www.linkedin.com/in/ryanlpeterman/
• Threads: https://www.threads.com/@ryanlpeterman
• Instagram: https://www.instagram.com/ryanlpeterman
• TikTok: https://www.tiktok.com/@ryanlpeterman
𝗥𝗲𝗳𝗲𝗿𝗲𝗻𝗰𝗲𝗱 𝗶𝗻 𝘁𝗵𝗶𝘀 𝗲𝗽𝗶𝘀𝗼𝗱𝗲:
• His PhD Thesis: https://www.usenix.org/system/files/conference/atc12/atc12-final118.pdf
• Masters paper: https://www.cs.princeton.edu/courses/archive/fall19/cos418/papers/vr-revisited.pdf
• Papercuts writing he mentioned: https://medium.com/@jamesacowling/embracing-papercuts-e6390055dfc4
• "Don't lead by example": https://medium.com/@jamesacowling/dont-lead-by-example-4f86b1174e64
• His writing about orienting teams around missions: https://medium.com/@jamesacowling/your-system-is-not-a-sports-team-e17f9eb16b94




