In short
Podcast Notes: The Changelog - Episode on Tiger Beetle
Episode Overview Title: The 1000x Faster Financial Database (Interview) Host: Jerod Santo Guest: Joran Dirk Greef Release Date: Not specified in transcript Key Themes: Database Design, Transaction Processing, Open Source, Performance Optimization
Episode Description In July 2020, Joran Dirk Greef discovered a critical limitation in general-purpose database design for transaction processing. This led to the development of TigerBeetle, a distributed database tailored for financial transactions, boasting performance three orders of magnitude faster than traditional databases. The episode explores the innovations behind TigerBeetle, its resilience and durability claims, and the intersection of open-source development and business.
Key Highlights
- Background and Genesis of TigerBeetle
- Initial Problem: Joran encountered limitations while consulting on a central bank switch that relied on MySQL for handling transactions.
- Performance Issues: Initial optimizations (e.g., using NVMe storage) failed to improve transaction speeds, which were capped at 76 transactions per second (TPS).
- Key Insight: Performance bottlenecks stemmed from row locks and the high number of SQL queries required for basic transactions.
- Design Principles of TigerBeetle
- Transaction Design: Focused on simplifying the transaction model to reduce the number of SQL queries needed.
- Performance Claims: Achieved significantly higher throughput by allowing for batch processing of multiple transactions in a single query.
- Unique Features:
- Concurrency Control: Utilizes prefetching and minimizes locking to enhance performance.
- Memory Management: Designed for direct memory access and static allocation, ensuring efficient use of resources.
- The Importance of Open Source
- Open Source Decision: Joran emphasizes that open-source licensing (Apache 2.0) is crucial for trust and community engagement.
- Trust and Brand: Joran argues that building a reputable brand and trust is essential for business success, especially in the fintech space.
- Balance with Business: While open-source projects can be perceived as "cheap," they can also be structured to provide valuable support and services, creating a sustainable business model.
- Technical Innovations
- Deterministic Simulation Testing: Implemented a testing system that simulates various fault conditions in the database to ensure reliability and robustness.
- Resilience to Failures: Discusses the development of a storage fault model to detect and handle issues like misdirected writes and disk errors.
- The Role of Zig Language
- Choice of Programming Language: TigerBeetle is built using Zig, chosen for its suitability for low-level memory management and performance-oriented designs.
- Community Impact: The use of Zig aligns with the trend of crafting high-performance software for specialized tasks.
- Anticipated Challenges
- Market Competition: Joran acknowledges the presence of industry giants like AWS but maintains confidence in TigerBeetle’s unique value proposition.
- Emphasis on User Experience: Plans to provide a push-button experience for users looking to implement TigerBeetle in their operations.
Conclusion The episode concludes with reflections on the future of database technology and the importance of leveraging open-source principles to foster innovation and trust in financial technology. Joran expresses enthusiasm for the potential of TigerBeetle in transforming transaction processing at scale.
Call to Action Listeners are encouraged to explore the capabilities of TigerBeetle, engage with their community, and consider the implications of open-source development in their own projects.
Related Links
- [TigerBeetle Documentation](https://tigerbeetle.com)
- [GitHub Repository](https://github.com/tigerbeetle/tigerbeetle)
- [Zig Programming Language](https://ziglang.org)
- [Sim Tiger Beetle (Simulator)](https://sim.tigerbeetle.com)
Sponsors
- Fly.io: Public cloud infrastructure for developers.
- AugmentCode: AI coding assistant tailored for professional software engineers.
- Depot: Build acceleration platform for Docker images and CI/CD workflows.
- Notion: Collaboration tool for project management and organization.
---
These notes provide a comprehensive overview of the podcast episode, encapsulating the key discussions, insights, and takeaways relevant to software development and database innovation.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:05Welcome everyone, I'm Jared, and you're listening to The Change Log Log. where each and every week we sit down with the hackers, the leaders, and the innovators of the software world to pick their brains, to learn from their failures, to be inspired by their accomplishments, and to have a lot of fun along the way. In July of 2020, Euron Dirk Grief stumbled into a fundamental limitation in the general purpose database designed for transaction processing. This sent him on a path that ended with Tiger Beetle, a redesigned distributed database for financial transactions that yielded three orders of magnitude faster OLTP performance over the usual general purpose suspects.
0:50On this episode, Euron joins me to explain how Tiger Beetle got so fast to defend its resilience and durability claims as a new market entrant and to stake his claim at the intersection of open source and business. Plus the age-old question, why Zig? But first, a quick mention of our partners at Fly.io, the public cloud built for developers who ship. You ship, don't you? Then you owe it to yourself to check out Fly.io. Okay, you're on from Tiger Beetle on the changelog. Let's do it.
1:29Well, friends, I'm here with Scott Dietzen, CEO of Augment Code. Augment is the first AI coding assistant that is built for professional software engineers and large code bases. That means context aware, not novice, but senior level engineering abilities. Scott Flex for me, who are you working with? Who's getting real value from using Augment code? So we've had the opportunity to go into hundreds of customers over the course of the past year and show them how much more AI could do for them. companies like Lemonade, companies like Codem, companies like Lineage and Webflow. All of these companies have complex code bases.
2:07If I take Codem, for example, they help their customers modernize their e-commerce infrastructure. They're showing up and having to digest code they've never seen before in order to go through and make these essential changes to it. We cut their migration time in half because they're able to much more rapidly ramp, find the areas of the code base, the customer code base that they need to perfect and update in order to take advantage of their new features. And that work gets done dramatically more quickly and predictably as a result. OK, that sounds like not novice, right? Sounds like senior level engineering abilities.
2:41Sounds like serious coding ability required from this type of AI to be that effective. 100%. You know, these large code bases, when you've got tens of millions of lines in a code base, you're not going to pass that along as context to a model. right? That would be so horrifically inefficient. Being able to mine the correct subsets of that code base in order to deliver AI insight to help tackle the problems at hand. How much better can we make software? How much wealth can we release and productivity can we improve if we can deliver on the promise of all these feature gaps and tech depth? AIs love to add code into existing software.
3:20You know, our dream is an AI that wants to delete code, make the software more reliable rather than bigger. I think we can improve software quality, liberate ourselves from tech debt and security gaps and software being hacked and software being fragile and brittle. But there's a huge opportunity to make software dramatically better. But it's going to take an AI that understands your software, not one that's a novice. Well, friends, Augment taps into your team's collective knowledge, your code base, your documentation, dependencies, the full context. You don't have to prompt it with context. It just knows.
3:54Ask it the unknown unknowns and be surprised. It is the most context-aware developer AI that you can even tap into today. So you won't just write code faster. You'll build smarter. It is truly an ask me anything for your code. It's your deep thinking buddy. It is your stay in flow antidote. And the first step is to go to augmentcode.com. That's A-U-G-M-E-N-T-C-O-D-E dot com. Create your account today. Start your free 30-day trial. No credit card required. Once again, augmentcode.com.
4:42i'm joined today by yaron creator of tiger beater welcome to the changelog hey jared so great to be here thanks for having me excited to talk to you i've been excited ever Ever since you were mentioned by Glauber Costa on our Terso episode, he spoke very highly of you, very highly of Tiger Beetle. And I thought, we've got to get this guy on the show. Yeah, so great to hear that. Glauber is sort of part of the mix in some of our design decisions for Tiger Beetle. So I've always held him in high regard. And yeah, so I'd love to dive into that later. But yeah, thanks again. Absolutely. Well, we'd love to hear what you've been up to.
5:20I read on your Tiger Beetle blog that you said in July of 2020, you stumbled into a fundamental limitation in the general purpose database design for transaction processing. Can you tell us that story? Sure. So it was the strangest performance optimization challenge that I've had to do. So I was doing performance engineering at the time, consulting on a central bank switch with the Gates Foundation. They had created an open source, non-profit central bank exchange. And I was brought in to see how can we make it go faster? So how can we do lots of transactions between banks? How can Alice pay Bob?
6:04And those are little payments. So this is the classical database problem, a transaction between one or more people. And if you look at the Postgres docs for Postgres transactions, the canonical example there is Alice Pays Bob. Databases need transactions because they need to record people transactions. And so the central bank exchange was like a good example of why you need a database, because it's literally money moving from one person to another. And so I came into this project and, you know, I had done a lot of performance optimization work and specifically in Node.js. I'd been involved in Node.js since Rindal announced it.
6:57I'd kind of already been writing JavaScript on the server before Node.js using Rhino on the JVM. And I was convinced that JavaScript was going to be server-side. So I was already writing it. And then Rindal came, you know, and then I was in Node.js. for 10 years and doing a lot of performance. Yeah. So this switch was written in, you know, in actually Node.js to make it really accessible, open source project. And so it was like, how do you optimize Node.js? And, you know, I had experience with this. But the surprise of this whole thing is that, you know, how do you make these transactions go faster through the switch?
7:37Inside, it's essentially just a general purpose database. 20, 30 years old. MySQL was the one. And it's very basic. They're just doing like, if you look at those Postgres docs on transactions, they're doing classic debit credit in a SQL transaction, nothing more. And that's it. And then I thought, but like, how do you make it go faster? So my first experiment was, well, let's give MySQL NVMe. And it didn't go faster. The system could only do 76 debit credit financial transactions a second and it didn't go faster. I thought, okay, that's odd because normally it makes the database go faster, Jared.
8:24Yeah, throw a little hardware at it. It'll go faster. So I thought, well, maybe I did it wrong. Maybe it's not NVMe. Maybe the database is memory bound. Let's give it more cash. Okay. So we gave it a lot of RAM and And nothing changed. It stayed 76. And then we thought, okay, there must be some CPU-intensive algorithm. So we profiled flame graphs, everything. And there was nothing. The CPU was like 1%. Even the database, everywhere you looked in the system, there was just no bottleneck. And yet the system could only do 76 TPS. Network also. So in what we call the four primary colors, network, storage, memory, compute, we gave them all hardware, and there was no change.
9:12And this was so puzzling for me because always with performance work, there's always some bottleneck somewhere, and you fix that, and then you fix the others. And here it was the strangest problem because you had to optimize a system where the hardware was doing nothing. It was idle. and what what i what what you know people who had worked on other payment systems they sort of took me aside and helped me along they said you know it's the rolox and what they meant was when you do alice pays bob through a sql transaction to represent a real world transaction you go to the database from the app and you say select me alice and bob's accounts i want to get their balances.
10:00And then in the app, you check, you know, does Alice have enough money to pay Bob? Yes. Then you do another query back and you say, okay, you know, increment the debits or the credit columns and also insert a record of the transfer. So you've only done two SQL queries within one SQL transaction. Right. Yeah. And this is what I saw. In reality, the switch was doing 20 SQL queries within a SQL transaction, but they were all necessary and essential. And I've, you know, since then chatted to lots of other fintech people. And usually it's like a, you know, there's about 10 SQL queries per debit credit.
10:43If you want to do it really well, you could do it just like two, as I described. And so actually, this was the problem. The SQL transaction was holding row locks on the account rows across the network latency. And even if that's, say, one millisecond round trip and you're doing two round trips, then you've got locks for a millisecond. And that means if your transactions are all sequential, your database tops out at a thousand tps a thousand debit credits a second and obviously you know the next question we're going to say is well let's go horizontal um let's just but it doesn't change anything because you have these row locks and the other problem that we saw is that there's like a power law of contention like the Pareto principle 80 % of real world transactions are through 20 % of like hot accounts.
11:43So for example, you know, on the stock exchange, anyone, you know, who's bought shares has probably bought Nvidia or sold Nvidia, you know? Sure. And so like 90 % of trades are Nvidia or the super stocks, Microsoft. Right. And some of the others, you know. And a bunch of other ones just sit there untraded. Exactly. Then there's the long tail. But if you're the brokerage behind the stock market, if you're powering like 30 % of NASDAQ's volume, then most of your SQL transactions are hitting like only 10 rows, your 10 super stocks. I mean, how many are there? Right. So they're permanently rowlapped.
12:28There's seven of them. Oh, there's seven. That's it. Yeah. And so that's like the brokerage. But the central bank was much the same. They only have like typically eight or so big participant banks around the table. And so all the money movements are literally going through eight SQL rows. But even for like, you know, if you're, say you're building like Earth's biggest bookstore, you've got the same problem. You know, you've got the bestseller list of books. And so all your Black Friday transactions are hitting that. But so then we started to see like this is many businesses, many different kinds of industries.
13:05But let's go back to the beginning, back to the question. And so it was Rolex. And we realized it's a fundamental limit. And it's got nothing to do with the SQL database. Whether you pick Postgres or MySQL or something horizontal or cloud, it actually doesn't matter. The problem is the interface of Rolex across the network. and that was when we realized like actually we can't go faster we need a transactions database not a general purpose so that that was the that was the impetus yeah okay so you're in 2020 you are trying to basically squeeze more transactions per second out of my sql4 kind of a specific type of transaction right like you are literally meaning financial transactions not sql transactions which could be a bunch of different things that kind of roll out into a single unit of work we talk about a sql transaction that could be a lots of different things and then it gets rolled back or it gets committed but you're actually saying like we're tracking and doing debits and credits like financial transactions and we can't do these particularly simple type of sql transaction any faster.
14:23We can scale horizontally, but we have the network row locks. I'm sure there's a way in MySQL to say, well, let's just fly closer to the sun and turn off row locks or something. But maybe that's, you're also a central bank. Yeah. You know, in Postgres, it's almost the default read committed isolation. You fly close to the sun just by default, you know? You actually have to, there's a lot of like tweaking you need to do just to get it safer. but yeah no i mean at the end of the day you decided this requires a different kind of database like we should have a transactional database is that how you think of it where it's like the core unit of work is a credit and a debit combined together is that right that's right yeah and and the re i i think you know i had a i had a little bit of luck in seeing this because i've i've been coding for 30 years since i was 11 so now i've given away my age uh date yourself Yeah.
15:22And I've been coding for 30 years and it's been my passion, you know, and self-taught. So I didn't actually study computer science because I was already like reading papers and I thought, well, I'm more like an autodidact. That's going along, you know, already. Actually, I've always loved business too. And people told me, they said, you know, a great subject to study if you want to understand the schema of business. across all sectors, you know, and anywhere in the world, it's double entry accounting. That's sort of, you know, Warren, there's a quote from Warren Buffett, accounting is the language of business.
16:01And so that's why, you know, like many, well, yeah, it's just, that's why I went and studied financial accounting. I majored in that at university because I had a love for business and wanted to understand the schema, you know, journal entries, debit credits, because with debit credit, you are doing transactions. You're doing multi-row transactions between multiple parties. A debit credit transaction could actually have many debits on one side and many credits. You're expressing something essential in life across – and it's not only businesses. You could use this not only for money but for inventory, stock counts, or just counting things, counting API usage or counting kilowatt hours.
16:45It's basically accounting. You're just counting. So any domain where you need to transact with counts and its quantities, it could be like valuable or not valuable. But it's just things are moving, moving around, and that's transactions. But, yeah, and that's sort of what Tiger Beetle was meant for. Sort of the canonical, you know, general purpose example would be debit credit. And so often you see this example given, but we thought, wouldn't it be great if you just got a database that did this out the box, you know, that let you in one query to the database execute 8 ,000 debit credit transactions instead of having to do 80 ,000 SQL queries to do the same.
17:35You know, so that's one query and you do 8 ,000 instead of 80 ,000 queries. And so that's just Tiger Beal. You get these debit credits. you know it's so we thought let's build it in you it's kind of like financial asset not not only row column consistency but multi-row column consistency so like transactional consistency double entry consistency yeah interesting now i have not been in financial technology as much as you have i've definitely dabbled and i've been a i was a contract software developer for many years. And so I worked on contract and I got into lots of different industries. And one thing I found is lots of different industries have their little niche databases that people on the inside know, people on the outside don't, or sometimes they're homegrown.
18:25I imagine inside these large financial institutions, there's probably many, maybe homegrown, maybe there's a proprietary thing that's like, this is the database we use because of the exact specific use case that you are speaking to now am i right about that are there tiger beetles living in large financial institutions and you decide we want one for the world or why did you decide to build this thing besides maybe go find one or yes i think i think we're like we're mostly right uh and and the reason is because yes everybody has to build this database but it's um they what what is what what you usually see and the, the, I also had no experience in, you know, in payments specifically as I worked on the central banks, which I came in from the technical performance angle.
19:19I had the accounting experience. I could see that, yes, they were doing debit credits and, you know, but it was the payments people that also took me aside and they said, you know, every single like FinTech or this is only FinTech, but it applies to any business that needs to record transactions, any money, you know. But they said, you know, they're all reinventing these debit credit databases. And typically what they do is they do it over a general purpose database. So I think that this is also interesting is that it's kind of, again, lucky on timing because the, for, you know, SQL is like 50 years old.
20:00and it's been able to power OLTP, transaction processing, business transactions for 50 years. And so this has always been a latent problem, but we were starting to hit it now on this system. And that's because things I think are becoming more, the volumes are increasing. But yeah, back to your question. So I think, yeah, all these fintechs were reinventing these databases, but they were building a database and didn't know it. So they were building a database in the application around. So that's why, you know, there was this central bank switch. And you can see the whole thing is a debit credit database.
20:44It's like tens of thousands of lines of code around MySQL. But if they just had an OLTP database with a debit credit interface, the switch would be so, so simple. And that was, you know, so that's how this work started. There was no desire to build a company, build a product, sell something. I was also passionate about the mission because the switch was open source, nonprofit. I was told, you know, and this was the mission, and it is, the cheaper, the more performance you make the system, you will reduce fees. the developing countries that run this will be able to you know lift a few billion well a few million people above the critical poverty line and it isn't literally actually you know i don't know on the order of about a billion people around the world that don't really have access to banking like we take for granted in the west because you know they just don't don't have like you know in those areas.
21:45They have to walk miles just to move money. But everybody has a cell phone. So these systems would be built that you could power people to move money instantly on cell phones. But the performance is so important because as you make things faster, you make them cheaper. And now for someone who's got, you know, they're sending like$10 back home. If you're taking, you know, 5 % fees or 1 % fees, that's life-changing because that 4 % means so much to them. They're going to maybe have an extra meal and it's going to compound. So this difference in fees, that is why they are doing this. So for me, it was really like, we do want to make it faster.
22:32I'm not just consulting and I'm not... There are like a thousand MySQL knobs I could have tuned. We did double performance. We could have made it a little bit more, but it was kind of like, we actually really want to make it really fast out of the box.
22:56Well, friends, I'm here with a brand new friend of mine, Kyle Galbraith, co-founder and CEO of depot.dev. Your builds don't have to be slow. You know that, right? Build faster, waste less time, accelerate Docker image builds, get up action builds, and so much more. So Kyle, we're in the hallway at our favorite tech conference and we're talking. How do you describe Depot to me? Depot is a build acceleration platform. The reason we went and built it is because we got so fed up and annoyed with slow builds for Docker image builds, GitHub action runners. And so we're relentlessly focused on accelerating builds.
23:30Today we can make a container image build up to 40 times faster. We can make a GitHub action runner up to 10 times faster. We just rolled out Depot cache. We essentially bring all of the cache architecture that backs both GitHub Actions and our container image build product. And we open it up to other build tools like Bazel and Turbo Repo, SC Cache, Radel, things like that. So now we're starting to accelerate more generic types of builds and make those three to five times faster as well. And so in simple terms, the way you can think about Depot is it's a build acceleration platform to save you hundreds of hours of build time per week.
24:07no matter whether that's build time that happens locally, that's build time that happens in a CI environment. We fundamentally believe that the future we want to build is a future where builds are effectively near instant, no matter what the build is. We want to get there by effectively rethinking what a build is and turn this paradigm on its head and say, hey, a build can actually be fast and consistently fast all the time if we build out the compute and the services around that build to actually make it fast and prioritize performance as a top level entity rather than an afterthought yes okay friends save your time get faster builds with depot docker builds faster get up action runners and distributed remote caching for basil go gradle turbo repo and more Depot is on a mission to give you back your dev time and help you get faster build times with a one-line code change.
25:03Learn more at depot.dev. Get started with a seven-day free trial. No credit card required. Again, depot.dev.
25:16Well, to just spoil the end of the story, it's not the end, but it's further down. You did end up designing something that's a thousand times faster. So I think you said three orders of magnitude. That's a big win. But let's talk about getting there. You decided like, okay, what we want is a database that's designed from the ground up for this kind of transaction. Then what do you do? You didn't want to make a business out of it. Here you have a business now. But did you decide I'm going to code up in my free time, an open source database for the world? Or I need to go raise money? Where'd you go from there?
25:54Yeah, so I didn't have really any of those thoughts. I think the first thought I had was, I figured we could fix the interface of the database. Instead of having row locks and multiple queries, you know, we pack a lot of debit credits, first class, one database query back again, and you're done. And you actually do get that 1 ,000x performance. And you have to only improve on 76 logical finished product TPS. I think people sometimes get confused. They think we're talking row inserts per second. We're actually talking logical transactions, debit credit with the contention row lock problem. If a database is only able to do 100 to 1 ,000 TPS, depending on the latency of the network and the contention, those are the two variables.
26:41You can plug them into Amdahl's Law. Typically, if there's a round-trip time, 10 milliseconds looking at 100 TPS, 50 % contention looking at 200 TPS. That's your max general purpose limit. And so if you want to go 1 ,000 times faster, but you are now packing 10 ,000 transactions in one database round-trip, it's only one meg of information, you're amortizing so many costs, it is not actually hard to do 100 ,000 a second. 200 we actually did a million a second with primary indexes only tiger beetle ships with 20 secondary indexes turned on so it does between 100 to 300 000 a second so it actually is like you know i hate benchmark wars that's why we never really do it we just try to explain to people first principles why it makes sense um but that that was sort of the first step like that was my first feeling like and we we only thought that the thousand x because that was what the we you keep saying we but you haven't described your partner okay yeah yeah so so actually it was it was me and then don chankfoot was working with me like guiding me around the system i i was doing the work and then don was helping me and he was like my you know my bridge and and and manager There were also other people at COIL, my managers there, and surrounding team, really bright people.
28:09My manager was the co-chair of the W3C web payments group. He was the co-chair, but there was Visa and MasterCard and all of them also on that. So it was people who really understood payments. So I say we because I always like to include the team. And, you know, so Tiger Walls today is there's a team. But I did just create it, you know, and it came out of this experience. So the first feeling was just we need to go a thousand times faster because that is going to drive down costs. And because the scale is actually a thousand X more countries that deploy these kinds of systems, you tend to get very small value transactions and massive volume.
28:59In India, for example, they've done four orders of magnitude scale in 10 years only. So in other words, 10 years ago, if you picked MySQL today, you need to – you're telling MySQL go 10 ,000 times faster, please. There's no cloud database designed – most cloud databases were designed 2012 or 2015. None of them were designed for like four-order magnitude increases. So there's no LTP database on earth that can power India or like Brazil's PICS. These systems are already past three orders. So we were like, okay, we need at least a thousand X because that it's not, this is something I think people also, it's helpful to understand.
29:45It's not as impressive as I'm impressed by what you're saying. Like there's people who've done 10 ,000, you're doing a thousand. I'm impressed, but India is not impressed, for instance. Yes. Well, you could actually take Tiger Beetle today, run it on your laptop, and your laptop could power India's transactions volume going through the cloud. You really could, and it would be pretty easy. So it resets the experience, which I think makes sense because we've got 30 years of hardware and software research advance, how you could build a database today. A lot has changed. So if you just work with the grain, you know, in those days it was spinning disk.
Read the full transcript
30:24That was the bottleneck. Today it's memory bandwidth. So all the parameters of the design have been inverted. But again, it's actually not that hard because there's just so much seismic differential waiting to be, you know, just, I was just excited that we could now do something. Yeah, because all the major ones, the general purpose ones, MySQL, Postgres, SQLite, etc., they're 30 years old, right? I mean, they're 90s, sometimes in certain cases, 80s. And so they were built with a different world of constraints. And they've shown amazing malleability and resilience and the fact that they're still good enough for lots of things today is an amazing feat of engineering by all those teams.
31:10but to start fresh and to start with a is it a limited domain i mean you're not also bolting on general purpose on top of tiger beetle you're saying use tiger beetle and for these transactions and then also use something else for the rest of your workloads and that's quite right yes so i think and i agree with you jared it's so nice that you said because i i was i was wanting to say it too it's a testament to yeah a testament to their designers that these you know postgres powers the world of general purpose transactions, you know, or general purpose, you know, if you need to build an app, whatever you, but I think the interesting thing here, like to your point about fixed domain, for many years, what we call OLGP, like general purpose database, that was used for OLAP up until 93, 96.
32:04and then OLAP actually became a term, you know, online analytics processing. And then because there was this whole paradigm shift that look, you know, Postgres is great. It's row major. That's how inside the elephant anatomy, that's how it stores the, you know, the information it's designed. It's sort of designed for a 50-50 read-write split. And so it has like a B tree that's maybe a little bit better for OLAP. But the OLAP people came along and said, no, no, we need to go column major. It's like a paradigm shift. So the anatomy of DuckDB is totally different to the elephant. And today we wouldn't say that Postgres is an OLAP database.
32:48It's a general purpose database. It's not Snowflake or BigQuery. And that's fine. And vice versa. OLAP is not always OLGP. But I think today, for the same reason, because of increasing scale and because you can specialize, there's like a paradigm shift. You know, OLTP is not just rows and columns, row major or column major. It's multi-row major. You know, you're doing things across rows with contention. So the concurrency control of the database inside looks totally different. The anatomy of a tiger beetle, I would say, if I may. Sure. You know, it's different. So it's a lot similar to Postgres, but it's also you can think of like the group commit that MySQL or Postgres has and just dial it up 10x, make a much bigger group commit, much bigger batches.
33:44It's also, again, like you said, fixed schema. So you don't need all the overhead of serialization. There's no inter-process communication. Just so many things. it was just incredibly fun to design Tiger Beetle because it was such a simple problem, just debit credit at scale. So in other words, I think this is it. Tiger Beetle is not an OLAP database. It's also not a general purpose database. It's just an OLTP database. And that's all. Right. And then Postgres is fantastic. So you never replace Postgres, just like if you're using OLAP, you don't. But what you would do is OLGP is like your control plane in your stack.
34:25So you put all your entity information. People call it master data or reference data. So it's the information, your users table, your usernames and addresses. If you're building Earth's biggest bookstore, usernames and addresses, names of your catalog, your book titles. Those are not really OLTP problems because that's just you update your top 10 every now and then, or you know what I mean. And OLTP is like when people move a book into the shopping cart because that's adjusting inventory that's held potentially for a shopping cart. After a certain amount of time, that debit credit times out and rolls back.
35:08Then if people do check out their shopping cart, those goods and services are moving through logistics, supply, you know, all of that to warehouses and delivery. That's all debit credit. You know, quantity is moving. and it's moving from one warehouse to another, to a driver, to the home, back again, okay, back again, all debit credit. And then you've also got the checkout transaction with the money. That's also debit credit. And so like that would be OLTP and that's sort of the Black Friday problem. The Black Friday problem isn't, you know, how do we, you know, store the book catalog or update that because it doesn't change often.
35:46So that's a great general purpose problem. Just like users, you know, they don't change the names often. So your database that's great for variable length information and a lot, you know, that information is actually very different to transactional information. Transactional information is very boring, essentially, just multi-row debit credit. Right. And so this multi-row major that you describe. I made it up. Yeah. But I would always say multi-row, but let's call it multi-row major. Yeah. Multi-row major versus row major or column major. Yeah. that makes sense to me because every single transaction has you know you're going to assume that over here there's an addition over there there's a subtraction there is this double entry thing where it's like there's going to be more than one row yes and pretty much anything that matters in tiger beetle right and so that's an interesting way to like think about it and fundamentally different way like you said versus thinking about it column or based on rows how does that fundamental primitive manifest itself in your decision making i assume there's storage concerns maybe like memory allocation maybe there's protocol i don't know where all that works its way out as you design this system how does that affect everything that tiger beetle is great i'm so excited to dive in please do let's do it yeah yeah so let's apply this like let's take the concurrency control, for example.
37:11So let's say we've got 8 ,000 debit credits. So one debit credit would be like, take something from Alice and give it to Bob. Then take something from Charlie and give it to Alice. Take something from Bob and give it to Charlie. And you've got 8 ,000 of these. Some of them might be contingent on the other. And you can actually express a few things around this. like, but let's just leave it like that. So you've got 8 ,000 debit credits in one query. So the first thing that comes off the wire into the database, the database is going to write that to the log, and it's going to then call F-Sync. What's great there is you've called F-Sync once for 8 ,000.
37:55So it's F-Sync once for one query, but that is amortized across 8 ,000 logical transactions. And F-Sync usually has a fixed cost. So it has a variable cost, but there's also always a fixed cost. And it's like half a millisecond or a millisecond. It depends all on the hardware. But there's always a fixed cost. But now you've amortized that massively over 8 ,000. Typically, the group commit for MySQL might be around 15 things. So it's much smaller. It will amortize F-Sync, but not by so many orders. That's the first thing. Now we've got the D in ACID, you know, atomicity, consistency, isolation, durability.
38:40Before the database processes it, it makes the request durable on disk calls fsync. The next thing it does is it'll take these 8 ,000 and apply it to the tables. You know, it'll actually update the state on disk. And if it was something like a general purpose database, what it'll do is it'll take the first debit credit, it'll read Alice's row and Bob's row for their accounts. Then it will update the rows and then it will write them back. Then it'll go on to the next one, read the two rows, update them, write them back, so on and so on. So you're looking at about 16 ,000 accounts that you read in and write out.
39:26And that typically takes what is called latches, little internal row locks also inside the database. So 16 ,000 little micro cuts, you know, and you're also contending there. And then, but you see, here's the catch is that the domain is usually, again, it's everybody's buying NVIDIA. So if like 80 % of your 8 ,000 are all NVIDIA, you're reading NVIDIA, writing NVIDIA, reading NVIDIA, latching NVIDIA, latching NVIDIA. And you're doing it 16 ,000 times. And so what Tiger Beetle does instead, and this is, again, with anatomy changes. Now, we first look through all 8 ,000 and we prefetch all the data dependencies in that request.
40:09So all the accounts, for example. And so we load NVIDIA once. Then we load the other six hot super stocks. And then there's a long tail that we load. But actually, there's not many accounts. and they're usually hot in L1 cache. So they don't even go to disk because we've got a specialized cache just for these accounts. Everything in Tiger Beetle is CPU cache line aligned. And we think of that these days as like 128 bytes. Cache lines are getting bigger, like with M1 started there. And everything is cache line aligned. We don't straddle cache lines for false sharing. Everything is zero copy deserialization, fixed size.
40:53very strict alignment, you know, powers of two, and we don't split machine words or just, it doesn't always make a difference on the hardware, but it can. So these are all the little edges, you know, the cache is optimized. But okay, let's go back. So now we load just NVIDIA, the super stocks and the long tail, but actually they're all in the L1, you know, or in the L2. And then we've got the data dependencies cached, then we push all 8 ,000 through and then we write them back. And so it's just, I think you've got it now. It's just drastically simpler. It's kind of like you just do less and that's how you go faster because you're doing so much less.
41:32You don't have SQL string passing. Right. It almost feels like you're cheating, but you're not. You're just doing exactly what needs to be done and nothing more because you're not general purpose, right? Yeah, exactly. I often like to say, you know, we didn't do anything special. It's kind of embarrassing. It's so simple. The 1000x trick, it's, yes, we use IO Uring, we do direct IO, you know, DMA to disk, a zero copy, all the stuff. Actually, it doesn't make a performance difference. It makes, it does. That's why we do it. It makes a 1%, 5%. It all adds up. But that wouldn't get you 1000x. Just like stored procedures in a general purpose also wouldn't get you 1000x stored procedures will get you you know you get you get those 10 sql queries down to one but now you're still doing one for one and you still got so you only went 10x faster or 10x cheaper if you really want to go 1000x you actually have to just have first class debit credit in the network interface change the concurrency control um we even you know Now, Tiger Beetle has its own LSM storage engine, LSM tree.
42:45We designed it from first principles, again, just for OLTP. So it's actually an LSM forest. We have an LSM tree for every object that the database stores, so transfers and accounts, accounts and transfers between them, debit credit transfers between accounts. Those are two trees. And then all the secondary indexes around each object, there's like 10 trees and 10 trees. And so for every size key and every size value, it's in a separate tree. And again, there's just things you can do now, like RocksDB or LevelDB, which is what you find in a lot of general purpose databases, they use length prefixes for the key.
43:29So it's a four-byte length prefix or it's variable length, but now you've got the CPU branching and costs, et cetera. And then there's again a length prefix for the value. But if you know that your secondary indexes are only 8 to 16 bytes and you then add the cost of length prefixes, 4 plus 4 is 8, you know, 8 bytes of overhead just to store 16 bytes of actual user data, you're burning a lot of memory bandwidth, a lot of disk bandwidth, and you're going slower than a database that doesn't have length prefixes at all. And so Tiger Beetle, each LSM tree stores pure user data. There is literally we put the key and the value.
44:15There is no length prefix because each tree knows at comp time what the size of the key is. So, yeah, we can go on and on. And then the consensus protocol we did also similar optimizations. Yeah. So hyper tuned for this specific type of workload, which also happens to be one of the most important workloads as well. So let's imagine that I'm an e-commerce shop and I'm not going to roll out Shopify. I'm going to do it myself because it makes sense. And so far, I'm a Postgres guy. And personally, I respect my SQL, but I just use Postgres as my example because that's what I've been using for 15 years.
45:04So let's say that I've got my Postgres database. It's been running everything just fine. It's got my users in there. It also has all my transactions. And I'm hitting up against scale issues. This is a good problem for me because I'm doing more sales, right? And so I'm selling a lot of books. And I'm hitting scale issues. And someone says, you should really look into Tiger Beetle for your transactions specifically. And I think to myself, Postgres hasn't failed me for 30 years. Tiger Beetle hasn't existed until 2020 at the earliest. You can tell me when at 1.0 or when you guys actually shipped a product because conception to now we're five years in.
45:40I'm going to trust my most precious, my sales. You know, I want to trust that to something that's this new. I'm sure you face this a lot as you go out and try to get people to try out Tiger Beater. What's your response to that concern? Because it's a valid concern that, you know, not super, not super, what's it called? Battle hardened yet. Or maybe it is. Tell me about that. Oh, yeah. Great. I love the question. That was my second thought as Tiger Beaver was created. First thought is we're going to need… But no one's going to trust us. Why are they going to trust us? Yeah. And the second thought was this question, how can you possibly be as safe as 30-year-old software that was created, Windows 95?
46:22How could you possibly be as safe? And I think the answer to that is actually a question. in the same way that how could you be a thousand times faster than something that is 30 years old. The question is, what has changed in the world in the last 30 years? So much has changed from a performance perspective. And then when we look to safety and especially mission-critical safety, so much has changed. So 30 years ago, consensus didn't really exist in industry. Brian Oakey had pioneered ConsenSys a year before Paxos view stamp replication in 88. That was his thesis at MIT with Barbara Liskoff. So ConsenSys did exist.
47:11How can you replicate data across data centers for durability to actually survive loss of a machine or disk? But that wasn't really an industry 30 years ago. And so you don't really have first-class replication. Yes, you do have it these days, but it wasn't designed in from the start. And I think there's more examples around testing. So deterministic simulation testing, what FoundationDB did. And we actually got it from the people at Dropbox, not Foundation. James Cowling and them at Dropbox were also doing DST. The idea is like you design your whole software system that it executes deterministically.
47:57So you don't use, if you use a hash table, the hash table given the same inputs always gets the same layout. You don't use randomness in the logical parts that users would see. So you can think of it basically like fuzzing your software or property-based testing. Given the same inputs, you get the same outputs. And you design all your software like this. then you can test it. But if the test fails, when you replay it, you'll get the same result. And so distributed systems are so hard to build 30 years ago. So not many people did. And then when they started really building them, they actually started with eventual consistency because people were still figuring out how to do strict serializability.
48:42So there was a lot of fashion around eventual consistency, which has gone away, I think. but back then to build a distributed system you just kind of assumed like it'll just be eventual um yeah but so so how to build distributed systems was quite hard and because it was hard because it was hard to test them because you know the failure of one system over there causes the failure you know of another system here and you can never when you find a rare bug you can never replay it and you need so many machines and these bugs you know before tiger beetle i worked on systems that were distributed like full duplex file sync hence my interest in dropbox but those systems were incredibly hard to test there would be bugs you know they're there they take a year to find and fix and and then you realize it's a buffer overflow in the buv and that was like some of my first c code that i ever wrote you know years before tiger beetle it was fixing a buffer overflow and wv um it was the windows you know windows event notification some interaction with multi-byte utf-8 and different normalizations of nfc nfd that was this distributor systems bug and it also needed long file paths so it took a year like we knew it was there it took it was that your nojs days or that was prior to nojs days oh that was nojs days yeah so around 2016 and that we were using like Jepson style techniques where you've got fault injection, like chaos engineering.
50:12That's what we were doing. So you could get the bug, but to even, you could get this amazing bug, but you knew it was there. It was like the tease and you could never find it, reproduce it or fix it. And it literally took a year. So coming to Tiger Beetle, I knew from Dropbox, there were newer ways to build distributed systems. Just like you don't need to use eventual consistency anymore, there's proper consensus. No one should give that up lightly. You don't need to because you can get great performance first, you know, and so much performance. Fundamentally, you fix the design, then you can add consensus.
50:49It's not expensive when the design is right. I mean, consensus is literally just, you know, you append to the log in F-Sync and in parallel, you send it across the network to a backup machine and f-sync. That's replication, you know, 99 % of the time. And then consensus does the automated failover of the backups, you know, of the primary. So consensus doesn't really have a cost. That's a common myth, you know. But all these things have changed. But coming back to it, you know, it's testing has changed, how you test distributed systems. Because you can actually model a whole cluster of machines in a logical simulation that's deterministic.
51:31Now you test it just like Jepsen, but the whole universe is under deterministic control. You can even speed up time. So you can like fast forward the movie to see if there's a happy ending. If it does, great. Watch another movie. Each movie is like a seed, you know how you can generate a worms level. You know, the worms game or scorched earth games would be like randomly generated all from a seed. And so this is just classic fuzzing, property-based testing with seeds. But you're taking these distributed systems and they're all, the database was born to run in a flight simulator. This is deterministic.
52:08And if you can do that, you can build systems that, you know, you're kind of doing Jepson, but you're also speeding up time. You can reproduce. And then you have, once you've done, had that experience, you have to ask now, you know, do the 30-year-old systems have this level of testing. Yes, they have 30 years of testing, but with DST, you can speed up time. So we've actually worked it out in Tiger Beetle. One tick of our event loop is typically 10 milliseconds. We collapse that into a while true loop. We get a factor of 700x speed up when you take into account the CPU overheads of just executing our protocols.
52:49So we flatten everything and we execute it and you get a 700x time acceleration. So we have a fleet of 1 ,000 CPU cores dedicated at Tiger Beetle. We got them from Hetzner. Thank you very much, Hetzner. Nice. They're in Finland, so nice and cool. They're burning clean energy. 1 ,000 CPU core fleet. They run DST 24-7, and that adds up. I mean, it is roughly 700x. it's a thousand cores because we pay for it you know dedicated and they run we do a lot of work to like optimize how much we're using those cores but it does add up to on the order of like 2 000 years a day of test time and but but you see now again we're simulating like disk corruption and we're simulating things like we write to the disk and the disk says yes i f-synced but the firmware it didn't.
53:47Or we write to the disk and the disk writes it to the wrong place. And so Tiger Beetle has like an explicit storage fault model. So we do assume that disks will break, but they don't only fail stop. They actually are like kind of, we call near Byzantine. So disks do very rarely, but 1 % in a year or two, a disk will have some kind of corruption or latent sector error. Then you get a latent sector error. A little bit less common is like silent bit rot. A little bit less is misdirected I.O. where it actually just writes to the wrong place. And this can be the hardware or the disk firmware, even the file system.
54:29So two years ago, XFS actually had a misdirected write bug. And if your database was running on that particular version and you triggered that, your database would write to the wrong location on disk. And now the question is, well, who tests for this? and you almost can't unless you're using these new techniques. So I don't know. Yeah. That's cool. I guess the question was, you know, it's not enough actually to be as safe as what was safe 30 years ago. We've got new techniques, and there's a few more of them in Tiger Beal. Yeah. No, I think that's super cool. It reminds me of light bulbs for some reason.
55:07You know, LED light bulbs, they say they're supposed to last 25 years or something, and I'm always like, you don't know that because, you know, you haven't been using them for 25 years in fact the house that i currently live in we're going on 10 years now and they sold us on all led light bulbs and i remember as the installer was putting them in he's like you're never going to have to replace any of these and you know what i've replaced a whole bunch of them so whoever did their their testing can't do what you guys can do which is they can't just fast forward time and prove that this seems going to last for 25 years because they haven't been building these for 25 years but what's super cool with this what's it called deterministic simulation testing yes yeah what's cool about it is you guys can actually just through cpu power and and uh design you can actually simulate all this time and so you can claim even though you've only been around for five years i'm giving you all five even though i'm sure it's technically less than that that you've actually tested for hundreds of years right like you could just say that because you've done that work through this you know three-dimensional simulation that i'm imagining that you put the system through at all times i think that's pretty cool yes at least we're trying to get there you know so we we would also say it's only as good as our coverage you know and our combinations of coverage and our state space exploration but we invest so much into that as well you know um and so it does because because we also know how many bugs we find and how rare they are like yeah require you know if you think of it like a hacker they require like eight exploits to be chained and each exploit is like tough and then you just know that like traditional software there must be so many like millions of bugs but but yet we don't find them in the real world because they're so so rare but with with the dst you do actually and we find them pretty quickly um so yeah so we wouldn't claim but it but it is it's very very strong it gives you some confidence that you wouldn't have otherwise and i think that's very valuable um but like you said or maybe you didn't say it but i was thinking it while you said it there's no accounting for the real world so like uh mike tyson's famous you know statement everybody has a plan until they get punched in the mouth or something like that it's like which is how it works you can simulate all you want and it's amazing what y 'all are doing but then there's the actual reality and there's always going to be something and so yeah is tiger beetle out there is it in production in reality yet are people using it what's the state of that side of it very much so so we had a customer reach out to us they needed to i mean we've got a few customers but just an example of one end of last year they reached out they said look you know some regulations are changing they need massive throughput because they can it can put their business ahead we had to get them for you know to for like like sheer business strategy they needed to migrate within 45 days a workload of like 100 million transactions a month logical 100 million logical transactions a month they needed to migrate within 45 days from the old system we we migrated them and they saturated their central bank limit and they were happy you know And we pulled it off and the system just works.
58:31That's got to feel good. It did, yeah. That was great. And there's like national projects, three different countries, whether it's for the whole national transportation system or the central bank digital currency or another central bank exchange. Tiger Beetle is going into the current production version of that gate switch now as we speak. But I think just to go back, Jared, to like the DST, it's a lot like formal methods. The difference is that formal methods check the design of a protocol so that you know that the protocol could possibly work. It's formally proven that it could work. But for me, always the challenge was, well, how do I know that I've just...
59:17Because I always feel like I'm getting slower every day. How do I know that I coded this correctly? I know the protocol is correct. I know, you know, these times replication, Paxos, Raft, they're formally proven, but the implementation is like thousands of lines of code. And so how do you check that? And so DST is actually checking the actual code. And you, you know, the simulator is checking for split brain. It's checking linearizability, strictizability of which linearizability is a part. and it's even checking things like FsyncGate. So cache coherency, Tiger Beetle's user space page cache is coherent with the disk at all times, even if there's FsyncIO faults.
1:00:00So Postgres, they are fixing this. They're adding direct IO. If people want to find out, it's called FsyncGate 2018. But most databases still can't survive FsyncGate. They were actually patches where databases patched like MySQL, et cetera, to panic. The problem is when they start up, they still read the log from the kernel page cache, which is no longer coherent. So actually they have to use direct IO. So Postgres has been on a long journey, notably to add async IO, direct IO. It's in, I'm not sure yet. It might already be in as the default, but those are kinds of the things you need to survive.
1:00:41And that isn't even an explicit storage fault model, But Tiger Beetle simulator is actually checking. Your simulator can reach in and check so many invariants. But then also, I think back to what we were saying about claims and coverage is you want to have a very buggy view of the world. So you take your four primary colors, network storage, memory, compute. You assign like explicit fault models. So compute, we say, look, that would be in the domain of Byzantine fault tolerance. So we don't solve that. memory and that's explicit memory would be ecc that's our fault model and then what tiger beetle focuses on is the network fault model so packets are lost reordered replayed misdirected classic jepson partitions all of that most you know that's what makes consensus so hard is just solving network fault model is is almost impossible then tiger beetle adds the storage fault model explicit it.
1:01:40So, you know, you write to the disk, wrong place, forgotten. You read from the disk, you're actually reading from the wrong place. So you'll get a sector as a valid checksum for the database, but it's the wrong. Now you need to daisy chain your checksums. And then what we do is with these, this is sort of the buggiest view of these four. Okay, the first two, we've been explicit that we don't solve those because they have different levels of probability. You know, the probability of a CPU fault is astronomically more rare than a memory fault, which in turn is astronomically more rare than network fault.
1:02:15Maybe it's the other way around. Sorry. The rarest thing is the CPU, then it's the memory, then it's the storage, and then it's the network. So most people are just solving network. Tiger Beetle solves storage. But storage is, when you're at scale, it happens more and more so you know one percent of discs you know you know around the two year window at enough scale you're going to hit that enough times that it seriously matters versus at small scale where you're like ah we don't ever have to worry about that yeah that probability thing is one of the reasons why i think logically as humans we fail often to reasonably consider scenarios because we think about worst case scenarios but we don't pair that with probability of scenarios.
1:03:03And so it's like, what's the worst that could happen? And for some reason in our, this is like off topic, but in our hearts, we give that like a hundred percent chance, right? Like that, if that happens, well, let's assume it does happen. So a hundred percent chance on that, then it's terrible. But we don't also think like pair that with probability. So that's kind of what you're saying with the CPU stuff is like the odds of a CPU failure. Okay. Catastrophic, of course, but like exceedingly rare versus the likelihood of a drive failure, for instance. And so exactly drive or network yeah right network is probably the most unreliable at this point i would assume since our drives don't spin anymore that's it yeah the network is incredibly unreliable unfortunately yeah
1:03:55Well, friends, I love Notion because Notion lets me do everything I want in a single application that lets me invite others to get involved in those organizational work. workflows, processes, collaboration, whatever you want to call it, right? The cool thing that I love most about Notion is that you can make it your own, meaning you can make operating systems, workflows, you know, processes, standard operating procedures, the way you do things, the way you work, the way statuses work for you. And I don't mean that you got to build this thing yourself from the ground up. No, everything is for the most part, push button, templatable, that you can start from somewhere and end up somewhere else that fits your model.
1:04:40For me, I could be in the middle of doing something, thinking how I could add one more property or one more status to the flow I'm doing things and make the change in real time while doing the work to enable the future work I'm doing to be better, to be easier, to fit me. That's why Notion is cool. And if you're not using Notion, well, now is the time to do it because there is no shortage of the way AI has helped many, many people. I love Notion AI. Notion for me is big. A lot of stuff in my workspace. I've got a personal one. I've got one for the changelog. And all these things fit into different places in my personal life, the way I personally use Notion.
1:05:23But Notion AI lets me search across everything in one single AI interface. And it's the coolest. But being able to combine your notes, docs, projects, all the things you want to do into a single space that's simply beautiful, easy to use, well designed on all the platforms, whether it's web, desktop application, iOS application, Android application, you name it. And Notion is there for you. And Notion is used by over half of Fortune 500 companies. Now, I don't know about you, but I'm not anywhere near a fortune 500 company, but they're also used by many, many teams. And we're one of them. We're one of those teams.
1:06:07These teams that use Notion send less email. They cancel more meetings because, hey, no meeting needed. They save time searching for their work using Notion AI, and they reduce spending on various tools because they consolidate it. And this helps everyone to stay on the same page and have the business stay more focused. So try Notion today for free. When you go to Notion.com slash changelog, that's all lowercase letters, Notion.com slash changelog to try the powerful, easy to use Notion AI today. And when you use our link, of course, you're helping us. So do that. Use our link that lets Notion know, hey, changelog is impressing people.
1:06:47They're sharing what we do well. And you know what? I wouldn't tell you if I didn't. I love Notion. It's awesome. You should try it out. Again, Notion.com slash changelog.
1:07:02While we're talking about the simulation stuff, I do want to take a quick aside and talk about something I found that was really cool on your website. Sim Tiger Beetle. Sim.tigerbeetle.com. This is a simulator for a distributed database scenarios. It's like a game. Well, for the video folks, we'll put it up for you guys to look at. This is really cool. you can play these different games like the mexican standoff the prime time and the the radioactive hard drive what's going on with this this is like out of left field when i saw it this is not i wouldn't expect this from you you're on what what is this thing i i thought you didn't expect it jared uh yeah i didn't expect it i saw it i'm like okay i started playing it and it's like it's polished the graphics are awesome there's sound there's music and it's like a a game all about distributed system failures.
1:07:58Tell me more. Maybe we, you know, our team, we just enjoy Tiger Beetle too much. You know, it like kept me company in lockdown and COVID that, that was July, 2020, March, 2020, I started on the switch a week later. It was my birthday, you know, and a week later, the world went into lockdown and I was locked down solving this problem July 2020 I sketched Tiger Beetle created the prototype mini versions and never stopped I don't know maybe it's because it's out of that experience that we as a team we have so much fun or we really enjoy it but some Tiger Beetle was kind of an expression of that because the feeling of our DST simulator The name of it is the Vopper, Viewstamped Operation Replicator.
1:08:48It's a homage or ode to War Games, you know, that's got the Vopper. Oh, yeah. That classic simulator in War Games. And so Tiger Beetle's simulator is called the Vopper. Nice. But the Vopper for us was such an experience, you know, switching it on for the first time. I was listening to Kings of Leon Crawl as I switched it on. And we had been building Tiger Beetle for a year. With the whole design, all careful, the fault models, we had the interfaces designed. I knew we would do DST, and I planned it like that. A year in, it took about a week to build the first version of the simulator, and then we switched it on.
1:09:31And just the bugs came falling out of the trees, and you'd fix like five a day. And each of these was like one of those one-year bugs from before. Now you're fixing like five a day. And that experience, it was just like special as a programmer. I wish it for everyone. And I think people are getting into this style. But this runs in our terminal. So I did a demo to a friend for the very first time, Dominic Torno, fantastic of Resonate. I really look up to him in distributed systems. He's become a mentor and friend. And if I have a hard problem in distributed systems, I ask Dominic. But I had just met him, and I showed him our simulator running in the terminal, and I'd never shown anyone before.
1:10:23And he was like, wow. He'd done formal methods. He was like, this is formal methods on the code. And I showed him the probabilities of the fault models for each simulation run. And he was blown away. Like he said, no, you've got to tell people about this. And then I thought, well, how do you show this to people? And, you know, I'm a dad and I thought, how do I encourage my daughter, you know, just not encourage her, but just, you know, how do you encourage the next generation to get excited about computer science? Yeah. And to me, this was the most magical part of programming in my own journey.
1:11:01So I thought, well, let's make a game. Like let's take our simulator and let's put a graphical front end on top to hook into all these events. And then we can create different levels for people showing them how consensus works if there are no network failures. So everything's perfect. You'll simulate the network latencies and disks. Everything's simulated, the clocks, all of that. But there are no faults. And so it's perfect. And now you can actually just teach this is just normal replication through the consensus protocol. And then the next level is like, okay, now we're going to introduce probabilities of partitions and network faults, but the disk is perfect.
1:11:45And then the next level is, okay, disk is radioactive. Yeah. Well, I played it for five or 10 minutes and I had a blast. It's like, here's a hammer. You can start, you know, start hammering stuff and see how it reroutes and really, really cool. Yeah, and you actually, you know, each of those beetles that you hit, they're running, each of them is running real tiger beetle code compiled to ASM. Yes, it's the real code in your browser for the cluster. When you take a hammer and hit the beetle, you're actually physically crossing the machine and it's restarting. And you're actually getting to touch a simulated world, but against the real code as a human.
1:12:27um yeah and that's amazing yeah i feel like uh as a tiger beetle engineer you could just be playing that game and and you know you're on asking what you're up to and you're just like i'm working you know i'm uh i'm simulating some crashes here come on well i have done that i i do do that jared it's like sometimes you know you instead of doom scrolling you just sim scroll yeah my daughter says to me you know papa can we play the tiger beetle game and then but i'm glad you call it a game because we we meant it like in the walking sim genre so it's it's a game that you can't win because no matter what you do things recover right and you can't knock the whole system out it's going to go back to good it's going to find its way back yes you actually can if you're really lucky when i do it in live presentations it's it's never failed me but theoretically you could crash the cluster um if you because you see the those human tools allow you to inject more faults than than the f tolerance and then then the system is designed to shut down safely um so you you might run into that but you would you would have to be very quick yeah well i was too slow in my five minutes of playing around I'm also always I didn't know it was possible so now that I know it's possible maybe I'll commit myself yeah to shutting that system down pretty cool oh but there is a game Jared at the end I don't know if you've played the credits in Radioactive I have not no is this an easter egg yeah after Radioactive at the end the credits like go and that becomes like a little platform jumper game and you can jump and spin a beetle like and you do get a score and that one is quite hard so that sounds like that sounds amazing I'm going to go do that after we hang up here.
1:14:19Yeah. You know, maybe I should just, if I can just quickly add, you know, that game, it was created part-time. I met Fabio Arnold in Milan, the very first ZIG meetup in Europe. He was there, and he created it part-time, just a few hours every week with Joey Mach's illustrator from Rio de Janeiro, Brazil, just the two of them. And they created it, like, very, very low budget. but they had such passion and we then carried on and we put more into it so it did become more polished but yeah um that's just a the skill you know tribute to them cool yeah i think for those who go out and give it a play you will notice immediately and i did just like how much love is put into something like that like you you mean it when you say we just love tiger beetle and this whole system because you know there's custom sounds and there's music this is like this is a labor of love for sure and a really cool way to show off uh what y 'all have built in a way that is difficult just with our brains you know as you talk about it for me to map it onto my brain and make sense of it but when you put it out there in that visual way um it's very compelling so shout out to them uh for their for their labor and for you to keep polishing that thing up i did want to touch on the open source side and kind of the business end.
1:15:42You mentioned some customers, you know, you raised some money now. So this is like a serious business, but it also is an open source database. Can you talk about the give and take there, the push and pull, the decision-making process? Because we talked to a lot of people in your circumstance and some of them have made other decisions. Some of them have made the exact same decision you've made and we're all trying to figure it out. How can we make this thing work? So tell us your side of the open source slash VC story. Yeah. Thanks so much for asking that, Jared. That was my third feeling. So my first feeling was, I thought, wow.
1:16:16I'm hitting all your feelings here. Yeah, yeah, yeah. I mean, but really like July, 2020, as the project was started, I remember clearly there were three moments. The first moment was like, wow, this prototype is fast. It's like the design works. I'm like, wow, like this maybe could change things for the open source switch. You know, that was it. And maybe it could change things beyond, but I didn't know. The second moment I remember was the DST, switching it on Kings of Leon crawl. The third moment was this wondering of like, what actually happened was we designed it to be so much safer that, yes, it gained adoption within the Skates Foundation project.
1:17:01We won trust and they are integrating it today. You know, it'll power countries at the national level, you know, who knows in how many years, but as it gets deployed. So we did solve the trust problem of being so much safer because we designed it like that as well. But then the third moment was people then within that project saying, this is all well. but you know open source is um is too cheap um where's the where's the company to support it where are the the you know the people that are going to be available to service and really work and you can't you need open source so this this system is apache to open source and all the software it uses it it cannot use software that isn't open source because otherwise you know it It just would be a non-starter.
1:17:58So Tiger Beetle was also created like contracting then, like it had to be Apache 2.0. Like that was obvious to me, open source, because otherwise, you know, you don't fulfill the mission, which is what inspired the performance and the safety, is actually to make this really safe because it is people's money. So then the question was, you know, people were saying, where's the company to support this open source is too cheap? but I still, I didn't have a clear vision of like business model at that stage. And then it became clear as startups said to me, open source is too expensive. So on the one hand, on the one hand, you have countries saying it's too cheap.
1:18:40And then you have startups saying it's too expensive. And I'm like, this is Goldilocks. We're just trying to make some great open source porridge. And it's either, you know, too cheap or too expensive. But nobody's saying it's free as in puppies. You know, it's too cheap. And then I realized, okay, that's it. You know, business model is orthogonal to open source. Business is about trust. People trust you, you know, at the national enterprise. And it's always about trust you. That is what you sell as a business is trust your brand, your reputation. It's actually, I use the word, it's brand. And I think people, startups talk a lot about go-to-market.
1:19:20I think it's more interesting to talk about brand. Do you understand? Do we all appreciate the value of brand? Because brand is trust. And I must thank my auditing professors. They used to ask us, what do accounts sell? Trust. That's the only thing you sell is trust, that the numbers are correct. Yeah, so business is about trust. And it's also about value. and so startups you know they need someone they want a push button experience who will run tiger beetle for me because that's that that can make it cheaper for me than if i had an engineer do three months of work around the sre you know of a database right so you can actually have a business and sell something that's going to make something cheaper for startups and similarly for enterprise you can have a business and sell something that is going to provide the value they need which is now they might have sre teams but they need the infrastructure you know to support massive scale like petabyte how do you connect tiger beetle to object storage s3 like oltp data lake not only olap data lake but like let's just connect the oltp direct to s3 And so this comes to your question about the tug of war and licensing and all of this.
1:20:45And I think the big mistake that we can make, and I used to make this until it became clear for me that third moment, was that an open source license affects the implementation, not the interface. But when it comes to competition in the business world, that doesn't happen at the implementation. Typically, it happens at the interface. so if you think of like some of the you know the fiercest competition the most um you know when things were really on a knife stage for the web it was the interface not the browser implementations it was the interface that the war was fought you know mozilla like fought that war and we needed other browsers to fight it because the interface was being embraced extended and extinguished triple e you know and then then you think of like android and java and again um it wasn't about the android implementation it was the interface and that was that oracle google lawsuit you know so and then again you think of like well confluent you know kafka is apache 2.0 open source then red panda came along and i'm a huge fan of red panda because very similar design decisions around being direct with hardware, static communication, all of this very, very similar.
1:22:05We came from a similar time period, but in that time period, the things were changing how you build things. But Red Panda came along and they saw the open source implementation of Kafka and said, well, thank you, but we don't want it. But that interface, that is great. That's where our business will be also, that interface. Thank you very much. And so they built a massive of high-performance implementation. And then Warpstream came along and they said, well, Red Panda, you are business source license, not open source source available. Apache Confluent or Kafka, you are Apache 2 open source. But that's all implementation.
1:22:46I'm going to do my own implementation, thank you very much, of object storage, but the interface is great. OK, now we're all competing. And so I think the myth is that source available is kind of the thing that always, I always feel that, you know, something inside of me dies when I see a great company relicense or where I see a young startup follow that lead. Because to me, source available says that we think it's going to stop competition. It doesn't. You know, you may as well be on the beach building little moats and sandcastles. But innovation technology is like a wave. it'll just find a way around you you know it'll warp stream around you you can't you can't legislate competition away and we shouldn't be trying to build companies where we think the success of the company is us creating a monopoly the world's too big you know there's too much for you to be you don't need monopolies to do really great and then that doesn't build trust you know to say to your customers you can only buy from me so i think people think it stops competition and they think it helps them sell.
1:23:50And it actually defeats both of those because you get complacent and you actually fail to build trust. So you, you burn trust when you relicense and if you start source available, you're going to be doing diligence with enterprise. I'm going to say, but sorry, you're not open source. There's confusion, license confusion, you know, it's, and maybe some people get it, but there's a little 1 % headwind and it's actually, you know, you're, you're, It's a category error because you're spending so much effort chatting to people about implementation licensing. And the rest of the world is competing on interface.
1:24:26And, you know, Tiger Beetle's interface is quite simple, you know, very simple. So we could, I don't know, whatever license we apply, it doesn't matter. Debit credit is where we compete. There are companies that offer debit credit as an API, and it's very similar to Tiger Beetle. But we compete on trust. You know, we didn't just take a general purpose database and slap on debit credit. We went deep. You know, we really cared and we built the whole thing. And people pay us. So we, like, before we were a company, we had an energy company reach out and we landed, you know, quite a good contract very quickly.
1:25:05And I think it came down to trust. So open source builds trust. So open source is great for business. It's also orthogonal. all and yeah i think and i think the other thing is like if you so the there's a lot of things in tiger beetle that are like the i had done many experiments of my own you know my my passion projects they're all in tiger beetle many parts of the design of tiger beetle come from these various you know experiments that i did and so i was never going to put that all into a project if it's not open source because it's just too valuable you know i want to always be able to play with it no matter what happens to it and i think we all feel like that like our critical infrastructure it just has to be open source and so yeah i think that's kind of how i think of it no i think that's a great perspective and a specific explanation that i haven't heard previously so i definitely appreciate i'm i'm mulling on it i think i agree with most what you said and the implementation versus interface dichotomy is one that I hadn't considered that explicitly and I need to think through it more.
1:26:14So I appreciate your thoughts on the matter. Question is, what happens, you know, Apache 2.0, your heart and soul is into this thing. What do you do? How do you respond if and when AWS comes by and says, Tiger Beetle, Tiger Beetle by AWS, you know, for sale now. Like, does that scare you? Does that threaten you? Like, what do you think about that scenario? because that's what a lot of people are concerned about most specifically yeah so i think one can try and stop it you know you can build you can build the the castles in the sand or you can just say look the wave is inevitable it's coming like let's prepare for it uh and then so what we do is we we get the surfboard ready and we're on the beach we're waiting for it to come okay and as actually we're already paddling out and there's the swell the wave will take its time you know to catch up but we're already surfing and so there are there are going to be um like we just get into the water and and it's like you know we could have the cavalry in the castle or we could get them out into the field and have great cavalry and great user experience like let's compete on let's actually add value and serve serve the community honorably at a profit let let aws catch us in that you know And great, if they decide that they couldn't build a debit credit database as well as Tiger Beetle so that they offer it, you know, as their OLTP, like flagship database, Tiger Beetle, well, that builds trust.
1:27:46And then rising tide lifts all boats. And we're still, now we're surfing the wave with the database. Swimming, we're out there surfing. I love it. I love it. I think it's a great way to think about it because, like you said, the wave is inevitable. So you might as well prepare for it. You might as well ride it, you know, ride that wave. And I think also we've seen this play out a few times that source available doesn't stop AWS. So if it's valuable enough, they'll write the implementation. If you relicense, they will immediately fork your community. Now you've got two problems. They're still competing.
1:28:17And now they lead the interface. And that's when it's fatal for a company. When I see them relicense and you see the classic AWS, you know, I think, oh, gee, you know, they did. Yeah. that's actually the thing you don't want is when you know your open source clients are being bought up um and and being led now yeah but i love aws i love their work i've learned so much from the distributed systems i've got friends that work there so really to me that isn't the threat the threat is you know what's the problem with the world i i am you know we we are as a team so So the threat is really that we stop investing in performance and safety.
1:29:02We stop being trustworthy, you know, building trust. And so maybe we should say, Jared, too, is that your product is more than the open source. So like there's a principle here, too, is that I think it works both ways. is if we connect Tiger Beetle to a proprietary API, proprietary interface, our principle as a company is that's viral. So if someone wants to license their interface as proprietary, our connector will be proprietary. And for example, S3, if we connect to S3 for massive scale, then we charge for that. You know, fairly, we make an honest profit. So serve the community honorably at a profit.
1:29:46But there must be a profit because there's entropy in the world. Sure. Yeah. But if something's open source, well, then we connect it. And there's lots of, you know, just like people would pay Amazon for Aurora because there's a great value in all the management around Postgres. Again, porridge is too hot or cold. So we'd be curious to get your thoughts once you've thought about it. Yeah. I can sharpen the argument because I really think we need more founders need to stand up. And we've all been given a lot from the previous generation of open source. And I think it would be great if we all say, okay, we're going to pay it forward as well, like make a technical contribution.
1:30:30And there's no reason not to. You know, I think it's better for business, actually. Yeah. no definitely uh appreciate everything you said right there and i will be listening to it back as we produce this and consuming it more i love the surf the wave analogy that's where you really sold me so so far i'm amenable i'm amenable to your argument but you know i'm very easily convinced on the air here is there anything else uh tiger beetle or otherwise that you wanted to touch on that i haven't brought up yet it's been a great conversation yeah no i've loved it too i mean I guess we should say, I should say, we wrote it in Zig, this new system.
1:31:13Oh, yes. I didn't bring up the language wars yet. We have to get our clip, you know, because the flame wars must rage on. You wrote it in Zig, I assume from the day one, the day one decision. And you are probably happy with it since you just brought it up now. So Zig for the win, it sounds like. Is that your overall message? You love it? yeah it was also just a big wave you know you could see the swell and you're like i'm going for that wave like you're hopping on that wave yeah the same reason yeah when no js came out like i jumped on because i you know it made sense and then i was really happy that i did and when i saw zig cam come i thought well it's not often that you've had these moments rust was another one but with tiger beetle the timing meant that you know we could have caught the rust wave but there were so many thousands on already.
1:32:03We would have been a drop in the ocean. Also, we wanted to do static memory allocation, and Zig is actually really ideal for that. It's really perfect for the Tiger Beetle design. It would have been much harder to do our intrusive memory patterns and direct hardware access, iUring, zero copy. Two of our team, one of our team is actually the co-founder of the Rust Language with Graydon. He was the project lead, Brian Anderson Breeson at Mozilla. His desk was next to Brendan Eich. And he's writing our Rust client in ZIG. Well, in Rust. Sorry, he's writing it in Rust. I was like, wait a second. But he writes ZIG normally.
1:32:48But he's writing our Rust client. And then Matt Clad is the creator of Rust Analyzer and IntelliJ Rust. He's also, he joined the company. he's basically like a co-founder. So there were a few of us. And my senior partner from Coil, not many people know, but he came with, and Matt Clad, Federico, Rafael, but they're sort of the core team. And Matt Clad joined. He was trying to write, he wrote a blog post called Hard Mode Rust, trying to do static allocation, very similar patterns. And then Jamie Brandon introduced us. And MatCloud was like, I've been trying to do this in Rust. And they're like, but TigerStyle, you know, the way of coding in TigerBeatle or TigerStyle, we've got our engineering methodology written up.
1:33:42Oh, really? Can you link me up to that? We probably don't have time to go into detail, but I'd love to read it. TigerStyle, you call this? TigerStyle, yeah. And that's sort of all the safety techniques. So assuming that we do still have bugs, we also have like 6 ,000 assertions that check the operation of Tiger Beetle in production. If any one of them trips, it's like 6 ,000 trip wires, then there's probably maybe a CPU fault or memory fault or a bug. And then we immediately shut down the whole database. You want to operate safely or not at all. And that way you get total quality control and the system becomes safer and safer.
1:34:23latent bugs. So Zig really suited not only the performance, but also some of the safety decisions. Right. It wasn't merely the trend. There's also technical reasons that you chose it. Yeah, I think we picked Zig before Bun. The only other major Zig project at the time was River by Isaac Freund, a Welland compositor. Amazing programmer, Isaac. He contracted on Tiger Beetle for quite a while. And then it was Tiger Beetle and then Bun. Also Mach Engine was around the same time as Tiger Beetle. But we really just picked it because I was doing a lot of C code and Zig was just a perfect replacement for C and for all these new memory patterns.
1:35:11Yeah, that was it. Very cool. Ahead of the wave on that one, you were an early adopter on Zig and probably one of Zig's, I don't know, largest code bases, but maybe most production grade and out there, like successful projects to date. Would you say that's fair? I would say, yeah, Bun is also pretty massive. For sure. And, yeah, also there's also Ghosty by Mitchell Hashimoto, which is like incredible code. You know, like his performance work there is very similar, trying to get as close to pure memory bandwidth as you can. So he's making a terminal. You know, how can you make that as close to memory bandwidth performance?
1:35:54Right. Which is also what we think with Tiger Beetle and same as Jared with Bun. Yep. Yeah. And I think that goes back to what we were saying, you know, the love for that sim that you see. We're actually trying to show what it feels like if you're coding in open source because we've really crafted everything. it's just we you may as well you may as well enjoy and this sort of came from anti-res you know his craft of redis impacted me a lot and yeah well speaking of all these things we just put a clip out today as a record of our conversation with anti-res which was just a couple weeks back and did you know he's hard at work trying to get redis to be open source again inside of redis inc He's advocating.
1:36:44He's returned to Redis and he thinks he can get the company culture moved in a place where they'll get switched off of that proprietary new license and probably AGPL. But I think you'll find that good news considering your stance on open source. I'm excited about it. Hopefully that happens. That's great. And again, I should be clear. I love open source also because I love how it enables businesses. So I actually think it's great for business. I don't just like open source because I like it, you know. But I actually think it is better for trust, for sales. It makes everything easier. But yeah, that's fantastic news.
1:37:24I can't wait to listen. Cool, cool, cool. While you're on, I appreciate you coming on the show. I'm fascinated. I don't have any use cases for Tiger Beetle in my life, but I respect it. And I'm sure our audience will enjoy this conversation as well. So appreciate your time, appreciate your work and your perspective on open source and business, which is refreshing, especially in a sea of people who are kind of moving away from your perspective. But maybe back again. We'll see. I mean, Elasticsearch is back. Maybe Redis is coming back. We'll see what happens. But I appreciate you coming. Yeah, there's always a new wave to come.
1:38:01Yeah, and I appreciate you too, Jared. Thanks so much for this. It's been really, really special.
1:38:09Tiger Beetle sounds pretty amazing, and I'm always impressed by what you can accomplish when you laser focus a solution on a narrow problem space. That being said, general purpose solutions are amazing too, because they're useful in so many different problem spaces. It really does depend at the end of the day what you're trying to do when you decide to build or buy a solution and what to buy if you go that route. What did you think of this conversation with Euron? Let us know in the comments. links to all the things are in your show notes that includes tiger's architecture the simulator so cool and tiger style which i'm interested in checking out also a direct link to continue the conversation in our zulip community hop in there hang with us no imposters totally free why not right let's thank our sponsors one more time fly.io you know we love fly and a shout out to augmentcode at augmentcode.com.
1:39:07To Depot, build faster, waste less time, depot.dev. And of course, to Notion, Adam's beloved collaboration tool, notion.com slash changelog. Please do use our links and discount codes when you kick the tires on our sponsor's wares. That lets them know we're helping spread the word and that helps us put food on the table, which we like to do, as I'm sure you know. All right. That's all for this episode, but we'll talk to you again on ChangeLog and Friends on Friday. Bye, y 'all.
1:40:24Game on!
From the publisher
In July of 2020, Joran Dirk Greef stumbled into a fundamental limitation in the general-purpose database design for transaction processing. This sent him on a path that ended with TigerBeetle, a redesigned distributed database for financial transactions that yielded three orders of magnitude faster OLTP performance over the usual (general-purpose) suspects.
On this episode, Joran joins Jerod to explain how TigerBeetle got so fast, to defend its resilience and durability claims as a new market entrant, and to stake his claim at the intersection of open source and business. Oh, plus the age old question: Why Zig?
