Turing Award Winner: Postgres, Disagreeing with Google, Future Problems | Mike Stonebraker

20 Apr 2026 · 57 min · 24 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Mike Stonebraker (Turing Award winner) discusses the origins of Ingres and Postgres, why “one size fits all” databases fail, disagreements with Google’s MapReduce and eventual consistency, and what future database problems look like (LLM text-to-SQL limits; agentic AI needing atomic read-write workflows). He also explains DBOS, a database-centric “operating system” for durable transactional workflows.

Guest backgrounds

Mike Stonebraker is a Berkeley professor and database pioneer; he helped create Ingres (academic/commercial) and Postgres (extendable types, flexible data modeling). He later worked with/around Informix and is involved in DBOS (with Stanford/Databricks founder Matei Zaharia).

Key claims

Query optimization is the hardest database problem. Oracle’s sales practices were “shady” (e.g., referential integrity “not yet implemented”). Postgres’ extendable type system was driven by GIS and bond-time/date arithmetic needs. Indexing doesn’t parallelize well with SIMD/GPU execution. LLM text-to-SQL benchmarks show 0% on his Beaver benchmark; RAG helps to ~10%, and providing FROM/JOIN boosts to ~35%. Eventual consistency breaks integrity constraints (e.g., stock can go below zero).

Notable examples

Ingres failed for Arizona State due to no COBOL on Unix. Bond instruments required custom date subtraction semantics. GIS needs points/lines/polygons, which Ingres standard types couldn’t efficiently support. Google’s eventual consistency example: overselling leads to “minus one” inventory. DBOS example: atomic workflow to transfer $100 between accounts.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Origins of Postgres

0:45 to 2:44

Mike shares the story of how he got into building database systems and the beginnings of Postgres.

“Then, as well as today, you're way ahead if you get adopted by a mentor who knows the ropes.”

The Development of Ingress

2:44 to 5:58

Mike discusses the development of Ingress and the challenges faced in its early adoption.

“So it was clear that was a horrible hack.”

Ingress Corporation and Market Competition

5:58 to 6:40

Exploration of the formation of Ingress Corporation and competition with Oracle.

“So unsupported operating system, unsupported database system, no COBOL, doomed us to, you know, irrelevance.”

Technical Innovations and Transition to Postgres

6:40 to 9:06

Discussion about the innovations in Ingress and the transition to developing Postgres.

“I saw that Ingress was competing with Larry Ellison's offering at Oracle.”

Postgres Features and Flexibility

9:06 to 13:24

Mike explains the key features of Postgres that distinguish it from previous systems.

“So you created Ingress and there was a lot of technical innovations in it so that it was better than the incumbents.”

Identifying Extraordinary Software Engineers

13:24 to 14:00

Insights on how Mike identifies extraordinary software engineers during hiring.

“And as, you know, in business data processing, most people were happy with the standard data types.”

Innovations in Postgres: Features and Challenges

14:00 to 15:06

Explore the unique features of Postgres, including time travel and inheritance.

“And so Postgres, that was the big thing in Postgres.”

Identifying Extraordinary Software Engineers

15:06 to 18:09

Learn how to recognize exceptional talent in software engineering through specific criteria.

“You said, I can't stand people who aren't really smart.”

Database Systems: The One Size Fits None Argument

18:09 to 19:16

Understand why tailored database solutions are more effective than one-size-fits-all approaches.

“So I think it's just as true today as it ever was.”

The Impact of GPUs on Database Optimization

19:16 to 21:13

Discover the challenges and opportunities GPUs present for optimizing databases.

“Probably, but I think the big challenge is that GPUs are, you know, SIMD, you know, single instruction multidata.”
Show all 24 chapters

Disagreeing with MapReduce: A Different Perspective

21:13 to 22:54

Analyze the limitations of MapReduce and alternative approaches to distributed databases.

“It's still, if you ask most any senior database programmer what's the hardest part, they'll still say the optimizer.”

The Fallacy of Eventual Consistency

22:54 to 25:56

Examine the pitfalls of eventual consistency in databases and its implications for data integrity.

“But that wasn't the only thing Google was stupid about.”

Critique of Major Tech Companies' Database Strategies

25:56 to 28:00

Explore Mike Stonebraker's criticisms of how major tech companies approach database management.

“So referential integrity in a sales system is integrity constraint is stock is greater than minus one.”

Critique of Database Systems Support

28:00 to 29:15

Discussion on the inefficiencies of supporting too many database systems.

“And I said, you're supporting too many database systems.”

Academic Influence vs. Industry Work

29:15 to 31:00

Exploration of the choice between academia and industry for making an impact.

“And my one thought that I had is why not work directly in industry?”

Introduction to DBOS Project

31:00 to 33:56

Overview of the DBOS project and its innovative approach to scheduling.

“I just thought it was a really interesting technical model.”

DBOS Programming Language Features

33:56 to 36:33

Description of DBOS's programming language capabilities and workflow support.

“you're dreaming if you think you're going to displace Linux.”

Future of Read-Write Applications

36:33 to 39:40

Insights into the shift towards read-write applications in AI and databases.

“So the company is selling and innovating in this area.”

Potential of Databases in Operating Systems

39:40 to 42:00

Discussion on incorporating databases into operating system functionalities.

“with things, with people wanting stuff to be read-write.”

Unsolved Database Problems and Future Challenges

42:00 to 47:27

Explore the challenges faced by large language models in database queries and future technological directions.

“I think we talked a lot about the past of databases, and I'm curious your thoughts on unsolved problems in databases and what you think the future might look like.”

Benchmarking and Human SQL Proficiency

47:27 to 51:15

Learn about the discrepancies between LLM performance and human SQL proficiency in real-world scenarios.

“And so if you think you're really good at doing text-to-SQL, try a real benchmark, not a fake one.”

Advice for Future Generations in Computer Science

51:15 to 55:28

Discover insights on career choices, education, and following one's passion in technology.

“Wow, I'm surprised that the LLM scores so lowly on this kind of benchmark.”

Introducing an Innovative Ergonomic Keyboard

56:03 to 56:21

Learn about the development of a new ergonomic keyboard prototype.

“I'll put a link to the keyboard in the description.”

The Impact of Listener Feedback

56:27 to 56:45

Discover how audience feedback shapes the podcast and guest selection.

“for me about the show, I'd love to hear it.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Mike Stonebraker:Computer science may well not be a growth industry going forward. This is Mike Stonebraker. He's a Turing Award winner, famous for his fundamental contributions to database systems like creating Postgres and more. What was the hardest part of that implementation? A query optimizer. It's just algorithmically difficult. How do you identify the people who aren't smart? Well, I mean, it's very easy. He shared interesting technical takes from his experience. On our benchmarks, large language models get 0%. Why did you disagree so much with MapReduce? That wasn't the only thing Google was stupid about.

0:38I'm curious your thoughts on unsolved problems in databases and what you think the future might look like. Here's the full episode.

0:52the first thing i want to go over is the story of how postgres got started but for that i kind of want to start the beginning how did you get into building database systems when i graduated

1:04Mike Stonebraker:i had the good fortune of being hired at berkeley and i it was clear i had to you know continuing what I did for my PhD was not going to go anywhere. Then, as well as today, you're way ahead if you get adopted by a mentor who knows the ropes. So Gene Wong, who is still alive and still kicking, took me under his wing and said, well, let's do something together. And this was 1971, which was the year after Ted Codd wrote his pioneering paper in CACM. Gene Wong said, well, let's take a look at database stuff. And at the time, the competitors were a thing called the CODISIL proposal, which you're probably too young to have ever heard of.

2:05Mike Stonebraker:And so it was a low-level spaghetti network proposal where you executed queries by following pointers. And then the alternative was the IBM proposal, which was a thing called IMS, which is still available. And it's hierarchical data. It's a tree. You organized your data as trees. And even at the time, IBM realized that trees were not general enough to solve many people's problems. So they hacked on a way to make it a limited network structure. So it was clear that was a horrible hack. The CODISIL proposal had all kinds of bad properties besides being low level and really hard to debug. It also had the property that if anything changed in what's now called your schema, You basically had to throw away everything and do it all again because it was absolutely rooted at the physical level.

3:18Mike Stonebraker:Whereas Ted Codd's stuff made perfect sense. And so Gene said, well, let's build one of these puppies. That's clearly the next thing to try. So he started building Ingress in 1972. I was an assistant professor at Berkeley. As you know, if you're an assistant professor, you have to – you have about – you get five years to prove that you're a big shit. And they fire you or they give you tenure. So Ingress was my ticket to getting tenure, which happened in 1976. That was where it started. And then again, you know, happenstance. at the time, a lot of people would build prototypes, which were sort of student-y like code, which means you could get it to run, but if you gave it to anybody else, they couldn't.

4:15Mike Stonebraker:So we put in the first 90 % to get something we could run. And then for whatever reason, and we put in the next 90 % to get it to where it really worked. So the University of California version of Ingress really worked. And so over the next couple years, about 100 universities started running it because Unix became the big thing. And so this was a free database system that ran on Unix. And so it was quite popular in the academic world. And so we started getting lots of visitors at Berkeley who would say, gee, this is really nifty-looking stuff. What's the biggest ingrist application you have? And we'd be forced to say not very big.

5:15Mike Stonebraker:And so this was brought home in spades when Arizona State University considered running Ingress on their student records data, all 40 ,000 students worth. And they could get over that they had to get an unsupported operating system from Bell Labs. They could also get over they had to run an unsupported database system from these guys at Berkeley. But the project went down in flames when they realized there was no COBOL available for Unix, and they were a COBOL shop. So unsupported operating system, unsupported database system, no COBOL, doomed us to, you know, irrelevance. And it was clear the only way out of that was to start a company.

6:12Mike Stonebraker:And so in 1980, we got venture capital as it existed then and started Ingress Corporation to move Ingress to DexVMS, you know, a real operating system. and we had a real company that would support Ingress. And that was the start of the commercial journey. I saw that Ingress was competing with Larry Ellison's offering at Oracle. Yes. I saw that Ingress was certainly better than what they were offering, but they were still competing somehow. How did they compete? Larry Ellison is a fabulous salesman. And at the time, he made present tense and future tense indistinguishable. And so he basically lied to customers.

7:16He would ship stuff that didn't work and have his initial customers help him debug it.

7:25Mike Stonebraker:So I think he engaged in what I consider very shady business practices. But lying to customers, I think, is unconscionable. So, for instance, there was a thing called referential integrity, which is if you fire an employee and he's the last person in a given department, do you want to delete the department or do you want to have it be a department, a ghost department? It's all that kind of stuff. And so Ingress Corporation implemented referential integrity. Oracle Corporation wrote two manual pages that said, here's the definition of referential integrity, which everybody agreed to. And then down at the bottom, it said, not yet implemented.

8:25Interesting. Yeah, I had interviewed someone who worked at Sun Microsystems, and they had a similar opinion that Larry Ellison was a little bit shady. So it seems to be a commonality. I also saw somewhere else in something that you had said was that when Oracle acquired MySQL, that everyone kind of got afraid of that and moved to Postgres.

8:54Mike Stonebraker:That was the genesis of Postgres replacing MySQL as the preferred open source relational database system. So you created Ingress and there was a lot of technical innovations in it so that it was better than the incumbents. But ultimately, it went away and you developed Postgres. What was the thing that Ingress didn't do that Postgres would do? The big thing that guided us at the very beginning was the original reasoning for the academic version of Ingress was we were going to support a geographic information system that the neighboring professor Praveen Varaya wanted. And so to support a GIS system, you need points, lines, polygons, line groups, that sort of stuff.

9:56Mike Stonebraker:And it was clear that Ingress couldn't do it because the data types we put into Ingress were the standard ones, integers, floats, text strings. And you couldn't efficiently support GIS types on top of that. So as a GIS, the academic version of IGRS was a complete failure. And that was in the back of our mind. The other thing that happened, this is a little out of chronological sequence, but it helps make the point, is that the commercial version of Ingress, I think around 1985, you know, there was ANSI had just proposed a date and time standard for relational databases. And so commercial Ingress implemented date and time, using the standard Gregorian calendar.

11:05Mike Stonebraker:And so I was associated with the commercial version of Ingress as well as I was still at the University of California as a professor. So I got a call from an Ingress customer who said, you implemented date and time wrong. And I said, huh? We implemented the Gregorian calendar and you can subtract. And, you know, if it has, you know, days have 30 or 31 months except for February, except for leap years. So subtraction on dates works exactly the way you would expect it to. But he said, that's not what I want. In his particular world, he said he was dealing with bond financial instruments. And for some reason, I mean, you got the same amount of interest on his financial bonds during each month, no matter how long the month was.

12:12Mike Stonebraker:So he had the date you bought the bond, the date you sold the bond, he wanted to do a subtraction, multiply it by the coupon rate, and say that's the interest we paid you. But of course, his version of subtraction was March 15th minus February 15th is 30 days, because that's the definition of his calendar. And so he had to retrieve two dates out to user code, do the subtraction in user code, put the answer back, and it cost him a factor of two or three in efficiency. And he said, why can't I just overload your definition of subtraction with what I want? And, of course, with Ingress, it was hard-coded.

13:02Mike Stonebraker:And the problem was, this is a case where you wanted bond time, just like you wanted points, lines, and polygons. And so Postgres was engineered to have an extendable type system. So you could have whatever data types you wanted. And they were very efficient. And that was the main gist of Postgres, was that it had that flexibility. And as, you know, in business data processing, most people were happy with the standard data types. But relational databases started to spread to all kinds of other places. what are called abstract data types or stored procedures, a bunch of names they're called, had great applicability.

14:00Mike Stonebraker:And so Postgres, that was the big thing in Postgres. Postgres also supported what the AI guys at the time wanted in the way of inheritance. We also supported time travel, but the implementation absolutely sucked, and it got taken out after a while. So there were a huge number of really nifty things in Postgres. You mentioned you want to hire extraordinary software engineers, and I think you've said before that you have no trouble finding those people. How do you identify those people in your hiring, that they're the extraordinary ones? It's usually pretty obvious. I mean, I have a good feel for how difficult stuff is.

14:56Mike Stonebraker:If they get 3x the amount done in school that I think is reasonable, then they're incredible. On the flip side, you had this interesting quote. I wrote it down. You said, I can't stand people who aren't really smart. It's challenging to talk to them. How do you identify the people who aren't smart? Well, I mean, it's very easy. you talk to them and you can rapidly surface whether they're smart or not. What was your master's thesis? What did you do? Well, how did it exactly work? Well, how did you deal with error conditions? How many processes did you have? Why didn't you use threads? I mean, you ask them technical, deep technical questions.

15:53You gave a talk, and I think there's also a paper behind it, of this idea that one size fits all database systems, not optimal. One size actually fits none. And that what you really want is database solutions that target specific needs. What database offerings do you see today that are one size fits all?

16:15Mike Stonebraker:Well, in 2004, when I wrote the paper, we had an academic project which was building what became StreamBase. And so a stream processing engine looks nothing like a relational database. And we had the gist of an idea for column stores for data warehouses, which was popularized by Vertica, looks nothing like a row store. So here were three wildly different implementations that had no resemblance to each other. And in each case, they were an order of magnitude faster than the other guys. So it's pretty clear that with those three instances, you give up an order of magnitude when you're running a database system that isn't architected for your kind of stuff.

17:12Mike Stonebraker:I think that's still true. I mean, I think ClickHouse is a column store. Pinecone is faster than user-defined types on text-based vector processing. And so I think it's still very much the case. And I think there's no difficulty putting a common parser on top of multiple implementations. Postgres has so far chosen not to do that. They don't implement a column store. And so I think they are not competitive, you know, on sizable data warehouses. They also don't have multi-node support, again, for people with big data warehouses. That's table stakes. So I think it's just as true today as it ever was.

18:14I think that what is true is that if you want to get going, you have a database problem.

18:24Mike Stonebraker:You know, the answer is choose Postgres. and there's a huge programming community, all kinds of data type implementations. It's free. And you can find Postgres talent easily and get going. And so I think it's a great choice for lowest common denominator. And until you're trying to do a million transactions a second, it works just fine. Until you're trying to support a petabyte data warehouse, I say at the low end, it's absolutely the right. One size fits all, at the low end, it's Postgres. At the high end, that's just not true. GPUs, do they make available some new opportunities to optimize databases?

19:18Mike Stonebraker:Probably, but I think the big challenge is that GPUs are, you know, SIMD, you know, single instruction multidata. And that's the anathema of indexing. And so whenever indexing is the right answer, they're probably not a good idea. And I think also you've got to architect them so that the bandwidth from storage is not the bottleneck. And so if they're an add-on to the CPU, as often as not, the bus connecting it to the GPU to the CPU is a bottleneck. Can you explain why indexing would be not as effective when there's SIMD? So let's say I'm looking for Ryan's salary and I have a B-tree. So you go to the root of the B-tree.

20:31Mike Stonebraker:You find the divider that has both sides of Ryan. And you follow the pointer. That's a memory access for sure. Then you do it all again. And you do this like three or four times. So that doesn't parallelize well. So the answer is indexing doesn't parallelize well. You mentioned B-trees. When you first implemented that first version of Ingress, did you write all of that by hand? Because I imagine there's probably not some existing B-tree library or something. Yeah, we wrote the original version of Ingress was all written from scratch. What was the hardest part of that implementation? Query optimizer.

21:18And why is that hard?

21:20Mike Stonebraker:It's tough. It's just algorithmically difficult. It's still, if you ask most any senior database programmer what's the hardest part, they'll still say the optimizer. MapReduce came out at some point in the early 2000s, and it kind of took the data world by storm. People were really impressed by it. They thought Google really knows what they're doing. This is the best thing since sliced bread. But it seems like when I look at the literature and what you thought at the time, you can kind of disagreed heavily. Why did you disagree so much with MapReduce? Well, I think there were a lot of not very enlightened people who said, Google is really smart.

22:11Mike Stonebraker:They must know what they're doing. And so we'll do whatever they say. And so they would engage in Hadoop or engage with Hadoop. But Hadoop is ridiculously inefficient. And so at the time, you know, others, you know, Dave DeWitt and others who were involved in our 2011 paper, we understood distributed databases and understood that you could beat the heck out of Hadoop. with a distributed database system, which is basically what that 2011 paper says. And of course, it's true. But that wasn't the only thing Google was stupid about. So Google also had the opinion that eventual consistency was the right way to do concurrency control.

23:16Mike Stonebraker:And so that was postulated from on high by Google all during that same period of time. And it wasn't, and all the database people said, you know, you're out of your frigging mind. Because it doesn't, it solves one particular kind of problem, but only, and that very rarely occurs in practice. Why did they pursue eventual consistency? Okay, well, the idea is that you have an East Coast database and a West Coast database, and they're replicas. So you want them to be the same. If you say, I'm going to do a transaction, I'm going to decrement by one the number of widgets in the West Coast warehouse.

24:03Mike Stonebraker:Then I'm going to, before I commit that transaction, I'm going to update the East Coast warehouse, pay a message over and back to update it. And then to make sure everything goes well, it takes another round trip of message to make sure that both of them actually do the commit correctly. So it's expensive to do a distributed commit. And it still is. And so the idea was, well, you do the West Coast update. You decrease the widgets by one. You just send a message asynchronously and not in a transaction. So that eventually the East Coast warehouse gets decremented by one. So meanwhile, if you're on the East Coast, you decrement, you know, foodstuffs by one, you send an asynchronous message.

25:01Mike Stonebraker:Eventually the West Coast gets it and eventually everything settles out. So if you're allowed to go below zero, then what will happen is if the East Coast guy and the West Coast guy simultaneously sell the last widget, then eventually the state of the warehouse will be minus one and somebody won't get their widget. it. And so if you're allowed, like Amazon, to say usually ships in 24 hours, then maybe you're allowed to oversell. But most enterprises can't do that. And so eventual consistency just doesn't work. So we talked a million hours ago about referential integrity. So referential integrity in a sales system is integrity constraint is stock is greater than minus one.

26:10Mike Stonebraker:And that fails with eventual consistency. And so Jeff Dean of Google finally figured that out. And when they did Spanner, Spanner had a conventional transactional system. And so Google completely abandoned eventual consistency and completely abandoned MapReduce. So the tradeoff's basically correctness for performance. So it's performance versus data integrity. And if you don't care about your data, then you're willing to deal with bad things happening. So did you ever talk to the Google team while they were doing those things that you thought were so wrong? We talked to them before the 2011 paper and said, why don't we partner up and do some stuff?

27:14Mike Stonebraker:And they weren't interested. So they declined. Have you seen other examples in other big tech companies where their databases or database solutions where you actively disagree with them, like maybe Amazon or Facebook? Well, I gave a talk at Amazon maybe three years ago, and I told them all the things I thought they were doing wrong. And I think Amazon's problem is that they are supporting, you know, 15 different database systems. And that's about 12 too many. So I think they have their own culture. And I said, you're supporting too many database systems. And at this point, they haven't chosen to retire any of them.

28:09Mike Stonebraker:Why do you say that the 15 should be three? Well, they're supporting a graph-based database system. And it's well understood that a graph-based database system is almost never the performant option. And so if you want a graph, if you like the idea of having a user interface that deals with nodes and edges, is that's fine, put a layer on top of a relational database system that gives you that user model. And so most of their database systems, there's some other of their database systems that's better at what it does than it is. And so the answer is you should retire you should retire any database system that isn't performant in a big enough market to justify the maintenance.

29:14You've influenced industry significantly from academia. And my one thought that I had is why not work directly in industry? Or why do you prefer the position of being in academia and having influence in the way that you have versus just taking a job at AWS or something like that, being a very distinguished engineer there?

29:42Mike Stonebraker:Because that gives you a boss. And that gives you company rules, limits your ability to publish, limits your ability to go talk at conferences, limits your ability to go poke at what various competitors are doing that they won't tell their competitors. But mostly, I really like being in startups. And after the commercial version of Postgres got acquired by Informix. You know, I was working part-time for Informix, which was a 2 ,000-person company, and I didn't feel like I could make a difference because it was bureaucratic and and whatever the president wanted, he got. So I think I'm just not cut out for, I'm not cut out for politicking.

30:43Mike Stonebraker:I don't do that very well. And I have a hard time interacting with people I think are dumb. And that, again, so I guess I have some problems with big companies. I want to talk a little bit about DBOS. I just thought it was a really interesting technical model. Can you explain what DBOS is? We started the academic project in

31:14Mike Stonebraker:2019, 2020, something like that. And the gist of it was, at that point, Matei Zaharia, who was on the faculty at Stanford, was also one of the founders of Databricks was the original creator of Spark. And so he said at the time, Databricks, you know, basically was running people's Spark jobs on the cloud. And so he said at any given time, we might be orchestrating a million Spark jobs. And so we have to write a scheduler that's going to decide who to run next at scale a million. And he said there was no, we tried all the schedulers written by the OS folks, and they couldn't, they didn't scale.

32:14Mike Stonebraker:So we put all the scheduling data in a Postgres database, and basically a Postgres application was doing scheduling. And then it sort of clicked that, by and large, most everything you do in an operating system is managing data at scale. And you should do that using database technology. So why don't we just replace at least the upper half of Linux with a database system? So that was the gist of the academic project. and we worked on it at Berkeley and Stanford in the early 20s. And it was very successful. It clearly worked. And in the process, the Stanford folks wrote an extension to JavaScript so that you could program.

Read the full transcript

33:16Mike Stonebraker:you need some programming world that can talk to your implementation. So if you're doing what amounts to a programming language, and you're running on top of what amounts to an operating system that is a database, then the obvious thing to do is put all your state in the database. And that's exactly what they did. And so we had an innovative programming language model and innovative operating system model. And of course, then the idea was, well, can we start a company? And so we talked to the VCs who, to a person, said, you're dreaming if you think you're going to displace Linux. However, this programming language stuff is really nifty.

34:12Mike Stonebraker:We had what amounted to extensions to JavaScript that would allow any program to have all of the nice features of the database system. You know, stuff was durable. You could have transactions. If it failed, you'd fail over. You know, it was all that nifty stuff. So we got funded to start a company in 2023, and that was DBOSS Incorporated. And we decided that that was the name of the project. It's always been the name of the project. But we were basically in the programming language business. And so at the current time, And DBOS has a version of TypeScript, a version of Java, a version of Go, and a version of Python, which are basically seamless.

35:16Mike Stonebraker:It runs what looks like vanilla programs. In the world of the cloud, there's every incentive for you to structure your application as a workflow. And so we decided that we would support a workflow system, period. And so the workflow that Dboss supports in those four languages is the steps in a workflow, the individual micro-ops, whatever you want to call them, are transactional. Workflows are durable. so that once you do a step, it's not forgotten. And it's clear that we can make workflows atomic if there was a market for it, which means the whole workflow would either finish or look like it never happened.

36:21Mike Stonebraker:So it has very, very nice properties and is a great deal faster and a great deal easier to use than the competition. So the company is selling and innovating in this area. And so the idea is that you want to make state of your application persistent when you put it in the database. And then you figure out how to do it fast. And I think their business model, as we were talking earlier, is very much get leaf-level programmers interested. So it's been very much, you know, tell us, leaf-level programmer, tell us what you need that we don't have. Get it quickly and convince people to try it. And we've been very, very successful with other startups who want to choose the best thing.

37:30Mike Stonebraker:And we're starting to be successful with the big boys. So it's an interesting market. And I think the key thing so far is probably two-thirds of the customers are doing agentic AI. which means that they have a large language model surrounded by a bunch of stuff that adds more signal. And so far, most of agentic AI is read-only, meaning you want to produce a prediction for whether Ryan is going to be a good customer or not. And so that just runs some stuff and then produces a new thing that's given to somebody. So basically read-only, which means that you're not actually updating Ryan's credit rating.

38:39Mike Stonebraker:And so I think fairly quickly the whole world is going to move to using agents to do read-write applications. And that's going to make them very, very database-y. And DBOS does that stuff really, really well. And so, for instance, if you want to write an agent or two agents that move$100 from my account to your account. And so you debit my account, you increment your account. And these two agents have to agree to commit or you have to back everything out. which is to say the workflow needs to be what I call the atomic, which is it all happens or it looks like it never happened. And so I think the demands in this market will escalate with things, with people wanting stuff to be read-write.

39:49Mike Stonebraker:And so I think that And that will bode well for the market and bode well for DBOS. And what's being offered in the market today to application developers differs from the original research project, where that was actually swapping out the guts of an operating system with a database. I see. I mean, that's really cool. I never imagined replacing all the state of an operating system with a database. What's the, there's got to be some trade-off there. Well, a file system written on top of a DBMS is faster than the Linux file system. The scheduling engine is competitive with other scheduling engines.

40:35Mike Stonebraker:You can make everything failover so you get high availability without having to do anything else. The answer is there's really no downside. Then why wouldn't Linux incorporate that and upgrade itself with this? You hope they would. In other words, you should keep all the device driver junk down at the bottom because there's a lot of it and no one wants to do that and replace everything else with a database implementation. Is that something that you've mentioned to Linux people, and what's their typical reaction? Back in the academic project, when I'd mentioned that to operating system folks, they would get very, very threatened, which is this is the database guys trying to take over their turf.

41:31And I think the programming language guys ditto, which is the way to implement the runtime for a programming environment is with a database. That's interesting. I mean, if it's objectively true, then maybe it will take over.

41:52Mike Stonebraker:Well, I mean, it took Java 10 years to become widely accepted. I just think the time constant is substantial. I think we talked a lot about the past of databases, and I'm curious your thoughts on unsolved problems in databases and what you think the future might look like. Okay, so I think two different things that I'd like to talk about. The first one is, like everyone else, three years ago, we started to look at what were large language models good for. So we've been trying to get what's now called text to SQL.

42:38We've been trying to make it work on real-world databases.

42:46Mike Stonebraker:especially real world data warehouses. So we've been trying the technology on four different production databases, warehouses, where we've gotten the workload, the actual workload that's run and, you know, from the actual users using the system. And we've gotten them to reverse engineer the text that corresponds to that SQL. So we have text and SQL for, we have four benchmarks. When you say text to SQL, you mean like a human prompting model or something? Like I would just, in English, that text would be, you know, everyone over four years old. Tell me all the professors at MIT who won the Turing Award.

43:42Mike Stonebraker:And so an LLM is supposedly good at that. And so the text to SQL benchmarks, there's one called Spider, another one called Bird. And the best LLM systems are pretty good at those benchmarks, like 80 % accuracy or better. So not superhuman. not superhuman, but they're pretty good. Like you would consider using them. And, you know, like the current leaderboard is something like 85 % accuracy, which, I mean, it's getting there. You say maybe it's not quite ready for prime time, but it's simply, it certainly looks pretty good. Well, on our benchmarks, large language models get 0%. And if you enhance them with RAG and all the tricks, it goes to 10%.

44:40Mike Stonebraker:And if you give as a prompt the FROM clause, in other words, all the actual tables that need to be accessed, and all the actual join clauses that need to be joined, then accuracy goes to about 35%. So the definition of this stuff is not ready for prime time and not going to be for a while, if ever. So what's the difference? Number one, LLMs are trained on the pile. Data warehouse data is not in the pile. And there's an adage that if you haven't seen the data a couple times before, you have no chance of regurgitating it. That's number one. Number two, query complexity on Spider and Bird is maybe 10 to 20 lines of SQL.

45:44Mike Stonebraker:Real-world data warehouses, it's 100 lines of SQL. Complexity is bigger. Number three, the schema in spider and bird is clean. You know, the table names are mnemonic, the column names are mnemonic, and there's no duplication. In data warehouses, people have materialized views all the time. it means there's redundancy. And column names are often underscore Z, upper score, blah. And so they're not mnemonic. That makes it a lot harder. And then they also have idiosyncratic data. So J-term is a popular thing at MIT. It's a one-month term in January. not unique to MIT, but not very popular. So not in the pile, idiosyncratic data, simple queries, schema is a mess.

46:56Mike Stonebraker:Make it not work. And those are true of every data warehouse I know of. And so I think the technology simply doesn't work and isn't going to work anytime soon. So we've been, so what do you do? So, well, first of all, we published our benchmark. It's a thing called Beaver, which is an anonymized and abstracted version of these four data warehouses. And so if you think you're really good at doing text-to-SQL, try a real benchmark, not a fake one. So number two, borrowing from what I just said, if you don't have all the join terms and you don't have the from clause, you're toast. What's more, if you don't break down the query into simpler pieces, you're toast.

47:56Mike Stonebraker:So that says to me that you want to give your retrieval system simpler pieces, which include the from clause and include join terms. That's number one. Number two, the minute you want to talk to two different structured databases, you know, like your data warehouse and your CRM system, then it's pretty clear to me that doing a structured data join using an LLM is a bad idea. It's just you're much better off leaving them as tables and doing a join in SQL. So our point of view is we are trying out turning everything into tables. You know, we're working with the Department of Transportation in the city of Munich, Germany.

48:58Mike Stonebraker:And they have six people full time who are answering citizens' complaints, queries, which are of the form, how come I don't have enough time to cross this intersection next to my house before the light turns? All kinds of stuff. How come the trolley doesn't stop for enough time for me to get on the trolley? How come the trolley doesn't come more than once an hour? I mean, it's all this stuff. Their database is the trolley schedule, that's SQL. The light sequencing, that's SQL. The intersections, that's CAD. The federal, you know, country of Germany regulations of this stuff, that's text. City of Munich regulations for this stuff, which is text.

50:00Mike Stonebraker:So you've got to join SQL, SQL, CAD, text, and text. So our point of view is turn it all into SQL, all into tables, and do a join with what amounts to a query optimizer. So that's what we're working on. I think other people will have other ideas, but I think it's an extremely fertile area because people really want to do it. So that's number one. Number two, we talked earlier about agentic AI. The minute this becomes read-write, it's a distributed database problem. And you want atomicity, consistency, all that stuff. I think a very interesting area. So that's pretty much what I'm working on now.

50:51On that benchmark where it's 0 % right now, what percent is human? Like if you took someone who really knows SQL, what would they score? Like the average human.

51:03Mike Stonebraker:So once you disambiguate the text, a knowledgeable SQL user programmer with the schema will get very high accuracy. Okay, like 90 % or something. At least. Okay, okay. Wow, I'm surprised that the LLM scores so lowly on this kind of benchmark. Maybe when this goes out, someone who works at Anthropic will reach out to you or something and say, let's. I'd love to find out because, I mean, it's a terrific success story. For people who want to deeply understand databases and they're looking for some material to study, is there a book that you'd recommend that's a top technical book to learn databases?

51:50Papers in the Literature.

51:53Mike Stonebraker:I think Joey Hellerstein and I published a red book, what's called the red book, which is called Readings in Database Systems. It's now eight years old. I mean, I think that would be a great set of readings for eight years ago. And beyond that, papers, popular papers from the literature. If you could go back to yourself when you just graduated, give yourself some advice, knowing what you know today, what would you say? Back when I first took the job at Berkeley, without thinking about it much, we said, let's write a database system. And we knew nothing about databases, nothing about implementations.

52:40Mike Stonebraker:We were not skilled programmers like Bill Joy. So starting off doing something that was that crazy was really pretty crazy. And, you know, you effort and you make stuff work and you learn along the way. And so I think the answer is think outside the box, think crazy thoughts, and try and do them. And I think to me, it's not at all obvious. The better question is, if you were starting out today, what would you major in? Because I think, you know, computer science may well not be a growth industry going forward. And I'm not sure I would recommend 18-year-olds to major in computer science. I mean, I think health care and the building trades are safe bets and everything else looks much riskier.

53:46Mike Stonebraker:If you're about to get your PhD and are trying to decide what to do, then I think life is pretty easy. Take the most prestigious job you can get and find a mentor who's willing to help you. And then pick some area that isn't, you know, like our stuff, you know, which is called Rubicon, is definitely not going with the flow. So choose something that isn't going with the flow and try and make it work. Both my wife and I said, follow your passion. Somehow the money will work out. And I don't believe that for a minute, but I think that's what you have to tell your kids and your grandkids. If you don't believe that, then why do you have to tell them that?

54:42Mike Stonebraker:And my wife is a good example. So she has a master's degree in computer science, undergraduate degree in computer science. And she wanted to be a teacher, you know, K-12 teacher. And her parents said, you can't do that. It doesn't pay enough money. And so I think she, ever since that time, has regretted that decision. She wasn't passionate about doing computer science. It was simply a trade. And so I think find something you're passionate about and you will, you know, either you won't starve. you may not make a lot of money, but I think chances are you'll be happier than if you do something you're not passionate about.

55:34Mike Stonebraker:Because I think a lot of people I know view their job as simply a job, and life is what happens between 5 p.m. and 8 a.m. I don't feel that way at all. I really like what I do. It wouldn't matter whether I made a lot of money or didn't. awesome all right well thank you so much for your time really appreciate it thank you for listening to the podcast it's a passion project of mine that i really enjoyed building another passion project that i've been working on kind of in secret is building an ergonomic keyboard that i wish existed and i finally have a prototype so i'd love to show you what we've built it's ultra low profile and ergonomic and i couldn't find anything like it on the market so that's why we built it.

56:19I'll put a link to the keyboard in the description. You can take a look and learn more about the project there. We could definitely use your support. Also, if you have any feedback for me about the show, I'd love to hear it. Comments on YouTube have led to guests coming on like Ilya Grigorik and David Fowler. I wasn't aware of them until someone dropped a comment. Also, feedback in the comments helped me learn to reduce the number of cliffhangers in the intros. So your comments definitely make a difference. Please keep letting me know what you'd like to see more of in the show and I'll see you in the next episode.

From the publisher

Mike Stonebraker is a Turing Award winner famous for his contributions to fundamental database technologies. We discussed the story behind building Postgres, where he disagrees with Google/Amazon on databases, and what he's working on now.


𝗣𝗼𝗱𝗰𝗮𝘀𝘁 𝗹𝗶𝗻𝗸𝘀:


• YouTube: https://youtu.be/YPObBOwIrHk

• Apple: https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835

• Transcript: https://www.developing.dev/p/turing-award-winner-postgres-disagreeing


𝗘𝗽𝗶𝘀𝗼𝗱𝗲 𝗹𝗶𝗻𝗸𝘀:


• Red book of database readings: http://www.redbook.io/

• BEAVER: An Enterprise Benchmark for Text-to-SQL: https://arxiv.org/abs/2409.02038


𝗧𝗶𝗺𝗲𝘀𝘁𝗮𝗺𝗽𝘀:


0:00 - Intro

1:03 - How he got into databases

6:43 - Competing with Oracle

9:07 - What made Postgres special

15:55 - One size fits none

21:37 - Why he disagreed with Google

29:14 - Why he chose academia over big tech

30:58 - Replacing state in an OS with a DB

42:02 - Future problems in databases

51:36 - Technical book recommendations to learn databases

52:20 - Advice for younger self

55:52 - Outro


𝗪𝗵𝗲𝗿𝗲 𝘁𝗼 𝗳𝗶𝗻𝗱 𝗠𝗶𝗸𝗲:


• His current company DBOS: https://dbos.dev/


𝗪𝗵𝗲𝗿𝗲 𝘁𝗼 𝗳𝗶𝗻𝗱 𝗥𝘆𝗮𝗻:


• Newsletter: https://www.developing.dev/

• X/Twitter: https://x.com/ryanlpeterman

• LinkedIn: https://www.linkedin.com/in/ryanlpeterman/

• Threads: https://www.threads.com/@ryanlpeterman

• Instagram: https://www.instagram.com/ryanlpeterman

• TikTok: https://www.tiktok.com/@ryanlpeterman

More from The Peterman Pod

All 60 episodes
Turing Award Winner: Postgres, Disagreeing with Google, Future ProblemsThe Peterman Pod · 57 min
Listen in VO