The future of AI and the law

11 Jul 2025 · 36 min · 16 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How AI can support legal work (research, document scanning, legal reform) while warning that large language models can hallucinate and may not yet replace lawyers; includes examples of using AI to redact racially restrictive covenants and to identify obsolete reporting requirements, plus debate over open vs closed AI models and prospects for legal chatbots.

Guest backgrounds

Dan Ho is a Stanford professor of law, political science, and computer science, focused on AI analysis of legal documentation and legal reform.

Key claims

General-purpose LLMs hallucinate 60–80% on basic legal facts (about 800,000 benchmark queries); retrieval-augmented generation reduces but still hallucinates 1/5–1/3 of the time. AI is promising for high-volume legal document tasks, but specialized, evaluable use cases are safer than end-to-end chatbots.

Notable examples

Santa Clara County redaction of racially restrictive covenants in deed records (5 million records processed in days; multimodal OCR improvements). STARA system scanning statutes/regulations (San Francisco Municipal Code/Resolutions: ~16M words; ~528 reporting obligations flagged; 351-page resolution to delete/modify over a third). “Regulatory sludge” examples: reports on fixed pedestal newspaper rack zones; Federal Reserve reporting on the presidential dollar coin program (ceased in 2011). Open vs closed risk assessed via “marginal risk” framing (RAND bio-attack study found no significant difference).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Legal Hallucinations and AI

0:48 to 1:50

Discussion on studies showing AI's hallucinations in legal facts.

“like ChatGPT, when faced with roughly 800 ,000 benchmark queries, actually hallucinate 60 to 80 % of the time on basic legal facts.”

Introduction to Dan Ho and AI in Law

1:50 to 3:15

Introduction to Dan Ho and the use of AI in legal contexts.

“Today, Dan Ho will tell us that large language models like ChatGPT can be very useful for law, but you have to be careful about a lot of hallucinations, and they might not be ready yet to replace your lawyer.”

Understanding Law's Complexity for AI Opportunities

3:15 to 6:45

Exploration of what individuals must know about law for AI applications.

“law, political science, and computer science at Stanford University, and he's an expert at the use of AI to analyze legal documentation.”

AI's Challenges in Legal Research

6:45 to 9:13

Discussion on how AI models perform in legal research and the associated risks.

“And again, right before we get into some of the really compelling applications, one of the things many of us have played with these large language models like ChatGPT and others.”

AI in Identifying Racial Covenants

9:13 to 14:08

Details on a project using AI to identify and redact racial covenants in property records.

“Yeah, we have, well, let me start, There's sort of two projects that are directly on this topic.”

Examples of AI in Legal Reform

14:08 to 15:27

Learn about the usefulness of AI in legal reform and the importance of data in legal processes.

“Of where the AI system really was useful.”

Introducing STARA: Statutory Research Assistance

15:27 to 19:21

Discover how STARA aids legal research by using AI to analyze vast legal texts.

“Yeah, and it was an incredibly time-consuming research process.”

Identifying Obsolete Reporting Requirements

19:21 to 21:09

Explore how AI can streamline governmental reporting by identifying outdated regulations.

“We're able to identify around 528 reports.”

The Impact of Regulatory Sludge

21:09 to 23:04

Understand the concept of regulatory sludge and its implications for bureaucratic efficiency.

“This is the future of everything with Russ Altman.”

The Cost of Redundant Reports

23:04 to 24:47

Learn about examples of unnecessary reporting requirements and their impact on resources.

“Those are quite literally pieces of concrete that were elevated to have to be able to sell copies, physical copies of the San Francisco Chronicle.”
Show all 16 chapters

Open vs Closed AI Models in Law

24:47 to 27:08

Delve into the debate over transparency in AI models and its relevance in legal contexts.

“And each year, the Federal Reserve has filed this, but also noted that the presidential dollar coin program ceased to exist in 2011.”

Legal Hallucinations and Their Consequences

27:08 to 28:00

Explore the phenomenon of legal hallucinations and their potential legal implications.

“judge might ask a lawyer on what is your claim based and can I examine it or can the opposing counsel examine it?”

Examining Legal Hallucinations in AI

28:00 to 29:49

Learn about the documentation and implications of legal hallucinations in AI-generated cases.

“There have been these incredible efforts, including by one of our own Stanford law lums, Peter Henderson, to track all instances of legal hallucinations in cases.”

Assessing Marginal Risk of AI Models

29:49 to 31:00

Understand the marginal risk framework for comparing open and closed AI models, especially regarding bio risk.

“And the evidence for that, at least at that time in terms of the bio risk, was not at all clear.”

The Future of Legal Chatbots

31:00 to 33:18

Explore the potential and limitations of general-purpose legal chatbots in providing quality legal advice.

“What's also tricky here is that just as you know, openness is a spectrum.”

Specialization in Legal AI Solutions

33:18 to 34:40

Discuss the necessity for specialized legal AI chatbots to minimize errors and better serve underrepresented litigants.

“In some of the earlier studies I mentioned, these general purpose chatbots have a preventative to hallucinate at alarmingly high rates.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Daniel Ho:This is Stanford Engineering's The Future of Everything, and I'm your host, Russ Altman. I thought it would be good to revisit the original intent of this show. In 2017, when we started, we wanted to create a forum to dive into and discuss the motivations and the research that my colleagues do across the campus in science, technology, engineering, medicine, and other topics. Stanford University and all universities, for the most part, have a long history of doing important work that impacts the world. And it's a joy to share with you how this work is motivated by humans who are working hard to create a better future for everybody.

0:37Daniel Ho:In that spirit, I hope you will walk away from every episode with a deeper understanding of the work that's in progress here and that you'll share it with your friends, family, neighbors, co-workers as well. We did a couple of studies of legal hallucinations and we showed that general purpose language models like ChatGPT, when faced with roughly 800 ,000 benchmark queries, actually hallucinate 60 to 80 % of the time on basic legal facts. Then we did another study of legal AI when sort of using systems known as retrieval augmented generation, where you try to retrieve the relevant legal document and then provide an answer.

1:17Those have substantial improvements, but can still hallucinate at non-trivial rates between one-fifth to one-third of the time.

1:32Daniel Ho:This is Stanford Engineering's The Future of Everything, and I'm your host, Russ Altman. If you're enjoying the show or if it's helped you in any way, again, my favorite bar, a low bar, anyway, please consider rating and reviewing it so that others can see what you think of it and can consider tuning in. We love to get a 5.0 if we deserve it, and we also love to hear the comments. Today, Dan Ho will tell us that large language models like ChatGPT can be very useful for law, but you have to be careful about a lot of hallucinations, and they might not be ready yet to replace your lawyer. It's the future of AI and the law.

2:11Daniel Ho:Before we get started, another reminder to rate and review the show, particularly if you found it useful or have learned something new.

2:25Daniel Ho:You know, large language models have gotten a lot of attention in the last couple of years. They are amazing at taking large amounts of text, summarizing it, searching through it for the pieces of information that you need. And many of us are seeing them in everyday life increasingly. Well, that includes lawyers. As you know, lawyers need to be familiar with a lot of text information. They need to know the law, they need to know regulations, and they need to review previous cases to make sure that they're making arguments that are consistent with precedent. Well, there's a great opportunity for AI to help in law, especially in finding useful laws or regulations that help your case, but also in finding out-of-date, even racist regulations and laws that need to be removed from the public record.

3:13Daniel Ho:Well, Dan Ho is a professor of law, political science, and computer science at Stanford University, and he's an expert at the use of AI to analyze legal documentation. Dan is going to tell us how his work has advanced legal reform by going through huge corpora of legal documents and finding the ones that just need to be changed. Dan, what led you to study the law and AI together as a major focus of your academic work? Well, gosh, I had long been interested in both of these talks, but honestly, it goes back even to growing up when my parents had grown up in post-war Hong Kong. That generation had fled from mainland China, my parents moved to Germany.

4:06I grew up in Germany at a time when you had a generation of folks question the choices of the earlier generation. And I have this vivid memory of actually being a young kid when the Berlin Wall came down. And so I'd always had the sense of how fragile our social institutions can be. And so it had this deep interest in really understanding our social institutions and government in particular, and then always had in mind, well, what can we do to really help make sure our institutions adapt and are able to change with the times?

4:52Daniel Ho:Great. So we're about to talk about a lot of really interesting legal applications of AI that you've dove into. And it occurred to me that I should ask you, what does the regular person need to know about law and the practice of law in order to understand what the opportunities are for AI and computation? So many of us know that it's adversarial. We know that there are cases that serve as very important precedents. We have a rough idea that there are statutes and regulations that are not laws, but that are still in some way part of the expectations of society. What would you say we need to know as we are about to dive into some of your work?

5:30Well, gosh, there's so many things one could say, but I would highlight two. One is, you know, we might, and this will maybe be too dated a reference, but the popular understanding of what lawyers do might be a Perry Mason kind of standing up in the courtroom. But so much of what lawyers really do day to day is legal research and writing, identifying the sort of operative case law and statutes that are relevant. And then the second is that law is inherently contested due to this adversarial system, as you note. And that has actually meant that over the past few years, as we've seen AI beat so many of the conventional benchmarks, that law has become a really interesting terrain for AI research because it is not a simple kind of question answering setting where you can find the right Wikipedia article, get the right answer, and you can know whether that's a correct answer or not.

6:33Legal reasoning and legal argumentation is much more complicated, and that has made really for fertile research for AI. time.

6:45Daniel Ho:And again, right before we get into some of the really compelling applications, one of the things many of us have played with these large language models like ChatGPT and others. And there's always a question about whether they're ready to go out of the box or whether they need a lot of work. So I guess a high level question is, if somebody goes in to ChatGPT and asks for legal research advice, either a lawyer or a non-lawyer, out of the box, How are these performing? Have they shown good, I don't know, legal instincts? Well, we've seen remarkable advances in the ability of large language models to actually provide answers and have a lot of legal knowledge encoded in them.

7:27And I would say when you're talking about querying a model like this for widely available legal knowledge, they actually do pretty darn well. The challenge is that it's precisely in the instances where you need lawyers, where the law might be uncertain because you're in a particular state that has a different rule or you have novel circumstances, that these are not things that can just be readily learned off given materials of the Internet. And so these systems can perform poorly. We did a couple of studies of legal hallucinations, and we showed that general purpose language models like ChatGPT, when faced with roughly 800 ,000 benchmark queries, actually hallucinate 60 to 80 % of the time.

8:18Oh, wow. on basic legal facts. Then we did another study of legal AI when sort of using systems known as retrieval augmented generation, where you try to retrieve the relevant legal document and then provide an answer. Those have substantial improvements, but can still hallucinate at non-trivial rates between one-fifth to one-third of the time. And so I think one still has to really watch out for the potential pitfalls of hallucinations. One theory paper, that's a general theory paper, has shown unless the answer is directly contained in the training data itself, well-calibrated AI systems will necessarily hallucinate due to the way that they are trained to generate.

9:09Daniel Ho:Okay, great. So thank you for that very helpful background as we now go into some of the really interesting things you've done recently. And the first one that came to my attention is looking at large corpora of legal documents for statutory statutes or regulations that are of great interest to a lawyer in a certain case, or are something that needs to be purged from the record because they're out of date, anachronistic. So can you tell me about that project? What motivated it and what'd you do? Yeah, we have, well, let me start, There's sort of two projects that are directly on this topic. Let me start with one that was really motivated by a well-meaning but difficult to implement piece of California legislation.

9:53In 2021, California enacted a legislation that required all of the 58 counties to establish processes to go through their property deed records and identify and redact them for what are known as racially restrictive covenants. What are those covenants? When we purchased our home here in the Bay Area, I had to sign a piece of paper that said this property shall not be used or occupied by any person of African, Japanese or Chinese or any Mongolian descent, except for in the capacity as a servant. And these have been unenforceable since the mid 20th century due to a US Supreme Court case, but they still persist in many of these deed records because they run with the land.

10:40And so California enacted this piece of legislation, which was well meaning, but in a county like Santa Clara, that could mean that you have to sift through 84 million pages of property deed records that date back to the 1800s. until we started our collaboration with the county.

10:58Daniel Ho:Was it simply to, forgive me, but was it simply to identify these or to expunge and rewrite them? Both, to identify and preserve them for the historical record so that people could understand this bit of local history and state history, but also to go through a formal legal redaction process that involved the recorder and legal counsel to re-record the deed record to omit the unenforceable racial covenant. Right. And this could affect house transactions every day because anybody who's bought a house knows that there's these title searches and it's all very kind of, I don't know, choreographed.

11:39Daniel Ho:Exactly. This well-meaning rule, I could imagine, might disrupt that choreography. Oh, the worry was that because it's such a huge volume of deed records, you're asking each of the county recorders to go through, it could lead to a, quote, near shutdown of county recorder offices. Until we started this collaboration, they had a team of two that read nearly 90 ,000 pages manually to identify 400 of these covenants. And that was the worry, is this is just going to take us forever. Los Angeles contracted with a vendor to do keyword searches for$8 million to conclude the process in a matter of seven years.

12:22And we felt this is exactly a place where we can help build out an AI system. And so we curated a kind of good benchmark data set of these kinds of racial covenants and managed to actually collaborate with the county to build out a model to be able to go through 5 million of these deed records and actually accelerate the redaction process and bring it down to be able to identify these things in just a matter of a couple of days.

12:52Daniel Ho:Can that be exported to the other 57 counties? We are very much interested in doing that and are currently working with several other counties, including San Francisco and Yolo County here to basically bring that technology to accelerate the process elsewhere in California. And was part of your job to digitize these or were you blessed with previously digitized documents? In Santa Clara County, we were blessed with previously digitized documents. But actually what we found is that the existing OCR, optical character recognition system, was not great. And that's been another interesting kind of discovery in building up the system is that actually the ability of multimodal models to do really good character recognition, what we found evidence of is almost a kind of leapfrogging of conventional custom OCR technology through multimodal models that both had lower error rates and were actually cheaper than conventional OCR solutions.

13:53So we actually built that into our pipeline. So we basically ingest just the page images, do the full OCR process, and then have a built out, a fine tuned large language model to actually be able to identify where on the page one of these racial covenants exists. Great.

14:08Daniel Ho:So that was a beautiful example. Thank you. Of where the AI system really was useful. It really addressed something that needed to be done. In fact, it sounds like it was mandated, but nobody had a plan. And now there's a plan. You said there was a second example. Oh, sure. Yeah. What we sort of realized after that is this is one example of a pervasive problem in legal reform, which is that so often what you're overwhelmed with if you're trying to engage in a legal reform effort or trying to litigate cases or do legal research, it's just the sheer volume of materials in front of you. I am reminded of then law professor Ruth Bader Ginsburg, who hired a small army of Columbia law students to search for 59 keywords across the entire U.S.

14:58code, which in present day is over 30 million words. And they went through each one of these to read every single code provision to identify all instances of gender bias in the U.S. code. And that formed the litigation plan for how Ruth Bader Ginsburg began to really litigate and reconceptualize equal protection in the United States.

15:26Daniel Ho:Yeah, so it really all started with the data. Yeah, and it was an incredibly time-consuming research process. And so what we built out in this other research project is a statutory research assistance system. We call it STARA, that actually is able to ingest something like the US Code, represent it in a way that we would teach statutory interpretation to law students, and then be able to apply the power of large language models to do systematic scans of statutes and regulations, much in the way that then law professor Ruth Bader Ginsburg did. But in this case, it's general, if I'm understanding, it's general purpose.

16:08Daniel Ho:So you're not specifying the kinds of statutes to look at. You're creating a platform with which a legal person could do many different searches. Exactly. Exactly. And what are the types of searches that are, well, you already told us about these covenants, these racist covenants. Are there other example searches that are kind of of contemporary relevance? There's so many. In the 1980s, the Reagan Department of Justice tried to enumerate all federal crimes. And the person who spearheaded that effort after two years of an effort gave up and said, like, it's impossible for a human being to do this.

16:47The example of present day is that, you know, there's a lot of reflection these days about why we have so much procedural bloat in a lot of our regulatory processes. So after we built out STARA, we partnered with the San Francisco City Attorney's Office, led by David Chu, really to explore these different use cases. We looked at a range of different things, but what we really zoomed in on were these legislatively mandated reports or reporting requirements that the Washington Post in one piece on this called a kind of congressional black hole in the sense that Congress has a propensity just to demand many, many reports.

17:34And of course, at a time when we're demanding more and more of the civil service, these can really weigh our civil servants down to complete reports that are not necessarily serving present day policy purposes. I think -

17:53Daniel Ho:So this is a situation, for example, where they're spending a big bunch of money on some activity and they say, oh, and we should, as part of our stewardship, we should have a report every year or every quarter on how things are going or what's working well and what's doing. But this accumulates now over a couple of hundred years of government. Yeah, the challenge is that you have things that were enacted like this 80 years ago and are not necessarily relevant in present day. And these things can consume a lot of time. In Neil Korsuch's book, he points to one report by the Social Security Administration that required it to write up a report on its printing operations.

18:33And he reports that that took 95 federal employees four months of time to compile a report that, for instance, included things like the serial numbers of forklifts and printers, all nicely curated for Congress to read. And what the Washington Post reports is very few of these reports are really good. And so what we did with the city attorney's office is really think about that process in San Francisco. We ingested the San Francisco Municipal Code and Resolutions. You might be surprised to find out that it's on the same order of magnitude as the U.S. Code.

19:08Daniel Ho:Oh, my goodness. It's about 16 million words. And so that's a really difficult thing, a difficult research process for any human being to undertake. And we ingested it into the Stara system. We're able to identify around 528 reports. And then the city attorney kicked off a consultation process really to identify which of these no longer serve a present policy purpose, which are obsolete, which are just not even useful to do, and proposed in a 351-page resolution to delete or modify over a third of these reporting obligations. And who is the decider on that? Is that something that has to then go to, I don't know, the mayor or the council or who says, yes, let's delete that stuff?

19:58Yeah, the resolution has been introduced and will go to the board of supervisors for consideration, you know, in the same way that sort of other pieces of legislation are considered.

20:10Daniel Ho:And before we move on to another topic, how did you know? You said something like 500 plus reports. How do you know if you didn't miss any? And did you find some that actually, when you looked more closely, they actually weren't reporting requirements? Yeah, it's such a great question. One of the things that we did in the underlying STAR paper is really benchmark this against known tasks. And what we were able to show, for instance, in the federal context where there have been these attempts to try to enumerate all of the reports, we're able to get a near perfect recall of known reports and then find thousands and thousands more.

20:49And then we sampled the additional ones to get a sense of the hit rate on those. And the hit rate due to the performance of large language models is actually quite high on those. And so it really allows you to focus the city attorney's time on the kinds of things that actually matter the most as opposed to being swamped with an array of false positives.

21:12Daniel Ho:This is the future of everything with Russ Altman. More with Dan Ho next.

21:31next.

21:31Daniel Ho:Welcome back to the Future of Everything. I'm Russ Altman. I'm speaking with Professor Dan Ho from Stanford University. In the last segment, Dan told us about the basic opportunities for legal AI, some projects that he's worked on, and the problems with hallucinations in law. In this segment, he's going to give us some specific examples about some of the regulations that they uncovered using AI. He's also going to tell us about the debate between open AI systems and closed AI systems. And he's going to end with some advice and the outlook for a chatbot that might replace your lawyer. Bottom line, not so fast with that plan.

22:10Daniel Ho:Dan, you told us about a really great study where you were looking through old regulations, old statutes, and you found all kinds of stuff. I can't help but ask, can you give us some good stories? Oh, sure. I mean, you know, so much of what the city attorney's office did after getting this list of 500 some of these reporting requirements is try to go through which one of these are obsolete and which ones of these are really still serving a present policy purpose. And we found a lot of what you might call regulatory sludge. Sludge. Regulatory sludge. I love it. The kind of stuff that really just exists.

22:50There may have been a good reason at one point in time to put it in there, but it just doesn't make sense.

22:55Daniel Ho:It's not lubricating the wheels of justice. It is doing the opposite. It is slowing down the wheels of justice. So I'll give you a couple of examples. One is that the director of public works has to regularly file a report on so-called fixed pedestal zones for newspaper racks. Those are quite literally pieces of concrete that were elevated to have to be able to sell copies, physical copies of the San Francisco Chronicle. Wow. That was a live issue at some point decades ago. But those have largely fallen into disuse. It's not a good use of the director of public works his time to create a long report on this.

Read the full transcript

23:39Daniel Ho:So let me ask about that because when you were talking about it, I had – so do they just ignore it realizing – I mean I'm guessing that they just ignore these requirements and they're not that worried about being taken to task even though in principle they're in violation of a regulation. Is that the way to think about how a responsible bureaucrat deals with these things? That is the tricky position you put bureaucrats in. I think the Washington Post had this fantastic piece on it where one agency official said, you know, we've been doing this every year. It's been taking us so much time. What if we just don't do it this year?

24:13And you have a ton of staff members that themselves just describe being inundated with these reports and using them essentially as doorstops. And so, nonetheless, you also have a lot of instances of agencies continuing to dutifully file these reports. Because it's their job. Because it's a statutory requirement. So the Federal Reserve, I'll give you one other federal example, the Federal Reserve each year has to file a report on the presidential dollar coin program. And each year, the Federal Reserve has filed this, but also noted that the presidential dollar coin program ceased to exist in 2011.

24:57Could you please relieve us of this reporting obligation? And, you know, those are the kinds of things that really should be by default sunsetted or at least up for reconsideration because it is just not worth the time of the Federal Reserve to keep filing something about a defunct.

25:12Daniel Ho:OK, so my assumption that they would just do benign neglect is really a big assumption. And that I shouldn't assume that that's that these reports are actually in many cases actually still being created and written. Yes. Thus creating the sludge. Was there anything else from San Francisco that you wanted to pass on? Oh, I mean, we see provisions that are 80 years old. There's an 80-year-old requirement that the redevelopment agency file quarterly reports on things. And, you know, those are the kinds of things that really should be updated. We have a lot of reports on entirely defunct programs.

25:46Right. But of course, there are also reports that still are quite important. And in some of those instances, what the city attorney proposed to do is to try to actually consolidate where it makes sense, rather than having a handful of different reports that are all about the housing stock. Would it make sense rather than having five different teams do that to have an annual housing inventory report that consolidates all of those reporting requirements? Okay.

26:13Daniel Ho:Well, thank you. Those examples were as satisfying as I was hoping that they would be. I wanted to move to actually a more theoretical issue, and it's a little bit technical, but I think that we all need to think about it, which is, as you know very well, there are debates about these large language models and whether they should be opened or closed. Open models would be ones where it's transparent, the entire model, its weights, its architecture, maybe even all the data it was trained on, and I know those are very different things, are publicly known and available and then can be scrutinized and evaluated.

26:47Daniel Ho:And then there are others who say, no, no, no, that is closely held intellectual property, the weights, the architecture, and the data. And this is, as you know very well, very common for lots of the most famous LLMs where we don't have full transparency. And I'm wondering as a lawyer who's using these all the time, and at some point a judge might ask a lawyer on what is your claim based and can I examine it or can the opposing counsel examine it? So does this open, closed thing, does it get on your mind and what's your attitude or thoughts about it? Yeah, I think it is probably less salient within the legal system as it currently stands, in part because these systems are being integrated into legal search provider systems that actually have a history of being fairly closed.

27:44If you look at systems like Westlaw or Lexis, they have just not historically made much available in terms of how search even operates.

27:53Daniel Ho:And these are the research tools used by practicing lawyers every day to find the Precedential cases and regulations. Exactly. And it's quite rare, at least in those instances, for judges to want to know more about that underlying formal technology. That could, of course, change. There have been these incredible efforts, including by one of our own Stanford law lums, Peter Henderson, to track all instances of legal hallucinations in cases. We have over 140 of documented instances where there are filings with legal hallucinations. And those are, of course, the instances that are really going to raise the eyebrows, potentially just judicial sanctions and the like.

28:33But the more general question you're asking about open versus closed, I think, is such a central one for the innovation ecosystem. In the last few years under the Biden administration, there was a really pronounced focus and concern that open models could really raise catastrophic risk. There's a particular concern in the Biden executive order on AI that really centered around bio risk. And there the concern is, well, if a model is capable of actually leading to the proliferation of bioweapons, maybe we should actually have an approach that favors more closed approaches. In a collaboration that involved quite a number of folks, but co-led by Percy Liang and Arvind Narayan, and here Arvind Narayan is at Princeton, we really sort of took a look at the sort of evidence base for open versus closed.

29:36And really, we think in a kind of white paper that resulted from several convenings, we really tried to state it as what we called a kind of marginal risk framework. What we need to know is what is the marginal risk of open models versus existing technology or closed models. And the evidence for that, at least at that time in terms of the bio risk, was not at all clear. There was a great study done by RAND where they actually had teams that were randomized to have access either to the internet and a large language model and to create a bio-attack plan. And then those plans were in a blinded way scored by experts.

30:28And what that RAN study found was, one, there was no statistically significant difference in the bio-attack plan between those with access to LLMs versus not. And number two, the kinds of information that LLM systems were able to provide was not distinguishable from widely available information that you get from the internet. And so we really have to be careful and have careful assessments of marginal risk. What's also tricky here is that just as you know, openness is a spectrum. So Irene Suleiman has a really good paper on the gradient of openness. And so what that means is you can have a Lama model on one end, which some would say is not even fully open because they're...

31:21Daniel Ho:And for background, Lama is the model that Meta Facebook puts out for free. Yeah. But there is a kind of license that people do have to sign that has restrictions. And then you can have something that doesn't provide the model weights, but does allow you to do forms of fine tuning on a kind of platform. And one study by Peter Henderson found that actually all of those safety sort of alignment efforts that come along with a more closed approach can actually very easily be stripped with just a few cents of fine tuning. And so, you know, as you're traveling that gradient, even something that doesn't have the weights exposed can still potentially actually present a fair bit of risk.

32:06And so I think what we really need in this ecosystem is much more analysis of both marginal risk and marginal benefit of open versus closed.

32:14Daniel Ho:Great. Great. Thank you. So in the final minute, I just want to ask you a basic question. People are thinking about their legal issues. And like Google, I'm sure people are hoping for a general purpose legal chatbot. Maybe they might even pay money for it. What is the future of a general purpose legal chatbot that could give people quality legal advice? Yeah, well, I think President Carter at one point said, we are over-lawyered and underrepresented. So many people are caught up in the civil justice system with insufficient legal representation. And that's, of course, the huge potential for this technology to democratize access to legal knowledge.

33:01I think there is a lot of potential there, but one of the misconceptions is that a general purpose chatbot will be an end-to-end solution to any and all legal problems. In some of the earlier studies I mentioned, these general purpose chatbots have a preventative to hallucinate at alarmingly high rates. And what we show in that study, too, is they hallucinate more frequently in exactly the types of cases that are likely to have underrepresented litigants at trial court, things that don't involve sort of appellate litigation. there are instances where these systems are really prone to error is if you ask a question based on mistaken premises, they exhibit a quality that we in that paper called model sycophancy.

33:59They're likely to just reify what the user said. And oftentimes, underrepresented litigants don't yet know what the right question is to necessarily ask. And so that is a particular risk. And I think in my view, one of the real paths forward is to focus on specific use cases that are more tailored to the kinds of problems that users need solved. Because in those specific settings, it's also much easier for us to evaluate both the benefits and the risk profiles of these kinds of systems.

34:39Daniel Ho:So the chatbots may still come, but they will probably be high. Tell me if I'm right. They'll probably be highly specialized and very narrow in their application domain. But in exchange for that narrowness, you might have less hallucination and more of a confidence that it's giving you good advice. Yes, that's right. Thanks to Dan Ho. That was the future of AI and the law. Thank you for listening. Don't forget, we have 250 or more back episodes of The Future of Everything, and you can listen to discussions on a wide variety of topics at the touch of a button. Please remember to hit follow in the app that you're listening to right now.

35:16Daniel Ho:That'll ensure that you're always updated about new episodes and you never miss the future of anything. You can connect with me on many social media sites at Russ B. Altman or at R.B. Altman on LinkedIn, Threads, Blue Sky, and Mastodon. You can also follow Stanford School of Engineering at Stanford School of Engineering or at Stanford ENG.

35:43Daniel Ho:If you'd like to ask a question about this episode or a previous episode, please email us a written question or a voice memo question. We might feature it in a future episode. You can send it to thefutureofeverything at stanford.edu, all one word, the future of everything. No spaces, no underscores, no dashes. the future of everything at stanford.edu. Thanks again for tuning in. We hope you're enjoying the podcast.

From the publisher

Law professor Daniel Ho says that the law is ripe for AI innovation, but a lot is at stake. Naive application of AI can lead to rampant hallucinations in over 80 percent of legal queries, so much research remains to be done in the field. Ho tells how California counties recently used AI to find and redact racist property covenants from their laws—a task predicted to take years, reduced to days. AI can be quite good at removing “regulatory sludge,” Ho tells host Russ Altman in teasing the expanding promise of AI in the law in this episode of Stanford Engineering’s The Future of Everything podcast

Have a question for Russ? Send it our way in writing or via voice memo, and it might be featured on an upcoming episode. Please introduce yourself, let us know where you're listening from, and share your question. You can send questions to thefutureofeverything@stanford.edu.

Episode Reference Links:

Connect With Us:

Chapters:

(00:00:00) Introduction

Russ Altman introduces Dan Ho, a professor of law and computer science at Stanford University.

(00:03:36) Journey into Law and AI

Dan shares his early interest in institutions and social reform.

(00:04:52) Misconceptions About Law

Common misunderstandings about the focus of legal work.

(00:06:44) Using LLMs for Legal Advice

The current capabilities and limits of LLMs in legal settings.

(00:09:09) Identifying Legislation with AI

Building a model to identify and redact racial covenants in deeds.

(00:13:09) OCR and Multimodal Models

Improving outdated OCR systems using multimodal AI.

(00:14:08) STARA: AI for Statute Search

A tool to scan laws for outdated or excessive requirements.

(00:16:18) AI and Redundant Reports

Using STARA to find obsolete legislatively mandated reports

(00:20:10) Verifying AI Accuracy

Comparing STARA results with federal data to ensure reliability.

(00:22:10) Outdated or Wasteful Regulations

Examples of bureaucratic redundancies that hinder legal process.

(00:23:38) Consolidating Reports with AI

How different bureaucrats deal with outdated legislative reports.

(00:26:14) Open vs. Closed AI Models

The risks, benefits, and transparency in legal AI tools.

(00:32:14) Replacing Lawyers with Legal Chatbot

Why general-purpose legal chatbots aren't ready to replace lawyers.

(00:34:58) Conclusion

Connect With Us:

Episode Transcripts >>> The Future of Everything Website

Connect with Russ >>> Threads / Bluesky / Mastodon

Connect with School of Engineering >>>Twitter/X / Instagram / LinkedIn / Facebook


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

More from The Future of Everything

All 67 episodes
The future of AI and the lawThe Future of Everything · 36 min
Listen in VO