How Thomson Reuters Built AI Agents That Think Like Lawyers

3 Sep 2025 · 59 min · 20 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Thomson Reuters’ Deep Research AI agents for legal strategy and research (not just document search), including how they differ from generic “deep research” web tools, how they reduce hallucinations, and how lawyers should use them with human oversight.

Guest backgrounds

Joel Hron (CTO, Thomson Reuters). He leads the Deep Research effort and discusses building with Thomson Reuters’ legal content/tools and thousands of domain experts.

Key claims

  • Legal “deep research” needs domain-tuned reasoning plus access to legal tools/content; generic web-oriented deep research is different.
  • Deep Research uses Westlaw-style features (e.g., Keysight flags for conflicts/overrules) as “breadcrumbs” for the agent.
  • Quality is high but not error-free; it augments attorneys, not replaces them.
  • Trust requires transparency (deep links, flags, reasoning trajectory) and validation loops/tools to mitigate hallucinations.

Notable examples

  • Customer anecdotes: surfaced new arguments/new ways to view the law.
  • “BS detector” / litigation document analyzer: checks briefs against case law to validate statements.
  • Testing with 1,200 customers; built over ~6–9 months with iterative refinement and expert-in-the-loop evaluation.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

What is Deep Research?

0:21 to 1:34

Discussion about Thomson Reuters Deep Research and its capabilities.

“I'm Corey Knowles, and with me as always is Grant Harvey.”

AI Reasoning in Legal Context

1:34 to 7:38

Exploration of how AI tools are designed for legal research and the complexities involved.

“So perhaps you could walk us through, like, how this is different and how it kind of works.”

AI Reasoning in Legal Context

7:44 to 8:05

Exploration of how AI tools are designed for legal research and the complexities involved.

“That's fiddler.ai slash A-G-E-N-T-I-C dash O-B-S-E-R-V-A-B-I-L-I-T-Y.”

Impact of AI on Legal Research

8:05 to 14:00

Discussion on how AI tools enhance legal research efficiency and quality.

“You know, this is something Grant and I have discussed a number of times regarding the need for like specialty AI tools within various fields and, and that, that Westlaw is a part of, uh, Thompson Reuters.”

The Quality of AI in Legal Research

14:00 to 15:00

Learn about the impressive advancements and quality standards of AI in legal research.

“also transparent to the user in terms of the answers that are being generated and also the trajectories that the model is following along the way.”

AI's Role in Augmenting Legal Professionals

15:00 to 17:00

Explore how AI tools enhance lawyers' capabilities without replacing them.

“And I don't mean that like blowing smoke or anything.”

User Experience and Feedback Mechanisms

17:00 to 21:00

Understand the evolving user experience and feedback processes in AI legal tools.

“And so more of the time is being spent making hard decisions, hard technology choices.”

Explainability and Transparency in AI

21:00 to 23:00

Delve into the importance of explainability and transparency in AI-assisted legal processes.

“patterns exist already in the UIs, but I think, you know, we're certainly eager to like continue to evolve more of them as we go forward.”

Mitigating Hallucination in AI Responses

23:00 to 28:00

Learn how AI providers address the hallucination problem in legal contexts.

“I think the other thing that we've done in terms of like explainability overall is really try to like include the context of the source of any conclusion like within the application itself.”

Understanding AI Training for Lawyers

28:00 to 29:10

Learn about the importance of training and change management for lawyers using AI.

“Here's how you can read like what the agent is doing and here's where you should spend your time validating or invalidating certain parts of this answer.”
Show all 20 chapters

Model Selection and AI Performance

29:10 to 30:58

Explore how different AI models excel in various tasks and their implications for legal work.

“there was one of the biggest advancements in GPT-5 that no one was talking about is that this way lower hallucination rate.”

The Promise of Open Source AI

30:58 to 32:59

Discuss the potential of open source AI models in enhancing legal tools and workflows.

“We acquired a company about a year ago called SafeSign that was really focused around legal small language model development.”

Building Enterprise-Ready AI Systems

32:59 to 35:05

Discover the challenges and strategies for developing AI systems suitable for legal professionals.

“And I think companies that kind of have some interoperability there are going to be able to build the best systems at the end of the day.”

Navigating MVP in AI Development

35:05 to 40:45

Understand the importance of focusing on the whole problem when developing minimum viable products.

“our domain experts and our technical teams.”

Agents in AI: Use Cases and Future Potential

40:45 to 42:00

Examine the role of AI agents in applications like deep research and speculate on future breakthroughs.

“Like that's when you maybe dive deeper and spend more time focusing on an individual component.”

Harnessing Agency in AI Systems

42:00 to 45:56

Explore how agency in AI enhances legal applications and workflows.

“And you're just scanning the universe to like figure it out.”

Navigating Rapid AI Advancements

45:56 to 50:04

Discuss the challenges of keeping pace with fast-evolving AI capabilities.

“Yeah, this is kind of what I was getting at when I said, like, solve the whole problem.”

Using AI Tools in Daily Work

50:04 to 54:42

Learn about personal AI tools and their impact on productivity.

“That's where like a lot of the legwork comes in, you know.”

The Importance of Adaptability in AI

54:42 to 57:20

Understand why adaptability is crucial for success in AI development.

“I don't think you design something that's future-proof.”

Deep Research at Thomson Reuters

57:20 to 58:22

Explore Thomson Reuters' innovative tools and where to find more information.

“I think it's a really fascinating tool that's going to get a lot of use for an awful lot of years moving forward.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Joel Hron:Lawyers, what if you had a legal research tool that also helped you strategize? Thomson Reuters Deep Research does exactly that. And today, we talk to CTO Joel Hron about it.

0:20Joel Hron:Welcome, humans, to the Neuron AI Explained. I'm Corey Knowles, and with me as always is Grant Harvey. Hello, hello. Today, we're going to talk to Joel Rahn, CTO at Thomson Reuters, about deep research, an AI agent platform built for legal strategy, not just document search. We'll unpack how it works, what testing with 1 ,200 customers taught them, the hallucination risks in legal settings, and what this means for lawyers and knowledge workers moving forward. Joel, welcome to the podcast. Thank you guys for having me. Nice to see you, Corey and Grant. Nice to see you too. We're sure excited to have you on here.

1:01Joel Hron:We know there's been, you know, a lot of talk in the legal community and around AI and how people are using it and some of the benefits they're finding and some of the limitations they hit along the way. So we thought this was just a really good conversation to have, and we're excited to have you here chatting with us. Yeah, excited to talk about it. It's been a lot of great work from the teams over the last nine months or so to get here. So it's an exciting topic to talk about. Yeah, so let's dive into it. So Thomson Reuters just launched Deep Research. Can you walk us through what, like, for example, I think a lot of our readers are familiar with, like, ChatGPT Deep Research.

1:40So perhaps you could walk us through, like, how this is different and how it kind of works. Like, for example, how is this different than, say, just putting all your docs into Chachupiti deep research? Why would they want to choose this instead? Yeah, that's a good question. So there's a number of implementations of deep research today, as you mentioned, like Chachupiti, Perplexity, Claw, Gemini have a variant of this, and most of them really orient around research via the web, right? So they're tremendous at kind of navigating all the deep links of search and sort of reading and learning and taking notes along the trajectory of learning and doing more searches.

2:22And as you think about, you know, let's say like a B2C use case of like planning a vacation, it's a very iterative thing. Like, you know, I'm going over the summer and like, what are the options? I could go here. Like, what's cheaper? Like, what's more family friendly? Like, there's a ton of just like trajectories of like search that you would do for that problem, which is why deep research is really phenomenal. The same is true in the legal context, though, and in legal research context as well, in terms of the various different trajectories of research that could be taken. And it's actually even more complex in many ways, because many aspects of the law, like, you know, say the same thing, but in a different way, or they might like overrule each other or the jurisdictional nuance, you know, might be particular to the context and where you are and sort of where the case is being tried and all this stuff and what judge is in front of you and all these things that kind of, I think, formulate from a human's perspective, like how they would go do that research.

3:31Totally. And so the sort of aspect of of how the agent reasons through that process needs to be tuned for the domain of law. And, you know, when I'm not a lawyer myself, but, you know, having worked with many for the last several years, like, you know, you go to law school to learn, like, what the patterns of good research look like and what sort of the, you know, way to formulate an argument is. And there's precedent for how to formulate good arguments or how to, like, argue against arguments. And so there is a structure to how good label research should be done. And then the second part, so that's all on the like AI reasoning side of the equation.

4:17The other side of the equation is like, okay, what tools and content and information does this agent have available to it to go do that research? And as Tom's and Royers, this is sort of one of the things that we've been exceptional at for many, many years. We have probably the world's largest and most respected repository of content for this type of thing. But we haven't just sort of like put it on top of a bunch of raw content. There are aspects of our legacy software tools like Westlaw, for instance, that really are tools that human lawyers have used for decades to do good legal research. An example of that will be something like Keysight.

4:58Keysight is like a feature of Westlaw that basically annotates whether a new case, you know, conflicts or overrules a prior case along a particular topic. And so this is a tool that a user in Westlaw would use to say, oh, I'm looking at this case. Like, I think it supports my argument, but maybe it's no longer good law because this new case came out. I should go follow that Keysight flag and look at that new case, perhaps. And we have trained our agent to use Keysight in the same way, for instance, that a human lawyer would use it. And so the agent now is able to see these breadcrumbs and follow these breadcrumb trails to do better and more comprehensive research than it would have done otherwise.

5:47And I think that combination of both injecting the AI with good sort of legal grounding and how this research is to be done and also arming it with world-class tools that are really kind of tuned for how the agent needs to interact with them is a key ingredient of building, you know, good agentic systems and in particular building good deep research systems. And so that's really, you know, what some of our main focus has been for the last, you know, nine months or so.

6:18Joel Hron:So your shiny new autonomous agents are trading stocks, booking flights, and chatting with customers at 2 a.m. Then suddenly, one of them goes off script. Traditional APM tools throw up their hands because agentic workflows aren't just requests and logs. They're reasoning, coordination, and a maze of dependencies. That's why I'm telling you today about Fiddler AI's agentic observability. Fiddler gives you a top-down and bottoms-up command center built specifically for multi-agent systems, you get complete and total visibility. Application session, agent, trace, all the way down to individual spans, so you can see exactly why an agent chose option B over option A.

7:04Joel Hron:Root cause analysis is baked in, slashing time to detect and fix issues before they ripple out across your whole tech stack. And it aggregates insights at every layer of the genetic hierarchy. And it aggregates insights at every layer of the genetic hierarchy, from apps to sessions, traces and spans, enabling you to just pinpoint the one outlier that matters. The best part? Fiddler AI plugs right into the frameworks you're already using. Open telemetry, lane graph, Amazon Bedrock, so onboarding is painless. If you're serious about reliable production-ready AI, head over to fiddler.ai slash agentic dash observability and request a demo today.

7:48Joel Hron:That's fiddler.ai slash A-G-E-N-T-I-C dash O-B-S-E-R-V-A-B-I-L-I-T-Y. So keep your agents autonomous and accountable with Fiddler AI. You know, this is something Grant and I have discussed a number of times regarding the need for like specialty AI tools within various fields and, and that, that Westlaw is a part of, uh, Thompson Reuters. It really, you know, I mean, every lawyer uses that, you know, just about it's, uh, it's, it's the tool. And, uh, I think it's really fascinating that, that you guys have dove in like this, you know, quick and early and, uh, saw an opportunity and to need. You know, I feel like this might be a thing that makes attorneys stronger than they are without it.

8:46Yeah. I mean, we've certainly heard that. I mean, like you mentioned, you know, the thousand plus customers we've been testing this with. And I think the first thing we've heard, which is maybe the obvious thing, is like, oh, you know, this took something that would have been a 10-hour research project into 10 minutes. And that's all well and good. But I think, you know, the other anecdotes that we've heard from some of the customers is that this surfaced new arguments or new ways of looking at the law that, you know, hadn't been surfaced before. And I think that's like one of the most profound and compelling opportunities here.

9:24That's interesting. Do you have like, do you have an example of that, of something that you know off the top of your head? Totally cool if you don't, but that's really fascinating. Yeah, even if I did have an example, I probably would butcher the legal interpretation of what it does. Fair enough. Fair enough.

9:41Joel Hron:That's okay. That's okay. But yeah, we have some good anecdotes from customers where they've been able to use this to, again, help think through what are good arguments for this case and how can I create new arguments that maybe I hadn't thought of before. That's amazing. Yeah, that's really interesting. That's, you know, we've talked before about, you know, the idea of that if you put the tool in the hands of a person who knows the skill, it's more like a superpower than it is like making something worse. Like there's often this idea that what it's going to do is going to be worse. And it's like not not if you put it in the hands of somebody who knows the right way to ask the right questions, who knows.

10:28Joel Hron:You know, specifically how to use it in this case, I think I think what you guys have done is really smart and will have a big benefit for a lot of years to come. And it really speaks to also, I think, the process of building it with experts in the loop. Like, you know, one of the things early on that we, and that's certainly something we have, we have, I think, 4 ,500 plus domain experts that like sit with our teams doing this work and helping them, you know, score answers, you know, do preference judgments, identify good rubrics for correctness or helpfulness or these other things about the answers.

11:10And one of the things early on, you sort of craft these gold sets of questions that you want to be able to have the system answer. And one of the obvious things is like as a rubric is to say, okay, an answer to this question should reference these cases as source material. And you could design a system that like evaluates, okay, did the answer generate references to these cases? And is that good or bad? What we found is that the system, if you give it enough agency, was able to get to the right answer along a number of different paths or trajectories. Maybe it would cite to different cases or different sources of secondary law or primary law or otherwise, but it would ultimately get to the right answer just in different ways, even if you maybe ran it multiple times.

12:02And so it was an interesting observation. And I think, you know, really speaks to the need to have like real human experts in the loop evaluating these systems because because it is not like a black or white, like binary type answer. There's there's a lot of subjectivity in the practice of law, a lot of like domain interpretation in the practice of law that is important to build into the system kind of from the ground up. Yeah, I think about that in terms of so, you know, like as a writer, right, like Corey and I can look at writing that they'll say Chachibiti gives us and we can assess whether or not it's good or bad.

12:42And we can assess, you know, like subjectively, of course, but like we kind of know, OK, this is not what I intended with this or I don't think this will go well with this audience. You know, this doesn't matter.

12:52Joel Hron:This is better than what I expected to write myself. sure yeah or or oh wow this was not what i was expecting and therefore i'm surprised and impressed and i get all the dopamine hits and i'm like i'm running with this um um and and so it's really fascinating so like the basically what you're saying is that the the amount of experts that you had on the team were able to kind of instill that that expertise into the agent so that the agent could then make those kind of decisions on its own more or less right that's certainly our intent and you know certainly not an easy thing to do in all cases because you really do want i think to strike a good balance here between uh giving the models the agency to go out and explore these paths um uh because that's really like i think what they're good at and something like like i mentioned the breadcrumb trail things like you want to give the models agency to follow of these paths that may not look obvious at the outset, but might be, you know, elucidating of something new or different.

13:54And so you want to balance that like agency with some structure as well, like some structure about like how good legal process should be done and how to make that also transparent to the user in terms of the answers that are being generated and also the trajectories that the model is following along the way. And so, yeah, we spend a lot of time finding what that right balance really is between domain orientation and agency on the model side.

14:25Joel Hron:Legal research, you know, and we touched on this a minute ago, traditionally takes hours. And there's this idea that this can bring that down to minutes. Do you feel like the quality that you're getting out of what you see is up to a standard that you're, I mean, I assume it is, or you wouldn't be releasing it. But have you been impressed with the quality of one it's been able to find so far? Have there been any surprises in that way? I would say I've been super impressed, to be honest with you. And I don't mean that like blowing smoke or anything. It really is super impressive. Uh, I think the, the level of quality that the systems are able to deliver.

15:13Um, I think, um, I think, I think overall, like when we, when it's not a case where we don't see any errors in any answers though. And so this question off comes up is like, okay, like, do I just hit this button and send it to, send it to the courthouse? Like I'm done. And I don't think that's reality. I don't think it will be reality for quite a long time. Yeah. I think the models are super impressive in what they're able to do. But, you know, they really are just, I think, augmentations to really good attorneys at the end of the day rather than like full on replacements for them. And I think, you know, I like to draw this analogy to software engineering in a lot of these conversations that I have in that, you know, if you look at AI tools and AI agents in the engineering space, like they are and they have like completely transformed how you look at like building software from the ground up.

16:18Right. And the the sort of pace of development of those tools has outpaced other industries by about 18 months, probably just because, you know, the models are better at writing code. They have been better at writing code. It's a it's a testable, you know, a thing in terms of like, does it work? Does it not? Does it compile? Does it not? and so and so I think I think engineers have kind of outpaced the rest of the market and you look if you look at what has happened in terms of like engineering workforce with AI dev tools it has amplified the hardest parts of being an engineer and in many ways it's made like the best engineers even better than they were before so if they were 10x people before they're 100x people now.

17:06And that's because, you know, the hardest parts of the job, the ones that require the most like human judgment and good decision making are the things that are being done more often, like all the boilerplate stuff and the common function writing and, you know, bug triage stuff is done much more easily now. And so more of the time is being spent making hard decisions, hard technology choices. Like I got in the fork, a fork in the road in terms of this architectural choice, like which one's better. And that's like the very human part, I think, that, you know, is even more important than it was before.

17:44And I think the same is true for legal or will be true for legal as these systems advance in that the system's not like just giving you an answer to go like hit send on or in an engineering context, hit deploy on, but they're sort of inviting you in to make the hard judgments at the right moments. And those are like the most human and the most challenging parts of the job. So I think what you'll see is that really the best attorneys are the ones who know how to use these tools the best will actually become 100x of where they were versus, you know, just kind of leveling up everybody at the same time, if that makes sense.

18:22Yeah, we have a minor UX question. So let's say, for example, I'm using this tool and I go out and I have it create, you know, I'm not a lawyer, so help me out with the example here if this doesn't work perfectly, like a brief on the case that I'm working on. And I get back a report from agent and I'm one of those like 10X lawyers, let's say, and I can kind of tell, okay, this is good, this is good, but this part here doesn't, how is the iteration process with the agent? Like, can you go in, check its work and then see where it went off and give it notes? or how does the actual feedback mechanism work there?

18:58Yeah, it's a good question. And I think it's a pattern that's still evolving, quite honestly. I think Andre Karpathy gave this talk at Y Combinator a few weeks back where he talked about the generation verification loop of AI in humans and really designing user experiences experiences that spin that flywheel of generation and verification as fast as you can. And that's really how you get tremendous speed up and value at the end of the day. And so I think, frankly, some of these design patterns exist in our products. Some of them are still being evolved at the moment. What I will say, like the current UX of research, for instance, you know, shows the user, like the trajectory of steps that the agent is taking, what searches it's doing, what it learned, why it's doing this next search, or why it's going to this content repository versus this one.

19:59And it sort of spells that out. And it does that in real time, but it also does that at the end. And you can look at like, these are the notes that the agent has written for itself over the course of doing this research. And here's how it got to this answer. um what i would love to see like in the future is like users being able to inject themselves either after the fact or even in real time to say oh no you know i've seen that case before like you know i think you should check this one out instead and like i think these little redirect nudges and things like that add context to the agent as it's going through its process and are good ways for, you know, a really, a really like 10X, you know, lawyer, for instance, to really, you know, apply what, what they know and their experience and maybe their broader context of the case and the client and the judge and all this stuff, uh, to help the model get to the right answer or the most right answer that it possibly can.

20:57And so I think, um, again, some of these patterns exist already in the UIs, but I think, you know, we're certainly eager to like continue to evolve more of them as we go forward.

21:07Joel Hron:You know, that's really interesting. And to me, you kind of touch on the explainability of it all, which I think is going to be a really big discussion in the AI space and every space that touches it moving forward. You know, when we're dealing with subjects like law, like healthcare, human resources, banking, finance, For example, there is this need to be able to to both go back and look to see how it arrived at a conclusion as well as to intervene at a step or at least go back and have the ability to re-trigger it from this step with a different decision. Maybe that sends it sends it back forward if it's not in real time.

21:54Joel Hron:But it sounds like, you know, that was that was a goal you all were baking in early. Oh, 100%. And we hear that loud and clear from our customers. Like, you know, if we just spit out an answer, I don't care how right it is. Like, there is very little trust of it, as there should be, I think. I think trust is earned over the course of time. And, you know, we need to do everything we can to earn that. But I think in the time between now and then, like explainability and transparency are super important. We actually just had a product review this morning with our tax team. And we were going through the same topic about what the right levels of transparency and human agency is necessary in the process.

22:38You know, and it is a balance between giving the user enough of that, but also not making it so they're having to validate every step that the agent takes because that's not really helpful either. Kind of doing its purpose. Yeah. Yeah. So trying to find the right balance of what's enough there is a learning experience as we sort of work with customers and clients on this type of thing. I think the other thing that we've done in terms of like explainability overall is really try to like include the context of the source of any conclusion like within the application itself. So, for instance, in like a deep research answer, like there's deep links to resolve to like the actual cases that supported this answer.

23:36There's flags. I mentioned Keysight earlier. There's like the same UI flags that exist in our legacy application exist in this answer to give the user like these visual clues about, you know, where to go next and how to kind of launch from here. And so I think those have been tremendously important. We've also built, as tools for the agent, ways for it to check its own work. So if it makes a statement, there's a tool for the agent to use to go validate, is this statement supported by law? and they can go in and like check like okay yes this is supported by law i feel good about it and so these like recursive uh checks if you will for the agents are an important way to help mitigate uh some of that risk again you i don't think you'll ever like remove it but you know we try to do everything we can to help mitigate it yeah that's really smart to do that um you know right at the end make sure like okay the ultimate conclusion sure it might be sound logic in the context of what you're doing, but when you look at it back in the context of the law, is there actually legal precedent for this?

24:47Exactly. Yeah. We've all heard the horror stories of lawyers citing fake cases generated by AI. I think they get headlines. I haven't seen headlines of them lately. I think lawyers are kind of buttoning themselves up, but we've definitely seen that. They'll cite a fake case. And I think there was like one a week for a little while when we were covering this. So, I mean, the hallucination factor, right, is really huge here. I guess just like in the legal realm in general and specifically what you did with the research, how are you thinking about hallucination problem solving when the stakes are as high as, legal ramifications for failure.

25:32Yeah. I mean, I guess I would start with saying I don't believe that humans are immune to these mistakes either. I'm sure many of these things happen even predating generative AI and this happened, but by all means I don't think that obviates the need as a software and technology provider to do 110 % of our best effort to mitigate as much as we can. And, you know, the way that we look at that is really probably three dimensions. The first is around like the user interface elements that I mentioned to you before. Like, how do we convey some level of like confidence about an answer to a user? How do we convey, you know, these visual clues that yes, this case does resolve to a real case because it's got a bright blue deep link to it?

26:24How do we put these flags in place in the answers to allow users to navigate through the trajectory easier? How do we give them UI clues of how the agent resolved and what paths of reasoning that it took to get to this answer? So that's one element. The second is things like building actual tools to do this validation in real time for the agent itself. You know, we have an application internally that we also released with deep research. In the market, we call it litigation document analyzer, which is a mouthful. But internally, we called it the BS detector, which is basically like a way for somebody to upload a brief and analyze the brief against case law and say, like, is somebody extending the law?

27:15Are they like, you know, using sound logic here in terms of what the law actually says in making this statement. And so if you think about that, that's actually a tool for a research agent to use. Like, hey, am I making up BS right now? And it can use these kind of things along the way to help, you know, control itself from a hallucination standpoint. So we actually build, I think, tools and logic into the process. And then the third thing is really just training, like I said. I guess, you know, if you do that one time as a lawyer, you remember like, okay, I need to check this next time, right?

27:52And so, you know, we hope that that doesn't happen to any of our users and we try to train them in a way to indicate to them, here's how you should use these visual clues. Here's how you can read like what the agent is doing and here's where you should spend your time validating or invalidating certain parts of this answer. So I think this this element of training and change management for attorneys. And, you know, to the extent we can, we, we try to lean in and really help them with that process. But I think it's really all three of those things in combination. Yeah.

28:26Joel Hron:Understanding the, you know, the whole need for a human in the loop with AI and the reason that is so vital. You know, it's one of those things that's a struggle to teach people, but they learn it real quick if they learn it the hard way. You You really, you only once have to accidentally embarrass yourself because you rushed, because you overtrust. And, you know, I think what you're talking about, about continuing education with attorneys and making sure they understand this, understand the tool and what it can do. I think that's very valuable. You know, something just last week was we saw this that I think could have implications for you all moving forward, especially, is that, you know, there was one of the biggest advancements in GPT-5 that no one was talking about is that this way lower hallucination rate.

Read the full transcript

29:19Joel Hron:And which leads me to believe we'll start seeing that follow in other tools as, you know, research comes about. Is that something you keep an eye on in like, you know, model selection you're working with or is it fully proprietary inside? If you don't mind me asking. Yeah, no, 100 percent. I would say like, you know, we have a pretty agnostic model philosophy internally, and we're constantly benchmarking the latest models as they come out. Smart. And, you know, the rubrics that we have for these different use cases, hallucinations are one part of that rubric. I mean, obviously, they have much more, you know, I think, breadth to them than just hallucination.

30:06But like hallucination is certainly one of the most critical aspects of it. And I think what we find is that like each of these models are, you know, really good at different things. Like GPT-5, really good at a lot of reasoning stuff. Like the anthropic models are exceptional at like tool use and these sorts of activities, code writing. the Gemini Google models are really good at longer context type tasks where you're dealing with huge documents and things like that. And it's hard to break up the context. And so we really find that like, you know, testing these models on a regular basis, having really good internal benchmarks that we lean on to give us good signal as to where these models might be stronger or weaker helps us refine our focus of where we put these models.

30:59We acquired a company about a year ago called SafeSign that was really focused around legal small language model development. And that's certainly an area of research for us as well. As you think about, again, just the volume of raw content we have, but also the number of experts we have in this field who, again, know this legal process, no legal reasoning. I think the ability to train some of these like finer aspects into smaller models, whether those be tools or actual components and planning process has real potential as well. And so that's a big area of research for us at the moment, in addition to the third party models we use.

31:43Like there's a debate, you know, between open source and closed source. Like how do you think about that in your benchmarks, like personally for you and also for the work that you do? Are you worried about open source models and using them in production or how do you feel about it? I'm actually quite excited about it, to be honest with you. I think in my shoes, like as a company with terabytes of proprietary content and thousands of legal experts, I think we have all of the ingredients to really excel with open source models. And so as that tide rises, I think it creates a lot of opportunity for us, creates a lot of opportunity for us to deliver models to our customers in different ways that help them manage the security or privacy concerns they may have in a better way.

32:37So I think we're really optimistic about the future state of open source. I also think sort of the positive momentum of open source drives the commercial models better. They constantly have a higher bar to leap over, which makes them better at the same time. And so I think the future, particularly as you think about agents, there's a very likely ecosystem of models that develops under any agentic system where you perhaps have some larger planning and orchestration models and you have smaller models that are doing more, perhaps like discrete or focused to things along the flow. And I think companies that kind of have some interoperability there are going to be able to build the best systems at the end of the day.

33:28Makes sense. Yeah.

33:29Joel Hron:You know, you're working with, you know, frontier AI models in a highly regulated industry. What do you feel like it takes to make cutting edge AI enterprise ready for legal professionals, like actually in the workplace. I know you've talked about, you know, bringing in experts, but I'm also kind of curious about, you know, what you can share of maybe the technical end a little bit in terms of like, you know, how big a team you're dealing with and how long it took to build this out, stuff like that. Yeah. I mean, I think, you know, So professional grade AI is challenging to build. And I think I talked about Thompson Reuters before.

34:12Like, I think this is one of our biggest challenges. We're 150 plus year old company. And I think if you look over that history, our brand is associated with trustworthiness and truthfulness and like these principles. And I think, you know, as as young startup companies and there's tons of them now and they're doing phenomenal work. but they have a lot of freedom to go out and experiment and explore and like try something if it doesn't work they pivot and it's like from an engineering perspective that's great that feels awesome but as a company with as much history as we have I think anytime we put something in the market it's definitely held to a bar I think above and beyond anybody else just because again of the history behind us and we take that to heart in everything that we do and develop, which is, again, why I think we put so much emphasis and effort on this relationship with our domain experts and our technical teams.

35:14I think, you know, you mentioned like from a technical perspective, getting a little deeper. I think a couple of things that are like really good ingredients like for success in terms of of building one would be um uh developing benchmarks that technical teams can focus on and so that they can spin their flywheel somewhat independently at the early days if everything you do has to go to a human evaluator every turn of the crank it really slows down the development pace. So we try to like kind of phase automated testing or automated evals with human evals so that we can take a lot of swings at bat in an automated way.

36:01And then as we get closer and closer to production, bring in more and more human eval before we ultimately ship. And I think if you can build really good kind of automated frameworks or automated evals at the early stages, is that it helps teams move more quickly in the early days, particularly. I think the other thing that we try to prioritize is perhaps like keeping teams small at the beginning and then growing them sequentially over time. So, you know, we've been working on deep research in earnest for probably the last six months. and I would say that team started quite small, like, you know, less than seven people and, you know, probably every sprint grew by a couple, right?

36:52As we started to sort of like bring in more context and bring in more parts of the application ecosystem and stuff like this. So we try to grow that over time, I think, a little bit. I guess the other thing is from a technical perspective, Like there's a lot of new frameworks that come out to have this like agentic scaffold and this thing. There's a real balance in like diving into one of those too early. Like the space is really evolving quite quickly. And, you know, being too dependent on one of these frameworks in such a fast evolving space can sometimes be a handicap, even if it helps you get started a little faster.

37:35It can.

37:35Joel Hron:They disappear as fast as they show up sometimes. They do. They do. And you get to this point of development where you're like, oh, man, this is actually like more of a handicap than it is an accelerant. And so, you know, we try to be judicious about our choices there. Again, this is where I think really talented, good engineers come in place. Like even though Claude Code can go write 10 ,000 lines of code, like you can't make a decision like that really well and have an intuition for that really well. And so I think this is where really talented engineers come into play and are exceptionally important in the development process.

38:16And then the last thing I'll say is like starting with the whole problem. So like there's this idea in software development of like the MVP, right? And I think oftentimes teams, they try to minimize the problem, like minimum viable product. Like, OK, I want to make the problem as small as I can to be viable. And so this is the hardest part. This is the most valuable part of the focus on it. And I think if you over index on being too minimal, you actually miss the system. And like really the way the AI system and the agentic system operate isn't apparent until you build the whole system. Right.

39:00And so I think I think really focusing on the whole problem as much as we can on day one, rather than trying to like slice it and solve this component too early, has been a shift for us from from like a product engineering perspective. that has allowed us to really think about these things as a system rather than as an individual discrete component.

39:24Joel Hron:I was going to say a thing about, and MVP is such an interesting subject, because it kind of gets contorted a little bit, and I really feel like a lot of people don't spend enough time on the viable. Quite often you see the minimum. What's the minimum we can shove out the door? And then things have a tendency to stall. And I think it's important that if you're going to subscribe to the MVP idea that you have to understand, you know, it's not MVP and move on. It's MVP and finish building it in public. Yeah, exactly. Exactly. Yeah, I do. Yeah, I do. I think we sometimes just human nature index on like the minimum versus the viable word in that acronym or in that pseudonym.

40:14So I think, you know, not to say you don't want to ship fast and like you don't want to try to like build simply and that kind of stuff. I think those are all good core like agile principles that teams should continue to follow. But I do think, you know, at the early days, trying to take the blinders off and really like scope a big, hard problem and and and see, see what like build all the components, see what the agent is capable of. And then when you see, OK, this is where it feels like it's fallen over. Like that's when you maybe dive deeper and spend more time focusing on an individual component.

40:50But like try to put all the pieces together first and then and then dive in. yeah that's awesome that is changing gears slightly because we just mentioned agents again um so deep research was definitely like i would say the first killer app uh for agents like someone asked me like what's the you know what is the the best killer app for agents like i'm still not figuring out like what they're good for and i was like well do you use deep research he was like yeah use it every day i was like well then it's deep research deep research is the killer app but what do you think, just from your own personal opinion, what the next breakthrough agent use case will be?

41:28Do you have any gut feeling about that? Any intuition? We can just wax about it. What's your thoughts? Whether it's possible or not today. Yeah, it's a good question. And I agree for what it's worth. I think deep research is one of the most profound examples of agents in an application today. I think there's a reason for that, though. And that is because that is the highest agency use case you could possibly imagine. Like, you know, a user's entering this situation with no preconception or maybe some small preconception as to like what the right answer is. And you're just scanning the universe to like figure it out.

42:09Right. And so the agent has full agency to to go out and like use its tools and use its content and use like Maybe it's system knowledge to figure that out. And so you're not like constraining the agent in any way. So I think that's why it's such a profound use case. You know, you asked the question about like hallucination risk and how do you build confidence in lawyers and things like that early. I actually think that the agency dial is one of those ways as well. Like there are other parts of our application, for instance, where we do like, you know, more workflow centric stuff like drafting policies or drafting purchase and sale agreements and things like this, where maybe you don't want so much agency.

42:54Like, you know, my company has a way that they do this. They have like standard templates or standard forms. And I don't want to just go draft some document from scratch every time. And so being able to turn this agency dial, I think, is actually not like it's not a cool use case. I think it's a necessity of different types of work that these systems need to do. Bring you some efficiency too. Yeah, absolutely. I think beyond deep research, which is a really cool use case, I'm really excited by some of the stuff we're doing in tax as well. It's like quite different. But I think it also speaks to some of the things like the models are good at.

43:41So a year ago, everybody was talking about LLMs and like they were like, oh, it's not good at math. Like it'll never do math things and stuff like this. And now all of a sudden, like it knows how to use calculators and it knows how to like write its own code and like this kind of stuff. And score well on the International Math Olympiad. Yeah, silver medals, gold medals, whatever it is. Right. So like, guess what? I don't need to train the model how to like do two plus seven. Like I just needed to know how to write a Python script that does X equals two and Y equals seven, like add them together.

44:15So I think that paradigm change is like really profound. And if you look at like what we're doing in tax, it really brings together this concept of tools and code, I think, in a really interesting way where you think about a tax workflow. Somebody starts with a bunch of like W-2s and 1099s and things like that. And I need to use AI and LLMs in particular in this case to analyze these documents, like resolve context between them. Like I've got this invoice, I've got this tax form. Like do these numbers like match with each other? Am I missing information? Then it involves like aspects of being able to take those numbers and look at the tax law in some way and say, OK, like, how do I treat these numbers?

45:01Like, should I deduct this or should I not deduct this? And then how, once I've done that, how do I then map that to a tax engine that like calculates actually the tax and puts it into a filing form? And that's like an end-to-end problem that involves like a ton of different tools and research and data extraction sort of things all in one. And we've got a beta out there on Ready to Review that does this, at least in some of the simpler use cases. But I'm really excited about those forms of agents because they do touch on so many different like capabilities of the model to be able to do math things, call tools, do research.

45:44at all in one system. And so I think if I were to extend deep research, that would be sort of where it's going next is like this broader context of operating. Wow.

45:58Joel Hron:You know, there's something I've heard mentioned by a few people along the way is a tendency to kind of over engineer these products more than they almost have to be sometimes, because more often than not, If there's the cadence with which new models, new features, new tools are dropping, I keep hearing that it's really easy to get caught in, you know, a loop trying to build something out that's actually going to be solved for you in 30 days or 60 days, most likely. Yeah. Yeah, this is kind of what I was getting at when I said, like, solve the whole problem. Take the tax example I just gave you, for instance.

46:39Like, you know, I might decide, let's say a year ago, if I was working on this problem to go spend a bunch of time on data extraction because like, hey, if I don't extract good data from these documents, then like none of the downstream stuff works. So I'm going to spend a bunch of time on that. And in reality, if you look at where we are today, like LLMs are phenomenally good at this now. Yeah. And they've gotten phenomenally better in just in just nine months or so at this particular task. Right. And so, you know, that would have been a case where you like spend a bunch of time on something.

47:14But at the end of the day, maybe that wasn't the most valuable use of your time, which is the point of why I think like built the whole system, like put put the agent system to the test and have it do all the steps. extracting data, looking at the tax law, inputting it into a tax engine, resolving errors, let it do everything, and then figure out where it falls over. I think I've heard some of the, you know, the large research labs and model providers kind of say, hey, look, when we design our applications, we're designing for what we think the model might be able to do in six months rather than what it's capable of today.

47:54And I think there's an element of that in how you approach building as well, where you try to kind of forecast where these models are going and what they might be good at. And if you forecast that today, honestly, I don't think the models are going to get really good at math, for instance. But what they're going to get much better at is writing code and planning and adapting and these kinds of things. And so you really need to think about, okay, well, how do I need to architect my system to make those components as flexible and profound as I want, but then go spend my time perhaps on the tools or the content underneath them, so that as this boat rises, these boats rise too.

48:34Joel Hron:Yeah, because if you think that if we're going to start a project today, we're going to plan it for two weeks, then we're going to go into an early development phase, we're going to do this, and before you know it, we've spent four months on a project that now could have been done in a half hour with a couple simple agentic workflows. And I think that whole idea of building forward a bit is going to be really important to, you know, and you may get there early. You know, that may be a thing sometimes. You may get there early. Yeah, it's not ready for you. You're ready before the agents are or the models are.

49:10Joel Hron:Yeah, but I think it'll keep a lot of people from wasting time and energy. I don't want to minimize the work to get something out the door, particularly in like professional use cases like this. You asked how long sort of deep research took. I told you like six months. You know, our first spike in the first month was like, yeah, this is it. It was very obvious. You're like, this is the path. And we know exactly what we need to do. Not to say there weren't learnings and nuances along the way. But from month one to month six, it was all about refinement of like these edge cases, like these little like feedback loops that need to be incorporated, new tools that need to be incorporated.

49:54Again, like trying to not be so minimum, like solving a bigger problem and making sure that it's robust against a broader array of use cases. That's where like a lot of the legwork comes in, you know. but really when you kind of see what these models are capable of, a lot of times it's pretty apparent just in the first couple sprints that you're onto something. You just got to keep running that. Well, I got to ask, just because I'm curious, beyond Thomson Reuters products, what is your personal AI stack? What tools are you using day-to-day to get work done? Are you using any of these coding AI tools?

50:36What are you using? I like Claude Code. I think Claude Code has been phenomenal for things that aren't even writing code at times. You know, I find it's been like a really like flexible tool for me to be able to use. We also have licenses to copilot as well. You know, internally, we use Microsoft a lot for a lot of things. I think Microsoft has really done a good job of like integrating it into the graph. I think context is in these models and the ability for Copilot to kind of reach into teams and Outlook and SharePoint and all these different systems that like a lot of my context is in already is really a profound shift.

51:27And I think some of the recent stuff they've done with like Copilot Researcher and stuff like that has been really interesting for those kinds of use cases. But just like personal kind of like things like Cloud Code has been been like pretty amazing. Yeah, I just I talked to someone who is not even an engineer the other day and I asked him, oh, what AI tools are you using? And they're like, yeah, I use Cloud Code with like my personal assistant. And I was like, you use COD code? So it inspired this whole rabbit hole of people who are using it as a general purpose agent. And we wrote a whole thing about it.

52:03It's awesome. If anybody, we need to get it on the pod. I'd love to come back and talk to you about that one day. I mean, yeah, if you can get past the fear of looking at a CLI as a non-engineer, then it's really, really cool. That's so awesome. Using it as like a chief of staff. Like I've seen people say that. Yeah. Yeah, that prompt, you know.

52:25Joel Hron:One thing I've been working on is I've been looking at this. I've got this N8N workflow I'm looking at right now that's a template. But it's pretty much a genuine personal assistant. Like it checks your Slack and has the various operations there and in your email and in your calendar and in your project management solution. And I was like, man, it would sure be cool to finally have all of that brought into one simple feed. Yeah. As opposed to a bunch of different places to go all day, every day, you know. You know, it's really hard. Even like I mentioned Microsoft and like we touch a lot of Microsoft systems today just internally.

53:06But even that's not comprehensive. Like just a lot of things that I do day to day that are outside that ecosystem. And so it's not particularly easy. Like I said, context is king for these models, but actually context is really hard. You carry a lot of context around in your brain that you probably don't even realize that you're carrying around. And it's hard to catalog the list of all of those things to give to an AI model because it's a longer list than you might expect if you really think about it.

53:38Joel Hron:And more than you even know, most likely. Yeah. There's so much. And, you know, it's been really interesting. You know, one of the things we see in looking even at just context windows is that, like, depending on the tasks and who you are, like, you know, we know people who are working in gigantic spreadsheets. And a context window boost is absolute gold to them. So is speed. where with what I do a lot of times on the more creative end, I care less about speed and more about getting kind of specifically what I want. And I think we're looking at a future where everybody uses the model that works for them, kind of, or the tool that works for them, you know, and that may be three different ones.

54:25Joel Hron:It may be five at some point. I don't know yet. Joel, for professionals listening who want to level up their research game with AI, whether legal or not, what is one mindset shift or approach you'd maybe recommend? end well i'll tell you like when i think about hiring and talent and building out with my team there's like one skill that or maybe like quality i would say i don't even know if it's a skill that i have indexed on more than anything and that would be adaptability like you talk about the pace of technology changing and you know designing systems that are future-proof i think that's impossible.

55:08I don't think you design something that's future-proof. I think you learn how to adapt really quickly. And I think people that have, whether it's natural or learned, like a mindset or a capability for adaptability, and maybe there's an element of humility there as well, which is to say, like, you know, I always need to be curious about learning the next thing. I think that's one of the most important, I think, skills for people in this day and age is, is like not being too fixated on kind of where you were, where you've been or what you've done before, but being really adaptable and open to kind of try and exploring new things.

55:50So that would probably be the one I'd focus on more than anything myself. That's great. It's like basically making yourself future-proof. Like you are the future-proof tool, not whatever you built. So as long as you're adaptable, you'll be future-proof. So before this role, I used to lead our AI research and development lab. And I used to tell like a lot of scientists have this sort of identity of like being a scientist. I'm like, I do science and like I don't do engineering. And I can appreciate that. Like I think there's a lot of value in that. But I tried particularly in our field, which is more of applied research than it is like academic research.

56:30I've tried to instill into my teams like you're an engineer I don't care front end back end you know QA science not science like you're an engineer and like engineers like engineer new ways of doing things like that's that's your job that's what you're trained to do I don't care if you're a mechanical engineer a software engineer a chemical engineer like you're taught the same foundations of like how to logic and intuit and you should use that skill and it will future proofread to anything, quite honestly. I'm a firm believer in that.

57:03Joel Hron:I agree. I agree. I think we're moving into an age where the smartest thing you can do is learn to get excited about change. Learn to find change exciting and be able to roll with the punches and adapt well. Joel, thank you so much for joining us today. This has been excellent. It's been a really fun deep dive into a very specific industry's use case and to help us better understand where you go, better understand where you all at Thomson Reuters are finding these wins and losses and what you've gone through in developing deep research. I think it's a really fascinating tool that's going to get a lot of use for an awful lot of years moving forward.

57:47Excited about it. Thank you guys for having me today. It's been a great conversation.

57:51Joel Hron:Excellent. Where can readers go to learn more about what you all are working on? Maybe take deep research for a test drive. Yeah, I would say, you know, most of the stuff we do, I try to put out on LinkedIn so you can see the latest. We also have a Medium blog that you can follow that talks a little bit more about hopefully some of the technical stuff that we do. I also try to post links there. So I would say check it out on LinkedIn and you'll be able to see some of the videos and links to how you go out and try it out. Excellent. We'll drop a link down below in the video description as well.

58:26Joel Hron:Thanks again to Joel, to Thompson Reuters, and to all of you for listening. I'm Corey Knowles here with Grant Harvey as always. And today, farewell for now, humans. We'll see you next time.

58:50FR difference!

From the publisher

Thomson Reuters just launched Deep Research—an AI system that doesn't just search legal databases, but plans and strategizes like an experienced attorney. In this episode, we explore how one of the world's largest legal research companies is using AI agents to transform how lawyers work, the challenges of building AI for high-stakes legal decisions, and what this means for the future of knowledge work. CTO Joel Hron shares insights from testing with 1,200+ customers, tackling hallucination risks in legal settings, and building professional-grade AI systems.


Resources mentioned: Thomson Reuters Deep Research: https://www.prnewswire.com/news-releases/thomson-reuters-launches-cocounsel-legal-transforming-legal-work-with-agentic-ai-and-deep-research-302521761.html 


Westlaw & KeyCite: https://legal.thomsonreuters.com/en/products/westlaw/keycite 


Claude Code for development: https://www.anthropic.com/claude-code 


LinkedIn: Joel HronThomson Reuters Medium blog: https://medium.com/tr-labs-ml-engineering-blog 


Subscribe to The Neuron newsletter: https://theneuron.ai

More from The Neuron: AI Explained

All 106 episodes
How Thomson Reuters Built AI Agents That Think Like LawyersThe Neuron: AI Explained · 59 min
Listen in VO