#290 Joel Hron: How Thomson Reuters is Approaching The Next Era of AI

29 Sep 2025 · 1 h

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Eye On A.I. Episode #290 - Joel Hron: How Thomson Reuters is Approaching The Next Era of AI

Host: Craig S. Smith Guest: Joel Hron, Chief Technology Officer at Thomson Reuters Sponsor: AGNTCY (Agency)

Episode Overview In this episode, Craig Smith interviews Joel Hron to discuss Thomson Reuters' innovative approach to agentic AI systems, highlighting the transition from traditional prompt-based AI to more sophisticated workflows capable of planning, reasoning, and executing complex tasks. Hron explains the importance of maintaining human oversight in high-stakes domains like law and tax to ensure trust and accuracy.

---

Key Topics Discussed

  1. Transition from Prompt-based AI to Agentic Systems
  2. Definition of Agentic Systems: Hron explains the spectrum of agency in AI systems, where higher agency allows for autonomy in planning and executing tasks.
  3. Agency Dials:
  4. Autonomy: Levels of decision-making capability.
  5. Tools: What tools the AI can use to perform tasks.
  6. Memory: Ability to remember past interactions to inform future actions.
  1. Infrastructure and Architecture for Multi-Agent Collaboration
  2. Importance of Infrastructure: Emphasizes that effective AI requires robust infrastructure beyond just models, focusing on how agents communicate and maintain memory.
  3. Future of Coding: The evolving role of engineers in AI, where adaptability and understanding of non-deterministic AI systems become crucial.
  1. Human Oversight and User Experience Design
  2. The necessity of human verification in high-stakes environments is crucial for building trust in AI systems.
  3. User experience design plays a significant role in facilitating interaction between AI systems and users, ensuring accurate communication of AI outputs.
  1. Thomson Reuters' AI Strategy
  2. Acquisition of AI Capabilities: Discusses the strategic acquisitions Thomson Reuters has made to bolster its AI talent and capabilities.
  3. Leveraging Domain Experts: The integration of human expertise from legal, tax, and compliance fields into the AI development process to ensure systems are grounded in accurate information.
  1. Multi-Agent Systems and Future Predictions
  2. Discusses the potential for future organizations to operate with multiple agents collaborating on tasks, likening it to how APIs allow applications to interact.
  3. The importance of context and memory in multi-agent systems and addressing the challenge of shared memory among different agents.

---

Key Takeaways

  • Agentic AI: Represents a significant evolution in AI capabilities, allowing for more complex interactions and autonomous decision-making.
  • Human Oversight: Essential for maintaining accuracy and trust in AI outputs, especially in critical fields like law and finance.
  • Adaptability in Engineering: Future engineers must be adaptable and comfortable with rapidly evolving AI technologies and methodologies.
  • Role of Infrastructure: Effective AI systems require a thoughtful arrangement of infrastructure and architecture to support multi-agent collaboration.

---

Quotes

  • “Agency really means the level of autonomy that you give the system to sort of plan certain actions.” - Joel Hron
  • “The goal is to get to 100%. But like I said, I don't think you'll ever get there.” - Joel Hron on accuracy in AI systems.

---

Conclusion This episode of Eye On A.I. offers valuable insights into how a historical company like Thomson Reuters is reinventing itself in the age of AI. By focusing on agentic systems and maintaining a balance between automation and human oversight, the conversation highlights the broader implications of AI in the professional world.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The ecosystem of tools and capabilities is pretty nascent. And it also from even nine months ago has changed quite substantially to what it is today. Like what you thought you wanted out of an agent, you know, nine months ago is very different than what you want today. We're trying to automate a tax return preparation. Input documents would be like W-2s and 1099s and these kind of things. Output would be like a tax return that's ready to file with the IRS. And, you know, the tools that would need to be used through that process would include tools like, you know, OCR. it would include tools like some document extraction, like entity extraction, APIs.

0:36In our case, what we've been able to do, we have a tax engine as one of our core products in the tax space. And so we've been able to expose that tax engine as a tool for the agent to go use and run tax calculations and validate tax calculations and inspect errors. And so in many cases, those tools that we're building are built on top of our existing products and applications. Build the future of multi-agent software with Agency. That's A-G-N-T-C-Y. Now an open source Linux foundation project, Agency is building the Internet of Agents, a collaborative layer where AI agents can discover, connect, and work across any framework.

1:23All the pieces engineers need to deploy multi-agent systems now belong to everyone who builds on agency, including robust identity and access management that ensures every agent is authenticated and trusted before interacting. Agency also provides open, standardized tools for agent discovery, seamless protocols for agent-to-agent communication, and modular components for scalable workflows. Collaborate with developers from Cisco, Dell Technologies, Google Cloud, Oracle, Red Hat, and more than 75 other supporting companies to build next-generation AI infrastructure together. Agency is dropping code, specs and services, no strings attached.

2:18Visit agency.org to contribute. That's A-G-N-T-C-Y dot O-R-G. Nice to meet you. Appreciate you having me on today, Craig. So my name is Joel Ron. I'm the Chief Technology Officer at Thomson Reuters today. I joined Thomson Reuters about three years ago via acquisition. I was the chief technology officer of an AI company that Thomson Reuters acquired. And shortly thereafter, I led our AI research and development lab, which we call TR Labs. It's a group of about 250 research scientists and engineers really focused on more, I would say, applied AI research. um and uh probably close to a year ago now took over as cto and and as thompson reuters uh we are we're really born as a content company more than anything but i would say um you know over the years we've been really invested in not just like delivering content but also selling the experiences of how professionals interact with that content.

3:34And a lot of that comes down to search and retrieval and these sorts of like technical disciplines as well. And we do that across a number of areas, like the legal industry certainly is one of our biggest, but also within tax compliance, risk, maybe most well-known for the news and media side of our business as well. But we really solve that problem across a wide spectrum of industry today. Yeah. And what was the startup that was purchased? Yeah. So we were a document understanding company. We really were focused on the legal sector in terms of understanding and interpreting legal documents. And that's what Thomson Reuters had acquired us for.

4:20In particular, one of the things that we had done, and this sort of predates generative AI, which has sort of like superseded a lot of those supervised AI models that were being built back in that time. but we had built an MLOps platform and an annotation platform that was really tailored to lawyers in particular to, you know, allow them to develop ontologies of certain types of contracts and curate data associated with those contracts that would be required for training those models and evaluating them and deploying them. And we built a really good proprietary operational platform for that type of thing.

5:02and thompson reuters um acquired us to really scale that into our practical law editorial teams to take you know the practical law editors who have experience across a wide spectrum of legal practice areas and have them do that kind of work uh across different document types uh in different practice areas yeah yeah can you can you say the name of the company i talked to a lot of companies in that space and the supervised learning yeah it was called thought trace at the time and uh we changed the name when we were required to document understanding uh as a product name but prior to acquisition it was called thought trace right well congratulations jumping to cto i'm sure uh that's a pretty pretty exciting thing for a company yeah it's it's been a big change over last few years for sure yeah so we're going to talk about what agentic ai really looks like in production and one of the things i'm interested in is what separates true agentic systems from prompt based assistance how ai is planning reasoning and executing multi-step workflows goes autonomously and what it takes to build that responsibly.

6:24That's a mouthful, but if you could just talk about that, I'm sure listeners would be interested. Well, yeah, I'm sure. And a lot of threads that we can pull on throughout that question, it's pretty broad. I mean, I'd say maybe the best place to start would be defining like what an agentic system is. And I think when I think about agency, I think about it on a spectrum. It's not a binary, it is either agentic or not, but it's really a spectrum of agency. And what agency really means in this case is the level of autonomy that you give the system to sort of plan certain actions, the tools that you give that AI system to then take an action, and the ability to coordinate memory across either one sequence of actions or maybe across a full history of a user's interaction with the system.

7:29So those are kind of three, if you can think about it, agency dials that you can move up or down to give an overall AI system more or less agency. So if you start on the very low end of agency, you know, this might be like a traditional, like you said, prompt-based RAG system. Like it's got no agency to write its own prompt. It's got no agency to select what data sources it uses to retrieve information. Like it literally just takes an input and processes a discrete step or sequence of steps even to go execute it. Whereas if you move to the high agency, these systems are more goal-oriented rather than a discrete number of steps.

8:18They have a goal in mind and they have the ability to plan a sequence of steps that they think will help them achieve the goal. They have the ability to go execute tools along that trajectory of steps. And then they have the ability to react through the middle of that. As they see they've uncovered more information, it might re-inform the plan and maybe they re-plan additional steps. And so that would be like a higher level of agency. And I think depending on the capability that you're trying to build into the application, you can move that agency dial up or down in some cases like particularly in certain aspects of law like people appreciate high precision repeatability consistency and so you might want to dial that down to be like less agentic if you will but in other areas like legal research for instance this is a this is a discovery problem it's often very open-ended and you want to follow the breadcrumb trails.

9:23And so you may want to dial up agency in that case to, to give the model the freedom to go out and sort of learn and explore, uh, multiple paths. Yeah. Although that, that dialing up and down, that's done, uh, at the creation of the agent, it's, uh, or are there agents today that you can, uh, you know, there that you can set parameters, uh, after the agent's built yeah so it's a good question i think in general yes it's done at the time of creation or how you design or architect the agent i think there are some good um corollaries of how this can be done in the user experience itself though like if you look at an application like a coding application like cursor for instance uh i can constrain the agency of of cursor by constraining the context over which it works so for instance i can highlight just a line of code that i want it to explain to me or i can highlight a whole function or maybe a file or i can give it access to the whole repo and tell it to just like go to town and so that would be a way to on the fly control the agency dial of context, and thereby you give the AI system more or less freedom to go out and create.

10:50And that's really a function, you know, in the coding example, is a function of the use case and the user, and are they working on a mature code base that has hundreds of files and, you know, tens of thousands of lines of code, or are they building something new and just trying to vibe their way through it like that would determine whether they want to go high or low agency in that in that case and uh and i think i think those systems and those user experiences kind of lend themselves to allow the user to adapt it to what they need at that point in time yeah uh and these are llm based agents we're talking about uh and they're making tool calls and giving instructions to various tools that then go out into the world and execute and then, you know, come back with information that the LLM either returns to the user or does passes on to another agent which at thompson reuters do you guys have a preference for which models you're using to build agents on yeah it's a good question so um they are llm based in general and in particular like the the orchestrator like sort of the the master agent that's sort of planning and reacting along the way is one of these, what I would say, frontier models in most cases.

12:32We evaluate that pretty consistently. You know, we're close partners with OpenAI, Anthropic, Google, in a number of areas, and we use all of those models and more in certain situations. the tools themselves that the agent system is using may or may not be large language model based. In some cases, they could be a web search. In some cases, it could be another LLM, maybe a smaller LLM that's fine-tuned to a particular task. It could be some other type of model entirely. So the tools themselves can be a variety of things. Generally, at the orchestration and planning layer, there's kind of like a master model that's bringing it all together.

13:19And again, depending on the use case that we're working with, we've found that we've had good success working across all three model providers and they each tend to be good at particular things. And we pretty continuously evaluate that as well. And do you build your own agent building platform or do you use uh uh you know commercial uh uh solutions i just i've been getting i've been having call after call with uh with companies that that have agent building platforms and it's very difficult as a journalist they all kind of sound the same to me but yeah yeah uh but but everyone i talk to says they have the answer uh and i'm curious a company like thompson leuters it's been a bit tricky because you know this development pattern or design pattern for ai has really only existed for nine months you know maybe 12 months at best but So I would say that the ecosystem of tools and capabilities is pretty nascent, to be honest with you.

14:41And it also, from even nine months ago, has changed quite substantially to what it is today. Like what you thought you wanted out of an agent, you know, nine months ago is very different than what you want today. So we have used, you know, particularly like MCP, for instance, is something that we leverage. like i think there's a lot of good industry gravity around uh around that um other sort of frameworks uh like lang chain has their own uh crew.ai etc have their own i think they're all good at certain things we found that like building some of those components ourselves has actually made things uh a little bit easier for us that you know 12 months from now when the space is a bit more mature and some of these design patterns are a bit more stable, maybe we would go to a third party.

15:32But right now, we've kind of built a lot of this ourselves for the moment with, like I said, the NCP and other examples where we do leverage industry standard. From a tool building standpoint, one of the interesting things that we've found as we've done this development is that like the, you know, the same way you would design a UI for a human user to use it, or you would design an API for like a computer to computer interaction, like designing tools is very much a thought process of how is an agent going to interact with this? And what is the sort of modalities of information that it is going to provide?

16:21And how does information need to be presented back to it in terms of a tool output so that the reasoning model can sort of act or react on it you know most appropriately so there's a whole another like level of kind of uh design consideration for how you build tools that allow agents to execute and exercise them uh more efficiently and more effectively and again that's a very nascent space as well that we're still learning every day on. I see. So you're building your own tools. You're not sending it out to call. Yeah. In some cases, you know, like if you were going to go do a web search, for instance, there's a lot of good third parties that provide like web search tools that will sort of leverage out of the box.

17:10In some cases, like a tool could be accessing our content systems. And so in that case, we will build a tool that leverages our search APIs in the appropriate way to search and retrieve content from our systems. In other cases, it may be like taking action in the application itself. So for instance, a good tax example would be if we're trying to automate a tax return preparation, like input documents would be like w-2s and 1099s and these kind of things output would be like a tax return that's ready to file with the irs and you know the tools that would need to be used through that process uh would include tools like you know ocr it would include tools like some document extraction like entity extraction apis in our case what we've been able to do we We have a tax engine as one of our core products in the tax space.

18:10And so we've been able to expose that tax engine as a tool for the agent to go use and run tax calculations and validate tax calculations and inspect errors. And so in many cases, those tools that we're building are built on top of our existing products and applications that we've been maintaining for many years now. Yeah. You know, Thomson Reuters is 170 years old or something.

18:44You know, how did companies move from content or software to delivering intelligent systems? I mean, that to me is, it sounds like an enormous task because it's not simply building the systems. It's getting the enterprise to change behavior. And this is something these guys that are building platforms for agentic workflows, I've been talking to them about. If I were young and had some smart friends, I would pick an industry and

19:30use one of these platforms to field, you know, agent after agent. And you could basically challenge the behemoths with a company of agents, couldn't you? Yeah. And I mean, look, I mean, competition in the space today is fierce and it's never been easier to kind of get started, as you say. uh like developer tools and things like that you know make the the time to ship something that's meaningful uh you know orders of magnitude shorter than what it used to be i think i think for us the content and not just content but the the human experts that generate that content we see as a real sustainable and durable advantage for us.

20:26Like it's one thing to just build a superficial agent that goes out and drafts an NDA. But, you know, quite honestly, I could just go use ChatGPT or Claude or Gemini out of the box to do that quite well. And it's seen a bunch of the internet and is pretty good at that. But I think what really separates us in that regard is the fact that we are not just using those models to generate this, but we're using those models with the context of what practical law, for instance, says about drafting an NDA in certain jurisdictions, in certain industries. And we've had lawyers for many, many decades now who literally draft know-how, draft sample documents, draft how-to articles on how to do this well, what pitfalls to avoid, et cetera.

21:19And that content goes into grounding and steering the agents to doing that task better. And then the second thing I mentioned, which is not just the content, but the people who have drafted that content for many years, they are critical in the alignment and evaluation of our AI systems as we build them and as we operate them. and their ability to provide feedback and steer and continuous critique of what's working well and what's not working well, make sure that we get the best performance that we possibly can out of these systems and that we're continuing to keep them up to date with current law and current industry standards and practices.

22:05So I think it's very easy to build a very simple AI system today. I think it's very hard to actually ground it in, you know, a true and verified information and also to provide the, the transparency to end users that you're doing that and they can pick and trust those outputs in that way. Yeah. But on the other question of how do you, how do you get the organization to adopt agentic workflows or is that easier than i imagine i wouldn't say it's um i wouldn't say it's easier than you imagine i mean it's a 170 year old organization as you said uh years of of of legacy uh behind us that certainly we're very proud of and for good reason i think what we've done though is one, demonstrated real success early on, even for those who have been in this world for a long time and may be slow to change.

23:16When they see this happen, it's almost undeniable that this is a better path. So part of it has just been really leaning in and showing success early. And I think that brings a lot of gravity. The other thing we've done is we've been quite acquisitive over the last couple of years, particularly in early stage companies like Case Text, like Materia, who we've acquired, Safe Sign. And I think bringing new talent into the organization who has kind of a fresh perspective on this problem and partnering them with people who've done it for a long time creates a really good outcome at the end of the day.

23:55And I think we've done that quite well and quite aggressively and certainly will continue to do it. And then thirdly, it's just been organically upskilling and integrating talent into the organization who has experience in these kind of skills. I think that's been something that we have been super focused on for the last two years. I mentioned I came from the TR Labs. Like that organization in the last two years has more than doubled in size. We were about 120 people two years ago. We're 250 plus now and continuing to grow. And so I think our talent composition of people who have experience building and architecting AI applications has grown as a percentage of our overall engineering workforce.

24:45And as you get more and more of this talent into the organization, I think it becomes more and more part of your DNA. And that's certainly something that has happened over the last two years and I think will continue to happen over the next several. Yeah. And then there's, I mean, you're talking about building the products, right? The part of the organizations that's building these agentic products for use by other organizations. But within Thomson Reuters, I mean, it's a very large organization. Is there an eager uptake of agents or is that not really your responsibility to build agentic products or workflows for Thomson Reuters more broadly to use?

25:39it is and it isn't so so my team also owns our internal ai platform which we call open arena and and i would say that has been one of the like we we built and shipped that within like two months of chat gpt you know coming on to the to the market you know two and a half years ago or so so we were very early in terms of building our own platform and driving experimentation into the company. And I would say we have huge adoption. So I think our weekly active users on OpenArena today are more than 16 ,000. It's probably more than 60 % of the company who uses that on a weekly basis. And so we've seen really strong adoption of just AI in terms of day-to-day usage.

26:28And Open is not the only tool we're very eager and and leaned in towards um you know using third-party tools in sales or in marketing or uh um in other functions of the business where there's best of breed kind of domain tools but we also have this platform that has just general usage uh that it enables um and i think we've seen a lot of uptick as you think about agents though that's really like what I would say the next level of productivity from an internal optimization standpoint. And I think we are still on the path of figuring out how to do that best. So today, if you think about where AI was a year ago, it was all about AI augmenting or being helpful to a human.

27:19You would go in there and you would like say, hey, help me rewrite this email. And it would like rewrite it for you and it was helpful well with agents maybe now you're saying go write this email and send it you know or maybe you're not even telling it to do it it just knows that it should go write that email and send it and so there's there's this level of sort of autonomy and automation that you're shifting to ai and and that requires i think more technical depth and certainly more attention from a process standpoint about where you want to do that. And again, what level of agency you want to give the system to do that.

27:56And those are areas that we're building around today. I can tell you, like for my teams on the software development side, we're not just sitting there letting AI write all of our code, but there are certain parts of the coding process maybe around test development or accessibility or other like uh aspects of the development process where we can achieve higher levels of automation and we're trying to drive that forward yeah and and in building products or or uh agentic ai for um production uh what skills You were saying that you're acquiring new skills in the workforce. What skills are you looking for?

28:43Because that's another question I get from a lot of young software engineers. Am I obsolete? What skills should I be developing to keep myself relevant? Yeah. I know a lot of people are concerned about that and what it'll have impacts to the workforce. And I'm sure there will be a workforce transition. One of the things I see is that AI amplifies the best engineers orders of magnitude more than it does the rest of the population. So for great engineers, it makes them like 10x greater than they were before, right? And I think, you know, and that doesn't mean that they are a scientist. It doesn't mean that they've trained and deployed AI models for their whole career.

29:33They could be a standard back-end software engineer. So I don't think there's a specific necessarily AI skill that's necessary to be really good at this. Certainly, you need some of those. You need some core engineers. But for all of those people, I think the most important skill is actually adaptability. the pace at which technology is changing right now we talked about like agentic patterns and frameworks like there's a new one every week how do you know whether you should use it or not or whether this is better like you need to be very adaptable in this environment you need to be comfortable like picking up something new and learning it and trying it and failing it and moving on and that level of adaptability to just pick up anything and get a very quick sense for it and be able to apply it if it makes sense uh is a really important skill i think the second skill that's super important in engineers uh is understanding the lack of determinism in ai systems i think traditionally for engineers they've thought about like here are my requirements like when i click this button like this screen better show up and it better look like this whereas ai systems are not deterministic in this way and and so building for non-determinism requires you to think about like, well, how do I eval something?

30:59And how do I like build validation into the application? Or how do I build rubrics that support the validation process that needs to happen that are stochastic and not deterministic? And I think that's also a skill that some people have. And certainly I think any engineer could develop, but is an important thing for engineers to develop, to get comfortable in building AI systems like this. Yeah. So you think coding is still an important thing for kids to study? Because in order to become a software engineer, you need to understand the structure of software and to understand the structure of software, you really need to be able to code.

31:47I think it is exceptionally helpful. And I don't like I said, I think the best engineers are even better than they were before with these tools. And I think that will continue to be true as we go forward. I think the the the population of people who know how to code because of these tools will also grow. So I think there's tons of people out there who didn't go get a computer science degree, but will go learn how to code in a way that they become great engineers themselves. Right. And and I think to do that well, you don't just sit there and like let the AI model run wild on everything. I think anybody that's done that, like written an application knows that like they can go off the rails pretty quickly.

32:36And and so understanding logically, like how do the applications get built? Like what does service architectures look like? Like how do I keep this thing on a leash as I walk it through this process? Like those are how I think really good engineers get a lot of efficiency out of these tools. So I don't think the engineering disciplines going anywhere. Like I said before, I think there are certain aspects of the discipline that lend themselves to really full or substantial automation in a lot of cases. But the core aspects of engineering, I think, will continue for a while. Yeah. And I've also been asked by software engineers that are not AI engineers, you know, what should they study to become an AI engineer?

33:31but like in the syngentic age that we're entering the ai uh is uh accessed through an api call i mean it's the building agentic system is really an engineering problem not an ai problem am i wrong on that no i think you're right it's an engineering problem what makes it an ai problem is this aspect of non-determinism like i mentioned like i think ml engineers historically would have been trained and well versed on like how to build uh good testing and validation sets and how to look at things statistically and say am i hill climbing this thing to get better or worse or on what dimension is it getting better or worse like they think uh statistically in that sense about the system.

34:23And I think that's what every engineer will need to learn how to do to be able to build AI systems well. I think today we call that AI engineering. It's probably the same analogy to 15 years ago when we thought cloud engineering was like a separate thing. Now it's just engineering. And I think probably the same will happen with AI engineering. Today is somewhat of a unique skill, But I think, you know, five, 10 years from now, it'll need to be ubiquitous for all engineers. Yeah. Why is infrastructure and not just models important going forward in AI? Yeah. Well, you know, I think I would separate that maybe for a company like ours or like the hyperscalers.

35:15you know infrastructure for the hyperscalers is like a fuel for them this is like if they're an electricity provider like a utility provider like this is their power plant like that that is what um right or wrong i think is necessary to create more advanced uh reasoning models more advanced versions of these models that keep coming out you know month on month and so i think that's where you see a lot of this infrastructure investment going is really around training infrastructure. But I think what you've seen more recently is this concept of like test time compute, where the inference in infrastructure actually gets quite large very quickly as you start to expand the cycles of inference that happen on any individual call.

36:07And I think this is another thing that's driving a lot of the big infrastructure spend that's going on within the industry. I think what you've seen over the course of the last several years is both the cost to run and the speed at which these models perform has just precipitously dropped, like orders of magnitude from two years ago. Who knows whether that, you know, curve will continue on, on that particular trend or not. But again, what, what, what that really allows you to do is not to necessarily deliver the same output for a lower cost. What it allows you to do is spend more cycles running the model so that you get a better answer.

Read the full transcript

36:54And that's really what this test time compute kind of concept speaks to is is being able to exercise the model over more cycles and get the better answers. And certainly as we look at areas like law or tax or otherwise, that's the most important thing for us. Like our users demand accuracy, they demand confidence, transparency, these kinds of things. And so anything that we can do to drive accuracy and confidence up in the answers that our AI systems are given is what we're aimed to do. Yeah, and in this case, when you're talking about infrastructure, you're not talking about data centers and chipsets and things like that.

37:41You're talking about architecture in effect, right? How you stack orchestration layers and proprietary tool chains and domain-specific logic. Is that right? Well, yeah, I'm sorry. So when I was talking about test time compute and the training side, I was talking about core chip infrastructure. But yeah, there's another side of infrastructure, I guess, which you would call architecture, which is how does the agent reason, how does it get access to tools? How does it maintain memory? How does it know when it's done? which is a hard problem, how does it verify information, like the citations it provides, for instance?

38:34How does it verify that if the AI generated this statement, where did that statement actually come from in its memory? So all these are hard problems that speak to the architecture of the agentic system itself in terms of how you build it, how you set up prompts, how you set up memory, these kind of things. And again, I think that that's a super nascent space right now. That's probably where we've spent the majority of our time over the last nine months building and designing these systems. Not on the core model itself. It's really on all of these kind of surrounding bits of infrastructure or architecture that really drive model performance.

39:16But if you're working with a probabilistic foundation model as sort of the core reasoning and logic engine, can you ever satisfy that demand for accuracy? or do you just keep narrowing the error? At this point, it's a great question. I think in an autoregressive system, you will never satisfy to 100 % accuracy. It's just I don't think that it's possible. And so the name of the game is to continue to narrow the error bars and to design good evaluation systems and to have good human in the loop interaction with those evaluation systems to make sure that you're doing this well. I think the other thing that's super important is this user experience design of how the user interacts with the agent itself.

40:33And Andre Karpathy gave a talk on this at Y Combinator. I thought it was really good. And he spoke about the speed of the loop of the generation and verification. So if you think about the agent generating something and the human verifying it, like the faster and tighter you can make that loop, the better the overall system is. And the user experience that you build in an application needs to facilitate a fast verification. And so, for instance, the way that we do that in legal research is, you know, our deep research system provides clear descriptions of what it's doing, what cases it's searching, what it found out of this case, what it's doing next and why.

41:20And a user can follow that if they want. And at the end of a research task, you know, we provide deep links into Westlaw. We provide key site flags as to the cases we're saying. So that experience itself is what drives confidence in whether the answer is correct or not, or what aspects of the answer might be wrong that they might want to do further investigation into. So the goal is to get to 100%. But like I said, I don't think you'll ever get there. And the job is to build as close to 100 % as you can, but to build user experiences that drive easy and fast verification by humans into the process.

42:01And, uh, and so I think both of those things together is what ultimately builds, builds user confidence in the system. And in your experience, law firms, for example, do they have paralegals going through and verifying the output, checking the references? there's this thing called automation bias you know that you know something works once it works twice it works 10 times pretty soon you're you're just taking it for granted that it works yeah yeah we certainly uh when we train users to use our application this is how we treat we we train them on how to verify how to how to look at something quickly and to know whether it is right or wrong.

42:55So for instance, I mentioned citations in our answers for deep research. So like a big thing that's been in the news is around like AI systems that hallucinate cases, like they make up court cases, right? So one of the very like easy things that we do in our UX is like we have the richest repository of case law in the world, right? If we get an answer and we're able to resolve it to a case, like we create a blue deep link to that case in Westlaw. You can very plainly see that this is a real case because it's tied to a real case that exists in Westlaw. If it's a case that doesn't have a blue link, like that would be a concerning thing for a lawyer and they should see that and they should do something about it, right?

43:43So there's very, it's a very simple like UX thing, for instance, that helps users to hone in on where they need to spend their time looking versus like maybe be more trustworthy of things. So these are the kind of ways that we try to train users in terms of using our system. I certainly hope that, you know, in all cases, users are applying some level of verification to the process. I don't think our intention is that people are just like sitting there clicking the button and send it to the judge at this point. Yeah. Yeah, that's interesting. So looking ahead, where are multi-agent, where is multi-agent collaboration going?

44:37I mean, well, first of all, just on that issue of, you know,

44:45providing as much accuracy as possible. I assume then you have sub LLMs that are doing what a paralegal would do. Look up the case law, verify that it's a real case. These are all tools, again, that the agent used. Like there's a case reader. There's like a key site investigator. There's like a verification. Like, so all of these, again, are tools that the agent is able to use to, you know, assess its answer and the correctness of it. Yeah. So on multi-agent collaboration, I mean, as I said at the beginning, people are talking about building organizations that are essentially agentic with, you know, thousands of agents.

45:38doing each step and then passing work back and forth among themselves and then structured memory and user feedback loops where do you see this going and and how do you see knowledge work of the future more broadly yeah again i think the analogy to to apis is a good one like in the same way i build apis so that different you know applications can connect to them connect to each other like i think uh we will build applications for agents to be able to to use and talk to each other i think the part that's somewhat unsolved to me at the moment particularly as you think about like multi-party agent systems so you know vendor a vendor b vendor c is this concept of memory and because at this point there's sort of a single orchestrator of a flow and generally that orchestrator owns the memory across and so you know as you think about that like context is really what drives the performance of these systems like as they run and as they build memory that is then re-injected into the context and the model becomes more aware of what's going on and what it needs to do differently and how to adapt and this kind of thing.

47:06And so memory becomes kind of the most critical, not the most, but one of the most critical parts of executing the flow correctly. And so who owns that memory? How does that memory get shared between these agents is something that I don't think is quite well solved yet. Again, like self-contained within Thomson Reuters and the agent we build, like we can solve that problem but but if you start like bringing in vendor a b c d into that equation i think that still needs to be figured out i think mcp is a good protocol in terms of like the communication pattern between them but memory is is another thing and certainly other vendors have tried to like throw their hat into that ring like a to a and things like that to help facilitate that but i think that pattern will evolve in the same way that like API standards evolved for how applications communicate with each other.

48:02I think you'll see similar patterns with agents. Again, I don't think that this verification loop is going anywhere anytime soon, though. I think this idea that you just have this autonomous systems of agents that is going out and doing all these things without any human ever seeing anything. is probably far off. If you look at self-driving, self-driving was really good a decade ago, but I was in the first Waymo in my life just a few months ago. So it's taken a decade or more for people to get really to the point of comfort of this. And I think the same will be for AI, that human verification will be an important part of the software experience for quite a while, I think, until we develop a level of comfort and consistency with this level of automation for some use cases.

49:00Again, for others where they're just maybe lower risk and there's less criticality, maybe you'll see those get automated fully more quickly. But for many things that really matter, I think this human verification loop will continue to be an important part of the software design experience. and and uh i mean that's interesting about context and memory because i just had a conversation with a guy talking about the problem with transformers uh that that the larger the context gets the slower the system gets and the more expensive it gets and um how do you guys handle um context and memory do you do you use like a scratch pad where you have everything that's happened in a kind of a knowledge base that then the orchestrating agent can refer to or is it being held in the context of the model yeah i would say this is also an area that is under development and kind of continuously changing as the models change.

50:20I think if you were to, I guess the short answer to your question is I think some external system is generally the best approach right now, whether that's a scratch pad or whether it's actually a database or an index or a graph or something like this, some external source of knowledge uh that the model can can go references is um is generally better than trying to like keep it all in the context window i think going deeper on that point if you look back a year or or a year and a half ago um when this was still quite early there was sort of like everybody was i remember i was on several of these interviews they were like what what's sort of the next big thing with ai and a lot of people were talking about context window and like at that point you recall context windows had grown from like 128k or even like 32k and 128k was like a big deal and then it got to like a million or two million and everybody thought context windows were going to go infinite and and um not a lot of people were talking about agents at that point i i remember having a couple conversations about agents and how profound they could be and i think what's happened is agents have like not removed the need for infinite context windows i mean big context windows help But there's certainly still diminishing returns with larger and larger context windows.

51:43And having agents in test time compute have allowed you to kind of minimize the level or the amount of the context window you use and store some of this information offline and organize it and let the model go retrieve what they need when they need it, like in a more targeted way. And certainly the model's better when it does that. If you fill its context window with a bunch of information that doesn't matter, then it's going to get lost at some point, right? Or it's going to follow the wrong path at some point. So I think agents allow a pattern where you're not just flooding the context window all the time, but you're being more intelligent about how you use the context and use it in a more targeted way.

52:29Yeah. I'm coming up to an hour. Is there anything I haven't brought up that you want to say to listeners that you guys are doing? Yeah, it's a good question. I mean, I think we covered a lot of ground today. I think one of the most important things that we touched on a little bit early on is around these validation exercises that our human experts do. So a lot of people talk about content and how important it is for AI systems. And that certainly is true. But I would say we have seen a tremendous shift. I mean, we have 4 ,500 domain experts, lawyers, tax analysts, et cetera, that we employ at TR who have spent their careers like generating content for the industry.

53:24And these people like know more about their domains or their practice area of law than maybe anybody else in the world. Right. And so being able to leverage these people to design and build these AI systems, I think, has been one of the most profound things for us because it's allowed us to sort of take their experience and their knowledge and their feel for the certain practice area that they work in and really align and train our systems to look, feel, behave like those people do. And so I think that's been one of the like really most interesting shifts that we've seen over the last two years.

54:02And we're still making more of those shifts as we go through time. But that's been just a critical part of our development. And I think a critical part of the quality of the systems that we're able to build is being able to really bring those people together with our engineering and science teams to build great products. And you're talking about using those people in an RLHF, is that right, loop? Or are you talking about them evaluating just as sort of as external evaluators that can watch the model and flag when it's wrong? Yeah, really all of the above. So if you think about what an evaluation is, I mean, in many ways, like it's a preference judgment that can be used in an RL routine, for instance.

54:56And so really you can leverage this data in a number of different ways. But from evaluating our systems to writing prompts themselves because they understand how they want the thing done to to to doing actual annotation or RL data for us. also to like designing rubrin rubrics for evaluation like so one of the things that gets a lot of attention is like llms as a judge where llms sort of assess their right answer and the best llm as a judge systems are really rubric based which means don't just like give me the right answer and then ask the llm is this answer right you need to design a rubric that says what about this answer makes it right or wrong?

55:43It references this case. It reaches this conclusion, blah, blah, blah. And so these are rubrics then that the LLM can evaluate an answer on and say, did it meet the criteria or not? So designing those rubrics also is quite a domain specialized task. And so we've been able to leverage our SMEs across these industries for all of these kinds of tasks. And it's been a really huge enabler for us, I would say. Wow. And I've got to ask, do you build, do you have any personal projects that you're building on the side? Do you have a Joel Huran chat bot that people can talk to or something like that? A digital twin?

56:27I wish I did. I'm sure there's many people on my team who wish they had that digital twin, too. They'd probably be pretty helpful. Not too much. I have nothing like personal project that I'm building for any material use. But I have done personal things just to try to understand where the AI dev tool market is going. So I've built things that interact with my email, for instance, and help me track personal finances, personal vacations, these kind of things. and and i don't do that because i like want to use them i i do it so that i understand where ai dev tools are going for two reasons one i think it's important for me like as i think about the workforce and the talent and like how they're going to be working in the future that helps me get a better sense of it um but secondly i think ai dev tools have been like 18 months ahead of the rest of the market in terms of AI, particularly because code lends itself to AI particularly well.

57:39The language is well-structured. There's a lot that's publicly available. It's testable. Diffs work really well. There's a lot of things about code that make it very conducive to AI. And so you've seen that space really move way faster than any other space. And so I like to look at AI DevTools as kind of like a foreshadowing of maybe what a lot of these other industries are going to look like in a couple of years as models improve. And so I spend a lot of time working with these tools just to try to get an insight and feel for that as well. Build the future of multi-agent software with Agency.

58:18That's A-G-N-T-C-Y. Now an open source Linux foundation project, Agency is building the internet of agents, a collaborative layer where AI agents can discover, connect, and work across any framework. All the pieces engineers need to deploy multi-agent systems now belong to everyone who builds on agency, including robust identity and access management that ensures every agent is authenticated and trusted before interacting. Agency also provides open, standardized tools for agent discovery, seamless protocols for agent-to-agent communication, and modular components for scalable workflows. Collaborate with developers from Cisco, Dell Technologies, Google Cloud, Oracle, Red Hat, and more than 75 other supporting companies to build next-generation AI infrastructure together.

59:25Agency is dropping code, specs, and services, no strings attached. Visit agency.org to contribute. That's A-G-N-T-C-Y dot O-R-G.

From the publisher

This episode is sponsored by AGNTCY. Unlock agents at scale with an open Internet of Agents. Visit https://agntcy.org/ and add your support.

 

Joel Hron, Chief Technology Officer at Thomson Reuters, joins Eye on AI to unpack the future of agentic systems and what it takes to build them responsibly at enterprise scale.

We dive into the shift from prompt-based AI to true agentic workflows capable of planning, reasoning, and executing complex tasks. Joel breaks down how Thomson Reuters is deploying generative AI across law, tax, risk, and compliance, while keeping human experts in the loop to ensure trust and accuracy in high-stakes domains.

Topics include:
- What separates agentic AI from simple prompt-based tools
- How “agency dials” (autonomy, tools, memory) change system behavior
- Infrastructure and architecture required for multi-agent collaboration
- Why human verification and user experience design are essential for trust
- The future of coding, engineering skills, and AI adoption inside enterprises

If you want to understand how a 170-year-old company is reinventing itself with AI — and what’s next for agentic systems in business and knowledge work — this conversation is a must-listen.

Stay Updated:
Craig Smith on X:https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI

More from Eye On A.I.

All 266 episodes
#290 Joel Hron: How Thomson Reuters is Approaching The Next Era of AIEye On A.I. · 1 h
Listen in VO