The path towards trustworthy AI

29 Oct 2024 · 52 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Practical AI - Episode Summary: The Path Towards Trustworthy AI

Overview In this episode of the Practical AI podcast, hosts Chris Benson and Elham Tabassi, Chief AI Advisor at the U.S. National Institute of Standards & Technology (NIST), discuss the critical theme of trustworthy AI. The conversation delves into NIST’s AI Risk Management Framework (AI RMF) and its alignment with the White House’s Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence.

Key Concepts Discussed

  1. Introduction to NIST
  2. NIST is a non-regulatory agency established in 1902, aimed at advancing U.S. innovation and industrial competitiveness through measurement science and standards.
  3. NIST's work involves developing standards for various technologies, including a focus on AI to improve its trustworthiness.
  1. AI Risk Management Framework (AI RMF)
  2. The AI RMF is a voluntary framework designed for the flexible and structured management of AI risks.
  3. Developed through multi-stakeholder collaboration, it seeks input from both technology developers and those impacted by AI technologies.
  1. Trust in AI
  2. Trust in AI is a multi-dimensional concept that includes characteristics such as:
  3. Validity and reliability
  4. Accountability and transparency
  5. Safety, security, and resilience
  6. Explainability and interpretability
  7. Privacy enhancement
  8. Management of bias
  9. These characteristics are interrelated, and trade-offs may need to be considered in their implementation.
  1. Engagement and Collaboration
  2. NIST emphasizes community engagement in developing guidelines and standards, ensuring diverse perspectives are considered.
  3. The importance of including experts from various fields, such as economics and psychology, in AI discussions to create a holistic understanding of trust.
  1. Executive Order Impact
  2. The Executive Order issued by the White House on AI aims to enhance the safety and trustworthiness of AI, building upon the foundation laid by the AI RMF.
  3. NIST is tasked with developing evaluations, safety guidelines, and consensus-based standards in response to this order.

Steps for Organizations Elham Tabassi provides a roadmap for organizations aiming to adopt the AI RMF:

  • Read the AI RMF Document: A concise 30-35 page guide that provides an overview of principles.
  • Utilize the Playbook: Offers detailed recommendations, categorized functions, and guidance for implementation.
  • Start Small: Organizations should begin with a few recommendations that suit their specific use case.
  • Continuous Monitoring: Risk management is an ongoing process that requires regular assessment and adjustment.

Future of Trustworthy AI

  • Elham envisions AI being utilized as a powerful tool for scientific discovery, impacting fields like healthcare, education, and environmental science.
  • The role of NIST is to ensure reliable measurements and evaluations of AI to foster trust among users.
  • There is a need for developing clear and robust standards to ensure the accountability and governance of AI systems.

Conclusion Elham Tabassi’s insights highlight the importance of establishing trust in AI systems through standards, collaboration, and continuous improvement. The ongoing dialogue between NIST, industry, and the broader community is essential for the responsible development and deployment of AI technologies.

Links and Resources

  • [NIST AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework)
  • [AI Resource Center](https://airc.nist.gov)
  • [Executive Order on AI](https://www.whitehouse.gov/briefing-room/presidential-actions/2023/10/30/executive-order-on-the-safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence)

Episode Information

  • Host: Chris Benson
  • Guest: Elham Tabassi
  • Published: October 2023

---

This summary captures the fundamental elements explored in the podcast episode, providing a comprehensive overview of the discussions surrounding trustworthy AI and the role of NIST in this evolving landscape.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:28Welcome to Practical AI. near your users. Learn more at fly.io. Okay, friends, I'm here with a new friend of ours over at Timescale Avthar Suothan. So Avthar, help me understand what exactly is Timescale? So Timescale is a Postgres company. We build tools in the cloud and in the open source ecosystem that allow developers to do more with Postgres. So using it for things like time series, analytics, and more recently, AI applications like RAG and Search and Agents. Okay, if our listeners were trying to get started with Postgres, Timescale, AI application development, what would you tell them? What's a good roadmap?

1:08If you're a developer out there, you're either getting tasked with building an AI application or you're interested and you're seeing all the innovation going on in the space and want to get involved yourself. And the good news is that any developer today can become an AI engineer using tools that they already know and love. And so the work that we've been doing at timescale with the PGAI project is allowing developers to build AI applications with the tools and with the database that they already know, and that being Postgres. What this means is that you can actually level up your career, you can build new interesting projects, you can add more skills without learning a whole new set of technologies.

1:45And the best part is, it's all open source, both PGAI and PG VectorScale are open source, you can go and spin it up on your local machine via Docker, follow one of the tutorials on the Timescale blog, build these cutting edge applications like RAG and such without having to learn 10 different new technologies and just using Postgres in the SQL query language that you will probably already know and are familiar with. So yeah, that's it. Get started today. It's a PGAI project. And just go to any of the Timescale GitHub repos, either the PGAI one or the PG Vector Scale one, and follow one of the tutorials to get started with becoming an AI engineer, just using Postgres.

2:23Okay, just use Postgres and just use Postgres to get started with AI development, build RAG, search, AI agents, and it's all open source. Go to timescale.com slash AI, play with PGAI, play with PG Vector Scale, all locally on your desktop. It's open source. Once again, timescale.com slash AI.

3:04Welcome to another episode of the Practical AI Podcast. I am Chris Benson. I am a Principal AI Research Engineer at Lockheed Martin. And unfortunately, my co-host Daniel is not with us today, But it is my pleasure to introduce Elham Tabasi, who is the Chief AI Advisor at NIST, which is the National Institute of Standards and Technology. Welcome to the show, Elham. Thanks for having me. You guys are doing so much in this area in terms of AI and kind of setting the stage. And I was wondering, for those of us in the audience who may not be familiar with NIST, If you could kind of start out with just telling us a little bit about NIST, what you do both in AI and maybe outside to give a little context and give us a little intro into what NIST is doing in AI and your role in that.

4:01Yeah, happy to. NIST, our National Institute of Standards and Technology, is a non-regulatory agency under the Department of Commerce. NIST was established in 1902, and our mission has not changed since then. NIST's mission is to advance U.S. innovation and industrial competitiveness. At NIST, we have a very broad portfolio of research, from building the most accurate atomic clock to modeling the behavior of the wildfire. But most importantly, we have a long tradition of cultivating trust in technology. We do that by advancing measurement science and standards. measurement science and standards that makes technology more reliable, secure, private, fair, in other words, more trustworthy.

4:46And that's exactly what we are doing in the space of AI. As I mentioned, this was established in 1901 to fix the standards of weights and measures. Our predecessors created and advanced standards to measure basic things such as length, mass, temperature, I don't know, time, light, electricity, all those that were essential for technological innovation and competitiveness at the turn of the 20th century. We are following the same course, working with and engaging the whole community in figuring out proper standards and measurement science for advanced technologies of our time, which is artificial intelligence.

5:27And then the way we do it is exactly or maybe an improved version of what we have been doing in the past century or so. NIST's day-to-day work is focused on helping industry develop valid, scientifically rigor methods. And one thing that I want to emphasize is that we do this through multi-stakeholder open, transparent collaborations. While we have a lot of really good experts and expertise at NIST, we also know that we don't have all of the answer. And it's really important and vital for us to foster a consensus and bind across our stakeholder community. So what we do is that we listen and we engage.

6:08We get the input. We distill them down. We develop a path for measurement to build up or bolster the scientific underpinning. And then we develop tools, guidelines, frameworks, metrics, standards, et cetera, to support industry and technology. And we have done that for development of the AI risk management framework. We have done that for quantum computing, for cybersecurity. And we are continuing doing that for improving methods and measures for risk management and trustworthiness of AI systems. That's a great introduction. I'm curious, you know, you talked about a couple of things about collaboration.

6:53You seem to be right at the center, kind of, you know, sort of an interface between government interests in these technologies and the issues around them and industry. And I know you work with a number of different organizations as, you know, that NIST does in these different things that you've talked about. And you specifically called out trust. I was wondering if you could talk a little bit about kind of how those different collaborations work, how trust in technology can evolve, and how does NIST go about that process of starting to, you know, in AI and in other adjacent technologies, how does it go about that process that it's been doing for so long?

7:35Yeah, thank you for that question. As I said, it's sort of, I think, the magic sauce for us to do stakeholder engagements, to work with the community and ask for their inputs, leverage the knowledge of the community and build on the really good works that the community has done. And by working with all of the experts, strengthen the scientific underpinning, building the right technical building blocks that are needed for development of scientifically valid guidelines and standards. In terms of the engagements, particularly in the space of the AI, we all know that AI is a multidisciplinary and in order to understand the concept of the trust, what make AI systems trustworthy and what constitute trust, that was one of the main questions as we were developing the AI risk management framework.

8:30And in the engagements that we were doing with the community early on, we recognized that as much as we need the input from the community that developed the technology, this is the community with expertise in math, statistics, computer science. We also need input from the community that studied the impact of the technology. That's economists, sociologists, psychologists, cognitive scientists. And we need to bring all of them together because AI systems are more than just data, compute, and algorithm. They are a complex interactions of data, compute, and algorithm with the human, with the environment, with the human that operates this, with the human that can be impacted by the systems.

9:16So that engagements with a very broad set of actors in the community bring different expertise and backgrounds became really important. In answering your questions about the trust and what constitutes a trust, as I said, that was one of the important and central questions in development of the AI RMF. The AI RMF or the AI Risk Management Framework, very briefly, it was directed by a congressional mandate, and it's a voluntary framework for managing the risk of AI in a flexible, structured, and measurable way. It was, as we do with anything else, developed in close collaborations with the AI community and engaging diverse groups of different backgrounds, expertise, and perspective to particularly get in and focus and hone on the concept of trust and trustworthiness.

10:09So on the side of what makes AI systems trustworthy, when we started the process, there have been very good high-level value-based documents that talk about AI systems to be non-discriminatory, ethical, and there has been a lot of different other papers and publications. Basically, there were many different views about what make an AI technology, AI systems trustworthy. And these views were not all aligned and on the same page. So it's not a property that can be defined with perfect rigor, but based on the collaborations and engagements and consultations that we did with the community, we understood that there are well-established key characteristics of trustworthy systems.

10:57With the help and consultation with the community in the AI RMF describes trustworthy AI systems as those that are valid and reliable, accountable and transparent, safe, secure and resilient, explainable and interpretable, privacy enhanced, and firm with harmful bias managed. It takes it a step further and for each of these characteristics provide a sort of a definition or bring the community on a shared understanding with expectation from each of these characteristics and also talks about how these characteristics interrelate and trade-offs involved in decisions about how safe is safe, how private is private, or how to enhance interpretability or transparency while at the same time, for example, preserving the privacy or ensuring the security and resilience of the AI systems.

11:51I'm curious, and as we go, we'll certainly dive into those topics, but one of the things around trust is, you know, folks like us who are in this industry and are living and working in and around and developing AI every day. You know, these are kind of work topics that we're going through and the guidance that NIST provides is invaluable, especially being part of that process of developing and as you described. But I would, before we surge into all that, there is so many people out there that are not in this line of work as we are and that are those that are you know curious about how I you know they see it in their in the news every day and they see you know they're trying to understand what are these technologies that we're working on and many people in the audience for this podcast are what we would describe as kind of AI curious as opposed to just you know we have practitioners but we also have AI curious people who are trying to understand how it fits into them.

12:52And I was wondering if you'd take a moment and kind of talk about the context of trust and AI for those of us who are not in this industry in a direct way like that. And, you know, like how does NIST try to frame it for the larger population or is it more for practitioners? How do you see that for the larger world? Yeah, thanks for that question. So Let me try to answer that with an example. We are seeing enormous advancements in the AI technology. Just in the past year, we saw a lot of release of the powerful models. We are also seeing that these technology, these AI systems, are being incorporated into a lot of the functions of the society and the way we do our work.

13:46I want to explain the concept of the trust in an examples in use of the AI systems in health domain. When we go, you know, for medical imaging, I'm coming from the computer vision. That's where my training was, and that's where they feel comfortable. So when we do a medical imaging, imagine some sort of imaging of the brain And the question is, is it a tumor there or not? So an algorithm can go and be employed to help the physicians to make that decision. So first for that systems, for that algorithm, we wanted to be talking about the IRMF, trustworthiness characteristics. We want it to be valid and reliable.

14:34We want to make sure that it has some certain level of the accuracy. So the false positive and false negative is low because there is going to be, you know, you don't want to scare a patient saying that, yes, there was a tumor when there was none. Or vice versa, a tumor is because of the errors of the systems is going unrecognized. So we want the systems functions, functions as intended. We want it to be valid and results being reliable. On top of that, we also wanted the systems to be secure and resilient because if it's not, And if the systems get hacked, there's a lot of the personal information that can get in the hand of non-friendly users.

15:17Talking about that, we want the systems to be privacy enhanced. We have read that particularly the large language models, they have the tendency to memorize the training data. Even before the large language model, there were papers that showed that with certain level of expertise, the training data can be inferred from AI systems. So if the system has been trained on real patient data, we don't want to have any hole that can give access to those private informations. Explainability and interpretability. So if it comes and say that, yes, there is a tumor, we expect it to give some reasoning, some sort of explanations of why you decide that there is a tumor.

16:06And then there's a lot of nuances there too, because that explanations, if it's been given to a physician versus a technician versus the patient, it's going to be different level of the technicality and different level of the information's being shared. And of course, we want it to be fair. We don't want AI systems that has been more accurate for certain demographics versus others. You know, this usually happens if the training data is uneven. So all of this at the end, we want to build the confidence in this, that this technology works. and the results, predictions, recommendations that the system is providing for better decision-making in this case, analyzing a scanning of the brain to see if there was any tumor was there or not.

17:00So all of these things are with the end goal of AI technology has a lot of promises. They are very powerful tools. They can transform the way we work for better, but make sure that at the same time it uplifts all of us and we get the maximum benefits while minimizing the negative consequences of the technology.

17:40What's up, friends? I'm here in the breaks with David Hsu. founder and CEO of Retool. So David, Retool has definitely cornered the market on internal tool software development. But zoom out for me. What's the big idea? Why did you start Retool? What is the big idea with internal software? Yeah, so Retool started at this point seven years ago. And when we started Retool, the core idea was that internal software is a giant, giant category that no one really thinks about. And what's surprising to most people is that internal software represents something like 50 to 60 % of all the code written in the world, which might sound pretty surprising.

18:20But if you think about it, most of us in Silicon Valley, we work at software companies, whether it's like an Airbnb, a Google, a Meta. These are all companies that are software companies selling software. And so most engineers in these companies of working on external-facing software. But if you think about most software engineers in the world, most software engineers in the world actually don't work at these software companies. There's not that many of them. There's maybe 10, 20 of them, big ones at least. Most of the companies in the world are actually non-software companies. So if you think about a company like an LVMH, for example, like a Coca-Cola, for example, or like a Zara, Zara's not selling any software.

18:55They actually have a lot of software engineers, actually. And all their software engineers, all they do day in and day out is basically build internal software. So that's, I think, one reason we started to retool. The second reason we started to retool is if you look at all this internal software that people are building, it is remarkably similar. So if you take a look at, you know, like a Zara, for example, versus Coca-Cola, two very different companies, obviously. One a clothing company, one a beverage company. But if you actually look at the software they're building internally to go run their operations, it is remarkably similar.

19:25It's basically forms, buttons, tables, all these sort of pretty common building blocks, basically, that come together in different ways. But then if you think about, you know, not just the UI, but also what's the logic behind a lot of this stuff, they're pretty much just hitting API endpoints, hitting databases. You care about authentication, you care about authorization. These are sort of a lot of common building blocks, if you will, to internal tools. And so for us, the insight was, wow, internal software is a ginormous category, and it's all so similar. and developers hate building it. And so could we create a sort of higher level framework, if you will, for building all this software?

20:02And that would be really cool. That would be really cool. Okay. So listeners, Retool is built for everyone, built for enterprise, built for scale, built for developers, and that's you. And if you found yourself nodding your head to what David was saying, then check out Retool at retool.com slash changelog. It's the fastest way to build internal software. Do yourself a favor, get a demo or start for free today. Again, retool.com slash changelog.

20:45so i know in uh in the early part of 2023 uh nist issued uh the ai risk management framework that we've been talking about. But a few months later on, or almost exactly a year ago, as we're talking, in late October, the White House issued its executive order on the safe, secure, and trustworthy development and use of artificial intelligence. I was wanting to understand that how the issue of the executive order might have either altered or accelerated or changed any of the work that NIST was already doing. You guys were already very much involved in artificial intelligence through the framework and other activities.

21:32Could you describe the impact of the executive order on the work you were doing? Absolutely. In answering your question, if I can just go back from the release of the AIRMF in January 2023 to release of the executive order end of October, October 30th, 2023. So the AIRMF was released January 2023. In March of that year, we released the AI Resource Center. This is a one-stop shop of knowledge, data tools for AI risk management. It houses AIRMF, its playbook in an interactive, searchable, filterable manner. And by the way, the AI Resource Center is definitely a work in progress, and we want to keep adding to that and adding more additional capabilities, things such as standards hub, repository for metrics.

22:25We want it to be really a one-stop shop of all of the informations, but also a place for engagements across the different experts. In June of 2023, just give a little bit of context, the ChatGPT3 was released in November 2022, to a month or so before release of AIRMF and CHAT GPT-4 was released in February or beginning of the March, a year after, a month after release of the AIRMF. So in response to all of these new development and advancement technology, we put together a generative AI public working group where more than 2 ,000 volunteers help us study and understand the risk of the generative AI.

23:12And then in October, as you said, we receive our kind of our latest assignment, Executive Order on Safe, Secure, and Trustworthy AI. This executive order really builds on the foundation works that we have been doing from the AI RMF to Playbook to Resource Center to Generative AI Public Working Group, and supercharge our effort to cultivate trust in AI, mostly by giving us some tight timelines of the things to deliver. The EU specifically directed NIST to develop evaluations, redeeming safety and cybersecurity guidelines, facilitate development of consensus-based standards, and provide testing environment for evaluations of AI systems.

23:57All of these guidelines infrastructures, true to the nature of NIST, will be a voluntary resource for use by the AI community to support trustworthy development and responsible use of AI. We approach delivering on the EO the same way that we do all of our work going to the community. We put a request for information out to receive input. Based on the input that we received, we put draft document out for public comment. Based on the comments that we received, we developed the final documents that we were very pleased that all of them were released by the deadline of July 26 that the EO had given us.

24:44Quickly, a quick overview of the things that we put out. One of them was a document on a profile of AI RMF for generative AI. The document number is, at NIST we like to refer to everything with the number. So that document is the NIST AI 600-1. It's a cross-sectoral profile companion resource to the AI risk management framework based on the input that we had and discussions that we had on the generative AI public working group, responses to the RFI and inputs that we have received. I think one main contribution of that document, if I want to summarize it, is its description of the risks that are novel or exasperated by generative AI technologies.

25:32These risks span from CBRN information capabilities, eased access to synthesis of materially nefarious informations that can lead to design capabilities for CBRN, confabulation, dangerous, violent, hateful content, data privacy risks, let me remember the rest, environmental impact, bias, human AI configuration, information integrity, information security, intellectual property, degrading or abusive content, and the concept of the value chain and component integration. With the generative AI, we are moving from the binary, the deployer, developer, kind of actors and dynamics. And now we are having upstream of the third-party components, including data that are part of this value chain.

26:32So one of the things that we're working in continuing that work is to work with the community to get a better understanding of the technology stack, of AI stack, if you will, for AI, understand the role of the different AI actors involved there so we can do a better risk management. As you're talking about that, could you describe a little bit, and this is just a question in my mind, like when we're talking about AI as risks, as a set of risks, and we talk about kind of that effort to create trust in technology, how do you tie those together? In this process, you've identified these risks and you just enumerated those.

27:22And with the purpose of ultimately kind of helping people get to a point of trust and being able to implement the technologies productively, how do you approach getting to trust through mitigation of risk? Does that, I'm not sure if the question makes sense or not. It certainly makes sense. I'll try to answer the way I understood this. So AI systems are not inherently bad or risky. And it's often the context that determines if a negative impact will occur and also what are the risks. So an example that I usually use is that if I use face recognition to unlock my phone versus face recognition as in the airport that now our faces are boarding pass to get on the plane.

28:13Or face recognition in the context of the law enforcement. It's the same technology, but in the different contexts, there is different risks and different level of the assurances that we want to have for the systems to work in a trustworthy manner. So what we have been trying to do as part of our work in approaching trust and trustworthy AI, the first one was to unpack the concept, try to get into the characteristics that make a system trustworthy. That helps to answer the question of what to measure. If I want to know if it's trustworthy or not, what are the measurements I need to do? So I listed the seven characteristics from valid and reliable, safe, secure, et cetera.

29:03So that gives a more of a systemic approach and structural approach to what are the dimensions, what are the characteristics that together can make a system trustworthy. And by the way, AIRMF talk about this, that not each of them by itself make a system trustworthy. You can have a system that is very secure, but not valid or accurate. So that's not going to be trustworthy. And a system that's 100 % accurate, but not secure, is also not trustworthy. So that gives, again, more structured approach on what to measure. Then the next step is how to measure methods and metrics for the measurement. Those type of measurement gives an information about limits and capabilities of the systems, the type of the risk that can occur, the magnitude of the impact if those risks occur.

29:57And then based on this information, then we can come up with mitigations and management of the risks. So AIRMF, its recommendations is really categorized in the four functions of govern, map, measure, and manage. The govern is giving recommendations on procedures and processes, roles and responsibilities that we want to have in an organization to do effective risk management. So what is the accountability line? What are the role and responsibility involved? The map functions provides recommendations on understanding the context of the use, going back to that examples of the face recognition, understanding the environment that AI systems are operating there, understanding the community that can be impacted by that, identify the risks in this particular context, understanding the laws, regulations, and policy that are effective in this context of use.

30:57The measure functions provides recommendations on the how-to measure. So for all of the risks identified in the MAP measure, provides quantitative or qualitative recommendations on how to measure them, how to take into account the trade-off between all of those trustworthiness characteristics. And all of this information is being used during the managing the risk part that the recommendations can go from safeguards and mitigations that can put in place to mitigate risk. to sometimes we cannot just mitigate risk and the risk should either be accepted or transferred or the systems is too risky that it should not be developed or deployed.

Read the full transcript

31:42So that is how the process in AIRMF.

31:54There's a lot of your personal data out there on the internet and you know this, anyone can see this stuff. There's more than you think, though. Your name, your contact info, your social security number, your home address, your various addresses, your past addresses. There's even information about your family members, maybe even the name of your cat. This is all being compiled by data brokers and is being sold. Now, these data brokers, they make a profit off your data, obviously. So they do it. Your data is a commodity and anyone on the web can buy your private details. They can identity theft you.

32:27They can fish you. They can attempt to fish you. They can harass you. They can send you unwanted spam. They can call you nonstop. And this is something I get lots. But now you're able to protect your privacy online with Delete Me. As a person who exists publicly for some time now, especially someone who shares their opinions online quite frequently, I'm aware, hyper aware of safety and security. And I take it seriously. And it's easier than ever to find personal information about anyone online, really. All this data is just hanging out on the internet and can have actual consequences in the real world.

33:04That's why I was excited about finding this recent solution and a sponsor of this show, Delete Me. Delete Me is a subscription service that removes your personal information from hundreds of data brokers online. When you sign up, you can provide Delete Me with exactly what information you want deleted, and their experts take it from there. They send you regular personalized privacy reports showing what information they found on the internet about you, where they found it, and what they removed. And Delete Me isn't just a one-time service. They are always working for you, constantly monitoring, constantly removing your personal information that you don't want on the internet.

33:40And to put it simply, Delete Me does all the hard work of wiping your data, your family's personal information, and all these things you don't want out there from those data brokers' websites. Now, the next step is to take control of your personal data and keep it private forever by signing up for Delete Me. Now, at a special discount rate for our listeners, of course, this is awesome. Get 20 % off your Delete Me plan by texting practical to 64000. Again, text the word practical to 64000. And of course, you may know this already, but message and data rates may apply. Check the terms, all that good stuff.

34:21Once again, text the word practical to 64000 and get 20 % off Delete Me. Enjoy.

34:44so that was uh very useful for me in terms of trying to frame and understand uh you know what you're relaying here in terms of govern map measure manage and and you you talked about something a moment ago that was really interesting in the sense of you have these characteristics, you know, that you're trying to measure toward trustworthy, but it's not just one and it's not just, you know, a black or white issue. You have a collection of them and they vary across different types of use cases, it sounds like. So you kind of have a, you know, characteristic profiles in a sense. How do you think about, if you are out there as a consumer of the guidance that you're providing from NIST, maybe you're in a small company that's doing some work in AI and you're trying to implement the guidance from NIST and you're kind of evaluating your own profile of characteristics through that govern map measure manage process, how does one frame that if you're kind of just getting into this and trying to implement the guidance, could you talk a little bit about how an organization that maybe had not done this before might go about implementing a particular, you know, whatever their use case is and, you know, how do they get started in the process?

36:08What's your recommendation there? The first thing I will say is that you don't need to implement all of the recommendations in the AIRMF to have a complete risk management. So our recommendation is that start by looking at and reading the AIRMF. It's not a very long document. I forgot. I think it's about between 30 to 35 pages. So get a kind of a holistic understanding of this. And then check out the playbook in the AI Resource Center, where for each of the recommendations, AI or MF is in high level for functions. Each function is divided into categories and then subcategories. So in a sort of a granular approach, we give recommendations on what to do for the govern.

37:00And then for each of those recommendations, get into a little bit more granular recommendations. The playbook for each of the subcategories, which is about, I think, 70 subcategories in the ARRMF, provides recommendations on suggested actions and informative documents that you can go read and get more information. And also suggestions about transparency and documentations for implementation of that subcategory. So we often suggest that get a better understanding of the AIRMF, spend some time in the playbook to get a better understanding of the type of the things that can be done. And then based on the use case, based on exactly what you want to do, start by simple, small number of recommendations in the AIRMF.

37:58and start implementing that. Govern or map functions are useful starting points. Govern provides recommendations about the setup that you need for a successful risk management. So it can give you ideas or an organization's ideas about the resources that's needed, the teams that needed to do this, so they can align it with their own resources and the teams that they have. And the map functions, as we discussed, gives recommendations of a better understanding of the context, getting answers to what needs to be measured. I will also add that the functions govern, map, measure, manage, there is no order on doing that.

38:48Depends on the use case, depends on what needs to be done. The starting point can be recommendations of any of the functions. We usually recommend start with govern and map and then start with as few number of the subcategories or recommendations that the resources and the expertise of the entity allows for their implementations, of course, prioritize in terms of their own risk management. And then the last thing I also add is also be mindful that the risk management is not a one-time practice that we just do at once. And you say, okay, I'm done with my risk management. AI systems, you know, there's data drift, model drift.

39:34These newer models can change based on the interactions with the users, with the environment. So we suggest a continual monitoring and risk management. So I think one of the recommendations in the MAP or govern is to come up with a cadence of repeating the assessments of the risks. So that would be my recommendations. Another thing that I would say is that I mentioned the AIRC. I mentioned the playbook. We also in the AIRMF talk about profile. So I keep emphasizing the context of the use and mentioning the importance of the context in AI system deployment, development, and the risk management.

40:20At the same time, ARRMF by design is trying to be sector agnostic and technology agnostic. We try to kind of come up with the foundations, the common set of the practices that's needed to be aware of and are suggested for risk management. But we also have a section on AI profile and recommendations on building verticals. These profiles are instantiations of the AI RMF for a particular use case or domain of use or technology domain so that each of the subcategories can be slanted or be aligned with that use case. So there can be a profile of AI RMF for the example that I used, medical image recognition.

41:07You can imagine a profile of AI RMF for financial sectors. That's something that we have been asked to work with the community on. That was a very long intro to say that there are a couple of profiles posted on the AI Resource Center. One is the one that Department of Labor did for inclusive hiring. Another one that Department of State did for human rights in AI. So that can give some sort of a window to or idea about where the organizations can start. In addition to the profile, we have also posted a few use cases and we will post more use cases. And that is how different organizations are using AI RMF that can hopefully be more practical examples of how to use AI RMF.

42:03No, that's a fantastic set of suggestions right there. And I'd actually like to kind of ask a follow up to that. And as a prelude to my follow up, if I'm understanding, kind of go to the AI RMF, read that core document. It's not very long. It's very consumable. Go to the playbook. Look at the subcategories. I believe you said there were about 70 of them. You know, it has suggested actions and references to other docs in that and then start to bite off kind of simple, small chunks in terms of how you're going to approach the functions that you mentioned, starting kind of with govern a map and then kind of how to put together resources and teams and then kind of cycling back with a cadence of repeated assessments that are also specific to the vertical that you're in.

42:53And as you're doing that, it's feeling really practical from my step. But, you know, we're practical AI. So that appeals to us. I'd like to ask, are there now or do you expect kind of tooling? You know, like if you look outside of AI, kind of the software industry at large is kind of a predecessor to that. as standards and workflows and kind of best practices arose in software development at large. Lots of tooling arose around how to do, you know, agile methodology and you name it. There are many different approaches to software development. Are you expecting tooling or do you have any thinking around what kind of tooling might help AI development teams that as they're building these teams and their resources so that they can be productive over time, how are you seeing that evolve going forward?

43:50Or do you think that there'll be a cottage industry kind of forming around this the way we've seen in software and other areas where there's a lot of tool support? Yes, we have already started seeing some of that. So there are entities that are putting tools for implementation of the IRMF and dashboards and all this. that they have developed those tools and they are having it on their websites. If I can just go back and thank you for your excellent summary of my very long, windy answers. No, it's very good. I'm learning a lot here. And I ask your listeners to, I think, start with the AI Resource Center.

44:31The URL is airc.nist.gov. So ARRMF is there and Playbook in an interactive, filterable way is there. So if their businesses is only, you know, they are developers, they can go and first filter all of the, you know, from the 70 recommendations, anything that is only applicable to the developer. So they're not overwhelmed with all of that. or if they only care about deployment and the issue of the bias for the deployment, they can go and say, you know, filter from the AI actors for the deployers and from the characteristics from the bias and that gives them, that saves them sometimes. So that is where they get information from our website and some hint about, you know, kind of we have it in more filterable way.

45:21And yes, there has already sorted entities that are putting more tooling in. And with the 600-1, that was the cross-sectoral profile of the AI-RMF for the generative AI. The work that we're doing with the community, we are focusing on, we use the word operationalization. So what are the tools that are needed for operationalizing and implementing AI-RMF? And going back and emphasizing the community engagements and the role that the input from the community plays in all of these things, some of the tools can be developed by us, but the majority of the tools are being developed by the community and shared by the community.

46:03And we hope that we see more of that. I hope so, too. It's fascinating. I love the framework that you've given us here that can be applied in so many different verticals and so many different ways, and yet is flexible in its guidance that way. As we wind up here and we have seen so much advancement in the development of AI, both as a technology and as the industries around it, and as you are kind of sitting there in the nerve center of kind of where this guidance and these standards come together, bridging both government and industry, as you look forward, what are some of the things when you're not in a particular position?

46:49meeting and you're just kind of winding down and you're, you're kind of thinking creatively about where things are going. What are some of your own thoughts about the future of this, both for NIST's role and for the, the industry and the technology at large, where we're going, because it's just, you know, it's going at such a rate, it's so fast and it's fascinating and is, you know, changing the face of business, changing the face of how we are as humans and stuff in terms of the tools that are available to us. I'd really love your insights into where you think all of this is going in the days and years ahead.

47:28I think what, and for me, the end goal, for me, what I'm hoping to see a lot of that is to use this powerful technology in the way, as a sort of a scientific discovery tool in the way that we are doing the science and discoveries there. I think that is where we are going to see a lot of really advancements into precision medicine, individualized educations, climate change, anything that's going to make life a lot better for all of us. I have to say this, that my heart was warmed by seeing a Nobel Prize for things such as AlphaFold. I keep saying for a long time that needs a lot more recognitions.

48:20But really all of the recognition that AI got through those prices. But I'm also very aware of very important things that NIST can do and the community needs to do. I think we all agree that there is a lot that we don't know about how these models work. And we ought to do something about it. We need to have a better understanding of how these models work, their capability and limits. That gets me to the important topic of evaluations and testing. We talked about it at the beginning of this podcast that it's important to unpack the concept of the trust into the things that needs to be measured.

49:04But at the end of the day, we need to have reliable measurements for assurance that the systems are trustworthy. At NIST, we are, as a measurement science agency, we are the big fan of this quote from Lord Kelvin that if you cannot measure it, you cannot improve it. So if you want to improve the trustworthiness and the reliability of the systems, we need to have a good handle on how to test them and how to evaluate for reliability, for validity, for the trustworthiness characteristics. And our knowledge on how to test AI systems is very limited. We need better evaluations, as we can see. Benchmarks are too easy.

49:41They get saturated very quickly. We need to have a better understanding of how they work. That gets to the assurance that can build trust into the technology and give users, everybody confidence that the systems works. And the third item that I put in, once we have built that knowledge base, once we have a good scientific foundation, when we have true did research and the work with the community, we have built the technical building blocks. Let's develop clear, understandable, technically robust standards that can help with global improbability of AI evaluations, AI assurance, and AI governance.

50:26Fantastic. Well, Elham Tabasi, thank you so much for coming on the Practical AI podcast. It was very, very instructive in terms of how to frame this. Certainly information that I'm going to be using going forward and really appreciate you taking time to talk with us today. I appreciate the opportunity to be here and talk and really enjoy the conversation. Thanks.

50:57all right that is practical ai for this week subscribe now if you haven't already head to practicalai.fm for all the ways and join our free slack team where you can hang out with daniel chris and the entire changelog community sign up today at practicalai.fm slash community thanks again to our partners at fly.io to our beat freaking residents breakmaster cylinder and to you for listening we appreciate you spending time with us that's all for now we'll talk to you again next time

From the publisher

Elham Tabassi, the Chief AI Advisor at the U.S. National Institute of Standards & Technology (NIST), joins Chris for an enlightening discussion about the path towards trustworthy AI. Together they explore NIST’s ‘AI Risk Management Framework’ (AI RMF) within the context of the White House’s ‘Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence’.

Join the discussion

Changelog++ members save 10 minutes on this episode because they made the ads disappear. Join today!

Sponsors:

  • Timescale – Real-time analytics on Postgres, seriously fast. Over 3 million Timescale databases power loT, sensors, Al, dev tools, crypto, and finance apps — all on Postgres. Postgres, for everything. 
  • Retool – The low-code platform for developers to build internal tools — Some of the best teams out there trust Retool…Brex, Coinbase, Plaid, Doordash, LegalGenius, Amazon, Allbirds, Peloton, and so many more – the developers at these teams trust Retool as the platform to build their internal tools. Try it free at retool.com/changelog
  • DeleteMe – DeleteMe makes it quick, easy and safe to remove your personal data online. 

Featuring:

Show Notes:

Something missing or broken? PRs welcome!

More from Practical AI

All 157 episodes
The path towards trustworthy AIPractical AI · 52 min
Listen in VO