In short
The TWIML AI Podcast - Episode #705: ML Models for Safety-Critical Systems with Lucas García
Episode Overview In this episode, host Sam Charrington welcomes Lucas García, a principal product manager for deep learning at MathWorks. The discussion revolves around the incorporation of machine learning (ML) models in safety-critical systems, emphasizing the significance of verification and validation (V&V) processes.
Key Topics Discussed
Introduction to Lucas García
- Background as a mathematician with 16 years at MathWorks.
- Transition from programming in finance to product management in deep learning.
Verification and Validation (V&V) in Safety-Critical Systems
- Verification: Ensures the model correctly implements the intended design and meets specified requirements.
- Validation: Evaluates the final product to ensure it meets its intended use and functionality.
The V-Model vs. W-Model
- Traditional V-model emphasizes verification on the left side and validation on the right.
- The W-model (proposed by EASA) adapts the V-model to better fit the requirements of AI models, ensuring continuous learning assurance throughout the development process.
Challenges in Applying Deep Learning to Safety-Critical Applications
- Use case example: battery state of charge estimation.
- Importance of data quality, model stability, robustness, interpretability, and accuracy.
- Discussion on formal verification methods and the complexities involved in applying deep learning neural networks.
Formal Verification and Abstract Transformer Layers
- Introduction to abstract transformers for model verification.
- Discusses the need for formal methods to ensure AI models avoid errors and maintain performance reliability.
Constrained Deep Learning and Convex Neural Networks
- Constrained Deep Learning: Incorporates domain-specific constraints to ensure desirable properties like monotonicity or convexity in outputs.
- Convex Neural Networks: Maintains properties that prevent getting stuck in local minima during optimization, emphasizing the trade-offs involved.
Applications in Safety-Critical Industries
- Challenges in integrating AI in regulated industries such as automotive and aerospace.
- Overview of specific standards (e.g., DO-178C) and the processes for certification.
Benchmarking and Best Practices
- Difficulty in establishing industry-standard benchmarks for safety-critical AI applications.
- Introduction to the MLLEAP project, aimed at guiding users in AI verification and certification processes.
Key Takeaways
- The integration of AI models into safety-critical systems requires a thorough understanding of V&V processes and the complexities introduced by AI.
- The W-model provides a framework to enhance the traditional V-model, addressing the unique challenges posed by AI in safety-critical applications.
- Constrained deep learning and convex neural networks present promising avenues for building reliable models that can satisfy safety requirements.
- Collaboration across industries, research, and regulatory bodies is essential for establishing standards and best practices for AI in safety-critical environments.
Conclusion The discussion highlights the critical need for robust verification and validation processes in the deployment of ML models within safety-critical systems. As the field evolves, continuous collaboration and research will be vital in overcoming the challenges and ensuring the reliability and safety of AI applications.
For complete show notes, visit [TWIML AI Podcast Episode #705](https://twimlai.com/go/705).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00There are techniques out there, there are observers, traditional observers that have been used for many, many years. The Kalman filter or the extended Kalman filter is perhaps the most popular or the most widely used technique out there. And one, of course, can think about, all right, why on earth would I use an AI model instead of the extended Kalman filter that is working correctly, right? It's doing its intended design. It's performing well. It's fast. It's reliable. Well, there might be a few reasons.
0:42All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Lucas Garcia. Lucas is Principal Product Manager for Deep Learning at MathWorks. Before we get going, be sure to take a moment to hit that subscribe button wherever you listen to today's show. Lucas, welcome to the podcast. Thanks, Sam. Happy to be here. I'm super excited to jump into our conversation. We'll be talking about incorporating AI models into safety critical systems. This is a topic you presented on at the last NURPS conference. As a way to lead into that, I'd love to have you share a little bit about your background.
1:21So my background, I'm a mathematician by training. So math is really deep into my roots. I joined MathWorks around 16 years ago. So I've been around for a little bit. Before that, I was a programmer, a C++ programmer in the space of finance. So I was working on pricing derivatives in the finance industry. And at some point, I got tired of that, decided to move on to something else. And I joined MathWorks. So I joined MathWorks as a training engineer. So my first role at Mathworks was to deliver training all around the country and teach people on how to use MATLAB. Over time, I've been in other customer-facing engineering roles, and ultimately I moved to the product management area.
2:21So I now work closely with the development team in the deep learning space. I thought we'd start out by talking a little bit about some of the differences in your experience between the way software is written for safety-critical systems like aircrafts and how it differs from how we might write code for a business or, you know, for your background financial type of an application. Yeah, yeah. And there are some differences. And of course, there are many commonalities as well. Well, to me, it all comes down to having an understanding of, a good understanding of the verification and validation process.
3:03And so perhaps just to provide all the context here, verification is a process of checking that the model is correctly implementing the intended design, right? And meets the specified requirements. So in the context of perhaps traditional software development, as you're building some engineered system, you would have requirements that pertain to the overall system, some of the subcomponents that that system is comprised of and so on. And so as an engineer, you're going to have to go through the process of understanding those requirements, deriving requirements from the system level requirements and so on, and then implementing the software, right?
3:47And then, so the verification aspect is all about answering this question, are we building the product right? So the verification tries to deal with internal consistency and correctness of the model within the design phase. And here, I guess I'm overusing the word model because we have the AI models and the models that we use for simulation that perhaps do not necessarily contain an AI aspect to it. But that's an important piece. So that's the verification aspect. And then we see the validation, which is basically the process of evaluating the final product and ensure it meets its intended use and intended functionality.
4:29at D. So here we're, instead of answering or trying to answer the question, are we building the product right? We ask the question, are we building the right product, right? And so validation is focused more on the system performance, making sure that it's able to deal with the real world data and so on. In software development practices, we've been following this V model or the V diagram that really emphasizes the process of verification and validation at each stage of the development process. So kind of you have the left side of the V focusing on verification, the right side of the V focusing on validation.
5:07And we want to ensure as we go through all the different steps in this workflow that each phase is complete and is correct according to the original requirements. Now, the tricky aspect here is when we incorporate AI. And it's important to note that the traditional V diagram or the traditional V and V workflow or verification and validation workflow, such as this V cycle, might be insufficient or might not always work when we deal with AI models. So there was this interesting contribution from EASA, the European Union Aviation Safety Agency, that deals with this situation. And instead of using a V diagram, proposes to have a W cycle or W-shaped development cycle.
6:00So basically what Iyasa did was to go through this traditional workflow, the traditional steps of putting together an engineered system. And think about what of these steps work well with AI and what of those do not work well. And so the whole point about this process is to make sure that we have a good learning assurance process so that we put together planned and systematic actions to make sure that the system is able to capture errors in a data-driven engineer process. And that if these errors have been captured, we are able to not only identify them, but also correct them. Can you give maybe a concrete example of a software system that you've worked on or seen and what some of those verification and validation steps might look like?
7:00Sure. Yeah. So if we're building, let's say, a battery state of charge estimator, right? I think we all have cell phones today, and the smartphones we use today, as well as any of the EV vehicles out there on the road, are unable to tell us with an adequate level of confidence, what is the actual truth with respect to state of charge? We don't have a sensor that estimates or that gives us the actual reading, the actual state of charge. And instead of that, we have to estimate it. We have to come up with an observer, with an estimation that tells us, okay, so the right amount of state or the state of charge is 85 % right now if I look at my phone.
7:40Right. So, for instance, one requirement for an AI model that is able to provide state of charge estimation values would be my state of charge has to be in between zero and one, for instance, or the battery should monotonically decrease as I'm off charge and should monotonically increase as we are charging. So these are some of the examples. Now, as we go through the W diagram, we need to make sure that we are appropriately addressing different steps of the workflow, data management, making sure that we are capturing the right amount of data for the purpose of the problem that we're trying to solve in the context that I was mentioning earlier on ADAS systems.
8:36We need to make sure that if the system is intended to be used a day at night, we need to have day and night data, of course, right? So we need to go through all of these steps, ultimately to come up with an architecture of an AI model. We need to train it. But after we train it, and before we put it on the embedded system, or we integrate it with the rest of the system, we need to do something else, which we didn't have in the traditional V-diagram, we need to have some guarantees that the emerging properties of neural networks are not impacting the performance of the system. So this is what we are seeing in the EASA contribution called this learning process verification.
9:22But many in the research space talk about AI verification, for example. So it's a key aspect of being able to put these AI models into safety critical systems, for instance. um yeah is it worth backing up and talking about if you've got traditional uh models like why you know folks are trying to use ai models like what are the cases that are lending themselves to folks wanting to do um an alternative like for example the you know the battery state of charge like there are traditional models for doing that why would i even want to use a neural network to do that what am i gaining yeah no excellent question i mean yes there are traditional models out there uh in the context of battery stage of charge or even more broadly virtual sensors right which is a battery stage of charges is you can think of it as a form of a virtual sensor um there are techniques out there there are observers traditional observers that have been used for many, many years.
10:37The Kalman filter or the extended Kalman filter is perhaps the most popular or the most widely used technique out there. And one, of course, can think about, why on earth would I use an AI model instead of the extended Kalman filter that is working correctly? It's doing its intended design. It's performing well. It's fast. It's reliable. Well, there might be a few reasons. Of course, with the hype in AI and a lot of interest coming into the field, that doesn't go unnoticed and customers and users want to explore AI. And that's a perfectly valid approach. But I'll leave that aside for now and focus on the more technical aspects.
11:23What we've seen is that on many occasions, the ability to generalize from these type of traditional base observers is not really working as customers expect. So there is this example from a battery manufacturer called Goshen. They presented in our MathWorks Automotive Conference, I believe it was in 2022. They basically created this onboard SOC estimation, and they were able to deploy it onto their ECU and making sure that it was compliant with the Autosar standards and everything in there. And so basically, it all comes down to being able to come up with a tool or a functionality that's able to generalize well enough.
12:20the extended Kalman filter does have an inconvenience which is that you require a mathematical model of the battery and that's not always useful or possible to have or achievable if you don't have a battery model that or a physical battery model that digital twin right of the battery it's sometimes challenging to build and so if you have access to data input output data it's It's extremely simple. You get access to, there are many public data sets out there these days, but if you have access to current voltage and temperature in the battery, all you need to do is have a controlled environment from the test lab being able to accurately measure current.
13:12and then you compute the state of charge as the integral of the current over the capacity. And that's the label, right? You've been able to very accurately label the data. You have current temperature and voltage, and now you have the labeled SOC values that you would use, right? In this context, it's going to be a regression problem, of course, but you have data that you can now work with. and basically it all comes down to choosing a good model for your needs. Look back at your requirements, never forget your requirements and see, okay, what are my performance requirements? What are my interpretability requirements?
13:57What are the physics requirements for my AI model? Naturally, if I'm putting together a battery of state of charge estimator, I want my outputs to be in between zero and one. So I might have to have some saturation at the output, some saturation activation function, some sigmoid function or some clipped ReLU or something along those lines. If I want to bake in something like monotonicity, then I need to look into what is the adequate construction of the network that gives me that monotonicity guarantee. So it all comes down then to the performance. And do I have very strong performance requirements?
14:46If I do have very strong performance requirements as I deploy, I might not be able to choose some of the more advanced... Okay, advanced these days, it's perhaps an overuse of the word. I would perhaps... It is advanced for the microcontroller, perhaps, in some sense. But some of the recurrent networks, if I use a stacked LSTM, I could get very good performance. But this is accuracy performance or very good accuracy, but might not be able to get a good execution performance as I deploy. So in here, I might want to consider different strategies, different alternatives, and be able to use a framework that allows me to deal with all these complexities and ultimately train a model that satisfies my earlier requirements.
15:39So you mentioned this kind of V model and the W model. Can you talk through some of the specific steps within this model? Are there 10 or 100 or what's the scope of the model itself? I'll go with the W-shaped workflow first and see how it resonates. It all starts with requirements that are allocated to the AI model, sometimes in the context of the AeroDev industry, the AI constituent. And here we need to make sure that the steps that follow the AI constituent follow this W diagram, whereas the non-machine learning related components follow the V cycle. So we have, as a first stage, to go through the data management process here, and we need to make sure that we are allocating the right amount of data, that we have a good representation of the operational design domain we're dealing with.
16:48We need to make sure that we are able to understand what are the edge cases with respect to that operational design domain. What are the boundaries under which my system would have to perform? So the next step involves the learning process management. And in here, what would happen is that you would typically have to go through an understanding of what is the most suitable AI model for your needs, for your problem. Is it a neural network? Is it a support vector machine? Is it a deep neural network? Is it a transformer convolution type of architecture? Is it a yellow network? What is it that you're trying to solve?
17:26and what is the most suitable architecture or set of architectures for your problem. And you would typically, in this learning process management step, deal with looking at different hyperparameters for your system. You would have to experiment with different models and different architectures, right? You ultimately next go to the model training and here you would need to get access to the right hardware. You would need to, it's quite obvious, but you would perhaps need to use a GPU or a set of GPUs or cluster or cloud, whatever it may be, to come up with a trained model. The next step is going up in the W.
18:12So it's kind of the peak of the W. So we're right at the center of the W here. And that involves this learning process verification stage where you want to taste your model against an independent test set. That's as obvious as that. But depending on the industry you might have, or the criticality level, you might have additional requirements to go through. Perhaps you have specific requirements that relate to robustness of the system or interpretability. you might have requirements to test specific scenarios or even so you know make sure that the boundary conditions or the boundary or the edge cases of the problem are taken care of in this phase and then you would move on to the right hand side of a W.
19:09The right hand side of the W deals with the implementation and integrating the AI model within the larger system under design. And so the first stage will be to do the actual implementation. And here you would see that you might see some differences between the model that you use for development and the model that you're using as you're implementing the final solution. So if your intent is to run this AI model into a resource-constrained processor, like a microprocessor or an FPGA or something along those lines, you would have to actually implement that AI model. If it's a neural network, you would have to implement the neural network for that specific target hardware.
20:02And so in that regard, what we do in one of our value propositions is to be able to generate code directly from MATLAB or from models that have been imported into MATLAB, if those are PyTorch or TensorFlow models. And then ultimately, you know, get the source code that you want to deploy into these resource constraint processors. So perhaps it's C or C++, or it's CUDA code if you have access to a Jetson board, or it could be HDL code if you're running it or you're intending to run it on an FPGA. Now, it's very important that now we have transformed the code, right? We had a model that had been developed in one framework.
20:50It could be PyTorch. It could be MATLAB. It could be TensorFlow. flow. And now we have a model that's implemented in C++, for instance. We need to make sure, and so we need to have specific requirements along this too, that we need to make sure that the translation, the code that we've translated now is also meeting all the previous requirements. And that by just translating the code into C++, we didn't introduce any unintended behavior. here. Then the last stages of the W diagram involve integrating the AI model, now the source code of this AI model, into a larger system under design. So if you are building, for instance, a system that is able to recognize stop signs or something along those lines, you would want to make sure that the model that is detecting those stop signs is well integrated with the rest of the system.
21:55Perhaps you would want to integrate it with some image, traditional computer vision tracker, and enhance the operational capability of this neural network, for instance, so that you don't have to detect the stop sign at every single frame, but instead detect it in a couple of frames and then track it all over, you would then want to make sure that it integrates with the other steps within the system. If the AI model is part of, let's say, an emergency braking system, you do want to make sure that it's well integrated into that and we do appropriate testing for the system level integration, making sure that different scenarios are tested and under control.
22:46Now, how often in these environments are you working with external specifications or regulations as opposed to design specifications? Very often. So it is frequent that our customers, especially in regulated industries, need to go through the process of certifying their code before it goes into their product. So, for instance, in the context of automotive or aerospace or medical devices, there are specific standards and guidelines that users have to go through. And interestingly, of course, now with the new government regulation coming up, the UAA Act and other regulation from other governments, we're seeing an increased interest from our users into this space and the VNV space, trying to understand what of the existing standards works for them and what doesn't work for them.
23:59And so new standards are being put together to cover for the gaps that the current traditional standards will not be able to address. When you're talking about these kind of verification and validation efforts, like how much of what customers are needing to do is formal proof of functionality or capability or some of these other aspects that are specified for the systems? So, yes, that's an interesting question. So definitely when it comes to formal methods and being able to actually prove correctness of the code itself, that's something that has been coming up frequently in traditional software development for these safety critical industries.
24:47Now, when it comes to AI, there is less of a regulation in that space, even though there are specific standards that are being put together for robustness of AI models and things like this. But it has less presence in the standards. But new standards in aerospace and automotive will be addressing this. and the basic idea behind it. I like how one of my colleagues in the development team puts this together. So when it comes to thinking about what the, let's say, what adversarial attacks are like, he kind of thinks about us in this way. Adversarial attacks are like finding bugs, right? And you're, with respect to your traditional software.
25:42and they allow you to reveal vulnerabilities that can be exploited. So for instance, if you think about traditional software best practices, you want to make sure that there's no integer overflow, right? That is definitely a bug in the code. And meanwhile, employing formal methods for AI is like providing or making sure that you're able to prove the absence of errors. So can I guarantee that there are no overflows in my code with respect to the traditional software engineering example? And so in the context of AI models, we're seeing a lot of interest in being able to offer a mathematical proof, a mathematical guarantee of the robustness of these systems, these models.
26:35So for instance, as you may know, the robustness is one of the main concerns that people have when deploying neural networks in safety critical situations. It's been shown and it's been shown publicly in many occasions that small input perturbations to the image can actually, or to the input can actually change the output significantly. And that's not good, right? We want to make sure that we have a model that we can trust. And by changing a single pixel, it doesn't look like so. So in this space, based upon the interest we were receiving from the field, we worked towards putting together a specific library to address this.
27:23We called it DeepLining Toolbox Verification Library. And it basically allows our users to test the robustness of this deep learning networks or test the stability of these deep learning networks. And the way it does it is using the formal verification methods that you were mentioning. So the idea is that we're able to mathematically prove, and it's a prove in the mathematical sense, that for the input space that you're perturbing, there is no change to the output to the network. So in the context of having image-based applications, which I think are the easiest to think about as we put together examples, you can think about some driving context or some situation where you would want to change or you would want to test the AI model against changes to brightness.
28:30for instance, right? And of course, you can create multiple tests and have a unit testing framework and change, you know, a couple of pixels or change the brightness to the entire image, or, you know, come up with thousands of tests and make sure that the AI model is able to capture that and is able to predict the right response. But that's not very smart. And because you would have to test perhaps an infinite number of images if you wanted to have a good understanding of how the model behaves. Instead, formal methods using abstract interpretation offer this approach where you can have a formal proof of the correctness of the system.
29:16And so the idea behind it is that you can have an input image in your test set. Let's put that concrete example. and define what a good perturbation for that input image could be. In the context of brightness, you might be changing pixels up, right? And you define what is the boundary or what is the bounds of that perturbation. And instead of actually doing individual tests, where you're actually, in fact, modifying the image, the original image and running it through the network, you define a volume or a polytope where you basically include all possible images within that contemplate this perturbation.
30:04And so what you would do is instead of running each image individually through the network, you would pass the entire volume through the network. And of course, that requires some transformation to the layers that you're dealing with. You need to rewrite the layers in a very specific way. They're called abstract transformers, but it's different to the transformers that we're all excited about in the GNI space. But you would have to rewrite the layers so that they're able to deal with these polytopes or volumes. And then as the volumes go through the different layers of the network, whether it's a fully connected or a dense layer, or it's perhaps a conf layer, and so on, it gets transformed.
30:53The volume will, the polytope will be transformed up to the point where we get to the output. If the output of the network doesn't change, so we know that the entire volume is on the same side as the image for which we wanted to do the prediction in the first place, then we get the properties verified. And so this is a very strong argument because you can formally guarantee that the entire set of images within that original volume will not change their output if you change some pixels within those boundaries. It's a very strong argument. You said on the same side, is that meaning in the context of like a classification boundary?
31:40Perhaps, yeah. In the context of a classification example, you would think of the decision boundary and you would say, right, so if my output volume, I'm trying to go with simple terms instead of polytope here, but if my output volume is to the same side of the decision boundary where I was getting my original prediction, then every image within that volume will get the same class, basically. Of course, that's not always the case. And there are some situations where you might get violations because the volume is crossing the boundary or perhaps on top of the boundary. And so you're getting some situations where you cannot always prove these properties, but you would have the ability to now go on a fine grained volume and try to better understand within your operational design domain, within the images that you're dealing with, how is it really performing, right?
32:44Is it perhaps you need to train differently because your model is not trained in a robust fashion, and that gives you some ideas into how to take the next steps. So this W diagram is far from being linear, as you can see, right? We're going to have to go back and forth multiple times, And that's part of the process. And so this model that you've created where you've substituted your traditional transformer layers with, I think you said abstract transformer layers. um are you then taking this abstract transformer and that's what you're putting into production or is it a modeling tool that you're using to determine uh verification but then you're using the you know the original it's basically a modeling tool uh we we need to make sure that we can have a way to verify the correctness of the network.
33:51And so for that matter, because the network is not able to work with bounds, we have to transform the way the network operates. And so it's very important that these layers are developed in the right way. There are, of course, some limitations. There are some layers that the research is not there yet. Yeah, that's kind of the direction I was going. Like once you introduce another model, then you have to do verification and validation of that unless there's like you can rely on a mathematical relationship between your kind of original transformer and these abstract transformers that assures you that that, you know, that other model isn't introducing its own issues.
Read the full transcript
34:40Yeah. And so, of course, like with a lot of things that we see in AI, there's no one size fits all here. And sometimes the formal methods will not satisfy the needs that you have from a verification standpoint. And you need to go with other approaches. You need to go with scenario-based testing or Monte Carlo testing. You would perhaps need to make, this is something that you would do anyways, right? Have some sort of safety monitor or runtime monitor sitting side by side with your network. So there are additional techniques out there to make sure that the AI models are behaving and do not go crazy as we deploy them to the safety or to the embedded systems.
35:28I think at the end of the day, it's important to know that it's not just about building a smart system. It's also making sure that it aligns with the values and the safety and that it fits the intended design. You know, thinking about kind of testing your way through these scenarios, I'm thinking of kind of an analogy to traditional code coverage metrics. Are there ways that folks try to extend that analogy to AI systems? Yes, yes, there have been. And it's a really interesting area of research. There is a technique called neuron coverage that has the basic idea to make sure that as you train your AI model, you want to make sure that you are covering enough neurons in the neural network.
36:32kind of the same idea as with the code coverage approach, right? If I have sufficient tests from my traditional code, I want to make sure that as I run my tests, I get 100 % coverage. That's kind of the ideal situation to be in, right? In a similar situation, moving to the AI space, you would want to have something along the same lines where given my test data, which in this case, it's perhaps images or it's perhaps signal data, I want to be able to run that through my neural network and make sure that I get sufficient neuron coverage so that we have some guarantees on the network itself. There have been some interesting papers in this space.
37:21We've implemented our own neuron coverage for MATLAB. We released it on GitHub a couple of years ago. So it's exciting to see how these techniques are being used first in research. They turn over to tool vendors like MathWorks. And, you know, ultimately there's a way to actually use them in the context of what our industry customers are working on. How confident are we that covering all the neurons actually covers the edge cases that we care about? it seems like a complex relationship potentially. It is indeed. And in fact, this is something that has been brought up at various working groups that I'm part of.
38:12I'm working with EuroK and SAE on this new standard for AI and aviation that will be called ARP 6983. And this has come up a couple of times. We are seeing that it's eventually challenging to connect neuron coverage to some of the objectives that one might have and requirements that one might have. Nevertheless, for some simple examples, simple shallow networks and applications, it is definitely worth a try. And I'm thinking particularly in the context of replacing large lookup tables, for instance. That's something that a lot of people have been interested in. You have a very high fidelity first principles model.
39:05And then this perhaps is coming from a third party simulation tool. And we have customers interested in integrating that high fidelity first principles full order. component into Simulink for system-level simulation. And it happens that, of course, as you do co-simulation with this third-party tool, it becomes tricky and it becomes computationally expensive. And if you are developing a controller to go along that plant, it's perhaps challenging to do this very effectively and to rapidly iterate over the design. So what many customers have been interested in is this these reduced auto modeling techniques or the surrogate models and definitely the easiest form of a surrogate model could be a lookup table um the challenging aspect with a lookup table is that if you would like to deploy this in in some way and you want to have a very accurate model then you have a very large lookup table and so memory footprint is going to be large This is something that we can mitigate with a shallow neural network, for instance.
40:19And in the context of this shallow neural network, you can perhaps easily look into the neuron coverage and have a good understanding of how perhaps adding an additional class to the test set or to the training data would actually modify the neuron coverage itself. For larger problems, though, I think that's quite challenging and perhaps it's still an open research problem. In talking about the polytopes and that whole conversation, we were talking about transformers. Are you seeing folks using transformer-based architectures in these safety-critical systems? Or is that more aspirational in kind of the direction that folks are going?
41:09Wow, great question.
41:14Yes and no. Okay, so let me go into that. Transformers, it's really exciting, right? Definitely a very exciting area. We have seen tremendous advancements over the course of the last six years or so in this space. when it comes to safety critical industries and the use of AI within safety critical situations, transformers could be yet another mechanism or tooling to be used. Now, I think it's important to know that we can use definitely transformers in the context of LLMs and Gen AI that's tremendously popular out there these days. We can also use transformers in the context of time series modeling and build a transformer instead of use a transformer architecture instead of an LSTM architecture or a GRU architecture or perhaps some shallow network that has some inputs from the past, which is how we were doing machine learning 20 years ago.
42:27And so in that sense, yes, we're seeing transformers definitely being used in the context of embedded AI. And it's specifically the use of smaller types of transformer architectures for time series modeling is the sweet spot that we're seeing. That said, of course, the use of larger transformers or LLMs or visual transformers, it's also relevant perhaps for other types of applications, not as much on embedded AI, but we're seeing an interest, of course, in gen AI, both for text and image. yeah it does seem like um you know trying to apply some of the things we've talked about to llm-based systems is a whole nother level of complexity yes it is yes uh it is definitely um in fact you know if we if we look into some of the latest competitions on neural network verification uh from from the past year there is this competition called vnn comp that happens on a yearly basis, it deals with, you know, you have different benchmarks and you have the ability to run your code through different situations and such.
43:51We have seen there for the first time, I believe it was last year, a benchmark on object detection for YOLO network architectures. So, yes, perhaps to answer your question, I believe that this is an interesting area of research. It will come. We're not there yet. But yeah, definitely these days we're seeing players focusing on GNI and looking into other ways to test the LLMs. But that said, perhaps formal methods are at this point in time more of a future thinking technique. thinking technique. You mentioned benchmarks. In academia, the kind of tendency is to standardize the problem, collect a lot of data, publish a benchmark, and then lots of folks share that same benchmark and use it to create comparisons.
44:55In industry, the problems are more one-off to a particular product or project,
45:05is there an analog for benchmarking or how do folks in industry in these kind of safety-critical environments, have they come up with benchmarking analogs for these kinds of systems? It's challenging. It's a challenging area. There are some benchmarks out there for AI verification that are publicly available. So these VNN-com benchmarks not only include participation from academia, but also research groups within corporate labs or industry. So they are definitely part of these benchmarks. Now, that said, certification bodies like EASA, for instance, are putting together concept papers, they're putting together white papers, putting together materials for people to look at.
45:58One interesting piece of information or one interesting collateral that was put together in the last year or so was the MLLEAP project. MLLEAP stands for Machine Learning Approval Process, if I'm not mistaken. And it's basically, you can think of it as a Rosetta Stone of these different steps in this W diagram. So what would be the different techniques and methods that one would use for AI verification? What would be the different steps for data validation? What methods are out there? What are the tools that one can use? So perhaps not benchmarks per se, but it does provide a good level of detail that industry players would want to go through, specifically if they're working in aerospace.
46:54And then as new standards get released, there will be supplementary materials that give you the ability to look into more specific use cases. So we worked on one of these use cases in the past year. That was the presentation I gave at NURBS. And that was basically making a case for building a runway sign classifier for the runway at the airport. So as a pilot, you would want to have some sort of assistant that gives you information about the signs that you're seeing in the runway. In a nutshell, when we look into when we're at the runway, I don't know if you've noticed, but there are two types of signs, pretty much.
47:44There are red signs that are mandatory signs, and then they are yellow signs that are informational signs. Now we want to make sure that that information gets processed and gets handled over to the pilot. It's all about having a complete use case that users can run through as they learn more about how AI integrates into these type of systems and applications. So in this sign example the The primary gaps are around explainability and those other more qualitative types of metrics as opposed to being able to accurately identify the signs and classify them. The intent that we had is being able to take an incremental approach towards certification.
48:41I'm not sure to what extent you're familiar with standards in aviation. The software standard is called... Not very. The software standard is called DL-178C, right? And the basic principle is that there are different criticality levels going from E, which is no criticality, D, low criticality, all the way to A, which is highly critical system. And so we worked on this end-to-end workflow to make sure that we were able to provide a blueprint of all the steps in the W diagram and give the user the ability to replicate this for their own use case. We started with a DAL-D, low-criticality application.
49:32So the idea, like I said, being able to build a runway sign classifier that would perhaps show up in your head-up display in the cockpit. And it's really not integrated within any of the controls or any of the control system in the aircraft. and this allowed us to basically reduce the number of objectives that needed to be satisfied from the ai perspective and so we could use the ai as a black box model in some way but we didn't want to stop there and we wanted to see if we could target the next level of criticality and so to do that we used a technique which is very frequent in this space which is called architectural mitigation.
50:20So basically, architectural mitigation deals with the fact that if you have two DAL-D components sitting side by each other, you can actually target a DAL-C certification if these two components are dissimilar from each other. Like they have been built with this principle in mind. So think about it in this way. Meaning redundancy, is that the idea? It is redundancy, but at the same time, It's redundancy with a dissimilarity in mind. So let's say you and I, we're working on AI. We get a data set. I get a data set. I label it. You get a different data set and you label it too. Perhaps we're using different labeling tools.
51:04I come up with an AI architecture. Perhaps it's a faster R-CNN. You use a YOLO architecture. I train MATLAB. You train PyTorch. There is independence of the tooling, there's independence of the data set, the modeling, and the modeling framework. And the idea is that if ultimately these two different components agree on the output, it is an indication that the output is the intended one in some way. So basically using this approach called architectural mitigation, we're able to certify a DALC component. And that's where we are at this point. We're able to give our users blueprints and templates so that they can go through the certification process themselves or collaborate with us if they need to and go through the certification process for this DO178C.
52:00And it's interesting that standards are not out there yet, but thanks to the research from the Technical University of Munich, from the working group, we have been able to mitigate these gaps in the existing standard and be able to put something together to certify these systems today. I forget if you mentioned if the data sets were independent. Are you collecting independent data sets for each of these efforts? Yes. So in this particular case, in this particular example, yeah, there's data independence, there's human independence, there's framework independence, there's architecture independence on the network.
52:45Yeah. you would want to bring as much independence as possible to make sure that the design is what it should be. The basic idea is that each of these systems independently, due to the statistical nature of the models, could only achieve a D-level certification, but with them acting independently or kind of this architectural mitigation, you're able to certify to the higher level. That's exactly right. And yeah, it's an interesting concept, interesting idea. Of course, you will also need, perhaps to complement this, you would also need some pre-processing step and post-processing step, right? These two, you know, DALT neural networks will not run in isolation.
53:42They would need something to go along with, right? And so these pre-processing and post-processing steps, especially the post-processing step is very important because that's how we are doing the safety monitoring, right? They have to be DAL-C. So these systems cannot be DAL-D, right? So they need to be DAL-C. They need to be of the higher level of criticality. That post-processing, for example, would be, you know, some kind of quorum mechanism or or whatever is looking to make sure to correlate the two decisions. Yeah, exactly. That's exactly right. So some quorum mechanism, perhaps you want to have some sort of voting system in the context of two networks.
54:22It's challenging. So in the context of two networks, you would want to do something perhaps along the lines of intersection over union of the output. So if the intersection over union of the bounding box is above, let's say, 80%, then you consider the output to be correct, right? That's one of the metrics we're looked into, for instance. So you would have to have the safety monitoring net, not a network, but safety monitoring component after the dial-d components. So, so far we've talked primarily about kind of existing machine learning models, neural network architectures, and maybe varying them a little bit for testing in the case of these abstract transformers.
55:13But are there efforts to build models that are more suitable for these types of safety critical systems kind of inherently? Right. So, yeah, definitely. So the space where we are, we've seen some interesting progress, perhaps not as much consolidated research, but it is very appealing to me. it's the constrained deep learning area. Constrained deep learning, it's quite simple. So if you're familiar with physics in foreign machine learning, it's a similar concept, but has a little twist. It's basically an approach to train a deep neural network by incorporating domain-specific constraints in the learning process.
55:58So the idea is, for instance, that you integrate a constraint into the construction of the network that guarantees a desirable property of the network. For instance, in the context of some of the examples we were mentioning earlier, battery state of charge, right? So you could perhaps guarantee that the network is monotonic or the output of the network is monotonic by design. In battery state of charge estimation, you would like a good property could be to have the network to be monotonically increasing as the battery is charging, right? Or monotonically decreasing as the battery is off charge.
56:40There are other examples. So for instance, you would be not only interested in making sure that the network is monotonic, perhaps you want the network to be bounded, right? Or the network to be convex. I like the convex example because it's quite relatable. So in simple terms, you can think of that. So what is a convex function? So if we look back into what it is, so basically a function is convex. If you are able to, within the function, connect any two points and the line segment connecting the two points,
57:30lies above the curve, or the curve lies below on the curve itself. So you can think of a function that has a U-shaped, that's a convex function. And now this is perhaps interesting in the context of optimization problems, because convex functions have only one minimum. And so it makes it easier to find the best solution without really getting stuck in local minima. So for instance, when minimizing a convex objective function numerically, you don't have a risk of getting stuck. And having a network that is a convex function is very useful in some fields. So what I mean by a network that is a convex function is that, okay, so the output of the network is convex, right?
58:25So for instance, in model predictive control, where the goal is you perhaps have an optimization problem, you would want a neural network to predict, let's say, energy consumption, right? and you can guarantee that the network is convex if you build the network in some specific way. Building a convex network starts with some basic operations. So if you use convolutions or fully connected dense layers, these operations are basically convex because linear functions always satisfy convexity. So if you add convex functions, the output is convex. Now, the problem is when you add the nonlinear activation functions.
59:17So with ReLi functions, for instance. So what you do is to make sure that whenever you introduce activation functions, you keep convexity in place. you need to use activation functions that are both non-decreasing and are convex. And so the hit-reli function that I was mentioning is a good example. It fits well here. However, a function that is like a sigmoid function, it doesn't preserve convexity, right? Because they are convex. They aren't really convex across the entire space. So that is an interesting problem, right? So, for instance, let's say, try to think about a problem here where, yeah, so if you're trying to perhaps estimate energy consumption or fuel consumption, right, you might have some input to the network.
1:00:23the input to the network might be something like perhaps it's the steering angle, the speed, or the force that you're putting on the throttle at some point in time. If you are able to create a convex function, you would be able to minimize the output of the network given that the function is convex and has only one global minima. and so that i kind of see this as a very appealing space because um instead of having to formally prove a property of a neural network you're you're baking in that property into the design of the network itself um by making sure that you know you're you're using an aggregation of of convex uh functions or layers as you put them together and then finally when you want to train the convex network, of course, there's something that you need to take care into there.
1:01:23You need to make sure that you force positive constraints to the weights, right? So basically, the idea is that with each training iteration, you need to apply a non-negative constraint similar to a Raleigh function to kind of keep the weights positive. If the weights are kept to be positive, then definitely you will be training a convex network. There might be some challenges, of course, there's no free lunch, which is, for instance, in the context of classification problems, softmax functions, right, don't preserve convexity. And so when you work with classification problems, we cannot use the softmax layer, and so we lose the probabilistic interpretation of the outputs.
1:02:12But that's perhaps something that is okay to have as long as we're able to train the convex network. It's okay to have because there are acceptable alternatives. Yeah. At the end of the day, you're going to lose that probabilistic interpretation of the output, but you will still be able to determine what the right output is. So you basically look into the last fully connected layer and you have an idea of what is the class, what is the right class for the problem itself. Assuming that we can do this, like in this, you know, what we've been talking about here using convex neural networks as opposed to more traditional alternatives to the architecture, you know, kind of pulling on that no free launch thread.
1:03:03Like, what is the, what's the catch? Like, why aren't we just using those? You mentioned the probabilistic interpretation for classification. Are there other areas where it's, you know, is it performed equally well, or do you have to give up some degree of performance? Are there other non-desirable properties? What do you think there? Yeah, there's definitely a trade-off, right? You got it right. Right. I think the most significant trade-off is that convergence slows down. And so because you're introducing, so when you're basically casting all the weights to be positive, you're introducing some noise into the training process that doesn't go unnoticed, right?
1:03:54And so convergence definitely slows down. Indeed, it might require a more complex network architecture to be able to deal with the nonlinearities in the dataset, the nonlinearities present in the dataset. But that said, it is perhaps a good trade-off to look into. I mean, the alternative, in the context of, say, critical scenarios and situations that you might have to deal with, the alternative is worse, perhaps. So it's part of the techniques and the methods that one should have available or at their disposal when putting together a model and training it. You should think about not only performance-wise or not only accuracy-wise what the BIS model is, but perhaps also to what are some of the other good properties that my system should have and try to figure out if there's an opportunity or there's a way to build the model in such a way.
1:05:07because ultimately we will get a more reliable model, right? I guess I'm trying to get past the feeling that it's harder than you're saying. Like, you know, why isn't this approach more dominant? All of the issues that you're, you know, that we've been talking about in the context of safety critical systems, like people want the same things for non-safety critical systems, like robustness and explainability and these kinds of things. is there a lot of research? I guess, why isn't there more research being done into using these convex neural networks? Why aren't we using them more? I'm just trying to kind of wrap my head around that.
1:05:55Yeah, no, it's a valid question that I've gone through myself, I guess, as well, at some point. It's just like we glommed onto something that worked and then we're running with it and kind of not putting enough energy to these things that are, you know, a little bit out of that strike zone or are there, you know, drawbacks that are, you know, or, you know, for example, you mentioned it takes longer to converge. Does that mean it just takes a lot longer to converge or, you know, it takes longer and it only converges one out of 10 times? So in my experience, yes, it does take longer to converge.
1:06:37It's not a drastic change in the process itself. So I've seen, yeah, going from toy examples to more complicated examples, convergence or the time to convergence increases somewhere between 20 % to 50%. So it's definitely impacting the training process in a good way. But because ultimately we're, yes, we're taking longer to train, but we are getting a network that better represents what we're after, right? I mean, it may seem simple, but I guess with the physics-informed neural networks, we've been more often trying to bake that physics into the networks themselves. This isn't really necessarily...
1:07:35Well, there's some form of physics that we're adding into the design of the network, but it's more in the sense of thinking about the safety or thinking about the reliability of the models. When we did our research and we put together this GitHub repo on constrained deep learning, perhaps a year ago or so, we noted that, yes, there is research out there. There's good research, but it's very scattered and it's hard to find a good... Sometimes there's a good central point for looking into these different techniques. And there are libraries out there. We didn't find a good set of libraries that would cover the space of constrained deep learning, monotonicity, boundedness, robustness, lipsets, continuity, all these good properties that you want a neural network to have.
1:08:37For some examples, we didn't find a good source, a good repo. And so we decided to build one, right? So I think this is interesting. Is it the case maybe that I'm not catching on to use case constraints in the sense of do you only get these properties if your data or problem match a particular pattern? like are linear or convex themselves in some way, and most of the things we want to do don't have those properties? Or is it just that it takes a little bit longer to train, but then you have a network that it just doesn't, it doesn't seem like, again, I'm trying to just poke at like, why are we doing this more?
1:09:28Like, for example, like I'm, I'm thinking of it a little bit analogous to, you know, another type of like gauge equivariance and stuff like that. And like those networks have great properties, but the data has to, you know, fit in, you know, has to exhibit this type of symmetry that, you know, not doesn't apply to all problems. Is there a similar kind of thing here where, you know, it works great, but you have a narrow band of application. So, yeah, I think it's a good analogy. So what I'm seeing is that, um,
1:10:07it's so being able to one of the things we noted is when we put this together is that being able to come up with a tool that allows you to build these type of or bake these type of properties into the networks themselves it's it's hard to find because you need to build this on a case-by-case, right? So for instance, if we are putting together like a convex network, we want to put together a convex network and I already have a network that I've trained, how do I make that network convex, right? It's not obvious, right? Yes, if of course, in principle, the sum of convex outputs would be convex and forcing this positive weights would allow it to continue having convexity and making sure that you don't use things like Sigmoid or Softmax.
1:11:03And that will guarantee that the output continues to be convex. But the thing that I haven't seen out there and which triggered our interest in this space and putting together some tooling around it was being able to say, all right, build me a convex network with this number of features, with this number of hidden layers in the first layer, second layer, third layer, and so on. And that gives you an architecture right away that you can train with the data that you have available. I think the gap has been that there's been some very good research for these different areas, but it's been challenging to come up with a framework that captures all this together into a consumable way for the end user, the machine learner practitioner.
1:12:00Instead, I believe that what's been happening is that perhaps as a machine learning or a data scientist, you've had to, on a case by case, every time you were thinking about putting together these networks, you would have to build a network with these ideas in mind. Maybe the question is, like is there are there inherent reasons why you know sometime in the future you know i couldn't have the opportunity within my you know on hugging face or someplace to download you know yellow v12 and yellow v12 convex like is it just not that simple and you have to you know manually convert it for every particular use case and there can never be you know a library of convex models like we have for non-convex models uh or are we just not there yet but we could be there if we spent enough time on it yeah i would i would think so i think i think we're not there yet but but definitely we could be and and like you said i think well i'm yolo v12 well who knows who uh what that would look like, but perhaps, well, it's not that far away, right?
1:13:22But definitely what would make sense is, okay, so look into the YOLO V12 architecture in that case into the future and see, okay, so what is making this network to be non-convex or what is making this network to be not having a specific property. And try to understand if we can add that property to the network itself by changing some of the weights or by changing some of the layers to the network, the activation layers. Perhaps we need some skip connections. Perhaps we need to connect the layers slightly different to make sure that this property is preserved by the sign or is baked by the sign into the network architecture.
1:14:09And then the whole story on AI verification becomes much simpler, right? Because you just have that property by design. And if you're able to come up with, you know, the accuracy results that your requirements are asking you to get, then it makes sense. to choose that as a preferred option. But yeah, this is an interesting area of research too. I think there's definitely more that needs to happen. We will continue to make progress in this space and hopefully see some good advancements that allow users not to have to spend a lot of time in the verification phase. Well, Lucas, thanks so much for jumping on and going through some of your experiences with validation, verification and safety critical AI.
1:15:13Yeah, thanks so much. Thanks for having me. Thank you.
1:15:30you
From the publisher
Today, we're joined by Lucas García, principal product manager for deep learning at MathWorks to discuss incorporating ML models into safety-critical systems. We begin by exploring the critical role of verification and validation (V&V) in these applications. We review the popular V-model for engineering critical systems and then dig into the “W” adaptation that’s been proposed for incorporating ML models. Next, we discuss the complexities of applying deep learning neural networks in safety-critical applications using the aviation industry as an example, and talk through the importance of factors such as data quality, model stability, robustness, interpretability, and accuracy. We also explore formal verification methods, abstract transformer layers, transformer-based architectures, and the application of various software testing techniques. Lucas also introduces the field of constrained deep learning and convex neural networks and its benefits and trade-offs.
The complete show notes for this episode can be found at https://twimlai.com/go/705.




