In short
Podcast Summary: Generative Now | Episode with Adam Wenchel
Episode Overview Title: Adam Wenchel: How Arthur AI is Making LLMs Trustworthy Host: Michael Mignano Guest: Adam Wenchel, Co-Founder and CEO of Arthur AI Description: In this episode, Adam Wenchel discusses the challenges of ensuring the trustworthiness of AI, particularly focusing on large language models (LLMs) and the solutions provided by Arthur AI, including firewalls, validation tools, and benchmarks.
---
Episode Chapters
- (00:00) Intro to Adam Wenchel
- (05:31) Adam's background and early work in AI at DARPA
- (08:39) Transition from cybersecurity to Capital One
- (14:20) Arthur's role in AI transparency
- (21:45) Impact of ChatGPT on Arthur AI's direction
- (26:18) Guardrails for trustworthy AI
- (31:46) Discussion on competition in LLMs
- (37:00) Observation of the AI landscape post-ChatGPT
- (44:22) Arthur AI's firewall and hallucination management
- (47:11) Current state of AI regulation
- (51:22) Predictions for the future of AI
- (52:39) Hiring opportunities at Arthur AI
---
Key Concepts
Trustworthiness in AI
- Hedging and Hallucinations: These are critical concerns in LLMs that Arthur AI aims to address, providing assurance to users in sectors like military and finance.
Adam Wenchel’s Background
- Experience with AI: Adam's journey includes work at DARPA and leading AI at Capital One, where he scaled an AI team and developed models impacting millions of consumers.
Arthur's Mission
- AI Monitoring and Measurement: Arthur AI focuses on delivering transparency in AI decision-making, ensuring models function effectively and justifiably.
Transition to Generative AI
- Impact of ChatGPT: The advent of ChatGPT reshaped Arthur AI's roadmap, pushing them to develop new tools and solutions to address emerging challenges with LLMs.
Guardrails for Trustworthy AI
- Implementation of Firewalls: Arthur AI's solutions include firewalls to intercept hallucinations and ensure the accuracy of AI outputs, particularly in sensitive applications.
The Future of LLMs
- Competitive Landscape: Discussion on which companies will dominate the LLM market, with an emphasis on the balance between open-source and proprietary models.
Regulatory Environment
- Need for Regulation: Adam discusses the importance of light yet effective regulation, emphasizing accountability for organizations deploying AI models.
Predictions for AI's Future
- Continued Evolution: Predictions include a mix of compelling applications and potential mishaps as AI matures, particularly in light of the upcoming election.
---
Key Takeaways
- Importance of Explainability: Providing insights into model decisions is crucial for trust, especially in high-stakes environments.
- AI's Rapid Development: The landscape of AI is evolving swiftly, necessitating constant adaptation and learning from professionals in the field.
- Diverse AI Ecosystem: A blend of proprietary and open-source models will likely shape the future of AI technology.
- Hiring Opportunities: Arthur AI is actively seeking talent to expand its team, highlighting the growing demand for skilled professionals in the AI sector.
---
Conclusion The episode with Adam Wenchel underscores the significance of creating trustworthy AI systems. With rapid advancements in generative AI, organizations must navigate challenges like hallucinations and regulatory requirements while fostering transparency. Arthur AI is positioned to be a key player in ensuring the integrity of AI models across various applications.
---
Stay Connected
- Arthur AI: [Website](http://www.arthur.ai)
- Generative Now Podcast: Available on various platforms including Apple Podcasts and Spotify.
- Follow Lightspeed: [Twitter](https://twitter.com/lightspeedvp), [LinkedIn](https://www.linkedin.com/company/lightspeed-venture-partners/)
---
*Note: The content discussed in this episode is not intended as legal or investment advice.*
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:04Hey, everyone, and welcome to Generative Now. This is a podcast where we talk to the builders who are creating the world's most exciting AI products and companies. I am Michael Magnano. I'm a partner at Lightspeed. And for today's episode, we're tackling a topic that people are talking about a lot lately, and that is making AI trustworthy. For this subject, we've got an expert for you, Adam Wenchel, co-founder and CEO of Arthur. Now, Adam's been studying AI for decades. He studied it in college, did AI for DARPA after college, And then he founded a startup, which he sold to Capital One, where he then ran AI for them before starting Arthur.
0:44So I really think you're going to enjoy this conversation. Check out my talk with Adam Wenchel, co-founder and CEO of Arthur. Thanks so much for doing this. Been looking forward to it for a while now. I love to start with these things and sort of helping our audience really, really understand the guest, who this person is, why we're talking to them. So why don't you give me your background? Tell me all about you. Yeah, absolutely. So Adam Wenshaw, CEO and co-founder at Arthur, which is an AI company. We focus on measuring and monitoring AI, including generative AI and LLMs, which is very much top of mind for a lot of people right now.
1:22But going back a little further, I actually started when the School for Computer Science, followed one of my professors over to DARPA, where I was doing AI research when I graduated in 99, building actually a lot of really cool agent-based systems that are in some ways very similar to what we've seen emerge in the last year. And then from there went into the startup world where I've been pretty much the whole time ever since. Founded my own startup back in 2013, which was acquired by Capital One a year and a half later, and joined Capital One. and shortly after I joined, they ended up asking me to start their AI team.
2:00So I started their AI team, first time working inside of a large company, and scaled that up to a little over 300 people. And so even though I've been doing AI for a while, it was the first time I had done it at that kind of scale, where you had tens of millions of consumers who were being affected by the performance of these models and who got credit and what was tagged as a fraudulent transaction and customer servicing and money laundering and all these different aspects. And obviously on the business side, hundreds of millions of dollars and more at stake in terms of making sure that these models made good decisions at all times.
2:33And of course, heavily regulated. So you had to justify everything you were doing to regulators. And so that was a real educational experience for me. And so I spent three years there and then left in 2018 and really started thinking a lot about some of the challenges we had there in terms of putting guardrails around models and making sure that we had sort of modern proactive tools for alerting and for monitoring so that we could find performance problems before they resulted in a hole in the balance sheet. And so that really was how I came to start Arthur with my co-founder, John Dickerson. And so we got going early 2019 and we've been building ever since then.
3:13Maybe just to go all the way back to studying computer science. You know, I don't know how old you are, but just by looking at you, I think you and I are probably roughly the same age. How did you decide you wanted to major in CS? Like, did you learn to code growing up before that? Or, you know, what led you there? And where'd you go, by the way? Yeah, absolutely. So I actually got into coding in like, I don't know, fifth or sixth grade. Our classroom had a couple Apple IIe's in the back of it and kind of got obsessed with them a little bit. And then we bought an Apple IIe at home. And actually, I think even before then, my father worked in the Navy and was in logistics.
3:46And so was responsible for bringing some large scale computer systems into the Navy to modernize their operations. This is back in like the 80s when it was kind of pretty revolutionary. And I guess it started actually, he had a buddy who had an Atari 800 that we would go over. he'd hang out with his buddy and I would sit there and program on. And so I was really into it like fifth, sixth, seventh grade. And then I kind of fell out of it. Didn't really, didn't, didn't really do much in high school. And then when I got back to college, actually started as a double E major because at the time, you know, the top companies in the world were chip companies, which in some ways that's like coming, coming back to, to happening again.
4:24But then, you know, just found myself really drawn to the computer science side. I was at the, worked at the campus radio station. That was kind of my big club at school at the University of Maryland. And at that time, we were like streaming audio was a brand new thing on the internet. And so we were one of the very first radio stations to stream 24 seven. And so I with the help of some alumni was got us kind of up and streaming and yeah, like probably one of the first first five kind of streaming radio stations, which was good because we only had a 10 watt transmitter. So if we wanted to be heard, we had to we had to think a little bit outside of the box.
4:58And, and so really just did a lot of programming and automation as part of that. And then just was like up at 2am one night, like kind of programming out of passion. And was like, why don't I just do this for for a major. And so that was, that's what I did. And started focusing in on AI at that time, which was, this was in the late 90s, was very out of fashion at the time. And, you know, lots of jokes like, oh, you're, you're an AI, you're great at predicting the past. And, and so did a lot of lisp, wore out parentheses keys on a few keyboards and then went from there, like I said, over to DARPA.
5:31I saw on your LinkedIn, you were doing AI work at DARPA 20 years ago. I mean, talk a little bit about that. How does the Department of Defense think about AI all the way back then when, you know, obviously like it's not the coolest thing in the world. And as you said, like a lot of people almost look down at it in a sense. Yeah. So I'll tell you, it's kind of amazing because the work we did there actually just in the last year has like reemerged in the the generative AI phase. And so we were doing the two programs I worked on were, one was called, first one was called Command Posts of the Future.
6:00And I was working with, you know, Stanford and Carnegie Mellon researchers and a bunch of really top minds. And a lot of what that involved was sort of multimodal type models, right, where you could take inputs that were video combined with gestures, like people circling things on screen, combined with sort of verbal commands, and assimilate that and then affect action, have a model that would have that would kind of action off of like different kinds of inputs coming in. And so, you know, obviously all that multimodal type of work in generative AI is very, very hot right now. And then the second one I worked on was called Control of Agent-Based Systems, led by a professor that I followed over, Jim Hendler, who's now the chair of the computer science department at RPI.
6:42And so that one, you know, agent-based systems are like really coming roaring back with a lot of the generative AI with agents. And so, you know, back then it was all a science project. Like if it worked in a demo, we were psyched. Like the idea of, you know, actually using it in the real world was was not, you know, it was clearly not there yet. But you could see the promise. Like if you could get to this point to where it worked, it would be really transformative. And so, you know, kind of a lot of that stuff sort of fell off the radar a little bit for for the intervening 20 years. But in the last year, it's been amazing to see a lot of those ideas actually start to work and actually start to, you know, create real world value for the first time.
7:19And now you're seeing all these companies pop up that are selling to the government and the Department of Defense that are very AI focused. And so that must be really fascinating for someone like you to see. Yeah, it is. It's great to see the government embracing it. I think there's a lot of reasons why it's really important that the government do it for competitiveness reasons and other reasons. And so it's been amazing to see. And I think there's a huge opportunity. I mean, we worked on a project with the Air Force on modernizing their supply chain. And just the complexity of keeping thousands of planes spread all across the world, a single engine, this one engine program that we were helping to modernize with AI, the engine had 29 ,000 parts.
8:05So you can imagine if you have these planes scattered all around the world operating in like sandy environments and things like that that cause high wear and a limited number of parts that you can stock, like where you place them, it's actually super impactful. And so the fact that this is in many ways still managed by tribal knowledge by humans is is there's a huge opportunity there to really kind of gain an advantage. And logistics is definitely what wins battles. So, you know, you see all the all the press around killer drones and things like that. But there's such bigger opportunities there.
8:38So, I mean, it's crazy to me that you were so ahead of the curve at the Department of Defense in college. And then I would say even when you sold your company to Capital One, I mean, again, like you're you're building, scaling and selling an AI company before anyone, you know, at least the mainstream really cares about AI. Talk a little bit about that and kind of the work you did at Capital One. And in what ways like was Capital One leveraging AI? I mean, I imagine it's probably for things like, I don't know. I don't know. I'm just this is probably like simple, like calculating credit scores and things like that.
9:11But I'm sure it's much, much more involved in that. Give us some examples of that journey and then what you worked on specifically at Capital One. Yeah. So I guess, you know, getting back into AI, I would say in like 2012, 2013, I was leading the engineering team at a company called Endgame, which was a cybersecurity company that actually was ultimately was acquired by Elastic and kind of the foundation of their security division now. And that team started out actually doing sort of actually on the offensive side, doing a lot of work for the intelligence community. And that was sort of leveraging that knowledge from building a lot of offensive cybersecurity tools into defensive capabilities.
9:52And one of the technologies that we really got good at was using machine learning to spot attacks and viruses, things like that. So that company was sort of going in a little bit of a different direction. So I left to start the company Annex that I started with some co-founders to really fully explore some of those ideas about applying AI to cybersecurity, which nowadays every single cybersecurity company claims to be using AI. But back in 2013, it was a little more novel, right? And so that's kind of how I sort of got, I would say, back into AI. And that was around the time when people were really starting to take advantage of GPUs and kind of some of the acceleration you could get from the parallel processing through them.
10:39And so these things that had kind of worked in the lab all of a sudden started to kind of work in the real world and just created a ton of possibilities. And so when I joined Capital One, originally the mandate was sort of rebuild our cybersecurity and technology analytics like Data Lake and kind of the analytics that go on top of that, which we did. And then around that time, one of the board members who was from Amazon started talking in the board meetings in early 2016 saying, hey, this AI thing is coming. You guys should get ahead of it and talking to the leadership there. And so they ended up tagging me to help them kind of get that going.
11:21And the applications are kind of endless. But if you're in a large consumer bank, the leverage points are you talk about credit decisioning. There's billions of dollars in getting credit decisioning right. Totally. And who you give a credit card to or an auto loan to or a mortgage to, how much you give them, what rate. All these variables are massively impactful.
11:45And it's not just like obviously a single model. Like there's dozens of models that go into that decision that look like maybe aspects of you as a consumer, but also analyze things like the macroeconomic trends, like their portfolio risk. Are we overweighted in a particular category of consumer or things like that? And so making sure that you understand not just the performance of a model, but of like a really complex system of models that are maintained by different teams and that have different goals and sort of optimizing is a really fascinating issue. And again, like similar theme there, even though the sophistication of the models and the model building increased a lot, a lot of the kind of management of them was still very much tribal knowledge type of things.
12:27And so, you know, there was a real opportunity to kind of get better at that and build better tools. because, you know, as the models, you move from like simple linear regression type models to more state of the art machine learning models, that sort of human oversight, manually reviewing process just is a tedious and B doesn't work. And so, you know, leveraging the same kind of sophistication on the oversight and guardrail side is critical. Totally. So talk about, you know, how did that experience end? I think you were there for a couple of years, right? You know, talk about that transition out of, out of Capital One and how you decided to start Arthur.
13:02Yeah. So I was there almost three years and, um, yeah, all my, all my startup friends were like, all my Capital One friends were like, Oh my God, you're leaving already. And, uh, all my startup friends were like, Oh my God, you lasted three years at a large company. So, uh, it's the way it goes. But, um, you know, I took a couple months off during this, uh, left in like July and, and, um, spent the rest of the summer just hanging out with the kids, um, up in Maine and get some time out in the outdoors and then really got back to when school started and got back to work and just started saying like, yeah, during that time that was kind of thinking through like, hey, I think there's a real need in the marketplace for this kind of observability and the same kind of visibility you get from, say, like a Datadog or Grafana or Splunk into a lot of your tech operations.
13:46How do you bring that into the data science world? Because they just didn't have that kind of operational maturity at all. and with the rise in importance of AI and AI starting to underpin more and more key leverage points in these businesses, that ends up actually becoming really existentially important because, you know, now these models are starting to more and more drive the P &L of your business. And so if you don't know what they're doing, then you're kind of flying blind, right? And so, yeah, that was kind of what got going. And then we, you know, formally incorporated in March of 2019, raised around and then just have been building since then.
14:20So let's unpack this a little bit. So, you know, I understand that Arthur helps with observability and security and things like that of large language models. And by the way, correct me if I'm getting any of this wrong. But back in 2019, I mean, so many businesses were not yet really thinking about large language models. I think, you know, I think that, you know, the paper had only recently been written. So so talk talk a little bit about what Arthur was doing for ML models back then and sort of how it's changed to now what's happening today. Yeah. So you're right. Back prior to a year ago was, you know, pre pre let's say the current generative AI era.
15:00Yeah. And so we were focused a lot on, you know, what we'll call traditional AI, which is classification and regression models, computer vision, NLP. a lot of the, even with the computer vision, NLP and tabular data, a lot of them use for either classifying things or a regression model, which is where you give like a numerical output. And so whether it's like a model, should I give this person a credit card or not? Or in a supply chain context, should I order, how many of these parts should I have on inventory? What's the right number? Should I order more? How likely are they to go on back order and run out?
15:33Answering those kind of more traditional questions that models have been used to predict. Yeah, making sure that you get that right. So it's a bunch of things like, you know, false positives, false negatives, precision, recall, accuracy. There's a whole suite of metrics that you can use to track those. And then also on top of that, providing some really important oversight beyond performance around what a explainability. So like, if you don't order parts, and a critical defense system gets taken offline, and you get hauled in front of Congress, like you need to be able to explain why the system told you not to order more or why these planes were grounded.
16:07And so providing that kind of level of transparency and explainability into the operation, into the model decisioning. So you can, for any given decision, you can say like why the model made that decision. And then the other one is around bias. There's a lot of bias in the world. In a lot of these systems, you're automating human decision-making, which because it's human, has some degree of bias in it inherently. And so if you're not careful, you end up automating it. But you also have an opportunity to kind of measure it and correct for it to some degree, which is which is exciting. Got it. So basically, these models, you know, even pre LLM, these things are obviously making countless decisions inside of the models, countless calculations that we as humans can't really understand unless we sort of peel the whole thing back.
16:55But then then we would have to do this millions and millions of times. So Arthur is basically digging in and helping us understand what's actually happening in the model. So then we can explain it back to ourselves, you know, like you said, in Congress or for our own auditing purposes. Is that is that a way to think about it? Absolutely. Yeah. OK. And you can also do things like look for is it keying off the right. You know, you can look at it at the aggregate level as well to understand, like, which which features you're putting in are important, which ones don't really matter. like those kinds of questions.
17:26Got it. So, okay, this totally makes sense. And, you know, I could totally see this being useful. Um, you know, my, my, my previous job, uh, at Spotify, Spotify, obviously for years has been leveraging machine learning to do algorithmic content discovery and personalization. I can totally see you using something like this to better understand how it's making the decisions that it's making. Is that a way to think about it? Absolutely. Yeah. And so you can, you know, you can understand like, why was I recommended this podcast or this song. And even, you know, whether it's to the data scientist consuming it or to the people being who have to like, you know, action off of that, like click on something to listen to.
18:03And maybe, you know, even more impactful example, we work with a large manufacturer that does some large chemical manufacturing, right? And they use AI to optimize their plant operations. And so when they're like turning, turning valves and adjusting the flow of different chemicals, Those situations, if you get it wrong, you can cause a large-scale explosion. As our sponsor there said, you can take a town right off the face of the map. And so what happens sometimes is these systems, whether it's advising a doctor or a plant operator, will say something that seems counterintuitive to the human who has to action off of it.
18:40So it could be to a surgeon. It could be to a plant operator. It could be to someone ordering more parts of their airplane. And if you have to action off of it and it seems unintuitive, one of two things is likely the case, right? Either you're right because you built up a lot of tuition and the model's wrong and it's making a decision for bad reasons, or it's seeing something that you don't see. And if you kind of like overrule it, you're potentially not doing a great job, right? And so by providing that explainability, it gives the human who has to action off of it a lot more ability to reason about, like, why is my intuition differing from what the model is telling me to do so that you can make the right decision for everyone?
19:19Got it. And the bias capability you mentioned is also really interesting, you know, to go back to, say, content recommendations. Obviously, lots of systems are doing this now. But as you said, they often might seem biased towards a certain political slant or something like that. How exactly does Arthur dig in and understand where these biases are that the teams can then act off them? You know, in some ways, bias in the kind of demographic sense is, you know, just taking performance, but segmenting it by whatever you want, but typically like a protected class. Right. And so it's not always that, but that's that's the most common scenario.
19:54So, you know, if you're if you're segmenting by gender or by race and you want to make sure the couple things, one, that the outcomes are relatively equitable for for for different categories. Like I'm getting, we're approving credit cards at the same rate or similar rate, things like that. And then the other thing you want to do is make sure that the model is equally kind of accurate or inaccurate at different segments. So, you know, it's not okay to kind of do a great job at making decisions about which males get credit cards, but which females you're just throwing darts at a dartboard, right?
20:27And so there's a really broad set of metrics that people use, but it kind of, those are the two big categories, I would say. And it turns out you can't optimize for all of them simultaneously. And so it's really, there's a lot of nuance. I think we all grew up thinking bias and fairness is this black and white grade school type of issue. And it's not. It's very use case specific. And it's very nuanced. And so what we do is just give the tools to our users so that they can make informed choices about it and kind of find the right balance for them. And what format does the output take for any of these decisions, whether it be the parts for the airplane or, you know, algorithmic bias in, you know, crediting?
21:11Like, how does one interpret what Arthur spits out? Yeah. So if you have something like where there's like a plant operator or human loop, then typically that, you know, there's an application. It's part of an AI driven application. And there's an interface for that, that it gets all the information gets presented to the to the end user. But also we have a platform where people can go in and log in. And, you know, it's really useful, I think, for the data scientists and the people building the systems, product owners and things like that to go in. And they can do all sorts of reporting and segmentation and analysis on it in our platform.
21:42Got it. Now, jumping ahead to LLMs, obviously, like you said, everything kind of changed about a year ago. How has that affected Arthur's focus? I mean, I'm guessing you've had to rebuild a lot of your offering to target towards LLMs and, you know, obviously want to get into the different capabilities that Arthur has for LLMs. But, yeah, talk to us a little about the transition over the past year, which has been so crazy for so many people in so many companies. And I'm especially assuming Arthur has as well. Yeah, you're right. It's been a huge change, you know, massive change to our roadmap like this.
22:13At the beginning of the year, we basically ripped up our roadmap and rebuilt the whole thing and have launched a couple of new products. and augmented our existing observability product. And I think there's some interesting things, right? Like the way you think about performance with these systems is wildly different than traditional AI tasks. And so if you think about things like false positives and false negatives are kind of meaningless when you're talking about producing text or images that are fairly subjective. And so I think for the first year, what most people have been doing is as they're building these systems, they're sort of like manually going in there and typing in prompts and then looking at what's generated out of it and sort of giving it the sniff test, which is A, tedious and time-consuming, and B, again, it's not very effective because you don't end up covering, like you do your handful of favorite prompts and it's not necessarily representative of what your end users are gonna be asking and things like that.
23:11So just qualitatively, you're saying like, I just like put in a couple of prompts and be like, yeah, this looks good. Exactly. I mean, ideally you would do a lot, right? But at some point, after doing that for two hours, you're like, all right, this is really boring. I'm not going to, you know, I think it looks good. Now it definitely looks good because I'm bored doing this. And so we've come out with a number of new tools. But a big part of that is our R &D team. So my co-founder, John Dickerson, he's a Carnegie Mellon PhD, tenured faculty member of Maryland, and is with us full time. But he and his AI team really had to kind of rethink, like, what are the new metrics that we need for generative AI?
23:49And there's some really basic statistical ones, like, you know, what's the length of the response? How concise is it? There's statistical measurements for readability, believe it or not, token count, things like that. But then there's a whole layer of more complex ones, including doing things like using LLMs to evaluate the output of other LLMs, which is a really fascinating area. And so that is actually really useful for some of those seemingly more subjective type of things, like how good is this response? But you can kind of segment even further and have one like, how helpful was this? How friendly was this?
24:23How fully did it answer the question? And you can actually create, you know, through a combination of special code and prompts and things like that, you can create a system where you can actually do that same kind of human-like evaluation with LLMs, but do it at scale and do it for a really comprehensive suite of prompts so that you can be sure that it's not just sort of like, I checked my 20 favorite prompts and it looks fine, but you can actually have a lot more confidence about how it works. So the use case for something like this would be a company is interested in, you know, building on top of an LLM, whether that be open AI or Anthropic or maybe, you know, one of the open source models.
Read the full transcript
25:04And basically as a way to evaluate the LLM before integrating it, they use Arthur to do that. Is that right? Yeah, absolutely. So we kind of work across three major areas of the lifecycle, right? So free production validation, which is like what LLM should I use? If I'm using, most people are using RAG, Vector Data Store, Retrieval Automated Generation. And there's a lot of choice of which RAG database do I use, which Vector Store, what's my retrieval strategy, how do I chunk up the documents when I put them in there. There's all these kind of knobs and dials when you're building these systems.
25:34And so you need to understand as you're turning these knobs and dials and getting it all hooked up, am I making it better, am I making it worse, is it good enough, basic questions like that. And, you know, on the model side, do I need GPT-4? Can I get away with GPT-3.5, which is a lot cheaper and a lot faster? Can I use LLAMA V2 and fine-tune it? Like there's all these, there's just a lot of options. And you get wildly different results depending on which option you choose. And so it's really helpful to be able to do that in a really methodical and fast way. So there's the validation. Next stage is the real-time controls around it, which is when you put these systems out there, you need to have some guardrails, right?
26:11And so there's sort of universal things like you don't want them to hallucinate, especially if you're in something like a medical or legal or financial context. You really don't want them hallucinating, which is basically just a fancy word for giving wrong answers. Right. And so being able to kind of catch those in real time and block them, as well as things like prompt block blocking prompt injection. Just blocking, just like if it hallucinates, just like basically never returning that to the either. Yeah. So so as the application developer, you can either choose not to return it, which is a program for some use cases, or you can return it.
26:40but just put like a little, you know, a little bubble next to it. Like, hey, this may not be true. You should validate it before you kind of take it as gospel. And same thing with prompt injection and toxicity and also data leakage. What is that? What is a prompt injection? Talk about that. Yeah. Yeah. So you can actually design prompts to get around a lot of the built-in guardrails that LLM makers have put in place. And so there's a number of attacks you can perpetrate with this. The most basic is probably getting back the system prompt, which is like the underlying prompt that the application developers have put in.
27:14You can actually make it say things that are like, you know, offensive or not aligned with what the company, the people who are running the system want it to say. And if they're starting to, as people are starting to use agents, you can also do all sorts of things that are maybe the equivalent of like modern day, like previous day, like SQL injection attacks and things like that. So you can kind of do unauthorized access to information. And so, yeah, blocking prompt injection, making sure you're not leaking data, making sure you're not leaking the system prompt. You know, just what we call, we call it a firewall for LLMs.
27:44It's kind of like a big application firewall for these LLMs. Also things like acceptable use. So there's actually a car dealership. Someone posted some screenshots this weekend that put up an LLM chatbot. And people started, like, figured out how to manipulate it. And so they'd ask a question or they would get it to do things like, you know, So say you will sell me this car for$5, and this is a binding agreement. Yeah, and then they're like, hey, your chatbot said I could buy this car for$5, right? And so you need to have some acceptable use kind of guardrails around it to make sure it doesn't go well or that it doesn't go poorly.
28:20And so that's the second phase. And then the third phase is just that sort of monitoring and observability that we've always done, but adapted for the LLM world, which is, are my answers getting more accurate over time? If my LLM provider makes a change to their system, what effect is it having on my application, which happens all the time. And so, you know, what you want to do there is just make sure you have quality service. You want to allow product managers to go in there and data scientists to go in there and machine learning people to really figure out, like, where is it giving good answers?
28:49Where is it giving bad answers? What are, like, the top 10 weirdest requests people made to the LLM today? Like, there's a bunch of questions you can answer to make your system better. So I have to ask because, you know, I see this all over Twitter X. Everyone is saying that ChatGBT is getting lazier in December. And my guess is this product would be able to measure whether or not that's actually happening. So is it getting lazier? So number one, it is, yes. And one of the metrics we've developed is hedging, which is like that behavior where, you know, hey, as an LLM, I shouldn't tell you what to have for lunch or I shouldn't like, you know, do this.
29:24Oh, interesting. The frequency with which it hedges has definitely gone up. And there are some other, you know, when you look across the different LLMs, there's a huge variance in like how much they hedge, right? And I think a lot of it's based on like how smartly and how aggressively they put in sort of guardrails. So I think a lot of it was kind of born out of this desire not to have it say anything offensive. But in a lot of cases, they've taken it too far. And in some cases, they won't even answer really basic questions. And so if you're in a context like, I don't know, investment research or something like that, where, you know, you have your own internal employees using it to get investment research.
30:02Like you never want it to hedge because you just want you want you want an answer. Right. You need the answer. And so that that can inform your choice of which LLM. And so you may not want one of the big mainstream LLMs, which is very powerful, but doesn't always give you an answer. So why is it getting lazier? Do you have any and is it just ChatGVT or is Claude getting lazier? or Lama? Like, are they all getting lazier? It seems like ChatGPT hedges the most. And I think, again, it's just, it's a byproduct of the, basically people will also post on Twitter and acts like, hey, look at this kind of offensive thing or this problematic thing that I got ChatGPT to say.
30:41And so then the team kind of like goes in and puts in guardrails, which are like, hey, don't answer these types of questions because you may say something offensive or that's just inappropriate. And so then the byproduct of that is like those guardrails maybe are, are, uh, overzealous sometimes and, uh, lead to it then not even wanting to answer certain innocuous questions, um, which, which is frustrating as an application developer. Got it. So it almost sounds like maybe we've just reached a tipping point in terms of the guardrails being implemented such that it's, it's creating like a bit of a restrictive environment for the LLM.
31:16Um, it's not that it's actually getting lazier. Yeah, it's all, I mean, this is all, this is all really early technology. Yeah, of course. Yeah. We're everyone's learning as I'm sure open AI is too, right? I mean, Oh, big time, big time. Yeah. So I think it's a lot of like people are seeing in a very public way, just sort of like the, the throes of a very powerful, but very early technology. That's, you know, it's not, it's not hard to see the possibility, but, uh, it's also, there's just so many things that can go wrong right now, uh, given the early nature of it. Yeah, totally. Speaking of open AI, like, how do you think about, um, sort of the, the big kind of LLM wars.
31:52I mean, is open AI kind of walking away with the market? Is Claude still in there? Are we actually moving more towards an open source future? Like, how do you see this playing out? Yeah, I think that the future is going to be pretty diverse, actually. And I think that's good for everyone. Open AI is certainly a really powerful model and system, right? And definitely kudos to them. They've been really kind of attacking this problem in a bigger way than other people for a long time now and I think that you see that. But I also think that they're building for kind of the mainstream and there's a lot of like the open source models that we see already have huge advantages in areas where people are doing really cool things like taking a smaller like 7B model and fine tuning it for certain specific use cases and they can get wildly better performance.
32:41And so I think that there's, we're starting to see more of that. We're starting to see people begin to train their own models more often. And so, you know, that stuff's a little bit early. I think it'll be another year before you really start to see that in the real world a lot. But we're starting to see it now. And so I think that there's going to be a really nice mixture of commercial providers as well as open source offerings, everything from kind of large general purpose models like GPT-4 to these really kind of focused models like, you know, fine tuning a small LLM that just detects prompts for prompt injection, you know, things like that, that are very sort of purpose driven ones.
33:18And as a consequence, they don't have that horrible latency. They're much cheaper to operate. They've just got a big, big advantages that way. Yeah. And it seems like every week there's sort of a new model coming out that is claiming to be better than the last previous model, you know, coming out with their own benchmarks. Gemini, obviously a couple of weeks ago i mean how how should we um as customers of these models think about these benchmarks that are getting released every other week i mean are like are they accurate can anyone just kind of uh tweak their measurement in such a way that their benchmark always looks the best i mean there's certainly some of that that goes on that's always really what's what what's there's like a name for that effect is it obviously is it dunning kruger or different effect Like once you sort of like start optimizing for a particular metric, it becomes useless.
34:06And there's definitely some of that. I think, look, those kind of general purpose metrics are a good, you know, first cut at like understanding the performance of a model. But what we preach, and so we have an open source tool called Bench, which allows people to understand. What you want to do is understand not just how does it perform against some generic benchmark, but how does it perform against my exact kind of workload, right? So if I'm using it for customer service or for investment research, I want to take like, you know, 500 of my real prompts and responses and run it through the system and see how it how it does on that.
34:37Because I don't really care about these like generic benchmarks. And they're not right. They're oftentimes not that indicative of how it's going to perform on my application. Right. And so that's what we have, like basically a testing tool that allows you to test it with your with your prompts on your data. and it's allowed you to know, like, all right, this big headline about how this is like revolutionizing that this is like the latest and greatest state of the art LLM. Like, does that actually matter for me or not? Right. The generic benchmarks don't matter. It's my benchmarks and how this product is actually going to perform for me that I should care about.
35:09Exactly. With my vector data store and my data and like all these variables. Yeah, that makes so much sense. So, you know, my my understanding is that Arthur is mainly focused on LLMs, but obviously as we're talking about sort of the future of models, we're seeing all these different types of models, right? Image models, video models, models for generating music. How do you think about the opportunity for Arthur in, you know, other, other formats and other forms of media? Yeah. So we, I mean, you know, prior to the, the generative AI boom last year, we like, we support a computer, we support computer vision to this day.
35:39You know, we have really robust, um, support for computer vision, for NLP, um, for, for tabular data, all, um, time series, these recommender systems. And so for us, I think we see a ton of business value being created or about to be created through text, generative AI for text, so LLMs. I think that the business value creation part of these multimodal models is still earlier. I don't doubt that it'll happen. But I think one of the things we're looking closely at and talking to our customers about is like, you know, what kind of value are you going to get out of being able to create these videos on the fly or create images on the fly, like in a business context, like, is this going to help you automate some sort of back office process?
36:23Or, or is this going to help you like service your customers better? And so I think that it's still kind of early in those conversations, like people are still figuring that out. And so we're watching it closely. I think because we do have like, you know, good expertise in computer vision and some of these other areas, we'll be ready for it when it when it really starts to have an impact in the world um so we're watching it closely uh but it's still early days yeah i can see what you're saying like especially in sort of the b2b use cases which is where you know i think arthur plays most like the the value is still not super clear where it obviously very very much is with llms and language right you know i think we need to see how it plays out outside of the consumer use cases for some of these generative media models yeah talk a little bit about like you know we've talked about the products we've talked about what they do, but talk about the impact of the past year on your company and your team.
37:13What's that been like? Again, at the beginning of this conversation, we were talking about how you have been working in AI for literally a couple of decades. Your team has been working on AI at Arthur for five-ish years. But now in the past year, it's like the whole world wakes up and decides that this is the most important technology in the world. What's that been like for your your cadence, your culture, everything. Yeah, it's been a crazy year. So I think, number one, don't go into working in AI if you're not the kind of person who thrives off of constant learning and change and adaptation because it's just wild.
37:50Every few months, there's something new coming along that just completely changes your understanding. And I've been working in AI for 20 years, and it's just one of those fields where every day I learn more and more and feel like I know less and less because there's just so much new stuff happening. But it's been, you know, it's obviously been a huge boost for the team, for hiring, for like talent, like so many talented people now are really just, I think, just moved like in a really existential way by the power of this technology. And they're like, oh, my gosh, I want to go work at a company that's building, that's like helping kind of define the future with generative AI.
38:23And so that's been really exciting. you know it's funny I think when Chatsby was first released a little over a year ago the I think there was a part of me who has been you know from having been in the world I was just like alright this is like the latest like hype du jour and I was almost sort of like dismissive of it for the first couple months but then you know by the time we got into January and we were just spending time with our customers and they were like talking not just that they were excited about the technology but here's all the ways we're going to use it and here's how like all the areas that we've identified that this could really be a game changer I started to get past my jaded nature and really get excited about it.
38:57And that's, I think, when we really started to say, hey, we've got to do something. This isn't a matter of just supporting generative AI and our observability. We actually want to build out this validation tool because that's one of the biggest impediments we've heard from our customers to getting started with generative AI. And then the second one being these real-time controls. And so there's a real opportunity for us. I think anytime you have this sort of just massive change, it just it creates opportunity. Right. And there's like winners or losers created out of it. So we just really wanted to go back to first principles and not be beholden to kind of what we what we were already doing and really just think about what the what the market needed and what our customers needed to be successful.
39:38For sure. And how do you think those needs are going to evolve? I mean, as an example, I mean, could you see a lot of these LLMs, especially the big ones like OpenAI and Anthropic? Like, could you see them building evaluation tools, firewall tools, performance tools directly into their models and offerings? And what would that do to Arthur? We already see like little pieces of it for sure. Really? And I think it's not that different from when we started just AI observability and people were like, well, AWS is going to build that, right? And I think so. I think there's a couple of things. One, those kind of first-party offerings are tied to their stack.
40:16And even when they build it in a way they say it's not, really, you're only going to use it with their stack, right? So if OpenAI comes out of tools, you're only going to use OpenAI. What we find is our customers, especially large enterprises, they don't just have OpenAI. They're agnostic. They have all of them, right? And they want to be able to choose the best tool for the job. and so they want that kind of single pane of glass across all of them that gives them that real-time visibility into what's going on. And then the second thing is, if you're like Amazon and you're selling compute or Azure, you know, compute and storage, like it's a lot, as you get up the stack, sort of the imperative to do a great job becomes like, you know, less and less strong.
40:57It's less and less existential for you. And so what we find is like, you know, They may have basic versions of observability or basic versions of guardrails, but a lot of times it actually helps us because people will start down the path of using those basic built-in tools, and then they realize they're not sufficient, and then they start looking around and they find us. There's still opportunity for software companies, even though obviously those companies have created a lot of momentum and they've created some really powerful engines, but we still like where we're at. What about companies that are sort of doing their own model development?
41:33You know, sort of starting with OpenAI, then maybe graduating to being a little bit more model agnostic and then starting to build and train their own models. Is the need still there for something like Arthur? Or are they just building this evaluation directly into whatever they're doing? Oh, no, those are that's a huge part of it. I think both for the validation, because there's like, again, there's like the amount of when you're training your own model or fine tuning the number of knobs and dials and also just the potential that like you can actually make things way worse if you fine tune a model, if you don't know what you're doing.
42:04And so you need to be you need to like make sure that you're you have a really good bead on what affects the changes you're making are having. So it actually creates even more opportunity for us. And the other thing is like where some of the guardrails like toxicity is a good example. Right. I would say the major LLM providers have done a pretty good job of like getting rid of like the most harmful like that. You know, you can't get it to cuss up a storm the way you could like in the early days, things like that. But if you're building your own, it's like, are you really going to invest like a hundred person team in like, you know, in building in those kind of toxicity filters?
42:39Are you just going to buy something off the shelf for for a relatively small amount of money? You'd much rather just that's not the strategic part of what you're trying to do. And so it's better to kind of work with someone like us who can just bring that to you right away, like a best debris product. How does your product get more challenging to build and maintain as the capability of the model improves? You know, for example, we've seen information leak that, you know, GBT 4.5 may be dropping soon. GBT 5 will be here, you know, whatever, before we know it, I'm sure. What happens each time a new model drops?
43:13Like, what does your team have to do in that moment? Yeah, I think what we do, so we're kind of agnostic to the LLM, but where it does affect us is, as I mentioned, a lot of our evaluation routines do use LLMs. And so like hallucination detection, we have kind of a layered approach to techniques. But one of those techniques is actually using LLM to help in the kind of determination of where the hallucinations are and what they are. And so, you know, every time a new LLM comes out, we benchmark our own internal tools against it. Like, hey, and there's typically it's not like, is this better? Should we worse?
43:48It's typically a tradeoff, right? Like, oh, hey, we got a better score out of like GPT-10, but it's also costing$7 million a day to run. Like, maybe that's not what we need. Or it takes 30 seconds to come back with an answer. And that's like not acceptable. Right. And so it's really like every time one comes out, like we ourselves use our own tool bench to understand, like, what are the tradeoffs with this particular LLMS? Like, how does it perform? How fast or slow is it? How much does it cost us to run? like what is it worth it is it not which of our own internal kind of like routines is it is are those trade-offs good for which are they bad for things like that got it we've talked a bunch about hallucinations how exactly does firewall cut down on hallucinations i mean this might sound simple but like is it doing sort of real-time google searches to like validate claims like how like how does that actually work so there's like a whole taxonomy of hallucinations and wrong answers and so there's some that are intrinsic to the model the kind that we focus the most on because this is the most helpful in the world.
44:47So like 99 % of the LLM systems that go in the real world are these RAG retrieval automated generation systems where you're taking all of your proprietary knowledge and putting it in a vector data store. And then when I go to ask a question like, hey, what's this drought in Brazil going to do to energy stocks there? Something like that. It'll use like all this investment research that's been loaded into this vector store. And a vector store is a way of like, it's kind of a way of indexing knowledge by, in a way that like makes sense from a language perspective. And so like, you know, cat and kitten are actually like, even though they're spelled very differently, they don't like say anything.
45:26If you put them in a vector data store, they're actually like right next to each other because they're very close in meeting. Right. And so, so if I say like, what is this drought in Brazil going to do to energy production in that country? What it does is it doesn't just give that to the model and answer the question because the model won't have a lot of information. What it does is it goes to the vector data store and it pulls back pages and pages of relevant kind of research that maybe my analysts have done in the past that goes back decades, that's maybe seen similar trends. And it'll append that knowledge base to my question.
45:58So it'll say to the LLM, here's Adam's question and here's some information you should use to answer it. And then it generates the answer that way. And that way, you know, you can get a much richer answer about data that the model has not been trained on. And it tends to be much more accurate. Now, what we do is we take the response. There's a number of techniques, but one of the more common ones is we take that response, break it down into a series of claims. And for each of those claims, we validate that against the underlying knowledge that's available to make sure that each claim is well supported.
46:33Because oftentimes it's not. And some of our clients see some pretty high rates of hallucinations depending on what they're doing. And so being able to spot them is huge. It's a simple concept. There's a lot of tradecraft that goes into, you know, you can get like 80 % accurate pretty easily. But getting them to like the high 90s, like we're achieving, it takes some iteration. So the team's done some pretty good work there. So much of what Arthur is doing seems like it's geared towards keeping LLM safe inside of organizations. And so, you know, I have to ask about sort of regulation and maybe the climate in general.
47:08Obviously, it's a very, very hot topic these days. Where do you think we as a country are headed in terms of regulation? And what do you think is needed versus not needed? It seems like this has become a very, very sort of tribal discussion very quickly. Help us make sense of it and kind of cut through maybe somewhere in the middle to figure out what's the actual ground truth. Yeah. So, I mean, we could do a whole podcast just on this topic. It's a fascinating one. But it's, you know, generally speaking, you want the regulation to be light enough that you can maintain room for innovation and companies can innovate.
47:41But you want it to be heavy enough that people aren't getting hurt in the process. Right. Like whether it's like getting run over by a driverless, literally getting hurt, like run over by a driverless car or just being negatively impacted, like can't get a credit card or something like that. And so I think the U.S. has done a pretty good job balancing that. Compared to the EU, the EU has been much more pro-consumer, or I shouldn't say pro-consumer, but much more focused on consumer protection and less on innovation. The U.S. has probably balanced more towards innovation. But I think the executive order that came out last month was a good start.
48:14And in the executive order, it's a good framework. If you read it, a lot of the pieces are sort of like charging, like, all right, we need to do something in this area. know, we're going to give this agency, you know, six months to figure it out and report back. And so it kind of created the framework and a lot of the actual details are going to be filled out in the next year. And so we'll see how that all shakes out. I think, you know, from my perspective, the thing that needs to happen is there just needs to be clear accountability for the people who are deploying the models. And so I think one agency that's done a good job of this is the EEOC, the Equal Employment Opportunity Commission.
48:51If I purchase an off-the-shelf resume screening tool or candidate screening tool and I deploy it, regardless of whether I built that model, like if I'm a large enterprise that employs tens of thousands of people and it's shown that I use this tool that's discriminatory, I'm liable for it because I'm the one running it. I'm the one using it. So it doesn't matter whether I built it or not, the fact that I'm using it. And I think that kind of clear accountability for you're responsible for any sort of bad repercussions that come of this is really important. Yeah, for sure. What are you maybe as a consumer most excited about with respect to AI looking forward?
49:28Yeah, I mean, I've had a ton of fun with it. Just like even silly thing, like I use ChatGPT almost every day and just doing fun stuff, like helping my kids with their homework. My daughter's old enough now that she's starting to get into pre-calc and things like that, that I don't, chances off the top of my head, remember I have to go back and kind of refresh And so, you know, we were asking for to help her with a trigonometry problem the other day. And it didn't get it right, but it did get it actually broke it down into these really clear, concise steps. And the thought process was actually right spot on.
50:03The arithmetic was like slightly off, but that was easy to correct. And so I think it's been pretty awesome just to see that. Like, I don't know, it's just one of many examples of being able to kind of use it. Another fun one is, if you look at our website, we have a specific illustration style that we've kind of made our own. And we worked with an illustrator who, and this was all done sort of in an ethical way, but to have her produce a set of illustrations. And then we use different foundation models to actually create different ones in that same illustration style. So if one of our, you know, someone's writing a blog post and they want this abstract illustration that represents, I don't know, measuring hallucination rates or something like that, they can feed all this into this foundation model and actually get an illustration in that same style that is, but that's like completely custom.
50:55And so there's just so many fun applications of this as a consumer I'm doing. What about yourself? Do you have any favorite ones? I mean, I use ChatGPT every day as well for, you know, all sorts of use cases. And, you know, my background has a lot to do with media and creator tools. So I'm, you know, I'm very, very interested in in some of the the other gender models around video and imaging and things like that. I'm constantly always playing with that stuff. So maybe looking ahead, you know, I don't think any of us could have predicted what the past year has been like. But as we record this, you know, near the top of 2024, like what are your predictions for the year ahead?
51:35Yeah, I mean, to your point, I'm almost like, like I know whatever prediction I make will be like, it'll be wildly different, right? Because it's the nature. But I do think that, you know, we're starting to see the hype turn into reality a lot more. And so I think it's going to be what I think is there's going to be some really compelling applications of this technology that that hit the real world. And there's also going to be some high profile kind of like unintended effects of it that were people kind of bad things happen or funny things happen, depending on the scenario. And so I think that is definitely going to happen.
52:11I do think that the multimodal models like it's going to be very interesting, especially with an upcoming election. you know, what, what effect being able to just really quickly generate realistic videos and, and, and images will have on things. So yeah, that's also the, the, the coinciding of, of kind of a very contentious election, foreign interference and, and, and generative AI is going to be, it's going to be interesting to watch. For sure. For sure. I got to ask, is Arthur hiring and where can listeners find out more about how to apply if they're interested? Yeah, we are absolutely hiring.
52:48We're at, just go to arthur.ai, our website. Appreciate asking and, you know, hiring across the board for machine learning engineers and a number of like customer solutions type of engineers are definitely top of mind. And we're based in New York here, down in Soho. So, you know, we have opportunities for both sort of virtual people as well as people who like to be in the office, as a lot of our team does. So definitely check us out. Drop a line. Awesome. And anything else we should tell the listeners about Arthur or anything you want to make sure they're aware of? I think we covered it pretty well.
53:21You know, just it's an exciting place to be right now. If you're the kind of person who gets bored easily, if you're not constantly having to learn new things, then it's the right place to be. For sure. Awesome. Well, Adam, this has been so much fun and really appreciate all the time you've given us. Thanks for coming on the pod and hope we can do it again sometime. Yeah. Thanks for having me on, Michael. Thank you so much for listening to Generative Now. If you liked what you heard, please do us a favor and rate and review this podcast on Apple Podcasts and Spotify. It really, really does help.
53:54And if you'd like to learn more, you can follow Lightspeed at LightspeedVP on YouTube, Twitter, LinkedIn, or any other social platform. Generative Now is produced by Lightspeed in partnership with Pod People. I am Michael Magnano, and we will be back next week with another awesome conversation. See you then.
From the publisher
Over the last year, AI and LLMs have experienced an explosion of use. But the tech isn’t foolproof - hedging and hallucinations are just a couple of the concerns. That’s where Arthur AI comes in. They’ve built out firewalls, validation tools, and benchmarks to make sure that an LLM’s answer is one you can trust, whether you’re a military contractor or investment broker. Arthur AI Co-Founder and CEO Adam Wenchel sat down with Lightspeed Partner and Host Michael Mignano to talk about just how Arthur AI is allowing for the safe, responsible, and trustworthy adoption of AI.
Episode Chapters
(00:00) Intro to Adam Wenchel, co-founder and CEO of Arthur AI
(5:31) Working on AI at DARPA 20 years ago
(8:39) From cybersecurity to Capital One
(14:20) How Arthur lends transparency to AI decision-making
(21:45) ChatGPT ripped up Arthur AI’s roadmap
(26:18) Arthur AI’s guardrails can make AI trustworthy
(31:46) Who will win the LLM wars?
(37:00) What the ChatGPT sea change looked like at Arthur AI
(44:22) How Arthur AI’s firewall clamps down on hallucinations
(47:11) Are we on the right path for regulation?
(51:22) Adam’s prediction for the year ahead? Unpredictability
(52:39) Is Arthur AI hiring?
Stay in touch:
LinkedIn: https://www.linkedin.com/company/lightspeed-venture-partners/
Instagram: https://www.instagram.com/lightspeedventurepartners/
Subscribe on your favorite podcast app: generativenow.co
Email: generativenow@lsvp.com
The content here does not constitute tax, legal, business or investment advice or an offer to provide such advice, should not be construed as advocating the purchase or sale of any security or investment or a recommendation of any company, and is not an offer, or solicitation of an offer, for the purchase or sale of any security or investment product. For more details please see lsvp.com/legal.




