In short
The TWIML AI Podcast - Episode #704: AI Agents: Substance or Snake Oil with Arvind Narayanan
Overview In this episode of The TWIML AI Podcast, host Sam Charrington speaks with Arvind Narayanan, a professor of Computer Science at Princeton University, about his recent research on AI agents and his book "AI Snake Oil." The discussion delves into the capabilities, challenges, and risks associated with AI agents, the urgency for better benchmarking, and the broader implications of AI in society.
Key Topics Discussed
- AI Agents That Matter
- Capability vs. Reliability Gap
- While AI agents have remarkable capabilities, their reliability is often inadequate for practical applications. For instance, a failure rate of even 10% can render an agent ineffective (e.g., ordering food incorrectly).
- Benchmarking Challenges
- Traditional benchmarks, which worked well for simpler machine learning tasks, fail to provide a clear picture for complex AI agents.
- The transition from models to agents requires new forms of benchmarks that accurately represent the real-world tasks agents are designed to perform.
- Importance of Verifiers
- Verifiers are proposed as a technique to ensure AI agents perform reliably. They serve as a guardrail by validating agent outputs against established criteria (e.g., unit tests).
- AI Snake Oil
- Problematic Claims in AI
- The discussion addresses overhyped claims about AI capabilities, particularly in high-stakes domains such as healthcare and criminal justice.
- Narayanan emphasizes the need for skepticism regarding claims made by AI companies and highlights the potential harms of misrepresenting AI capabilities.
- Taxonomy of AI Risks
- Different categories of AI risks are identified, including:
- High-stakes Errors: Misapplications in critical sectors can have severe consequences (e.g., healthcare algorithms).
- Discrimination: AI systems might propagate biases, affecting individuals negatively but not on the scale of catastrophic risks.
- AI Policies and Regulation
- Current Regulatory Landscape
- Narayanan discusses the challenges of keeping up with rapid AI advancements in policy-making environments.
- He advocates for a balanced approach that emphasizes the enforcement of existing laws rather than rushing new regulations.
- Future Policy Directions
- There is a call for policies that address specific harmful behaviors (e.g., non-consensual image generation) and longer-term strategies for managing AI risks.
Insights and Takeaways
- Understanding Agentic Behaviors
- The definition of AI agents is fluid; it includes several factors such as environmental complexity, task difficulty, and the level of autonomy allowed.
- Performance Metrics
- Emphasizing the need for multi-dimensional benchmarks (considering both accuracy and cost) rather than simple performance leaderboards.
- Real-World Applications and Challenges
- Practical use cases highlight how AI agents can be beneficial, but they also reveal the significant barriers to achieving consistent and reliable performance.
- The Role of Research in AI Development
- Academic research can fill gaps in understanding and developing AI technologies that are reliable and effective, contributing to the safe deployment of these systems.
Conclusion Arvind Narayanan’s insights provide a critical perspective on the current state of AI agents and the risks associated with them, along with suggestions for the future of AI policy and research. The conversation underscores the importance of rigor in AI evaluations and the need to differentiate between genuine advancements and exaggerated claims.
For further details, the complete show notes can be found at [TWIML AI Podcast Episode #704](https://twimlai.com/go/704).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00And here's the fundamental paradox with agents. The capability of agents, I think, is already in some ways mind-blowing. You know, if agents could do reliably in the real world, in the hands of consumers, everything that they're capable of, it would truly be an economic transformation. But even if they're going to fail 10 % of the time, it's a useless product because no one wants to have an agent that orders DoorDash to the wrong address 10 % of the time, right? These are the kinds of failures that consumers are actually reporting.
0:44All right, everyone, welcome to another episode of the TwiML AI podcast. I'm your host, Sam Charrington. Today, I'm joined by Arvin Narayan. Arvin is a professor at Princeton University. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Arvind, welcome to the podcast. Hi, Sam. Thanks for having me on. I'm super excited to jump into our conversation. We'll be talking about some of your recent work exploring AI agents and how we should be benchmarking and assessing their performance, as well as your recent book, AI Snake Oil, which I should mention you co-wrote with Sayash Kapoor, who we spoke to on the podcast earlier this year, talking about the risks of open AI models.
1:35And I think you were a co-author on that work as well. Yeah, we've done a bunch of research together looking at the risks of AI in various ways. Yes. So really looking forward to jumping in. And let's get started by having you share a little bit about your background. Sure. I'm a computer scientist. I'm also the director of a center here at Princeton called the Center for Information Technology Policy. So I split my time into a few buckets. One is technical AI research. These days, it's been primarily in the AI agents topic. In the past, a lot of it has been on AI bias. And then a second bucket is advising policymakers, you know, where AI is going, what sorts of guardrails are needed, policy on open versus closed foundation models, those sorts of things.
2:23And perhaps the third bucket is breaking this all down for a broader audience beyond AI researchers and policymakers. And that's kind of where the book comes in. And we have a newsletter as well, AI Snake Oil. So I spent quite a bit of time doing that as well. Yeah, talk a little bit more deeply about the way you kind of craft your research agenda. What are some of the problems that you've been digging into? What are some of the ways you formulated them? And how do you kind of pull all that together into a broad research agenda? So when we look at AI, one of the things that I think as academics we can contribute is doing things that companies are maybe not going to do, even if they're capable of doing, right?
3:15Because we don't want to be competing with Google or Microsoft or OpenAI. And so I think benchmarking is one area where academic research can really contribute. and more rigorous evals are needed now more than ever. So as we moved from traditional machine learning models to foundation models, evaluation became a lot harder, right? Because with a traditional machine learning model, it's built to perform one task and that task is often something really simple like handwriting recognition. So, you know, MNIST is a pretty decent benchmark for handwriting recognition. There are limitations, you know, it doesn't measure distribution shifts, et cetera.
3:55But that said, you know, it worked well for decades. But then you get to foundation models. These are models that are supposed to be able to do everything to some degree. And so benchmarking started to become really messy. I mean, we've all heard that story many times. You know, every time a company releases a model, they pick some set of benchmarks and their model looks good. But it turns out they prompted it in different ways or whatever, right? So, and then the vibes don't really match the benchmarks. So that was already a problem. And we're expecting it to get harder as the models are getting trained on training data that they've generated.
4:34Exactly, exactly. And another thing that makes it harder is when you get from models to AI agents, because agents are supposed to be able to do things in the real world. And, you know, if you have an agent that is supposed to book flight tickets, you can't actually ask it to book flight tickets, right? And so you have to create some simulated version of that. And now you have to worry about, does your simulated version actually model reality? Or are you encouraging developers to build brittle agents that will do well on your benchmark, but completely fail when you put it on a real website? So those are some of the challenges that have motivated our work.
5:10And this kind of ties in neatly with our other gig that I mentioned of calling bullshit on some of the overhyped claims. Right. So that's what AI snake oil is about. So can we can we actually take that premise and try to make it rigorous, build better evals and actually use that to push the community forward towards agents that work better in the real world and not just do well on benchmarks? Yeah, yeah. And it's interesting that kind of agents is one of the areas that, I don't know, I guess it strikes me as like this nexus for both promise and what many are hoping to enable with AI, even historically.
5:55And bullshit, because there's a lot of inflated expectations, I guess we can call it that. Um, what's your sense of, you know, where we are today with regard to kind of the capability of agents? And I guess I'm mostly interested in like, do you, when you look at the technology, do you see like fundamental or where do you see fundamental limitations such that, you know, we'll never catch up to the hype, you know, versus, you know, it's just being early and the tooling needing to evolve and the benchmarks needing to evolve and all that? How do you think about that landscape? That's a great question.
6:37And here's the fundamental paradox with agents. The capability of agents, I think, is already in some ways mind-blowing. And I'll put it the following way. If agents could do reliably in the real world and the hands of consumers, everything that they're capable of, it would truly be an economic transformation. And this is what I think a lot of developers have not fully recognized, right? This is what we call the capability-reliability gap. And when you look at a lot of the products that have flopped on the market recently, like the Rabbit and so forth, I mean, you know, I get the concept, right?
7:13And sometimes these magical experiences do work, but even if they're going to fail 10 % of the time, it's a useless product because no one wants to have an agent that orders DoorDash to the wrong address 10 % of the time, right? These are the kinds of failures that consumers are actually reporting. So that's one thing. The capabilities are really strong, but how do you get that last mile? And that's the same issue we had with self-driving cars, for instance, right? We had really good prototypes, you know, 20 years ago, and that's why developers and CEOs perpetually thought that self-driving was two years away.
7:49I mean, we're finally there now, right? Waymo is offering rights to the public in several cities, but it's taken decades to get there. And maybe, you know, maybe agents are going to take that long, but maybe we can accelerate that through better rigor in our research and development. There are, I think, a bunch of other challenges with agents, but reliability is the number one challenge. Mm-hmm. Yeah, when I hear that, I think a little bit about an argument that I see raging on LinkedIn and X about code generation. And, you know, it almost strikes me as like two sides talking past one another in the sense of, you know, anyone who's used it, particularly on something that they're not intimately familiar with, like is astounded because it's like helping you write some app in some language that you've never programmed before.
8:39And that's amazing. Um, but I think the calls for, you know, broad studies that look at the impact over many developers and, um, you know, the degree to which the bugs that it introduced, uh, you know, if you're not trying to build a production system, those bugs don't matter. But then when you're really talking about productivity on the economic scale, they do, uh, those are valid as well. Right. That's exactly right. That's exactly right. Yeah. And I think you really put the finger on it. There's a difference between coding and software engineering. And I think that's a big part of where the gap is.
9:14So interestingly for me, as a computer scientist, a researcher, I write a lot of code and most of it is not production software, right? So actually for me, agents for coding are extremely valuable because a lot of my code is something that only needs to run once because I'm just testing out some interesting idea that I had. Is this a promising research direction, right? And also even in my personal life, let me give you this example. This is one of my beautiful moments that I had with my kid, which was a positive use case for AI. She's five years old. I was teaching her to tell time. And you might be wondering, what does this have to do with agents and coding?
9:50But it does. And that's the crazy thing. Engineers always make things more complex than they need to be. Well, but yeah, but this wasn't complex. And it didn't require any engineering skills. So I drew a bunch of clocks for her to teach her how to do it. and she was really enjoying it, but it kind of got tiring to keep drawing these clocks. And, you know, I pulled up Claude and said, make a random clock generator app. Right. And so this is something that I call a one-time use app. I don't know if there's a standard term for it, but I find myself doing this more and more. And it produced the app within a minute.
10:23I gave it some modifications that did that. And then I played with her for five or 10 minutes using this app, having it generate random clock faces and asking her to tell the time. And she really got it. That's awesome. Yeah. Within those 10 minutes, and it would have been much more annoying for me to have to draw 50 different clocks on a piece of paper. And then I've never used the app again. So before AI, it would have been crazy to try to sit down and write an app for this. It would have taken two hours. But now I can do this with AI, and it's not a production software engineering scenario, so it doesn't matter if there are bugs, etc.
10:58But you mentioned self-driving cars as an analogy. And early on, one of the conversation points was always that the corner cases are too vast and we'll need to constrain the infrastructure, like build lanes for self-driving cars or that kind of thing, have them only operate as airport buses or what have you. I think we're a little bit beyond that to some degree, but I'm wondering if there's an analogy for agents that, you know, is the idea of a way of like, I guess, guardrails or like constrained environments for agents. I think in the conversation with Sayash, we talked a little bit about the idea of just how much risk is involved with allowing them to actually interact in the world.
12:01And I'm wondering how you see that playing out safely. Yeah, for sure. So there is a strong analogy here, and this is actually a big part of our next research project, building on the AI agents that matter paper. So we're making a bet on what people are calling verifiers. So the idea here is that if you have a coding agent, then a set of unit tests is a verifier. So you have the agent write code, see if it passes the unit tests. If it doesn't, you can throw it away and start again or give it feedback and ask it to debug it, et cetera. So a verifier is a kind of guardrail. And this relates to OpenAI's Strawberry model as well and other things that other companies are working on.
12:46from what I understand, the way that they have trained Strawberry is, or, you know, a one preview, yeah, whatever they're calling it, is having these kinds of domain-specific verifiers and then, you know, using reinforcement learning, if it didn't ultimately lead to a passing score from the verifier, then that's, you know, negative reinforcement, that sort of thing, that's great. So this leads to a general model that can work well in a bunch of reasoning domains, But still, it doesn't give you that 99.9 % reliability. So what we're asking is instead of a general model, if you want to make it domain specific, if you want something only for coding or if you want something only for web navigation, right?
13:28So can you build really strong verifiers that you can be confident with 99.9 % accuracy? It's going to be able to tell you if the code is correct or if it navigated the website correctly. And can you then use that to build domain-specific agents that can actually work reliably? It sounds like verifiers in this context is a design time constraint as opposed to envisioning some kind of runtime constraint. I would always say for a long time that agents is a political problem as opposed to a technical problem in the sense of for many years and to this day, companies actively try to restrict crawlers from accessing their websites.
14:15They are selective about what they make available via APIs and to who and their licensing agreements. and part of the promise is that I have this agent and it'll be able to do anything that I can do, but there are legal and other reasons why we haven't been that. In other words, part of what needs to happen is kind of changing the way we think about machines accessing systems. And I wonder if your research has kind of come across that idea. Do you think that the kind of momentum of AI will make some of those changes societally or will we still have these kinds of issues? There's a case that the momentum will overcome some of these barriers.
15:07There's also a case to be made that it will make things worse because everybody who's not a big AI company is just so terrified of AI companies capturing all the value basically. And every website or app now just becoming a sort of backend data provider or service provider. And I wonder if this is one of the reasons why OpenAI's GPT's feature didn't really take off where they wanted Expedia or whoever else to be essentially hooking into ChatGPT so that ChatGPT would be a universal interface to access all of these. Right. So that that, you know, that's that was an attempt to realize the vision you were talking about instead of having to build an agent that can navigate the web generally, you know, bring the web to the agent.
15:57Right. And I don't know why it hasn't worked. I think maybe it's on the user side. Maybe the user experience wasn't great or maybe it's because they didn't get as many providers to play nice with them as they wanted. And we're seeing that a little bit when it comes to web crawling, which, as you mentioned, there was a study recently showing that over the past two years, the percentage of websites using robots .txt to stop crawlers from accessing them has shot up dramatically. Interesting. I hadn't heard that. Yeah. Yeah, I think Shane Longpreay was the first author. Yeah, so I totally agree. It's as much economic and political as it is technical.
16:41I think for agents to work, we're going to have to figure out who's going to capture the economic value that they create and how to do that in an equitable way. You know, we jump right into this conversation as I often do, but defining agent is still an open question in many ways. When you think about agents or agentic behaviors, like what are all the elements that come to mind for you? So before I talk about those elements, let me say for a second, you know, one might ask genuinely, you know, is this even a meaningful term or is it just marketing hype, right? And I think that's a fair question to ask.
17:22But here's why, I mean, there is a lot of hype around agents, but here's why I do also think it's a meaningful term. Unless one thinks that companies are going to be able to scale models all the way up to AG high, and I don't think that's going to happen. Right. And so if that's not going to happen, then it must inevitably be the case that to take these models and to do more useful things with them than they can do just through zero shot prompting, you have to build more complex scaffolding. And, you know, you can call it whatever you want. You know, it needs some term. And I think for better or worse for these compound AI systems, as some researchers call them, which is also a fine term.
18:01but the more well-recognized term is agents. So I feel it is something meaningful. So in terms of how to define them, we looked at a bunch of attempts to define agents. And it seems like people have a certain set of factors in mind. One, how complex is the environment in which the system is asked to operate? So a chatbot operates in a really simple environment compared to a web agent or some other kind of agent. and how difficult are the tasks that it's asked to do? Are there multiples, take holders, whether people or other agents, is it asked to operate over a long timeframe? So that's one set of factors.
18:43A second set of factors is, is the user basically babysitting it or is the system given a certain amount of autonomy based on high-level goals to go off and do things in the way that it thinks is the best way to accomplish that goal. So that's the second factor. And the third factor is the design patterns themselves. Is this, you know, does it go just beyond an AI model? Does it do tool use, for instance, or reflection or planning or those types of patterns? So I think there's no strict binary dividing line, agent or not agent. And I think the more factors that a system has, the more agentic it is.
19:25Yeah, I think one of the interesting things and perhaps the central thing in the Agents That Matter paper is this idea that we can look at agent performance in different areas of capability and kind of law these agents. But unless we look at the cost and the efficiency of their operations, it's kind of meaningless. Is that what you're trying to communicate with that work? That's exactly right. So many years ago, Google's alpha code paper showed that if you can just repeatedly sample solutions from a code generating agent, and then you had some Oracle, maybe unit tests or whatever. and that let you, as long as one of the generated solutions was correct, you could somehow pick that solution, then you could just keep doing this and increasing accuracy.
20:21They stopped after one million attempts, but even up until one million attempts, the accuracy kept going up, right? Starting from a really weak model by today's standards, they could just keep increasing the accuracy. So if we don't measure cost, what does it mean to say that you have state-of-the-art performance, right? So you could always just invoke the model more and more times. It's like teaching a kid that there is no biggest number because it's always that number plus one. So it's just a basic insight, but somehow I think people were not willing to admit this because it's just very nice to have this abstraction of a one-dimensional leaderboard.
20:59And the minute you say, no, it's at least two dimensions, you have to consider cost. It's not a leaderboard anymore. You have to talk about what's the best model for this cost budget or whatever. And it's messy to talk about dollar costs, right, when you're trying to figure out what's the best model. So what we say is, no, we can, you know, let's plot these agents or models or whatever on a Pareto curve. One axis is the accuracy and the other axis is cost. And instead of a one-dimensional leaderboard, maybe these Pareto curves are what we should be measuring when it comes to benchmarks. and in the last three months, this has become much more common.
21:35It's almost become the default way to talk about performance now. And then when someone wants to look at this to see if they want to actually use an agent in a particular application, they can think about what their own cost budget is and based on that, how to go about things. And the other thing we say is if what you're trying to do is get the best performance for a particular cost budget, But often the simplest way to do that is just to retry until you get it right, instead of more complex methods like trying to debug the code and so forth. So don't forget brute force is one of the lessons of the paper.
22:10Yeah, I've had that experience personally working with some of the smaller models like Gemini Flash, for example. And I don't recall if currently they have, I think they actually, they do now have a structured output ability. But prior to that, it was like, you know, ask it to give you some JSON and try to, you know, confirm that you received JSON. And if not, just keep trying and eventually it'll give you JSON. And by the way, can I can I just point out my gratification in hearing someone else call it JSON instead of JSON?
22:52I have no idea how consistent I am in the way I pronounce that. And so as a follow on to this idea of kind of cost versus performance, one of the things you're working on now is some benchmarking around agents. Can you talk a little bit about those efforts? So we're taking it in a couple of different directions. One is related to, can we build agents for some very specific mundane tasks that people need to do as opposed to, you know, reasoning is very abstract and very tempting, very sexy. A lot of people want to work on that. But if you've solved reasoning, it's still a little bit unclear what you can actually do with that.
23:42On the other hand, let's take a real problem that many people are struggling with. And as researchers, our mind immediately went to scientific reproducibility. So this is just the idea that when someone puts out a paper and code, other researchers should be able to download that code, set up all the packages, all the messy stuff that's necessary, run that code, and verify that it produces the same results as are reported in the paper. So that's a basic prerequisite for being able to build on that paper, right? And unfortunately, it turns out that so often this basic notion of reproducibility doesn't work out.
24:20It's not because anyone's cheating. It's very rarely that, but it's just that, you know, sometimes it just takes too long and people give up. It's just hard to install a bunch of packages. Other times you have a slight difference in your library version. And because of that, there are minor differences in the code that actually translate to substantial differences in the output, things like this. So it's a very subtle problem. And worldwide, tens of millions of researcher hours are wasted on just reproducing others' code. So our question was, can we automate this? And so my team got together. Zach Siegel is the first author.
24:59And we've put out this benchmark called CoreBench. It stands for Computational Reproducibility Agent Benchmark. And so we sourced this from actual scientific papers, 90 different papers. And this is not just about machine learning. Our papers are across computer science, social science, and medicine. And what we say is, okay, so here's the code, and we give it to the agent at various levels of chaos, if you will, everything perfectly documented as it should be, or we remove the Docker file that's supposed to be there to make it easy to run the code or various things like that. So we simulate the way in which researchers very often don't do all the things they need to do to make it easy for others to reproduce the code.
25:47Meaning so for these 90 papers, you've gotten to the point where they should be machine reproducible and then insert kind of artificial noise to simulate deficiencies. Exactly. Exactly. I mean, they should still be reproducible. You know, a person can do it because they can figure out what's missing and how to fill that in. The code is there, write the Docker files, the requirement files missing, you know, figure that out. Yeah. And this turns out to be surprisingly hard for agents because I suspect, you know, it's hard to pinpoint an exact cause. I think it's a bunch of things, but it requires a lot of back and forth interaction with the shell.
26:26So that requires very long context. Agents actually get confused a lot. They forget exactly what the state of the shell is at any given time. And so that's one difficulty. And then after you run the code, especially in fields outside computer science, it turns out the result is not necessarily in a nicely formatted code output. It's in a PDF. And so now you have to visually inspect the PDF using a vision language model in order to extract what it has actually output. And these are all, you know, hard for a model. They're easy for humans, although mundane, time-consuming, and so forth, right? So that's the kind of task this is.
27:05And this is very different from some of the other tasks, which are more aspirational things like reasoning. And so two main differences, right? One is it's kind of mundane as opposed to some North Star challenge. And the other is if you solve this, you're probably immediately going to save people tens of millions of hours per year. So we think this is a nice challenge. Current agents don't do that well at it, not well enough to be useful. But we can see that even with some simple modifications to off-the-shelf agents like AutoGPT, we can push up the accuracy a lot. So we do think that with a good amount of effort on this, we can get it to a point where we can largely automate the problem of computational reproducibility.
27:50And to be clear, it's not a problem you solve. It's an open benchmark that you're offering to the community to kind of go after this problem. Exactly. I guess I'm envisioning like one possible use is, you know, when you check in your work at a conference, this thing does an assessment for you and gives you like red, yellow, green light kind of thing. But then if it can do that, like why not just fix the problems and put everything into a standard format? Like how do you see it being used? Yeah, yeah, absolutely. Yeah, that's that's the first step, right? So to be able to check if things are reproducible and the next immediate step after that exactly is to fix things.
Read the full transcript
28:33And I think, yeah, once you do the first one, the second one is not going to be that much harder. And we hope that this will lead to the integration of LLMs into, you know, packaging systems and so forth. So that for a researcher to actually take their code and make it reproducible, you can have a lot of AI assistance for that step as well. You asked yourself an interesting question in there in kind of thinking through what makes this particular task difficult for agents. Do you have a mental model of what agents are good at and what agents are not so good at today? Yeah, I guess not one that I can back up with a lot of evidence, but one mental model is that stuff that is well represented on the web, obviously agents are going to be better at that.
29:24And I think in this particular case, the reason it's hard is that it's a lot of back and forth. So that's not so well represented on the web. You know, there's a lot of shell commands on the web, right? So if it's a zero shot thing, what's the command for doing this? In my experience, at least playing with these models, they almost always get it right. But if you have to take a bunch of steps and kind of quote unquote remember what things look like at each point in time, then that becomes much harder. Understanding the state of a messy file system, all of that is harder. And just turn taking, I think, is hard.
29:58One of my favorite examples of this is anytime a new model comes out, I try to play rock, paper, scissors with it. And I say, you go first. And it'll say, you know, I pick paper or whatever. And I'll be like, I pick scissors. So scissors beats paper. And we do this over and over. I keep winning every time. And then I ask the model, why am I winning every time? and it's just amazing that the model doesn't recognize that it's because we're not going simultaneously right so i mean despite so much smarts in other areas despite you know terabytes of data from the web that's one thing the model's not going is not able to get this nature of turn taking unless it's been specifically fine-tuned for that these days models are better at that i suspect it's because developers have specifically um fine-tuned them to to understand understand So I think that's one of the difficulties of automating a long sequence of shell commands.
30:54Yeah. How do you think we get to models that are better able to reason? You know, some of that, quote unquote, emerged from, you know, LLM's training on the Internet. But we talked a little bit about O1 or Strawberry, and they've taken, you know, clear steps to build the model into more of a reasoning engine. Do you have a sense for the trajectory or what people are thinking about or the different ways that we're able to engineer in or focus on reasoning capability? Absolutely, yeah. So I think for LLM-based reasoning, there are three broad approaches. One is to just keep scaling up the models and hope that they continue to improve at reasoning.
31:43And I think that probably will happen if we see bigger models, they will be better at reasoning. But I don't think scaling is going to continue forever, and it's not going to get us to human level reasoning by any stretch. A second way to do it is inference time methods. And in particular, we've talked a little bit about these kind of brute force inference time methods of using a verifier and then trying it over and over. but we can certainly imagine much more complex stuff. And broadly, we can call this neuro-symbolic AI. So the model is the neural part, and there's a symbolic system, whether it's a verifier or a more traditional kind of planning engine or whatever it is, right?
32:26So you try to put a symbolic system and a neural system together and see what they can accomplish together. So for instance, when we look at benchmarks like, what's the Francois Chollet benchmark, the Arc AGI. So in Arc, so his hypothesis is that the best models are going to be neurosymbolic. So you have, you know, you're using an LLM as a clever way of doing program synthesis, for instance, but then you have to actually take the programs, you have to run the programs, you have to look at the output, you have to try to verify the output. So that's a more symbolic part of it, right? So that's the second big approach.
33:04And I think the third approach is what we're seeing with strawberry or 01, which is actually training or fine tuning the models to be better at reasoning. And so there you're using some symbolic system or something else during the training process itself. And I think that's, that's a, that's a new approach. I think this is the first time we've really seen that. And I think that's, that's what makes it exciting. I mean, these are three of the, I don't, I don't mean to say these are the only three approaches, but these are three that are really, as far as I've seen, being pursued a lot. And I think all of them are promising.
33:39And so I think it's exciting times for reasoning. And so do you think of the gains of Strawberry as being more attributable to this fine-tuning than inference scaling? For sure. Yeah. Has that been published? Or what do we know about the way it works? When I look at, and this could just be an artifact that it's displaying, I see echoes of an iterative process and more traditional agent behaviors. Like, I'm going to try this, I'm going to try that, I'm going to try that. as opposed to, you know, when I think of, you know, something that would be, you know, trained or fine tuned into a model, it's like a, you know, a single shot, but the, you know, maybe the time to first token is longer or the token generation time is longer.
34:31It seems more iterative. So I think what's exciting about this model is that it's, it's a true hybrid in the sense that, yeah, it is doing those inference time agentic behaviors, but it's trained to be better at that, right? Instead of just creating to be better at it zero shot. And I haven't verified it myself, but I have seen tweets from other credible people comparing it to what would happen if you took the same inference budget. It is, you know, using up a lot more tokens and inference, but still with that inference budget, if you did it with GPT-4 or whatever, it's really not reaching anywhere close to the same level of performance on benchmarks.
35:08So I do think that that this is something new. And also I would say that the kind of summary that it gives you, people suspect that it's a little bit bogus. Yeah. Yeah. You, I sugar or something like that to. Yeah. Yeah. Probably. Yeah. To, to, to make people feel like pressing the, the street crossing button. It doesn't really do it or changing the thermostat in an office building. Like it's just there to make you happy. Yeah. You know, how does all of this kind of translate to, you know, or with all of this in mind, like, you know, recontextualize this idea of snake oil and what it all means for consumers broadly, which I think is your target for that book.
36:00Yeah, definitely. So almost none of what we've talked about is snake oil, right? I mean, these are researchers and developers trying at very hard problems. Sometimes they fail, but that's perfectly okay. The only little bit dodgy product that we discussed is some of these agent hardware-based products that promised a lot more than they could deliver. And again, I don't think that was because the developers were trying to fool anyone. They just genuinely underestimated all of the practical difficulties, the capability reliability gap and just how hard the user interface is because agents are messy enough if you're looking at them on a desktop computer and you're babysitting them and making sure they don't do things they're not supposed to do.
36:44But then if you have a very limited voice based user interface, it's a million times harder to get the agents to work reliably enough. I've not been able to get OpenAI's advanced voice mode to come anywhere near the demo performance that I saw. Yeah. There. Yeah. So, you know, clearly a lot of hype, but it's, you know, again, there's not, you know, it's not a completely broken product. It's researchers really trying to solve hard problems, et cetera. Want to give companies the benefit of the doubt here, right? But I think there have been claims that are much more problematic. So let me talk about somewhat problematic claims and then some really problematic claims.
37:26So when GPT-4 came out, OpenAI bragged about how well it did on the medical licensing exam and the bar exam and so forth. And it's true in a very narrow sense it did well on those exam questions. But I think the way this was presented, the way it was amplified by the media and others overall, it portrayed a really, really wrong impression for the public as if these models were now capable enough to take over a lawyer's job or something like that. And that, as we know now, with the benefit of hindsight, was very, very far from reality. Anytime lawyers have tried to use these models in non-trivial ways, I think their results have been pretty disastrous, like hallucinating entire cases and then lawyers getting into trouble with the judge for submitting incorrect information and so on.
38:13And I think one reason for that is probably contamination on benchmarks. A more important reason for that is that lawyers' real jobs look nothing like the bar exam. The bar exam is a very constrained setting, a setting that in many ways is almost designed to let AI really shine. And there's a generalization gap between what it's being tested on and what it needs to do in the real world. And, you know, these are high stakes domains, errors have a very high cost and so forth. So I think companies, unfortunately, really freaked people out a year and a half ago. And, you know, working in the policy area, and on the societal impacts of technology, I saw the extent to which people were panicking and expecting massive and imminent job loss and not being sure how to react to this.
39:02So I think that was pretty problematic, that sort of thing we call out in the book. But the most problematic thing is actually not even in the generative AI space. It's companies who are using the hype around AI to take basic regression models that are just statistics that we've known how to do for 100 years, but then rebranding them as AI and selling them as something far more sophisticated than they are. So quote-unquote AI is used in all kinds of things, in hiring, in the criminal justice system, in various other domains. And so in criminal justice, for instance, what these algorithms might be trained to do is look at various characteristics of a defendant and predict what is the probability that they will fail to appear at trial.
39:51So when someone's arrested, there's this risk assessment, right? And someone is judged to be too risky, then bail might be denied or set at a very high amount or whatever, and they might be jailed until their trial, which might be months or years away. Now, just going back to the Compass work from years ago, are there more contemporary examples around criminal justice? Are other folks trying to do the same thing with the same classes of models? Yeah, I mean, this is happening all over the place. And these models are so brutal. Let me give a couple of examples. There's one by Arnold Ventures called Public Safety Assessment.
40:29I mentioned this because Arnold Ventures is an organization that I respect a lot, right? So these are, you know, this is not a company that's trying to fool anyone. This is a philanthropic organization really trying to make the criminal justice system better. And they've built the model with really good intentions. And yet, the problem is not that they are being irresponsible or something like that. It's just fundamentally that you can't predict people's behavior that well. And so the intrinsic, the irreducible error, as we often call them, of these models is very, very high. And so you're making decisions about someone based on little more than the flip of a coin.
41:09And so that's really morally problematic. So even the system PSA by Arnold Ventures, which was developed much more responsibly than Compass, right? But one of the failures here was it was calibrated on a national data set of defendants. But as we can imagine, crime patterns in various parts of the country differ dramatically. So when an investigative reporter looked in a particular county, for instance, in Ohio or somewhere, I forget, where the base rate of crime was relatively low. The model was not accounting for that. And apparently it was jailing 10 times the number of people that really should have been jailed, right?
41:44So that's obviously very problematic. Let me give you one more example from healthcare. Epic is a health tech company that is very ubiquitous. Yeah, exactly. And they built this model back in, I want to say, like 2017 or so. And it was for predicting which hospitalized patients might develop sepsis. Sepsis is deadly, and early detection really saves lives. Again, this is very well-intentioned, nothing wrong with trying to detect sepsis. But they claimed it had an accuracy in the sense of AUC of between 0.76 and 0.83. So that's pretty high. And so on the basis of this, it was deployed in hundreds of hospitals.
42:29but it's a proprietary system so the doctors or the you know the it teams in the hospitals couldn't examine these systems really and it was only much later that external doctors were able to publish a validation of the system and it turned out the accuracy was in fact much much lower in AUC of 0.63 and 0.5 is random right so this is only slightly better than random and it turns out what had happened was a classic error called leakage again this is not epic trying to fool people, but more fooling themselves. One of the features, one of the variables that was part of their set of predictors was whether the patient had been prescribed antibiotics by the clinician for treating sepsis, right?
43:11So the model is making the prediction once it's no longer useless and they had not even recognized this. And once we realized that the achievable accuracy here is pretty low, it's questionable whether we should be using this kind of system at all. And in fact, this blew up on Epic and they had to stop selling this one size fits all system. And now they're just saying, you know, here's our model, go train it on your own hospital's data set. And I think that's the more responsible way to do it because every patient population is going to be so different and there's going to be a distribution shift between those.
43:44And so these plug and play AI systems don't really work in these high stakes domains. And so those are the kind of more dangerous applications of AI that we call out of the book. And so from the perspective of advising policymakers, like how do you present that perspective to them? One of the things that we hear about, talk about frequently is that tech has gotten so far ahead of policymakers. They have no way of really comprehending what's happening. You're close to them. What do you see happening in the policy domain? Yeah, definitely. So yeah, tech does move very fast. And politicians are not technologists.
44:33I think there is a kernel of truth to that. But then going from there to saying tech policy is completely hopeless, I think that's a very big leap and involves, I think, a series of misconceptions. So I think the first thing to note is that the politicians that we see on TV, those are not the ones, you know, making policy right at the level of the details of the legislation. They might set priorities. They might say, oh, you know, everybody's concerned about AI failures in health care, so we better do something about it. But then it's up to the staffers and others to really turn that into legislation.
45:08But then again, it's up to the enforcement agencies, which are a whole different set of entities for actually putting that into practice. So if there's medical AI legislation, then the FDA will be tasked with specific rulemaking. And yeah, there's a whole complicated process. So there is technical expertise in many parts of the system. Maybe you could argue that it's not enough. But it doesn't have to be that senators need to be tech experts. That's not the problem, right? If anything, you know, it's more things like some of these agencies being under-resourced and not having enough budgets to do their jobs and that sort of thing.
45:45So that's one aspect. The other aspect of this is that regulation is not about the details of the latest models. It's not, you know, the regulation is not saying, you know, what should your metrics be for, you know, for GPT-4. That's not it's not how regulation operates. Right. It's in fact, usually it's not about the technology itself. It's about human behavior around the technology. Right. You shouldn't use it for in a discriminatory way and that sort of thing. And so in that sense, regulation doesn't have to even make any reference to the internals of the models. And so it's not as hard as people think.
46:26So I think if there is political will, if agencies were well-resourced and Congress were less dysfunctional, I think we could be doing a much better job with these things. Tech moving fast is, I think, a little bit of a problem, but it's not by any means the biggest problem with policy. And so how do you see policy evolving currently? What are the priorities and where are folks really digging in to try to figure things out? Yeah, so a big part of the priority right now is just enforcement. So enforcing long existing laws, right? So the FTC, for instance, I think it was just last week, came out with a sweep of cases against AI companies for misrepresentation.
47:09So we were talking a minute ago about AI replacing lawyers, and there is this company, Do Not Pay, that has been making those claims. They've built a robot lawyer, and they've been in the news a lot. Wasn't that a kid in the UK or something? Am I thinking of the same thing? I think it was a Stanford student who started it. They definitely operate in the US. one of their publicity stunts was saying they would pay a million dollars to any attorney who went and argued a case in front of the Supreme Court by having this robot lawyer talk to them in their earpiece, which you know they're bullshitting because that's first of all illegal.
47:52Electronic devices are not even allowed in the Supreme Court. Anyway, and of course they didn't have a robot lawyer. And so the FTC went after them and a bunch of other companies for making false claims about their AI products, right? So while the actual action here was about AI companies, the FTC's authority to police these kinds of false claims in the marketplace is, you know, 100 years old, right? And so that's one thing people miss. There's so much we can and should be doing even without new legislation. And yes, in some cases, this new legislation is needed. And I think there's two aspects to that.
48:30One is around the specific behaviors that we want to prohibit. So for instance, using image generators to create non-consensual nudes, I think that's been a really huge problem. It's affected hundreds of thousands of primarily women around the world, lots of teens. We've been hearing so many cases of this in high schools. So yeah, so on those specific problems, regulators need to be taking, legislators need to be taking action. But then there's like the big question of what to do about the industry, big problems like copyright, which don't have a simple solution in terms of prohibiting a specific action, what to do about the danger of catastrophic risks.
49:13Those are the areas where there are kind of bitter fights going on. And I certainly have my perspectives on that, But definitely a lot more debate is needed before we can get to clarity on those types of questions. And I guess on that point and your perspectives on that, I'm inferring from the AI risks work that you feel like a lot of that is overblown. at least from the perspective of LLMs teaching people how to make chemical weapons and those ideas that were thrown around that you... I guess I kind of took that paper as kind of a takedown of shallow thinking there. One of the things we recognized in pushing back against hype in terms of capabilities is that a lot of that also applies to the risks that people are alarmed about.
50:04We've been pushing back against those. Certainly, you know, researchers should be studying the risks, and I'm glad that there is a very vibrant safety research community. But I do think policy should still be evidence-based. It should, for the most part, focus on known risks. And certainly, we need to be making preparations to deal with, you know, more catastrophic risks if we have at least even circumstantial evidence to think that these are more realistic. So things we could be doing now are things like insisting on much more transparency from big tech companies who might know more on the inside about the potential for those risks than we might on the outside.
50:47So I think instead of doing that, trying to legislate against risks that currently are in the realm of science fiction might become reality, but it really depends on models behaving in a very different way than they do today. I think the potential for going really, really wrong with that and curbing AI development for little benefit and not even necessarily improving safety, I think that risk is too great. So while I do support action on AI safety, I think that action has to start with collecting a lot more evidence than we have now. Do you believe AI poses catastrophic risk? And sounds like no.
51:25But then what are the ways that you kind of taxonomize that risk? There are many ways to taxonomize it, but from a policy perspective, you can broadly think of two kinds of risks. One is, let's say, you know, AI that's used to discriminate against people or something like that, right? So that, you know, I would say is not a catastrophic risk. It might be catastrophic to a person who is affected, for sure, not minimizing that in any way. But from a policymaking perspective, we have a very strong presumption that we don't act unless we think those risks are real, as opposed to someone describing in a paper, hey, look, this might happen sometime in the future.
52:03So that's not a catastrophic risk. On the other hand, something that clearly would be catastrophic is an adversary taking over critical infrastructure. And the reason that's catastrophic is that it's not the kind of thing we can necessarily easily recover from. So it requires action even before such risks have materialized. It requires an application of the precautionary principle, right? And in the US, at least, we're generally very, very cautious. We don't rush into the application of the precautionary principle because we recognize that premature regulation also has a lot of downsides. Different countries or regions of the world might have different approaches.
52:46But I think in keeping with our approach before we can say this risk is so catastrophic that we can't recover from it, you know, bioterror or pandemics and so forth that could be aided by AI, that we should put limits on people's freedom because we're worried that this might potentially happen in the future based on models behaving very differently than we do today. To me, that's a bridge too far. I think, you know, let's gather more evidence on that. Awesome. Well, Arvind, thanks so much for taking some time to share a bit about your research and what you've been working on. Very great stuff.
53:27Thank you, Sam. This has been really fun. Thank you.
53:45Thank you.
From the publisher
Today, we're joined by Arvind Narayanan, professor of Computer Science at Princeton University to discuss his recent works, AI Agents That Matter and AI Snake Oil. In “AI Agents That Matter”, we explore the range of agentic behaviors, the challenges in benchmarking agents, and the ‘capability and reliability gap’, which creates risks when deploying AI agents in real-world applications. We also discuss the importance of verifiers as a technique for safeguarding agent behavior. We then dig into the AI Snake Oil book, which uncovers examples of problematic and overhyped claims in AI. Arvind shares various use cases of failed applications of AI, outlines a taxonomy of AI risks, and shares his insights on AI’s catastrophic risks. Additionally, we also touched on different approaches to LLM-based reasoning, his views on tech policy and regulation, and his work on CORE-Bench, a benchmark designed to measure AI agents' accuracy in computational reproducibility tasks.
The complete show notes for this episode can be found at https://twimlai.com/go/704.




