#156 Gordon Crovitz: Will AI SPREAD Misinformation & Fake News?

22 Nov 2023 · 48 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. - Episode #156 Summary

Episode Title

Gordon Crovitz: Will AI SPREAD Misinformation & Fake News?

Podcast Description Eye on A.I. is a biweekly podcast hosted by Craig S. Smith, featuring discussions on the implications of artificial intelligence with industry leaders and innovators. Episode #156 focuses on the growing issue of AI-generated misinformation, featuring guest Gordon Crovitz, co-founder of NewsGuard.

---

Key Points

Introduction (01:26)

  • Craig Smith introduces Gordon Crovitz, highlighting his background as a former publisher of the Wall Street Journal and co-founder of NewsGuard.

NewsGuard's Mission (02:24)

  • NewsGuard aims to combat misinformation through two main databases:
  • Reliability Ratings: Assessing over 30,000 news sources based on credibility and transparency using human analysts.
  • Misinformation Fingerprints: Cataloging significant false claims spread across the internet, including their sources and debunking information.

Challenges in Detecting AI-Generated Content (05:58)

  • Crovitz discusses the difficulties in identifying AI-generated content and the reliance on human analysts for this task.
  • Currently, there is no "silver bullet" method to systematically detect AI-generated misinformation.

The Impact of AI-Generated Misinformation (08:36)

  • AI-generated content is on the rise, posing risks to the integrity of information online.
  • Crovitz cites a specific example of an AI-generated false story about Israeli Prime Minister Netanyahu, illustrating the potential dangers.

The Future of AI Agents and Their Role in Content Generation (14:19)

  • The conversation shifts to how generative AI models can potentially enhance the spread of misinformation.
  • Crovitz warns about the ease with which malicious actors can generate and distribute false information using AI.

Use Purpose of NewsGuard's Databases (17:35)

  • NewsGuard's databases help AI models enhance accuracy by reducing the incidence of misinformation dissemination.
  • Collaboration with companies like Microsoft allows integration of reliability ratings into platforms, improving trustworthiness.

Tackling Misinformation Effectively (22:34)

  • Crovitz emphasizes the need for technical and regulatory approaches to combat misinformation.
  • Social media companies and search engines hold significant responsibility for curbing the spread of unreliable content.

Government and Regulatory Responses to Misinformation (29:38)

  • Discussion on existing regulations and the need for more comprehensive measures to address misinformation.
  • Crovitz notes that regulatory responses are currently limited but evolving.

The Role of Social Media in Misinformation (32:36)

  • The responsibility of social media platforms in managing misinformation is criticized.
  • Crovitz points out the reluctance of some platforms to inform users about the nature of news sources.

Impact of AI on News (39:22)

  • The conversation reflects on how AI-generated content can significantly alter the landscape of news consumption.
  • Crovitz expresses concern about the long-term implications of AI on public discourse.

Search Companies' Role in Combating Misinformation (42:32)

  • Crovitz discusses the need for search engines to prioritize reliable news sources over misinformation.
  • Collaboration with companies like Microsoft is shown to improve the reliability of AI in news generation.

Regulatory Measures to Stop Misinformation (44:13)

  • Crovitz believes in the potential for regulations to improve the situation but acknowledges the complexities involved.

---

Key Takeaways

  • Misinformation Challenge: The advent of AI has led to an increase in the scale and sophistication of misinformation, necessitating proactive measures from various stakeholders.
  • Human Oversight: Human analysts play a crucial role in identifying and countering misinformation, indicating a need for a human-centered approach to AI oversight.
  • Collaboration and Responsibility: There is a pressing need for collaboration among AI companies, social media platforms, and regulatory bodies to establish effective measures against misinformation.
  • Future Outlook: While there are significant challenges ahead, the potential for trust data to enhance AI safety and reliability offers a promising pathway for mitigating the spread of misinformation.

---

Conclusion In this episode, Gordon Crovitz provides valuable insights into the challenges and responsibilities associated with AI-generated misinformation. The discussion emphasizes the importance of collaboration, human oversight, and regulatory frameworks in the evolving landscape of digital information.

For more information, please visit [Eye On A.I.](https://eye-on.ai).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00One of the challenges for the large language models is that, with the exception of Microsoft, we're not aware of any of them that have used fine-tuning to differentiate between generally reliable and generally unreliable. news sources. They tend to go by the number of tokens that a site has. Hi, this episode is sponsored by Salonis, the global leader in process mining. AI has landed and enterprises are adapting, giving customers slick experiences and the technology to deliver. The road feels long, but you're closer than you think. You see, your business processes run through many systems, creating data at every step.

0:43Solonis reconstructs this data to generate process intelligence, a common business language. With process intelligence, AI knows how your business flows across every department, every system, and every process. With AI solutions powered by Solonis, enterprises get faster, more accurate insights, a new level of automation, and a step change in productivity, performance, and customer satisfaction. Process intelligence is the missing piece in the AI-enabled tech stack. Search Celonis, C-E-L-O-N-I-S, to find out more. Hi, I'm Craig Smith, and this is Eye on AI. In this episode, I talk to Gordon Krovitz, one of the founders of NewsGuard, an organization that helps identify misinformation, disinformation, and AI-generated content on the internet.

1:45In the few years that NewsGuard has been around, they've built some very significant databases that are helping large language model providers, search companies, and even social media identify and flag unreliable content. This is an issue that's growing fast with the advent of AI agents built on generative models that can handle the end-to-end workflow of disinformation creation and distribution. The conversation is troubling, but I hope you find it as informative as I did. I'm Gordon Krovitz. I'm co-CEO of NewsGuard and spent several decades as a journalist. I was publisher of the Wall Street Journal.

2:34And along with Steve Brill, another journalist, we started NewsGuard five years ago in order to counter misinformation. And our mission is to counter misinformation on behalf of news consumers, brands, and democracies. And we do that with two main databases. One is we have reliability ratings on more than 30 ,000 sources of news and information. We do this using humans, not using AI. We have journalistically trained analysts that use nine different criteria to rate news sources. These are apolitical criteria around credibility and transparency. And every source gets a point score from 0 to 100 and what we call a nutrition label write-up to explain the nature of the source.

3:26And those reliability ratings are made available to consumers through companies like Microsoft, which has integrated that into its edge browser and other parts of its services. We have licensees use those reliability ratings for functions like deciding which news sources to aggregate if they're a news aggregator. If they're a brand or ad agency or ad tech company, they use those ratings to decide where should ads go and which sites are not brand safe. Could be Russian disinformation sites or healthcare hoax sites or conspiracy sites. So those reliability ratings are one database. The other one relevant for our AI discussion is what we call our misinformation fingerprints.

4:14This is a large catalog of all of the significant false claims spreading on the internet. So that again is everything from Russian disinformation claims, healthcare hoax claims, conspiracy claims of various kinds. And those fingerprints include the statement of the claim, debunking of the claim, and then tools for both humans and machines, Boolean search terms, hashtags, et cetera. And we often find the origin of the false claim, the provenance of it, and how it spreads. And that too is done by humans. So for example, we have a team of journalists who look after 400 malign actors, sources we've identified, that account for virtually all of the Russian disinformation claims.

5:06There are different government offices, different websites, RT and Sputnik, many YouTube channels, social media accounts that have started and spread Russian disinformation, to take that example. So we have humans, again, journalists who have domain expertise in those areas. They're monitoring those sites all the time. When they see a new claim, they ask themselves, I wonder if that's really true, and do the reporting to see if it's really true or if it's not. And between those two databases, the reliability ratings database and the misinformation fingerprint catalog of false claims, that's used now by generative AI models greatly to reduce the hallucination or the making up of or spreading of false claims in the news.

5:58yeah and uh you know as i we were talking earlier the what brought me to newsguard and for listeners i've known gordon as an acquaintance for a long long time because we were both at the wall street journal i was a reporter um i was doing a piece on a paper uh about um generative ai models that degrade over time or their output degrades over time when they're trained on output from generative ai models because the you lose the long tails in the distribution you end up everything converges on the mean and it becomes not very interesting or accurate. And one of the concerns is that AI-generated content is the fastest growing segment of information on the internet.

7:05And while it's still very insignificant, it's growing and will someday represent a much larger segment of the Internet on which large language models and other kinds of AI models train. And that'll affect the accuracy of the models. And I cited NewsGuard, one of your reports, without knowing that you are one of the founders. And that report was about how many AI-generated sites you had identified. So these two databases were compiled by humans, I would imagine with some search or some filtering software. And I guess we'll talk about the disinformation. I'm certainly interested in that. But in terms of identifying AI-generated websites, how do you do that?

8:09Because there's a lot of talk right now about, you know, in education and other arenas, you know, being able to identify AI-generated content beyond just kind of the feel that you get when you read it. Do you have, is there any systematic way of identifying AI-generated content? So we have not discovered the silver bullet for identifying AI-generated content. We actually do it an old-fashioned way, which is that our analysts are on the lookout and using some search tools and other tools to identify what looks like it could be generative AI produced. And there are telltale signs. The occasional, you know, as an AI model, I cannot, for example.

9:05You know, literally that is included in some of the writing. And there are other telltale signs. It's not foolproof. And I know we're missing many, many, many, many sites because we're not able to identify them. But we started looking for them a couple of months ago, so in the early fall. And we're now up to 557 what we call AI-generated news sites. We call them unreliable AI-enhanced news sites. So 557 of them. They have names like Ireland Top News. So many of these take the names of what look to a consumer like a regular news site. Many of them get programmatic advertising revenue, and that may well be why they were created.

10:01There's a lot of money in it. NewsGuard did a study with Comscore that found there's$2.6 billion a year unintentionally being spent by advertisers on misinformation sites. So this is well beyond the AI-generated ones, but that is an incentive for people to create these websites. Your experience with AI models building on themselves and making, you know, here's the way I think about the internet before and after AI. On the pre-AI internet, there was already an enormous amount of misinformation on every topic that you can think of, healthcare included. And the Russians and the Chinese and the Iranians spend hundreds of millions of dollars, maybe more, on their operations, and they're very good at it.

10:54We now have the AI-enhanced misinformation on the internet. We discovered an example actually just this week, and let me tell you about it because I think it reinforces what you observed. we found on a website that was one of these 557 we had identified as ai generated this website is called global village space and this website published a story claiming that benjamin notanyahu the israeli prime minister that his psychiatrist had committed suicide And the story went on a great length about how frustrated the psychiatrist was and what a terrible patient Netanyahu was. Very well-written story. And what we determined was that it had used an AI model to rewrite a story from 2010 on a different website that was a satirical story making up this claim.

12:02But as an AI model, of course, it couldn't recognize satire or pass on satire. And it looked like a truly legitimate, well-researched story. And let me just give you a little bit of an excerpt from it to give you a sense of how well-written it was. So this AI written article claimed that the suicide note, which didn't exist, of course, written by the psychiatrist who didn't exist. This suicide note, quote, painted a grim picture of a man who had tried for nine years to penetrate the enigmatic mind of Netanyahu, only to be defeated by what he called a, quote, waterfall of lies, unquote. And then in this AI-generated article, there was a subsection called Shocking Diary Entries, and it said that Netanyahu had equated Iran with Nazi Germany and went as far as dubbing Iran's nuclear energy program a, quote, flying gas chamber, while also suggesting that all Jews were living perpetually in Auschwitz.

13:07So, unquote. So in other words, what this generative AI tool had done was it had been prompted by a malign actor to rewrite this story. It then was posted on this website. It was then picked up by numerous social media accounts across multiple languages. and it was featured on official Iranian broadcast news as an accurate story. So this is one of the fears of the generative AI models, which is that they enhance the work of malign actors, that if you're in the business of spreading false claims about a political leader or on any topic, at least in the old days, you had to do the work yourself.

14:05all you have to do is prompt a generative AI model and you end up with a news story that's so believable that many millions of people saw this story, which was entirely untrue. Yeah, and we're only at the beginning. I've been talking to people recently about AI agents, which have existed for a while, but when generative AI, particularly when GPT-4 came out, and this was one of the things that led to the so-called pause letter and people like Jeffrey Hinton starting to warn of the threat of AI, that this, the GPT-4, which is, you know, a question and response chatbot, can be turned with very little engineering into an agent where it is actually taking actions.

15:13And OpenAI has since introduced its own agents based on these foundation models. And once you have an agent that can take action, you no longer need the, you know, you can just turn it on. Eventually, it means it'll take some engineering, but it is only an engineering problem. A large language model can then create websites, register with ISPs, populate websites. and there will be potentially a combinatorial explosion of this deceptive content. And that's concerning for obvious reasons, for the political discourse and democracy. I'm particularly interested in what that does to the internet or as, as a data source for training AI.

16:27have you talked to I had a conversation with a guy who works on safety at Google he's actually a quantum physicist but he's on a sabbatical working at Google on safety and he's working on trying to figure out or embed statistical fingerprints in AI generated text so that even to a human reader, it wouldn't be obvious that it's AI generated to a computer trained to recognize those statistical anomalies. It would be flagged as AI generated and then you could remove it from data sets or flag it for the public and that sort of thing. As NewsGuard, are you following that research? Are you guys talking to people about those kinds of strategies?

17:35And we've also been involved in red teaming of all the major models. We've done some of those publicly. So, for example, with ChatGPT, which you mentioned, we red team ChatGPT 3.5 and 4. and our red team involved taking a sample of 100 of our misinformation fingerprints, the false claim, and looking to see out of those 100, how many would the models repeat, enhance, how many of them would they recognize as being false in some way. And with chat GPT 3.5, it recognized the false claim 20 times out of 100. In other words, spread the false claims 80 times out of 100. We were very excited to test chat GBT 4.0, which, of course, passed the bar exam.

18:28But it spread all 100 out of 100. Wow. But here's what we have learned. Microsoft has had an enterprise-wide license for NewsGuard data for some years, and it were used in many places, including Bing, traditional Bing search. And when Microsoft launched Bing Chat, it had had access to our ratings database for new sources and our catalog of false claims. The result of that is, you know, in contrast to ChatGPT, if people do a search or prompt rather on bing chat it's highly likely to treat different sources differently to identify a false claim in the news so i'll give you a real life example one of the russian claims about its war in ukraine is that there are americans fighting and dying in ukraine and that i guess is designed to dissuade the west from defending ukraine one of the false narratives involved a woman, the Russian disinformation site's claim was a mercenary.

19:40Her name is Rebecca Macirovsky. So if you do a prompt on was Rebecca Macirovsky a mercenary killed in Ukraine, you will get essentially the Russian answer on most of the generative AI models. On Bing Chat, in contrast, it says certain websites like RT and Sputnik, identified by NewsGuard as unreliable sources of Russian disinformation, say that she was a mercenary. On the other hand, sites that get a high score from NewsGuard, like Reuters in the New York Times, say that she was not. And then it often also says this was also identified as a false claim by NewsGuard. So, you know, that is delightful because it gives an answer to that prompt and it gives some context and some tools for consumers to decide what they want to believe or not believe.

20:42And it tells us that with, you know, different approaches to fine tuning and guardrails, that these generative AI models can indeed recognize a false narrative and take steps to mitigate. And so whether it's AI-generated misinformation from a malign actor prompting the AI model to create new content, or if it's an innocent person using ChatGPT or other chatbots kind of as a search tool, asking, is it really true? it will give an accurate answer instead of looking for the next likeliest word, which in the case of a Russian disinformation claim, the likeliest next word is coming from RT or Sputnik or TASS or Provda.

21:40So I think the challenge, the technical challenge of identifying what's AI-generated content and what's not, I think that's an enormous task. And perhaps people will figure out how to do it. I think we solve a simpler problem in a simpler way, which is, with the right trust data, can all the generative AI models differentiate and treat different sources differently? And can they be trained to identify a false narrative and take steps to mitigate the spreading of it? And so far, based on the work that Microsoft has done, it looks you know, quite promising. You know, we can't solve every problem, but, you know, we're focused just on significant topics in the news.

22:28But, you know, that is where so much of the misinformation is likely to come from. Yeah. And once you have these two databases, which I'm curious, maybe you can talk a little bit about building them. But once you have them, then you can use search or other tools to find sites that are using the information sites that you haven't discovered yet. Then you add them to the database, I would assume, and then it keeps growing. But how did you build the database initially? How many people? And are you guys nonprofit? I mean, how did you fund this? It sounds like an enormously labor-intensive exercise.

23:17It is an enormously labor-intensive exercise. But, Craig, once a newsman, always a newsman because you got right to the heart of how we are the process for this. So we've spent$20 million. We're a for-profit to create these databases. The reliability ratings database was the first one that we created. So we identified all of the news and information sources that accounted for 95 % of engagement, first in the U.S., and we now operate in 10 countries. And we identified nine criteria, basic journalistic practice. Is there a corrections policy? Is ownership disclosed? That sort of thing. And we had humans rate all of those sources.

24:05That's thousands of websites and now a total of about 35 ,000 sources, including YouTube channels and social media accounts. Very labor intensive, but scalable. So we've hit that 95 % figure in all the countries in which we operate. And that means that if somebody is in their Facebook feed or doing search and they see a story from a news outlet 95 % of the time, they'll be, if they have access to it, a little icon indicating the score from NewsGuard and one click to get a full write-up from what we call our nutrition label. We developed that second database, the misinformation fingerprints, off that process that you described.

24:55That is, as we identified new false claims, as we were rating news sources, we began to catalog them. And once we began to catalog them, we were then able to use tools and third-party data from companies like Meltwater and others to find all versions of a false claim anywhere on the open internet. We were then able to identify new sources of false claims. And those new sources of false claims had yet more new false claims. So it's a very effective and efficient cycle with humans at the center. So this is, we have AI solutions. We're not an AI company. Yeah. No, go ahead. I was going to say we use machine learning and other tools to scale the work of our humans.

25:53But, for example, we discovered that false claim about Benjamin Netanyahu having identified these unreliable AI news sites and monitoring them to see what they came up with. Very human thing to do. If you're a news person and you're tracking the news, you can pretty much instantly see what looks like a new claim. And you can ask, well, is that really true? And figure out, you know, quite quickly, is it likely to be true or not? Yeah, this reliability score badge or nutrition label, as you called it, that's available on Facebook? I've never seen that. It's available in one of two ways. We have a browser extension that's available on all browsers.

26:44So if you're on Chrome or Edge or Safari and you subscribe to our browser extension, you'll see it on Facebook and throughout the Internet. It's asking a lot of consumers to download a browser extension. So our preferred way of getting consumers access to our ratings is through third parties. I mentioned Microsoft. They're a large licensee of ours. And if you're on edge or in their discovery area elsewhere, you have instant access to news guard ratings as you're seeing stories from news sites. Yeah. Are you, as I was saying, as AI agents take hold and expand or extend the capability of large language models, Are you concerned that the generation of misinformation, disinformation is going to outrun our ability to flag it or combat it?

27:54I think the way I would put it is that as misinformation generated by AI is used as training data for yet more AI, the amount of misinformation is absolutely going to grow at some geometric rate. Well, that is, take somebody with a more advanced math degree than mine. um but having said that in the area that you know we're most focused on which is a spread of misinformation on topics in the news i think we've already shown that you can scale an answer to that problem in other words as ai is is is trained on ai that misinformation the number of false claims um is not going to grow geometrically because the ai models don't make up new claims they simply find new ways to repeat ones that already exist.

28:50So the fact that we have already managed to scale to handle the significant false claims across a number of different topic domains gives me hope that if there's a human will to use that kind of trust data to improve the trust and safety of the AI models, that the AI models would become safer and more trustworthy and maybe even could preempt the problem you've outlined. If the AI models recognize false claims and take steps to mitigate them, then maybe they won't spread them to one another. And maybe that will turn out to be one way to address the problem. Yeah. And you have buy-in from the major LLM companies beyond OpenAI.

29:49in using your databases to exclude things from the training data? Just to be clear, the AI model currently using us is Microsoft, and that's because they have had access to our data for some years. So when they were trying to make Bing Chat safer and more trustworthy, they had access to that data. We're in discussions now with all or virtually all of the large language models. They have a kind of, we have the benefit of the real-time case study of how much better Bing chat is on these topics than other AI models because it has access to trust data. So we're optimistic about it. And to us, the big question is, you know, do these AI models want to solve the problem?

30:36and you know we of course lived through the social media era when it turned out you know that facebook and others um at the end of the day didn't really want to solve the problem you know it finally dawned on us you know eventually that uh publishing misinformation that led to engagement and more eyeballs and more revenue was not a bug of that system that was a feature of that system. I think the AI models are in a different position, which is there's nothing in it for them to be spreading misinformation. They're not generating ad revenue. They're not trying to spread misinformation to make money off it.

31:15They want to be trusted by large companies and governments. And we know from conversations with trust and safety folks that the AI models, a lot of them are refugees from social media companies, you know, where they couldn't get done what they wanted to get done, whether it was at Facebook or elsewhere. So those discussions are, I would say, much more engaging and encouraging. There are a lot of reasons for urgency on their part as well. Sam Altman and others have been completely transparent about the risks on factual matters, and they completely understand the problem. The best academic work that I've read has been by people at OpenAI and the other companies.

32:04They know the risks better than anyone. So I think as they can see that there are solutions, at least in some domains, to the hallucination problem, I'm quite optimistic that they'll want to solve the problem as quickly as they can. Yeah.

32:24Well, I have two things I want to talk about. One is the filling up of the internet of AI-generated content, but I'll hold that for a minute. Social media, which continues to be the main vector through which this kind of content is spread, do you really think that the owners of the social media platforms are that cynical that they just don't care and that, in fact, it's a business decision to let this stuff run wild because it generates revenue? And are you hopeful at all that as the volume of misinformation, disinformation grows, which it certainly will, as I said, with the advent of AI agents that can sort of do the disinformation workflow end to end, that they'll get concerned enough that they will want to address it?

33:34I mean, I would imagine you're talking to these guys. What's the attitude that you get from them? Yeah, so I now understand why the social media companies have been so reluctant to give their users information about who's feeding them the news on their platforms. We work with a lot of researchers, academics, they license our data for their own work. And we've come to believe that something like 15 to 20 % of the users of the social media platforms are getting most of their news in one area or another from misinformation sources. So imagine if you're Facebook and you have those data too. They don't share the data, but they know.

34:20They have the data. They rate news sources themselves. They don't tell news sources what their rating is. Their criteria are secret and probably done by algorithms, not by humans. But they do the rating so they know the amount of misinformation on their platforms. Imagine providing their users with tools so that 15 % to 20 % of their users suddenly see that much of what they've been consuming on those platforms is misinformation. That would be embarrassing, probably not good for the share price. So I have become skeptical. I don't want to go so far as to say cynical. But they've had years to fix this problem.

35:00And with the exception of some platforms, again, including Microsoft, which has a different business model, they've really been reluctant to take responsibility for the known harms that they're causing or to give information to their users about the nature of the sources they're relying on, even though within the social media companies, they completely understand it. By the way, the people who work in trust and safety understand it. They're just not the decision makers. Yeah, that really surprises me. I mean, it's one thing to not want to intervene on freedom of speech grounds or claiming that they don't want to get involved in adjudicating right and wrong.

35:53But on disinformation where you can clearly demonstrate that something is false. They have your databases or access to your databases, and there are tools they could just delete anything. In fact, Craig, we have said since we started the company that we're an alternative either to the secretive ratings by the digital platforms or by government censorship, even worse. So our ideal is that people should have more information at their fingertips about the sources, not that anything be censored. I think the difference... The problem with that is that the average consumer is frankly not sophisticated enough or doesn't have the energy or time to do that.

Read the full transcript

37:07It's the responsibility of the platforms that are carrying this. It's absolutely their responsibility. But we have a lot of data based on people who have access to NewsGuard ratings that when they see the rating and they see a low score and a warning proceed with caution and they understand the kinds of claims that site has made in the past, they're highly unlikely to share that content or to believe it you know it kind of solves the crazy uncle willie problem crazy uncle willie keeps sending me these conspiracy sites i know they're not true but i you know can't prove it and you're right consumers don't have time to research every false claim that's why there need to be intermediaries third parties you know that will help i've come to believe that part of the problem which is not surprising i think craig to you and me as people who spent years as journalists.

37:54It turns out that people at the Silicon Valley platforms don't have great news judgment. They're great at some things. They don't have great news judgment. My favorite example of that is that when Russia Today, RT, the Vladimir Putin-funded website and broadcaster, when it became the first channel on YouTube, the first news channel on YouTube with a billion page views, video views. A very senior executive from Google went on to help them celebrate this. And he said, you have all this traffic because you're authentic, you're not propaganda. It was just the most absurd, misleading, and wrong assessment of RT.

38:40And had they been using our nutrition label and the low score for RT, they would have had a sense of the nature of what RT really is all about. And I think it's, you know, we've done a lot of research in this area. People do not trust digital platforms to make any news judgment. Consumers want to make the news judgment for themselves, but they welcome having information, apolitical information about the sources they're being presented in their newsfeed, so they can decide what to believe and what they don't want to believe. And the upside of generative AI in this content. Whether you have any thoughts on the volume of AI-generated content, whether or not it's just information on the growth of that volume and how that's, you know, what is the internet going to look like in 20 years, 100 years?

39:43yeah you know it's um machines uh creating and spreading misinformation can be done at a scale that humans could never do um i mentioned earlier we found 550 news in the first few months of looking for them and they're growing at a geometric pace so i think it's very hard to say how many that there will be someday, but it certainly looks like an enormous number. And again, there is a financial incentive for people to create these AI generated news sites, which is that if they publish crazy enough stories, they will generate an audience and they will generate programmatic advertising unless people take steps to keep ads off of these kinds of websites.

40:32And so it's incumbent on not only social media platforms, but search companies like Google and Bing to stop indexing sites that have been flagged as spreading misinformation. Would that be one of the avenues? And have you talked to them about that? Yeah, I think one of the challenges for the large language models is that, with the exception of Microsoft, we're not aware of any of them that have used fine tuning to differentiate between generally reliable and generally unreliable news sources. They tend to go by the number of tokens that a site has. There was a Washington Post article a few months ago using one of the models where they had access to the underlying training data.

41:26And they found, for example, that RT, the Russian disinformation source, was relied on more, had more tokens than Reuters. Infowars, as I recall, had more tokens than the Wall Street Journal. So it's no surprise that these generative AI models spew misinformation. It was garbage in, garbage out at great scale. So the first thing the generative AI models really need to do is to fine-tune to treat Reuters and the AP and the New York Times and the Wall Street Journal differently from Infowars and NaturalNews.com, you know, a crazy healthcare network, and Chinese disinformation and conspiracy sites.

42:11And it's not hard to do. There's data out there. We have those ratings. Others have done ratings. But that's the first thing that needs to happen. And it's not difficult, but it does mean that human beings at these companies have to decide that that is a priority for them. Yeah.

42:33And then for the search companies, have you spoken? You said that Microsoft has been using your data to improve Bing. What about Google? Google has taken steps on its own to try to improve the quality of the results. We think that they should be doing more. We're always delighted to license our data to companies, including to Google. but you know we've also red teamed bard the google search engine and the most recent time that we did that bard spread false claims in the news 80 times out of 100 80 you know fail rate so whatever they're doing is not working and again in contrast to bing which has taken steps to solve this problem, being shown that the problem can be solved with trust data.

43:32And I'm confident that the folks at BARD and Inflection and Anthropic and all the other large language models that are under pressure to do something about the hallucination problem will take steps in this important area of news and information. Yeah. And are you talking to regulators? Because ultimately it sounds like this needs to be top down from government that that there should be some penalty for spreading misinformation or is there a move on the regulatory side to do something what one of the reasons that we've done public red teaming of some of the large language models is to is in response to requests from regulators in the US and in Europe to understand better the scale of the problem.

44:25We're also a signatory to the European Commission's Code of Practice on Disinformation, which is a voluntary code, and all the big platforms have signed on to it. Twitter recently left it, but had signed on to it. And they just added AI as a new category of information. And so our contribution to the debate really is to explain the scale of the problem? How much misinformation is there? How likely are the generative AI models to create or spread misinformation? Answer, highly likely, without training. And to propose some solutions, which in this area, this fairly narrow area of news and information, the good news is that there's trust data that's available.

45:16There's steps that can be taken. The engineering's not terrifically complicated and the results get much, much better. Yeah. So in the end, you're optimistic or are you not? I'm personally very pessimistic on this topic, but. I'm optimistic that this narrow area of spreading of misinformation can be largely, not completely, largely solved. There are always going to be ways around it. There are always going to be malign actors using clever prompts to get around most anything. But there's a lot of resilience in what we have seen from large language models using trust data. So I guess I would say I'm as pessimistic as can be in the immediate term because these large language models were launched without any trust and safety protections in this area of news and information.

46:14but optimistic in the medium to long term, seeing how well the AI models operate when they're humans, take steps to provide them with access to trust data. Hi, this episode is sponsored by Salonis, the global leader in process mining. AI has landed and enterprises are adapting, giving customers slick experiences and the technology to deliver. The road feels long, but you're closer than you think. You see, your business processes run through many systems, creating data at every step. Salonis reconstructs this data to generate process intelligence, a common business language. With process intelligence, AI knows how your business flows across every department, every system, and every process.

47:11With AI solutions powered by Solonis, enterprises get faster, more accurate insights, a new level of automation, and a step change in productivity, performance, and customer satisfaction. Process intelligence is the missing piece in the AI-enabled tech stack. search Salonis, C-E-L-O-N-I-S, to find out more. That's it for this episode. I want to thank Gordon for his time. If you want to learn more about what we talked about today, you can find a transcript of the conversation on our website, eye-on.ai. Remember, the singularity may not be near but ai is changing your world so pay attention

From the publisher

This episode is sponsored by Celonis, the global leader in process mining. AI has landed and enterprises are adapting. To give customers' slick experiences and teams the technology to deliver. The road is long, but you're closer than you think. Your business processes run through systems. Creating data at every step. Celonis reconstructs this data to generate Process Intelligence. A common business language. So AI knows how your business flows. Across every department, every system and every process. With AI solutions powered by Celonis enterprises get faster, more accurate insights. A new level of automation potential. And a step change in productivity, performance and customer satisfaction Process Intelligence is the missing piece in the AI Enabled tech stack.

Go to https://celonis.com/eyeonai to find out more.

 

On episode #156 of Eye on AI, Craig Smith sits down with Gordon Crovitz, co-founder of NewsGuard and former publisher of the Wall Street Journal, to unravel the complexities of AI-generated misinformation.

In this episode, Crovitz delves into the challenges posed by AI in the spread of fake news, emphasizing the growing issue of AI-enhanced misinformation websites. He discusses the significant role of NewsGuard in countering this trend through human-analyzed databases and the importance of 'trust data' to improve the reliability of AI models. The conversation also touches upon the responsibility of social media platforms and search engines in this battle against digital deception. 

Tune in for an enlightening discussion that navigates the intersection of AI, journalism, and truth in today's digital landscape.

 

Stay updated:

Craig Smith's Twitter: https://twitter.com/craigss

Eye on A.I. Twitter: https://twitter.com/EyeOn_AI

 

(00:00) Preview

(00:19) Celonis

(01:26) Introducing Gordon Crovitz

(02:24) What is NewsGuard's Mission?

(05:58) Challenges in Detecting AI-Generated Content

(08:36) The Impact of AI-Generated Misinformation

(14:19) Future of AI Agents and Their Role in Content Generation

(17:35) Use Purpose of NewsGuard's Databases

(22:34) How Do You Tackle Misinformation Effectively?

(29:38) Government and Regulatory Responses to Misinformation

(32:36) What Role Does Social Media Play in Misinformation?

(39:22) Impact of AI on News 

(42:32) The Role of Search Companies in Combating Misinformation

(44:13) Regulatory Measures to Stop Misinformation

 

More from Eye On A.I.

All 266 episodes
#156 Gordon Crovitz: Will AI SPREAD Misinformation & Fake News?Eye On A.I. · 48 min
Listen in VO