In short
The AI Daily Brief: Episode Summary
Episode Title
Which LLMs Hallucinate Least?
Episode Date
[Insert Date Here]
Episode Overview In this episode, NLW discusses recent developments in AI, focusing on:
- New research regarding hallucinations in language models (LLMs).
- Google’s lawsuit against AI scammers.
- The potential agreement between the US and China about the use of AI in nuclear control systems.
Episode Highlights
- LLM Hallucination Rates
- Definition of Hallucination: Hallucinations in AI refer to instances where models generate factually incorrect information.
- Study Findings:
- New research published in *Nature* reveals significant rates of hallucination among various LLMs when citing works.
- Key Statistics:
- GPT-3.5: 55% hallucination rate on cited works.
- GPT-4: 18% hallucination rate on cited works.
- Overall accuracy:
- GPT-4: 3% general hallucination rate.
- GPT-3.5: 3.5% general hallucination rate.
- LAMA-2: Between 5.1% and 5.9%.
- Cohere Models: 7.5% to 8.5%.
- Anthropic Cloud 2: 8.5%.
- Mistral 7B: 9.4%.
- Google Palm: 12.1%.
- Google Palm Chat: 27.2%.
- Google's Legal Action Against Scammers
- Overview: Google has initiated legal proceedings against scammers misleading users into downloading fake versions of its Bard chatbot, which do not exist.
- Details:
- The scammers are reportedly based in India and Vietnam and deploy malware through deceptive Facebook ads.
- Google has issued over 300 takedown requests for these ads, emphasizing the importance of protecting consumers from such fraud.
- YouTube’s Policy on AI-Generated Content
- New Guidelines: YouTube is introducing policies to manage AI-generated content, particularly related to music.
- Key Points:
- Creators must label AI-generated content, especially if it impacts socially significant topics.
- There will be a moderation process for requests to take down AI-generated videos simulating real people.
- YouTube is investing in tools to detect AI-generated content but acknowledges current limitations.
- AI in Medicine
- Cardiac Health Study: A study from Oxford indicates that AI can predict heart attack risks from CT scans up to ten years before symptoms arise, showcasing the potential for improved preventative healthcare.
- Drug Development: In Silico Medicine is approaching late-stage trials for a drug developed using AI for idiopathic pulmonary fibrosis, marking a significant milestone in AI-driven drug discovery.
- Market Response to AI
- NVIDIA Stock Surge: There is renewed investor interest in AI technology, with NVIDIA seeing a significant rise in stock value.
- Geopolitics of AI: US-China Relations
- AI and Nuclear Weapons Control: The US and China are reportedly set to agree on not using AI in nuclear weapons control systems.
- Political Declaration on Military AI: The US has released a declaration aimed at promoting responsible military use of AI, signed by 45 states, which emphasizes ethical standards and compliance with international law.
Key Takeaways
- Understanding LLMs’ hallucination rates is crucial for their integration into professional workflows, especially in sensitive fields like medicine.
- Legal actions against AI scams highlight the necessity for consumer protection in rapidly evolving tech landscapes.
- YouTube's approach to AI content moderation may set precedents in managing digital rights and copyright issues.
- The potential of AI in healthcare signifies its transformative power in improving patient outcomes.
- The collaborative efforts between the US and China to manage AI in military contexts could pave the way for future diplomatic relations.
Closing Thoughts NLW emphasizes the importance of collaboration and regulatory frameworks in managing AI’s risks while harnessing its potential benefits. The episode underscores the necessity for ongoing monitoring and adaptation as AI technologies evolve.
---
Sponsor Message This episode is sponsored by Notion AI, a tool designed to enhance productivity by providing quick answers based on your existing documents and notes.
For more information, visit
[Notion AI](https://notion.com/aibreakdown)
Subscribe and Stay Updated
- Newsletter: [AI Breakdown Newsletter](https://theaibreakdown.beehiiv.com/subscribe)
- YouTube Channel: [AI Breakdown YouTube](https://www.youtube.com/@TheAIBreakdown)
- Community: [Join the Discord](bit.ly/aibreakdown)
---
Thank you for tuning in to this episode of The AI Daily Brief. Until next time, peace!
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:01Today on the AI Breakdown, the US and China are set to agree not to use AI in the control of nuclear weapons systems. Before that in the brief, which LLMs hallucinate least? The AI Breakdown is a daily podcast and video about the most important news and discussions in AI. Go to Breakdown.network for more information about our YouTube channel, our Discord, and our newsletter.
0:24Welcome back to the AI Breakdown Brief, all the AI headline news you need in around five minutes. One of the things that we have been talking about quite a bit here on the AI Breakdown is the fact that we are moving from a period that was largely characterized by experimentation and sort of first blush efforts to experiment with AI to a world in which organizations and professionals are increasingly integrating generative AI tools into their actual professional workflows. Now in that, one of the big barriers is of course the fact that AI models still hallucinate. For some professions and roles, this isn't such a big deal and it simply means that you have to double check facts that come through these LLMs.
1:03But for other areas, particularly think medical use cases as a for example, obviously hallucinations can have significant impacts. Well, somehow up until now, we haven't really had good information or research around which models hallucinate more or less. But we've just gotten a new report published in Nature called Fabrication and Errors in the Bibliographic Citations Generated by ChatGPT. The abstract reads, Although chatbots such as ChatGPT can facilitate cost-effective text generation and editing, factually incorrect responses or hallucinations limit their utility. Now, this particular study was focused on how often these different models saw hallucinations around the specific citation of works.
1:42And this is an area where there was a far greater rate of hallucination than there was in general. For example, across all works, GPT 3.5 hallucinated 55 % of cited works, and even GPT 4 hallucinated 18 % of cited works. And this is why, as Professor Ethan Mollick points out, it's important to understand how the hallucination rate looks not just in general, but for specific applications. Now going back to the overall accuracy and the general hallucination rate, this research from Victara has GPT-4 at the top of the heap, hallucinating just 3 % of the time. GPT-3.5 hallucinates 3.5 % of the time.
2:17The LAMA-2 models, based on their size, hallucinate between 5.1 % and 5.9 % of the time. Coheres models hallucinate 7.5 % and 8.5 % of the time, which is the same as Anthropics Cloud 2 at 8.5%. Hot new kid on the block, Mistral 7B, has a 9.4 % hallucination rate. And then Google Palm is all the way down at the bottom with 12.1 % and Google Palm Chat at 27.2%. Perhaps one more reason to be very eager for their forthcoming Gemini model to finally get out into the wild. Now, speaking of Google, the company is going on the offensive against scammers who are trying to take advantage of hype and excitement around artificial intelligence for their nefarious schemes.
2:56Basically what's happening is that there are a group of individuals and potentially companies in India and Vietnam, the subjects of the lawsuit aren't named in this case, who have been trying to trick small business owners into clicking Facebook ads that say that they will download a version of Google's Bard chatbot for mobile. The problem is that Bard is a web-based platform and isn't available for download, and so what these scammers are actually doing is installing malware that steals social media credentials. Google's general counsel said that the lawsuit is the first such lawsuit aimed at protecting users of a major tech company's flagship AI product.
3:28Now, while the overall size of the scam isn't clear, Google says that they've filed over 300 takedown requests to have the ads removed. According to Google, Facebook, and others have been generally responsive to the takedown requests, but with so many advertisers on the self-serve platform, there can be a bit of latency there, and of course, that still creates risk for these users. What's notable to me is the fact that Google is availing themselves of the legal system for this. In general, the response to these sort of scams, especially when they're international, tends to kind of just be focused on pressuring the platform, in this case Meta, to have better protections, and more broadly, just helping consumers get better educated about what is and isn't a scam.
4:03Now, speaking of things in the alphabet universe cracking down, YouTube is starting to implement policies around content that features AI clones of musicians. As The Verge describes, YouTube really has two sets of content guidelines when it comes to AI deepfakes. There is a looser general set of rules for the average user, and then a very strict set of rules when it comes to protecting the platform's music partners. Now, a blog post has just come out from the company, which shares how they are starting to think about moderating AI-generated content. Part of it is pretty common sense. YouTube is going to require that creators begin labeling what they determine as realistic AI-generated content when they upload videos, and those disclosure requirements are going to be more significant the more socially impactful the topic of the video is, such as elections or ongoing conflicts.
4:44Now, YouTube hasn't described what it thinks realistic means yet, but say there's going to be more detailed guidance when those requirements start rolling out next year. Now, one thing The Verge points out is that while YouTube has the ability to penalize content creators that don't label their AI-generated content, figuring out if an unlabeled video was actually generated by AI could be something of a problem. YouTube says that the platform is, quote, investing in the tools to help us detect and accurately determine if creators have fulfilled their disclosure requirements when it comes to synthetic or altered content, but tools for detecting AI-generated content simply don't work right now.
5:16Now, on top of that, there is going to be a moderation process whereby people can request that videos that simulate them get taken down. However, it's not a guarantee. YouTube says that it will evaluate, quote, a variety of factors when evaluating these requests, including whether the content is parody or satire, and whether the individual is a public official or well-known official. Once again, the vagueness of the definition of parody and satire could create a lot of problems when it comes to actually implementing these policies. However, when it comes to AI-generated music content from YouTube partners, there is no exception for parody and satire, and anything that quote mimics an artist's unique singing or rapping voice is subject to takedown.
5:52Now, one thing that is worth noting is that there won't be any automated detection of that, but instead there will be a manual request form that partner labels will have to fill out when they see violations of the policy. It also seems that YouTube is going to take a light hand when it comes to punishing the creators, especially in the early days of these policies rolling out. Now, part of why YouTube might be so concerned with their music industry partners and the copyright protections they're in is that their deals with those companies are very important for the way that they've set up their site and particularly the way they make music and sound available for YouTube Shorts.
6:22We've also heard that Google more broadly is in conversations with the music labels to create some sort of apparatus through which people can legitimately and in an above-board way make new music using synthetic versions of existing artists in a way that is approved and cuts artists in. Anyways, it will be an interesting case study to watch of how copyright plays out, not just in the courts, but in the business realm. Now, moving over to the world of medicine, a couple interesting stories there. One is a new study from Oxford that suggests that AI analysis of cardiac CT scans could accurately predict the risk of heart attacks even up to 10 years in the future, even before someone officially has heart disease.
6:58One of the doctors involved in the study said, Our study found that some patients presenting in hospital with chest pain, who are often reassured and sent back home, are at high risk of having a heart attack in the next decade, even in the absence of any signs of disease in their heart arteries. Here we demonstrate that providing an accurate picture of risk to clinicians can alter and potentially improve the course of treatment for many heart patients. Obviously, one of the big promises of AI is better preventative care when it comes to medicine that allows doctors to get out ahead of issues that their patients are likely to face in the future.
7:27Now, speaking of AI in medicine, Bloomberg also writes that in the race for the first drug to be discovered by an AI, a key milestone is soon to be reached. Bloomberg writes, The global push to use AI to find new medicines faces a crucial test as one front-runner starts approaching late-stage trials for a drug discovered by algorithms. In Silico Medicine, which has headquarters in Hong Kong and New York, used AI to develop an experimental drug for the incurable lung disease idiopathic pulmonary fibrosis. The treatment is in mid-stage trials in the U.S. and China, with some results expected early 2025.
7:59Now, the world is watching this one even more closely than other drug trials, because it's the first fully AI-based preclinical candidate. As Bloomberg writes, A string of other leading molecules that relied on AI have faced setbacks, and in silicoes could still fail in the process or take years to reach the market. At the same time, the implications of any success would be huge, opening the door for new and cheaper AI therapies that can save lives and cut costs for health systems. Now, even as the medical world watches that closely, moving over into markets, there are indications that Wall Street is falling in love with AI once again.
8:29As the Wall Street Journal writes, this year's hottest stock is regaining its momentum. The report is about NVIDIA, which has traded up for nine straight sessions and is up about 20 % over that period. Still, the big question will be what happens next week when the company shows off its third quarter results. Obviously, we will keep you informed about all those developments, but that is going to do it for today's AI Breakdown Brief. Next up, the main AI breakdown. And now, a quick word from today's sponsor. I am a huge Notion user. We're talking multiple accounts for multiple projects. I use it for everything from applicant tracking to note taking to project management to sharing public documents to frantically capturing ideas I have while out hiking or just driving around.
9:11Given that, and given the topic of the AI breakdown, I was excited to learn that they've launched a new AI tool called Q &A. It's like a personal assistant that responds in seconds with exactly what you need. Notion AI can give you instant answers to your questions using information from across your wiki, projects, docs, and meeting notes. For someone like me who makes dozens of notes per day around a huge array of topics, having a built-in AI tool to help recall that is incredibly useful. Now beyond that use case, think about this. Have an urgent question you'd normally turn to a co-worker to answer?
9:40Just ask Q &A instead. It'll search through thousands of documents in seconds and answer your question in clear language no matter how large or complex your workspace is. Plus, you can trust your data is secure because Notion AI is designed to protect your information. No AI models are trained with your information, the data is encrypted, and answers will never use information from pages you don't have access to. With Notion AI, it's even easier to do your most meaningful work. Try Notion AI for free when you go to notion.com slash AI breakdown.
10:22Welcome back to the AI Breakdown. When it comes to the geopolitics of artificial intelligence, there is quite obviously no more significant relationship than that between the US and China. Now, we have had numerous contexts where this has been on display over the course of the last few months. One is, of course, everything around the UK Safety Summit. Rishi Sunak's government made the controversial decision to have China participate in that AI Safety Summit, in spite of the fact that they were dealing with an active Chinese spying scandal, on the logic that if the world is really concerned with mitigating the biggest risks of runaway artificial intelligence, it needs the participation of everyone, not just some people.
11:02Now, at that event, there was a declaration around AI's potential for catastrophic danger that was signed by both the U.S. and China, among other signatories. But it wasn't really about anything more than acknowledging the risk. The so-called Bletchley Declaration was intended to be the first time that the governments of the world got together to collectively agree that, as the declaration puts it, there is potential for serious, even catastrophic harm, either deliberate or unintentional, stemming from the most significant capabilities of these AI models. Said UK Technology Secretary Michelle Donilon, for the first time we now have countries agreeing that we need to look not just independently but collectively at the risks around frontier AI.
11:40Now of course also happening recently is that the US has been tightening its export controls when it comes to AI chips. The Biden administration first put some rules into practice last year around this time, and this latest set of rules coming through the Commerce Department were effectively meant to close loopholes that had been identified over the course of the last year. This involved things like tighter restrictions even on lower-powered chips, as well as an inclusion of foreign subsidiaries that were owned by Chinese companies as part of the firms who were prohibited from getting access to these technologies.
12:09Meanwhile, as all of this has been going on, for anyone paying close attention, there has been a steady drumbeat of announcements and news and reports around both China and the US developing further AI capabilities when it comes to military power. Take for example this piece from Fox News on October 17th. China-US race to unleash killer AI robot soldiers as military power hangs in balance. AI technology is the new arms race pitting the world's power against each other, experts agree. Now, the details aren't super salient to the discussion that we're having today, but suffice it to say that while everyone is talking metaphorically about an AI arms race in the context of frontier models, there is an actual AI arms race happening between the world's biggest military powers.
12:52Now, the US quite clearly is thinking about the military implications of artificial intelligence not just in terms of a blank slate that they can do whatever they want with, but as something that needs to be managed on a global stage. This week, they released the Political Declaration of Responsible Military Use of Artificial Intelligence and Autonomy. This declaration was signed by 45 endorsing states and contains 10 what they call concrete measures to guide the responsible development and use of military applications of AI and autonomy. So what did the actual declaration say? Well, this is from the latest version that has been published on the State Department website, which comes from November 1st.
13:26It reads, An increasing number of states are developing military AI capabilities, which may include using AI to enable autonomous functions and systems. Military use of AI can and should be ethical, responsible, and enhance international security. Military use of AI must be in compliance with applicable international law, and in particular, use of AI in armed conflict must be in accord with states' obligations under international humanitarian law. So then, And what the endorsing states agreed to were things like the idea that military organizations should take appropriate steps to review their AI capabilities as relates to international and humanitarian law, that states should have systems for effective oversight of the development and deployment of military AI capabilities, that they should take proactive steps to minimize unintended bias, that they should ensure that the development of these technologies is done in a transparent and auditable way, that the personnel who approve and use this technology should be appropriately trained, that capabilities should have explicit and well-defined uses and that states should implement appropriate safeguards to mitigate risks of failures.
14:23Now, you see, these are very kind of common sense declarations, and they don't really limit what states can or can't develop. There's nothing here that says, for example, you're not allowed to use AI in such and such a way as it comes to military applications outside of already established norms of international humanitarian law. Ambassador Bonnie Denise Jenkins made statements around the launch event for the declaration at the UN in New York, saying, We cannot predict how AI technologies will evolve or what they might be capable of in a year or five years. However, we know that there are steps states can take now to put in place the necessary policies and to build the technical capacities to enable responsible development and use, no matter the technological advancements.
15:01We need, therefore, to come together as an international community around a set of strong norms for responsible development and deployment. Norms that will enable nations to harness the potential benefits of AI systems in the military domain, while encouraging steps that avoid irresponsible, destabilizing, and reckless behavior. Now, Jenkins' speech also noted that this was in many ways a foundation for deeper conversations. In that same speech, she said, It provides a basis for a much more concrete dialogue on what responsible means in practice. What does an effective testing and assurance process look like?
15:29How do you exercise appropriate care in a range of practical applications? She also said, We envision this collaboration among endorsers to be far more robust than simply committing to high-level principles. The declaration is a foundation for collaboration and exchanges, such as sharing best practices, expert-level exchanges, and capacity-building activities. Finally, she said that the broad terms used in the declaration was specific, and that the U.S. isn't trying to unilaterally decide for countries how to apply these principles, but to simply provide a starting point in a shared space of international agreement.
15:59Now, like I said, this declaration was signed by 45 countries, but the most notable absence was, of course, China. Said Sam Bresnik of Georgetown, it wasn't really surprising that Beijing declined to endorse this declaration. Bresnik said, Although Beijing likely supports many of the declaration's proposals, it is not enthusiastic about signing on to a US-led effort on responsible military AI. Instead, he said, quote, China seems more interested now in engaging in multilateral discussions surrounding the responsible development and use of AI, while unlikely to agree to binding agreements that might limit its ability to develop and field AI-enabled military systems.
16:33And indeed, that seems to be echoed in the fact that reports are that President Biden and President Xi will sign a deal this week focused on specific issues around the use of AI in military applications, most notably questions of keeping AI out of the control systems for nuclear weapons. At Wednesday in San Francisco at the Asia-Pacific Economic Cooperation Summit, US President Joe Biden and US President Xi Jinping are set to meet. Two sources familiar with the planned discussion say among the top items on the agenda is the proliferation of AI in military technologies. According to reports, the leaders will pledge a deal that limits the use of AI in autonomous weapons such as drones, as well as in the systems that control and deploy nuclear warheads.
17:10Indeed, one of the things that's interesting about this is that it appears that while AI is such a wedge issue in so many other contexts, here these presidents are using common ground around AI as part of an attempt to reduce tensions. Now, whether China agrees or not, it appears that the position of this U.S. administration is that AI should not be involved in the deployment of nuclear weapons. Secretary of State Antony Blinken was asked last week about this and said, I can't get into the specific issues that they would discuss in any such meeting about President Xi and President Biden, but, quote, I can say as a general principle for us that when it comes to artificial intelligence, we believe AI should not be in the loop or making the decisions about how and when a nuclear weapon is used.
17:48Now, I kind of saw two different categories of reactions to this. I didn't see anyone who is negative on it, but on the one hand, you have people like Max Tegmark who wrote, I'm delighted to hear that the US and China plan to agree on not empowering AI to launch nukes. On the other hand, you had folks like Matthew Pines, who said, I know the bar is low, but a U.S.-China agreement to avoid automating nuclear command and control systems is like the bare minimum of what functioning human beings interested in collective survival should agree to. I don't think that there's actually necessarily any sort of mutual exclusiveness between these two positions.
18:20Like Tegmark, I am very excited that this seems to be an agreement that the U.S. and China can make. And like Matthew, I feel like this is a low bar that we can all rally around. What I will say is that especially in as tense an environment as we have between the U.S. and China right now, getting to small agreements, even on issues that seem like they should be incredibly obvious, is a really important part of the diplomatic process. Getting to alignment between two nations, especially nations that compete as intensely as the U.S. and China do, where there is as much mutual suspicion as there is between these two parties, requires laying slow foundations of small alignments on top of one another, which have the potential to become strong enough to handle future breaks and cracks in that relationship that come along.
19:00History is littered with examples of countries that were extremely antagonistic or even outright at war, who could still come to agreement on certain involable principles, much to the good of the survival of the world. So perhaps, yes, this is not something to fist pump about in excitement, but it's not nothing either. Thanks for listening or watching as always. And until next time, peace.
19:28You
From the publisher
On today's episode, NLW looks at new research about LLM hallucination; Google suing AI scammers; and China and the US working towards an agreement not to use AI in nuclear device control systems.
Today's Sponsor:
Notion - Notion AI. Knowledge, answers, ideas. One click away. - https://notion.com/aibreakdown
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI.
Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe
Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown
Join the community: bit.ly/aibreakdown
Learn more: http://breakdown.network/
