In short
The AI Daily Brief: Episode Summary
Episode Title
The gpt2 Chatbot is Back as OpenAI Talks Data
Podcast Overview
- Title: The AI Daily Brief (formerly The AI Breakdown)
- Description: A daily podcast analyzing important developments in AI, exploring creativity, industry disruptions, philosophical questions, and ethical considerations surrounding artificial intelligence.
Key Topics Covered
- U.S. Government Measures Against China
- Expanded efforts to deny China access to advanced AI models.
- Preliminary plans by the White House for stricter export controls.
- Focus on AI models like ChatGPT, with potential regulations targeting closed-source models based on computing power thresholds.
- Geopolitical Context
- U.S. export restrictions on advanced AI chips to China and its implications for national security.
- Changing dynamics in the Middle East, with countries like the UAE and Saudi Arabia reassessing their partnerships with China.
- OpenAI Developments
- Search Product Announcement: Growing anticipation for a search feature within ChatGPT.
- GPT-2 Chatbot Resurgence: An unexpected return of a GPT-2 chatbot that generated significant interest due to its performance.
- AI Detection Tools: OpenAI's initiatives to develop tools for identifying AI-generated content, including partnerships with C2PA (Coalition for Content Providence and Authenticity).
- Data Management and Ethics
- Introduction of a "Media Manager" tool to allow creators to manage their content's usage in AI systems.
- OpenAI's shift towards clearer communication regarding their data collection principles and respect for content creators.
- Model Specification
- Launch of OpenAI's "model spec" to define expected behaviors of AI models.
- Emphasis on user feedback and clarity in distinguishing between inherent model behaviors and potential bugs.
Notable Quotes
- "The mere fact that OpenAI feels it needs to release a statement about training data principles... shows how powerful the movement against it has become."
- "Today's AI systems will seem laughably bad in 12 months." — OpenAI CEO Brad Lightcap
Conclusion The episode covered an extensive range of developments regarding AI governance, impacts of geopolitical tensions on technological advancements, and OpenAI's ongoing evolution. Listeners were left with a sense of anticipation for upcoming products and the implications these advancements will have on the field of AI and broader societal contexts.
---
Key Takeaways
- U.S.-China AI Relations: Ongoing restrictions and geopolitical strategies highlight the intersection of technology and national security.
- OpenAI's Innovations: The potential launch of a search feature and new detection tools signify OpenAI's commitment to evolving and addressing ethical concerns in AI usage.
- Data Rights and Management: OpenAI's initiatives to respect creators’ rights underscore a significant shift towards ethical AI practices.
---
Additional Resources
- [Superintelligent - AI Education Platform](https://besuper.ai/)
- [Managing the Future of Work Podcast](https://www.hbs.edu/managing-the-future-of-work/podcast/Pages/default.aspx)
- [Subscribe to The AI Daily Brief Newsletter](https://theaibreakdown.beehiiv.com/subscribe)
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Today on the AI Daily Brief, more confirmation that search is coming or even already here for open AI, but maybe a delay in the announcement. Before that, in our headlines, the White House might be moving into a next phase of trying to deny China access to advanced AI models. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. Check out the Discord link in our show notes to join the conversation.
0:26Hello, AI friends. One quick note before we get into today's episode. You may have seen that we just announced that the Breakdown, the crypto podcast, as well as the Bitcoin Builders show that I started last year, and a bunch of related media properties are being acquired and moving over to Blockworks. Blockworks is an excellent crypto media company. I'm really excited for the breakdown and the other crypto content to be over there. But the AI show is staying put right here. As part of the transition, we are shifting the name to the AI Daily Brief instead of the AI Breakdown, just to remove any brand confusion.
0:57The format will remain the same. The cadence will remain the same. Basically, if you like the AI breakdown, the AI Daily Brief is pretty much exactly the same thing. Anyways, just wanted to share why you had been hearing a new name, but there is a lot to discuss today, so let's get into it. Welcome back to the AI headlines on the AI Daily Brief. It's all the AI headline news you need in around five minutes. We kick off today with yet further efforts from the US government to limit China's access to advanced AI. For the last couple years, AI has been a centerpiece of the geopolitical tension between China and the USA.
1:32Since the end of 2022, the Biden administration has increasingly tightened export restrictions on advanced AI chips, while at the same time trying to create incentives to bring chip manufacturing back to the US. Now it appears that there is going to be a new front in that effort, with Reuters reporting that the White House has, quote, preliminary plans to place guardrails around the most advanced AI models like ChatGPT. Continuing, Reuters writes, the Commerce Department is considering a new regulatory push to restrict the export of proprietary or closed-source AI models. Any action they write would complement a series of measures put in place over the last two years to block the export of sophisticated AI chips to China in an effort to slow Beijing's development of the cutting-edge technology for military purposes.
2:10According to Reuters sources, any new export controls would likely target Russia, China, North Korea, and Iran. Back in February, Microsoft said in a report that it had tracked hacking groups affiliated with the governments of China and North Korea, as well as Russian military intelligence, in trying to use their software to improve their hacking. In terms of determining which models would be advanced enough to be under these controls, it seems like it might be based on the computing power it took to train a model. Reuters again writes, the sources said the US may turn to a threshold contained in an AI executive order issued last October that is based on the amount of computing power it takes to train a model.
2:43When that level is reached, a developer must report its AI model development plans and provide test results to the Commerce Department. That computing power threshold could become the basis for determining what AI models would be subject to export restrictions. If used, it would likely only restrict the export of models that have yet to be released since none are thought to have reached that threshold yet. According to the reporting, the agency is quote far from finalizing a rule proposal, but just the fact that they're considering it, Reuters argues, shows how seriously the US government is taking AI as a geopolitical concern.
3:10Now of course, one of the challenges if anything like this were to be imposed is how much open source just complicates things. Especially given how close to state-of-the-art open source is getting, it could be a real challenge. To reiterate, this is currently just reporting from unnamed sources. When Reuters reached out, the Commerce Department declined to comment, while the Russian embassy in Washington did not respond. The Chinese embassy, though, you know they were going to get a comment in, describing it as a, quote, typical act of economic coercion and unilateral bullying which China firmly opposes.
3:38Meanwhile, in terms of things that have actually happened on this area, the U.S. government has revoked licenses that allowed Qualcomm and Intel to supply chips to Huawei. The information writes, U.S. lawmakers and officials have been alarmed by Huawei's ability to make advanced chips for smartphones despite years of Western sanctions. Republican lawmakers concerned about national security risks had been urging the U.S. government to cancel licenses that allow U.S. companies, including Qualcomm and Intel, to continue to sell chips to the Chinese company. In many ways, this is just a continuation of policies that were already started, but shows that in general, this is not just idle talk.
4:10Another part of the world that is impacted by the China-U.S. battle around AI is, of course, the Middle East. For some time, the Middle East has been positioning itself as a literally middle player, maintaining relationships with both China and the US. However, the US has been increasingly uncomfortable about the Middle East relationship with China, seeing it as a way for China to getting around policies like the chip export controls. Recently, UAE company G42 explicitly started to move away from its relationships with China and took on a$1.5 billion minority investment from Microsoft, which was in part set up by the US Commerce Department.
4:42Now, Time Magazine is reporting that the head of Saudi Arabia's new investment fund for AI, said that if the US asked them to divest from China, they would. Amit Midha said, So far, the requests have been to keep manufacturing and supply chains completely separate, but if the partnerships with China would become a problem for the US, we will divest. In another interview, Midha said, We are seeking trusted, secure partnerships in the US. The US is the number one partner for us and the number one market for AI, chips, and semiconductor industry. What's more, the US government is not just prohibiting China from using AI, or trying to, but also figuring out how to use AI itself.
5:13For example, the Department of Homeland Security is now piloting using AI to train officers who review applicants for refugee status. Basically, the central idea of this pilot is that it can be really difficult for immigration officers to do interviews with refugees who have experienced significant trauma. Said Secretary Alejandro Mayorkas, refugee applicants, given the trauma that they have endured, are reticent to be forthcoming in describing that trauma. So then in the pilot, DHS is training the AI to act like refugees so that their officers can practice interviewing them. Meanwhile, while over the last few weeks, Microsoft has made a number of big investment announcements in Southeast Asia, they're now bringing it back home with an over$3 billion investment to build AI in Wisconsin.
5:51Today, President Biden will speak with Microsoft President Brad Smith in Mount Pleasant, Wisconsin, to announce a$3.3 billion investment in a new data center there. In addition to the data center, Microsoft said that they're also investing in a new AI lab at the University of Wisconsin-Milwaukee to train employees to use AI. Said Brad Smith in an interview, we have a huge responsibility to help ensure this technology serves people. Part of that is ensuring that it works safely and remains under human control. But another part of it is really supporting and aiding the transition of the economy.
6:19So lots going on in the world of AI and geopolitics. But for now, that is going to do it for the headline section of the AI Daily Brief. Next up, the main episode where we talk all about what's going on with OpenAI, or really, what isn't going on with OpenAI. As a listener of this show, I have a strong feeling you like to stay up to date on all things artificial intelligence, including its impact on the workforce, which is why I highly recommend checking out Managing the Future of Work, the chart-topping business podcast from Harvard Business School. HBS professors Bill Kerr and Joe Fuller talk to business leaders, technologists, and policymakers grappling with the forces like AI, globalization, and demographic shifts that are reshaping the nature of work.
6:58Recent guests include IBM CHRO Nicol Lamoureux on how Big Blue is adopting AI, Morningstar CEO Kunal Kapoor on how AI can raise the investment IQ, Microsoft Corporate Vice President Jared Spatero on how the tech giant is experimenting its way from AI assistants to autonomous agents, and many other prominent movers in business and the workforce ecosystem. So don't miss out. Follow Managing the Future of Work on Apple Podcasts, Spotify, or wherever you're listening now. Hello, AI friends. Today, I want to tell you about our platform, Super Intelligent. In short, it's a platform for useful, practical, immediately applicable AI learning.
7:34We have nearly 400 video tutorials, each of which comes with step-by-step how-tos. And the idea is to get you actually using these AI tools we talk about every day in a matter of minutes to actually solve problems, create new opportunities, and just do really cool things. To learn more and subscribe, go to besuper.ai. And if you do decide to subscribe, use code podcast for 50 % off your first month. Again, that's besuper.ai. Welcome back to the AI Daily Brief. Boy, is there a lot of OpenAI news today. None of it's huge or transformational. It's just really diverse. As you'll see, there is enough that it will easily fill an episode.
8:12First up, we got some more confirmation that OpenAI does appear to be getting ready to launch a search product that will rival Google and other AI-native companies like Perplexity. According to Bloomberg, quote, The feature would allow users to ask ChatGPT a question and receive answers that use details from the web with citations to sources such as Wikipedia entries and blog posts. One version of the product also uses images alongside written responses to questions when they're relevant. If a user asked ChatGPT to change a doorknob, for instance, the results might include a diagram to illustrate the task.
8:42Bloomberg also gave a shout out to the Twitter sleuths who found search.chatgpt.com last week, but we're still a little bit in the realm of unknown sources. The other OpenAI-related thing from last week was, of course, that mysterious GPT-2 chatbot that appeared on an LLM ranking site, and which seemed to many to be more performant and advanced than anything else out there. After it got all of that attention last week, it was taken off of LIMSYS, but then a couple days ago, it came back. Once again, people are really impressed. This time, it was called I'm a good GPT-2 chatbot, and much of the response I saw was people like Pietro Schirrano, who wrote, I'm a good GPT-2 chatbot is so good that it created a code interpreter that uses Claude Opus for me.
9:20Excuse me as I faint in ontological shock. Lior at AlphaSignalAI tweeted a network error due to high traffic that seemed to point to OpenAI as confirmation that it was behind the chatbot. There was also Sam Altman, who on May 5th tweeted, I'm a good GPT-2 chatbot. Siki Chen from Runway suggested that this chatbot also revealed that ChatGPT had already stealth-launched the search feature. Siki writes, Did ChatGPT already stealth launch search? ChatGPT used to reject prompts asking for current events. But now I get this without even a little browsing with Bing message. He shared an image where he had asked, what's the latest news on this I'm a good GPT2 chatbot model?
9:57ChatGPT responds, the I'm a good GPT2 chatbot model that has recently surfaced in the AI community is wrapped in a bit of mystery. This model appearing on the LIMSYS chatbot arena generated considerable interest due to its sudden appearance and impressive performance, sparking speculation about its origins and capabilities. There were initial reports that it performed better than GPT-4, but concrete details about its design or purpose were not disclosed and it was subsequently taken offline. What was interesting about this is that it also cited its sources, pointed to Daily AI and some other sources as where it had drawn this information.
10:27And again, this was not technically supposed to be the Browse with Bing version. All this is to say, there is even more evidence now that search with ChatGPT is coming, but it seems also like we might have to wait just a little bit longer to find out more about it. The information reported earlier this week that OpenAI was considering postponing an event that had been planned for this Thursday, where they had been intending to show off a set of new products, including presumably this search product. The information says the spokesperson from OpenAI declined to elaborate on the reasons for the change.
10:55If this alone were the OpenAI news slate, it would be a lot, but we are not even close to done. Yesterday, OpenAI announced that they were working on new tools to detect AI-created images. The blog post was called Understanding the Source of What We See and Hear Online. There were a couple things that were announced as part of this post. One was that OpenAI was joining the steering committee of something called C2PA, the Coalition for Content Providence and Authenticity. Earlier this year, they wrote, We began adding C2PA metadata to all images created and edited by DALI3, our latest image model, in ChatGPT and the OpenAI API.
11:26We will be integrating C2PA metadata for Sora, our video generation model, when the model is launched broadly as well. They also announced that they were working on new technology in this area, including what they describe as temper-resistant watermarking, i.e. marking digital content like audio with an invisible signal that aims to be hard to remove, as well as detection classifiers or tools that use artificial intelligence to assess the likelihood that content originated from generative models. As part of the announcement then, they shared that they were opening applications for access to OpenAI's image detection classifier to a first group of testers.
11:57They write that in their internal tests, the classifier correctly identified around 98 % of DALI-3 images and less than 0.5 % of non-AI-generated images were incorrectly tagged as being from Dolly 3. They write the classifier handles common modifications like compression, cropping, and saturation changes with minimal impact on its performance, but other types of augmentations, such as adjusting the hue or adding moderate amounts of Gaussian noise, can make a significant difference. Lastly, on this front, Microsoft and OpenAI also announced that they were launching a$2 million societal resilience fund, which is basically all about combating deepfakes.
12:29Then there was another blog post from yesterday, this time called Our Approach to Data and AI. The big TLDR of this comes in the section called We Respect the Choices of Creators and Content Owners on AI. They write, Decades ago, the robots.txt standard was introduced and voluntarily adopted by the internet ecosystem for web publishers to indicate what portions of websites web crawlers could access. Last summer, OpenAI pioneered the use of web crawler permissions for AI, enabling web publishers to express their preferences about the use of their content in AI. We take these signals into account each time we train a new model.
13:00That said, we understand that these are incomplete solutions, as many creators do not control websites where their content may appear, and content is often quoted, reviewed, remixed, reposted, and used as inspiration across multiple domains. They conclude, we need an efficient, scalable solution for content owners to express their preferences about the use of their content in AI systems. OpenAI's answer is something they're calling Media Manager. It's a tool that will, quote, enable creators and content owners to tell us what they own and specify how they want their works to be included or excluded from machine learning, research, and training.
13:28They didn't explain exactly how, but that's the idea. There was a lot of commentary here. Some people were pretty skeptical. But Brian Merchant noted that, quote, the mere fact that OpenAI feels it needs to release a statement about training data principles and to announce a program to let creators opt out of having their works included in training data shows how powerful the movement against it has become. And then finally, we get to today, where OpenAI announced something that they called their model spec. Sam Altman writes, we are introducing the model spec, which specifies how our model should behave.
13:56We will listen, debate, and adapt this over time. But I think it will be very useful to be clear when something is a bug versus a decision. We want to give users lots of control of AI with some hard boundaries that society eventually agrees on. This is another step. While this is definitely very different than Anthropic's constitutional approach to AI, in the sense that this is a public-facing document, not a document that specifically is meant to inform AI behavior, it still has some of the same ideas of articulating what they're trying to have their AI models be. So the approach includes one, objectives, broad general principles that provide a directional sense of the desired behavior, such as assisting the developer and end user, benefiting humanity, and reflecting well on open AI.
14:34There are tools, instructions that address complexity and help ensure safety and legality. These include things like follow the chain of command, comply with applicable laws, don't provide information hazards, respect creators and their rights, protect people's privacy, and don't respond with not-safe-for-work content. Lastly, there are default behaviors, or as they describe, guidelines that are consistent with objectives and rules, providing a template for handling conflicts and demonstrating how to prioritize and balance objectives. These include things like assuming best intentions from the user or developer, asking clarifying questions when necessary, being as helpful as possible without overstepping, assuming an objective point of view, encouraging fairness and kindness, and discouraging hate, and more.
15:10Joanne Zhang, who worked on this product, writes, I'm personally excited about this concept of a model spec for three reasons. One, there'll be more clarity on whether something is a policy or an RLHF bug. Two, principles are easier to debate and get feedback on versus hyperspecific screenshots or abstract feel-good statements. The way that she describes this is, it's easy for most people to agree on models should be something, but the more important questions lie deeper in thorny scenarios. For example, how should the model engage with someone who claims the earth is flat? Third, she writes, model spec feedback will help us steer our efforts and steerability.
15:40Unexplicit non-goal, she writes, for the model spec is to reach consensus on a one-size-fits-all model. That will never happen. We want to give users and developers as much control as possible while staying within hard boundaries that people understand. Hearing feedback on where and how everyone wants to steer the model is helpful in A, designing a more rigorous survey process, and B, informing the research and product roadmap. So just tons and tons of stuff happening. But when push comes to shove, what everyone's really waiting for is the next big products. OpenAI CEO Brad Lightcap told a conference this week that today's AI systems will seem laughably bad, his words in 12 months, after they ship GPT-5.
16:16However, Gurgelioros writes, OpenAI was amazing in 2022-2023 because they shipped a product that spoke for itself. Jaws dropped by those using it and seeing it for themselves. To see the company hype up future unreleased products feels like a major shift. If it's that good, why not ship it like before? It seems like we will get answers to that question sooner rather than later-ish. But for now, still plenty of open AI news to consume a cycle. That is going to do it for today's AI Daily Brief. Until next time, peace.
16:52Thank you.
From the publisher
Today on The AI Daily Brief, NLW looks at expanded restrictions the White House is considering to deny China access to advanced models. Plus, an absolute slew of news around OpenAI.
**
Check out the hit podcast from HBS Managing the Future of Work https://www.hbs.edu/managing-the-future-of-work/podcast/Pages/default.aspx
Join Superintelligent at https://besuper.ai/ -- Practical, useful, hands on AI education through tutorials and step-by-step how-tos.
**
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI.
Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe
Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown
Join the community: bit.ly/aibreakdown
Learn more: http://breakdown.network/
