GPT-4 More Creative Than the Average Person According to Recent Study

17 Sep 2023 · 12 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The AI Daily Brief: Episode Summary

Episode Title

GPT-4 More Creative Than the Average Person According to Recent Study

Podcast Overview Host: NLW Description: A daily news analysis show focusing on artificial intelligence, covering creativity, industry disruptions, and ethical considerations surrounding AI.

---

Key Topics Discussed

  1. AI Creativity Study
  2. Research Highlight: A study published in *Nature* suggests that GPT-4 outperforms the average human in creativity tasks.
  3. Study Context: This research utilized the Alternative Uses Task (AUT), where participants generate creative uses for everyday objects.
  4. Findings:
  5. On average, AI chatbots produced more creative responses than human participants.
  6. However, top human responses were still competitive with AI outputs.
  7. Implications: Raises questions about the nature of creativity and the evolving capabilities of AI.
  1. Methodology of the Study
  2. Participants: 250 individuals, aged 19 to 40, with diverse employment backgrounds.
  3. Judging Creativity:
  4. Responses were evaluated based on originality and subjective creativity ratings from trained human judges.
  5. The study emphasized that creativity is multifaceted, and this analysis is limited to divergent thinking.
  1. Future Directions of AI
  2. NextGPT: Introduction of an any-to-any multimodal LLM aimed at creating more advanced AI systems that can handle multiple types of inputs and outputs (text, images, audio, etc.).
  3. Purpose: Enhances human-like interaction with AI, leveraging various modalities to mimic real-world communication.
  4. Potential Releases: Anticipation of multimodal models from companies like Google and OpenAI.
  1. AI in Health and Sensory Functions
  2. AI That Can Smell: Dr. Jim Phan’s work on a neural network capable of identifying and categorizing smells.
  3. Method: Utilized a dataset of 5,000 labeled molecules to train the AI.
  4. Application: Could be useful in culinary fields and robotics.
  • RETFound: A foundational model for ophthalmology that analyzes retinal images using self-supervised learning.
  • Significance: Offers pivotal diagnostic tools for both eye health and systemic diseases, potentially reducing the need for expert human labeling.
  • BrainLM: A model trained on fMRI data for analyzing brain activity.
  • Impact: Demonstrates how techniques from generalist models can be applied to specialized medical fields, creating a feedback loop that accelerates AI development across disciplines.

---

Key Takeaways

  • AI Creativity: While AI has surpassed average human performance in specific creative tasks, exceptional human creativity remains competitive.
  • Advancements in AI: The next generation of AI models is focusing on multimodal capabilities, which could revolutionize interactions and applications.
  • Healthcare Applications: The integration of AI in health diagnostics is expanding, with models being developed to analyze complex biological and sensory data.

---

Conclusion The episode provides a comprehensive overview of the latest research in artificial intelligence, highlighting both the advancements in creative capabilities and the potential applications in healthcare. As AI technology continues to evolve, its implications for society and various fields are becoming increasingly profound.

Join the Community: [AI Breakdown Discord & Newsletter](http://breakdown.network/) Subscribe: [YouTube Channel](https://www.youtube.com/@TheAIBreakdown)

---

*This summary is designed to encapsulate the main points and insights discussed in the podcast episode for easy reference and understanding.*

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:01Today on the AI Breakdown, we're looking at some of the latest research in artificial intelligence, including one test at least that suggests that GPT-4 is more creative than the average human. The AI Breakdown is a daily podcast and video about the most important news and discussions in AI. Go to breakdown.network for more information about our Discord, our newsletter, and our YouTube channel. Welcome back to the AI Breakdown. Today we are doing a little bit of a research roundup. This isn't something that we've done for a little while now, and there has been, especially over the last week, a lot of research that I've made note of, I've bookmarked on Twitter, and so I wanted to share some of the things that I thought were the most interesting new papers to come out.

0:42I also thought that this would be good for a weekend episode when there's a little bit less news. So let's begin with a piece that I think has some pretty significant implications. Now, the title in the Nature article summing up this research is Best Humans Still Outperform Artificial Intelligence in a Creative Divergent Thinking Task. What I will call your attention to, though, is the use of that word best. Because this article could very easily have been framed very differently. Here's the key line. On average, the AI chatbots outperformed human participants. Outperformed them on what, you ask?

1:18Well, this was a study about creativity. As the authors write, creativity has traditionally been considered an ability exclusive to human beings. However, the rapid development of AI has resulted in generative AI chatbots that can produce high-quality artworks, raising questions about the differences between human and machine creativity. In this study, we compared the creativity of humans with that of three AI chatbots using the Alternative Uses Task, AUT, which is the most used divergent thinking task. Participants were asked to generate uncommon and creative uses for everyday objects. And then that's where we get back to this line, on average, AI chatbots outperformed human participants.

1:57While human responses included poor quality ideas, the chatbots generally produced more creative responses. However, the best human ideas still matched or exceeding those of the chatbots. Now, this paper is super interesting. Obviously, I'm going to include a link to all of these in the show notes for this episode. But for creativity geeks out there, this gets deep into different theories of creativity. This is, of course, one of those neural processes that we don't necessarily have the best understanding of. We have an intuitive sense as humans what is creative when we see it, and what people mean when they say they're creative or being creative, but it's not a term that has a strict scientific meaning, which of course makes it harder to test in a laboratory-style environment.

2:35So let's talk a little bit more about the methodology they used. Again, they write that the most used test of divergent thinking is the alternative uses task, in which participants are asked to produce uncommon creative uses for everyday objects. As they point out, divergent thinking has traditionally been assessed by tests requiring open-ended responses. In this test, there were a little over 250 participants, which included 108 males, 145 emails, two other and one that preferred not to identify their gender. The average age of the participants was 30.4 years, ranging between 19 and 40 years.

3:07142 were employed full-time, 37 were employed part-time, 30 were unemployed, and 42 were homemaker, other, retired, or disabled. The participants came predominantly from the UK and the US, and they were tested against three different chatbots, basically representing GPT-3, GPT 3.5, and GPT 4. For each of four objects, participants were prompted with, For the next task, you'll be asked to come up with original and creative uses for an object. The goal is to come up with creative ideas, which are ideas that strike people as clever, unusual, interesting, uncommon, humorous, innovative, or different.

3:39Your ideas don't have to be practical or realistic. They can be silly or strange even, so long as they are creative uses rather than ordinary uses. The objects tested included rope, box, pencil, and candle. Now, there are two ways the responses were judged. One was an attempt to systematize and mathematize it. They write, the originality of divergent thinking was operationalized as semantic distance between the object name and the AUT response. Basically, how uncommonly were the words written in the response associated with the word from the prompt? Now, on top of that, they also collected subjective creativity and originality ratings from six humans that had themselves been trained to rank creativity and originality in this case in a similar way.

4:18They were asked to rate each response on a five-point scale, with one being not at all creative and five being very creative, and they were not told that some of the responses were generated by AI. The conclusion the authors write? The results suggest that AI has reached at least the same level or even surpassed the average human's ability to generate ideas in the most typical test of creative thinking. Although AI chatbots on average outperform humans, the best humans can still compete with them. However, the AI technology is rapidly developing and the results may be different after half a year.

4:48On basis of the present study, the clearest weakness in humans' performance lies in the relatively high proportion of poor quality ideas which were absent in chatbots' responses. This weakness may be due to normal variations in human performance, including failures in associative and executive processes, as well as motivational factors. It should be noted that creativity is a multifaceted phenomenon and we have focused here only on performance in the most used task, measuring divergent thinking. So, of course, a couple caveats to the study. As the authors themselves point out, there is a whole lot more to creativity than this one type of test, so big claims and big implications should be taken with a grain of salt.

5:22Second, even within the context of a single test, this is one sample group. It's not at all to say that a different group of 250 humans wouldn't have a very different set of responses. And yet, in spite of all those caveats, it's hard to ignore entirely. AI writer Andrew Curran says, I've been saying this for some time. The average human was surpassed in March by GPT-4. It seems to me that there should have been more fanfare. Some of the bridges that are about to be crossed by the next iteration will be transformative for human society. Next up, let's move to something that is a little bit more about the future of LLMs.

5:56A paper called NextGPT offers an any-to-any multimodal LLM. Now, multimodal has been one of the big themes and one of the clear what's coming nexts in the world of chatbots and LLMs. Humans exist in a multimodal world in which we get our inputs from lots of different sources, be it text, images, audio, video, smells, something else, and likewise we output things in a variety of different ways. Now of course, the LLMs that we use today are much simpler. They're usually text in and something out, be it text-to-text like ChatGPT, text-to-image like MidJourney, or text-to-video like PicoLabs and Runway.

6:33It's been quite clear for some time, however, that adding multimodality and the ability to use different types of inputs was always going to be a major part of the next generation of LLMs. We've seen little nibbles against that goal, with some more recent chatbots allowing, for example, image-based inputs, but that's quite different than what this research paper is promising, which they call any-to-any multimodal. The researchers write, As we humans always perceive the world and communicate with people through various modalities, developing any-to-any MMLLMs capable of accepting and delivering content in any modality becomes essential to human-level AI.

7:05To fill the gap, we present an end-to-end general-purpose any-to-any MMLLM system, NextGPT. We connect an LLM with multimodal adapters and different diffusion decoders, enabling NextGPT to perceive inputs and generate outputs in arbitrary combinations of text, images, videos, and audios. Overall, our research showcases the promising possibility of building a unified AI agent capable of modeling universal modalities, paving the way for more human-like AI research in the community. Again, there is a link down in the show notes to this paper, go check it out. I think it's going to be especially relevant in the context of some of the upcoming releases that we're likely to see over the next few months.

7:38We've got Google's Gemini, which is expected to be a multimodal model. And OpenAI is also teasing its developer event in November, which among the various speculations of what they'll announce, some are thinking that multimodality might be a part of it. Our next set of research papers have to do with human-like health and body functions. Dr. Jim Phan from NVIDIA introduces an AI that can smell. He writes, A neural network can smell like humans do for the first time. Digital smell is a modality that AI community has long ignored, but may be one day useful for robot chefs. Here's how to do smell-to-text.

8:11One collected 5 ,000 molecules and asked humans to label creamy, chocolate, alcoholic, beefy, spicy, citrus, etc. This dataset is one of a kind and a huge contribution from the paper. Two train a graph neural network to map the molecule to each label. Each molecule is a graph of atoms described by valence, degree, hydrogen count, hybridization, formal charge, atomic number, etc. The GNN predictions match well with human experts on novel smells. The embeddings give us a, quote, principal odor map, POM, that faithfully represents hierarchies and distances among odorants. And now Jim did a pretty good job summing it up, but you can also check out the nature piece for a little further elucidation.

8:47Effectively, this is exactly what he described. It's a neural network that map different molecular compositions against human-labeled smells, and then can use that to provide descriptions for novel smells, including some that aren't from nature, and in the process creates a map where we see how related or unrelated different smells are based on their chemical composition. I think Jim's right to note that this is an area that hasn't maybe had quite as much development. But to his point, to the extent that there are robot chefs in our future, this could be extremely important. Another foundational model relating to the human body that was announced this week was RETFound.

9:21RETFound is a foundational model for ophthalmology. RETFound is a retinal image model that was trained using self-supervised learning. As Nature writes, that means that the researchers did not have to analyze each of the 1.6 million retinal images used for training and label them as normal or not normal, for instance. Instead, the scientists used a method similar to the one used to train large language models such as ChatGPT. The AI tool harnesses mirrored examples of human-generated text to learn how to predict the next word in a sentence from the context of the preceding words. In the same kind of way, RETFound uses a multitudinal of retinal photos to learn how to predict what missing portions of images should look like, says Pierce Kane, an ophthalmologist at Morsefield Eye Hospital, NHS Foundation Trust in London.

9:59Over the course of millions of images, the model somehow learns what a retina looks like and what all the features of a retina are. Now, there are a bunch of reasons that this is valuable. One is obviously it can just understand problems with the retina and the eye itself. But because of the retina's role in the human body, it can also be a good diagnostic tool for other problems. For example, Keene says, if you have some systemic cardiovascular disease like hypertension, which is affecting potentially every blood vessel in your body, we can directly visualize that in retinal images. Retinal images can also be used to evaluate neural tissue, which gives the model the ability to predict brain diseases such as Parkinson's.

10:31As Kean points out, the need for expert human labeling is a significant barrier to AI-enabled healthcare. By being much more label-efficient, RETFound opens the possibility of applying AI in rare disease. So really, when it comes down to it, this research is relevant not only for what it can do specifically, but also as an example for how future health-based models might be trained. And lastly, and not unrelated, is BrainLM. Van Deek Lab writes, Introducing BrainLM, the first foundation model for fMRI analysis trained on 6 ,700 hours of brain activity data. The abstract reads, Basically, this is another foundational model based on a particular part of the human body, in this case the brain, that can be used for a variety of purposes.

11:11Now, the implications for this one are extremely numerous. But for the purposes of this show, what I wanted to point out is just how much developments in one area of AI can influence how other areas of AI evolve as well. We're starting to see the techniques that were used to train generalist models, such as GPT, be used for these highly specific medical and healthcare uses and many other uses as well. This is creating a pretty powerful feedback loop that is, I believe, accelerating the way in which artificial intelligence is flowing to and shaping other fields of study. Hopefully today gave you a better sense of some of the latest and most interesting research in the field.

11:46And so until next time, peace.

11:52Thank you.

From the publisher

On this research recap, NLW looks at

AI creativity https://www.nature.com/articles/s41598-023-40858-3

BrainLM https://t.co/MUobqXULfb

RETFound retinal model https://www.nature.com/articles/d41586-023-02881-2

Any-to-Any Multimodal https://huggingface.co/papers/2309.05519

AI that can smell: https://twitter.com/DrJimFan/status/1701611251376497046

TAKE OUR SURVEY ON EDUCATIONAL AND LEARNING RESOURCE CONTENT: https://bit.ly/aibreakdownsurvey
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI. 

Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe

Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown

Join the community: bit.ly/aibreakdown

Learn more: http://breakdown.network/

More from The AI Daily Brief: Artificial Intelligence News and Analysis

All 1,099 episodes
GPT-4 More Creative Than the Average Person According to Recent StudyThe AI Daily Brief: Artificial Intelligence News and Analysis · 12 min
Listen in VO