Decline of the Titans: GPT-4's Deterioration Unveiled

12 Mar 2024 · 11 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI Today Podcast Episode Notes

Episode Title

Decline of the Titans: GPT-4's Deterioration Unveiled

Overview In this episode, the hosts delve into a study from Stanford that reveals a decline in performance of the AI model GPT-4, leading to discussions on user expectations and the future of AI. OpenAI's responses to the criticism and the implications for businesses using GPT-4 are critically examined.

Key Topics Discussed

  1. Study Findings
  2. Research Background: Conducted by Stanford and Berkeley, evaluating GPT-4's performance over time.
  3. Main Findings:
  4. Significant decline in GPT-4's ability to answer complex mathematical queries.
  5. Performance on prime number problems dropped to 2.4% accuracy, compared to near-perfect performance in earlier tests.
  6. Users report perceptions of deteriorating quality, sparking debates about whether AI is genuinely declining or if user expectations are rising.
  1. Comparisons with Previous Versions
  2. GPT-3.5 vs. GPT-4:
  3. GPT-3.5 has shown improvements, particularly in basic math, while GPT-4 struggles with complex coding tasks (only 10% success on LeetCode problems compared to 50% earlier).
  1. User Engagement and Expectations
  2. Decline in User Engagement: First reported drop in ChatGPT usage since its launch, linking to perceived performance issues.
  3. User Perception: Some believe the AI is becoming "dumber," while others argue that user understanding of AI limitations is growing.
  1. OpenAI's Response
  2. Counterarguments from OpenAI:
  3. Peter Welder (OpenAI VP) asserts that the AI's performance is consistent; rather, users are becoming more aware of its limitations.
  4. Emphasis on the importance of transparency in AI operations and updates.

Implications for Businesses

  • Risks of Relying on Unpredictable AI:
  • Businesses using GPT-4 may face challenges if performance worsens—potentially affecting customer outputs and satisfaction.
  • Need for Better Communication:
  • Experts suggest OpenAI should provide clearer updates on model changes to ensure users can adapt and test their applications effectively.

Expert Insights

  • Zaharia and Zhao's Positions:
  • Zaharia advocates for transparency regarding model updates, emphasizing the need for developers to understand changes and their impacts.
  • Zhao cautions against overcomplicating user experience with too much technical information.

Conclusion

  • The episode highlights critical concerns regarding the reliability and performance of AI models like GPT-4. It underscores the necessity for transparency from AI developers and the importance of user education about the capabilities and limitations of AI technologies. The discussion points to a broader need for understanding as AI continues to evolve.

Key Takeaways

  • Performance Decline: Evidence suggests GPT-4's performance has declined, particularly in specialized tasks.
  • User Awareness: Increased user scrutiny may contribute to perceptions of decline.
  • Business Risks: Dependence on unpredictable AI models poses significant risks for businesses.
  • Call for Transparency: Greater clarity about AI model updates is essential for effective usage and trust.

Additional Resources

  • Invest in AI Box: [AI Box](https://republic.com/ai-box)
  • Join the AI Box Waitlist: [Waitlist](https://AIBox.ai/)
  • Connect with the AI Community: [Facebook Group](https://www.facebook.com/groups/739308654562189)
  • Explore AI in Music: [Musical AI](https://musicalai.pro/)
  • Learn about AI Models: [AI Models Pro](https://aimodelspro.com/)

Privacy Policy For more information on privacy policies, visit: [Privacy Policy](https://art19.com/privacy) and [California Privacy Notice](https://art19.com/privacy#do-not-sell-my-info).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The wait is over. Dive into Audible's most anticipated collection, The Best of 2025. featuring top audiobooks, podcasts, and originals across all genres. Our editors have carefully curated this year's must-listens from brilliant hidden gems to the buzziest new releases. Every title in this collection has earned its spot. This is your go-to for the absolute best in 2025 audio entertainment. Whether you love thrillers, romance, or nonfiction, your next favorite listen awaits. Discover why there's more to imagine when you listen at audible.com slash best of the year.

1:03podcast we're going to be diving into if chat GPT and GPT-4 are actually indeed getting worse if this is just a figment of people's imagination is this just the fact that you know we're starting to expect you know and set higher bars for these AI models as it becomes less of a wow factor so we're going to be diving into all of that without further ado let's get into it. I think the first thing that is kind of making headline news right now is the fact that a handful of researchers out of Stanford have actually run an experiment since the beginning of ChatGPT till now, testing it on its ability to answer certain questions.

1:38And over the course of that amount of time, it was Stanford University and Berkeley, they found that indeed it did get worse at answering a number of questions. So essentially what happens is they showcased in the ARXIV preprint archive, And they suggested that GPT and its predecessor, which is GPT 3.5, are actually evolving backwards. So GPT-4's once really reliable ability to deal with complex mathematical queries, I mean really reliable, it used to be terrible, then they fixed it to be a little bit better. But even that has been on a downward slope, so it's now managing to correctly solve about 2.4 % of questions about prime numbers.

2:23it's some like prime number math questions and it can get it right about 2.4 percent of the time but back when they originally tested this which i believe was back in march um it was actually able to do it nearly every time so this is some serious um decline in its ability to solve certain types of math problems so i think interestingly enough while the modern version gpt4 has been giving you know not very great uh essentially not very good details for its problem solving steps the older version gpt 3.5 has actually been enhancing its skills so specifically in the realm of basic math um i still think that it is floundering when faced with more intricate code generation of course so that's not where it's great but the online tech communities have seemed to be you know debating this for quite a while and whether chad gpt's you know performance is indeed declining so in the case of AI getting worse, people are asking, is this the AI getting worse or the users becoming more discerning about its limitations?

3:26So some have even suggested that this perception might have contributed to the drop in engagement that we have. ChatGPT announced that from last month to this month, they did see a slight decline in engagement and usage of their tool. So a first for the application, that was the first time that this happened for the application since it was launched earlier. So GPT-4's most recent iteration seems to be really grappling with spatial reasoning and coding related queries. So for instance, researchers put it through its paces with coding problems for LeetCode, an online learning platform for coding.

4:02And they found that the recent version could only produce functional code 10 % of the time. So earlier this year, it was also tested and it was able to execute successfully 50 % of the tasks they gave it. So in a chat with Gizmodo kind of did an interview with the researchers Matai Zaria and James Zhao, and they expressed their concerns about the bot's declining performance. So they mentioned that GPT-4's response now features more base text and the code often requires more correction. So I think this really has a big problem that it poses for the business. And, you know, for any businesses that are intending to leverage ChatGPT to help with code generation, it's also kind of an underlying issue that is associated with depending on a proprietary AI whose functioning remains largely unscrupulous.

4:54Right. Like this is not very transparent how ChatGPT works. We can't see the code. It's not open source in contrast to projects like Llama 2 out of Facebook, which they have announced they're going to be open sourcing and offering for free. and so I think this is something that this kind of has a bigger issue which is the fact that if you go and get an API to an AI model because you test it and it can accurately you know do what you need it to do and all of a sudden this thing gets worse and worse over time this is a big problem for any developers and any organizations that have embedded GPT-4 into their business right like I personally have software where we are using the API to GPT-4 and we're getting into all sorts of really cool stuff.

5:36But if GPT-4 is getting worse and worse, all of a sudden my customers are going to be getting worse and worse outputs. And this is a serious problem, right? Like in reality, you should have one API to a tool. It should work. You should not be able to live update that tool. You should be able to make a new version, right? I think a lot of this kind of happened when, you know, OpenAI said, we're not going to be working on GPT-5. We're just going to work on fine-tuning GPT-4 and making it better and better and kind of refining that. But I think the issue is when you refine something that is live, if it's, you know, you make some sort of small tweaks and inevitably it starts negatively impacting the overall thing.

6:16I think a lot of this might be the fact that they're trying to do, they're trying to make this thing more performative, right? So it's obviously really expensive. If you have GPT-4, they only let you do 25 questions every three hours, I think, if you have premium. And the issue is, I'm sure they're like, well, let's make this thing as fast and snappy as GPT 3.5. But the problem is, inevitably, I'm imagining some of these performance enhancements they're trying to do are actually making the quality suffer. So this is a big problem. This is a big issue, especially for people that are embedding this into software.

6:52I think Zaharia, who's a professor at Stanford, he really emphasized the reliability challenges posed by integrations with these language models. But he also pointed out that the changes could be due to a shift towards more conversational patterns. So OpenAI's VP of product is Peter Welder, and he was countering these criticisms. And he said that ChatGPT is not dumber, but that its users are getting more aware of its criticisms, or I mean limitations. So, you know, you have this kind of, the researchers doing this are saying it's getting worse, opening eyes saying no, you guys are just becoming more aware of what it is unable to do.

7:30But I think, you know, from the perspective of the researchers, they did the same benchmarks and the same tests over a course of months, and it, you know, the tool performed worse. So I think that's, you know, a pretty smoking gun in this case. I think these recent findings don't actually indicate any major overhaul beyond fine-tuning, and I think they don't suggest any deliberate, preference of GPT 3.5 over GPT 4. But still, I think the researchers recognize the potential impact of even small adjustments, which could really significantly alter the AI's response pattern. So in response to this, Zaharia and Zhao are considering a broader study that encompasses challenges in other companies' language models as well, right?

8:13So they've done this for ChatGPT. it'd be really interesting to see this done on you know inflections ai and on anthropic and a number of others as well so i think despite the issues it is not all downhill for gpt like this thing isn't just a model that's gonna completely die because it's getting worse and worse i think the model actually displayed enhanced resilience to prompt injection jailbreak attacks since its launch so that's one thing that has improved it's also been um it's also reduced its response to you know some prompts that people would consider problematic which is just you know open ai trying to lock down their model more um and that again is uh you know that's something that people could argue about the pros and cons however i think there are some concerns remaining about its ability to generate responsives that are really responsive comprehensive and actually accurate it's not perfectly accurate um we know that i think everyone knows that you definitely have to fact check everything that comes out of it i don't think that is a surprise to anyone in any case zaharia is also an ai uh is also an ai consulting firm executive and he advocates for more transparency from open ai about the updates and changes to their models i think this is an absolutely must necessity right like if you are if you are supplying this to customers who are embedding this into their products and you're making tweaks and changes live you absolutely have to tell people what's going on they need to they need to be aware that they need to test this right coming from my own software background every time we make any sort of feature on our application we test anything related to you know any single button or feature or anything or bit of code that could touch the new feature we've done and then we also go through and test every feature on the entire app just to make sure nothing somehow got broken that we were unaware of and inevitably that does happen a lot right like we'll push out a new feature um that helps people do something better and then we don't realize but somehow it messed with the onboarding flow or some small aspect somewhere um and so you know you'll get a support request like typically you make one adjustment somewhere it doesn't affect something too far away but it does happen it is possible and i think the same concept applies to these ai models where you fine-tune uh one area for something specific and you don't realize that it may inadvertently have caused um harm or um yeah problems in another area of the model.

10:34So I think that, you know, Zahara specifically was saying that, you know, giving this kind of clarity to consumers on what changes were happening is going to help users understand the AI's behavior better. Zhao, on the other hand, doesn't think users would welcome that kind of complexity to their, you know, AI toy or whatever. He thinks it's just gonna be too confusing if OpenAI is giving updates on what they're tweaking. But in any case, regardless of their opinions, given the increasingly heated discussion surrounding AI regulation and potential harms, I think it would be wise for OpenAI to offer some insights into its workings to help, you know, its user base understand why their AI might be changing or doing things different.

11:17I think it's going to make a big impact, especially for businesses. I honestly think that's kind of a must so that we know what to test and if we need to make any tweaks. So this will definitely be something interesting to follow in the future.

From the publisher

In this episode, we dissect the shocking findings of a study revealing the decline in performance of GPT-4, while OpenAI deflects blame onto users, sparking debates on the future of AI.

See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

More from AI Today

All 897 episodes
Decline of the Titans: GPT-4's Deterioration UnveiledAI Today · 11 min
Listen in VO