Scaling Down: OpenAI's Choice to Avoid GPT-5 and Model Expansion

23 Feb 2024 · 11 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI Today Podcast Episode Notes

Episode Title

Scaling Down: OpenAI's Choice to Avoid GPT-5 and Model Expansion

Episode Overview In this episode, the hosts explore OpenAI's strategic decision to refrain from developing GPT-5 and increasing the size of its AI models. They discuss the implications of this choice and examine insights shared by Sam Altman, CEO of OpenAI, during recent interviews.

Key Topics Discussed

  1. OpenAI's Current Model Strategy
  2. End of an Era: Sam Altman suggests that the trend of creating larger models is ending. OpenAI aims to make AI better through methods other than increasing model size.
  3. GPT-4 Parameters: OpenAI has not disclosed the parameter count of GPT-4, leading to speculation about their data strategies and competition.
  1. Intellectual Property and Data Concern
  2. Data Scraping: OpenAI's previous reliance on data sources (e.g., Twitter API) raises questions about data acquisition and transparency.
  3. Protecting IP: Not revealing the model size may be a tactic to safeguard intellectual property and prevent competitors from understanding their data advantage.
  1. Diminishing Returns on Data Scaling
  2. Saturation Point: Altman implies that OpenAI might have reached a limit on the effectiveness of simply increasing data size, suggesting that advances will need to focus on improving model capabilities through other means.
  1. Future Advancements in AI
  2. Human Feedback and Tuning: The role of user interactions (thumbs up/down) in refining AI responses is emphasized. This feedback loop can potentially enhance model performance without the need for larger data sets.
  3. Transformer Model Improvements: Discussion on optimizing existing transformer architectures instead of scaling up parameters.
  1. Competitive Landscape
  2. AI Market Dynamics: OpenAI faces competition from other entities like Anthropic and Google, alongside emerging projects from Elon Musk.
  3. Regulatory Pressure: The emergence of calls from industry figures for a moratorium on AI advancements signals an ongoing struggle for ethical AI development amidst fierce competition.

Insights on Public Relations

  • Strategic Communications: Altman’s announcement about not pursuing GPT-5 may be a PR move to manage public perception while still working on enhancements to the existing model.
  • Perception vs. Reality: The distinction between product iterations (e.g., GPT-4.5 vs. GPT-5) is likened to consumer products where numerical upgrades do not always correlate with significant improvements.

Conclusion The episode concludes by emphasizing the importance of monitoring OpenAI's next steps and potential improvements to GPT-4. The hosts underscore that while the decision to halt development on GPT-5 may seem progressive, the underlying advancements could still be significant, albeit under a different naming convention.

Key Takeaways

  • OpenAI is shifting from scaling model size to enhancing existing models.
  • Intellectual property and data acquisition remain critical concerns.
  • User feedback will play a crucial role in the evolution of AI.
  • The competitive landscape is intensifying, with ethical considerations at the forefront.
  • Communication strategies are pivotal in shaping public perception of AI advancements.

Additional Resources

  • [Invest in AI Box](https://republic.com/ai-box)
  • [Get on the AI Box Waitlist](https://AIBox.ai)
  • [AI Facebook Community](https://www.facebook.com/groups/739308654562189)
  • [Learn more about AI in Music](https://musicalai.pro/)
  • [Learn more about AI Models](https://aimodelspro.com/)

Privacy Information

  • [Privacy Policy](https://art19.com/privacy)
  • [California Privacy Notice](https://art19.com/privacy#do-not-sell-my-info)

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Happy Monday, I hope you are all doing fantastic. I'm finally back in Arizona after being out of town for around a month, so it feels good to be back in the studio. And today we have a juicy topic. And that is why OpenAI, specifically as its training for chat GPT-4 and beyond, is currently not working on GPT-5. The reason for that, and also the reason why that might be a little bit of a deception. And why they say they are no longer going to be making bigger AI models. So let's kick the podcast off today. A bunch of the stuff we're going to be talking about today is coming from some interviews Sam Altman, the CEO of OpenAI, did recently at MIT.

0:43And he said that in this recent interview, he said that he thinks we're at the end of an era where it's going to be like these giant, giant models and that they're going to find ways to make them better in other ways. This is really interesting. We're going to dive into what that means. But essentially what he's saying is that this might be the end of when we just keep seeing these models get bigger and bigger and bigger. And this is interesting because they actually have not released how many parameters are in the GPT-4 model. when they came out with the first GPT-2, it was around 1.7 billion parameters.

1:20And when they came out with GPT-3, which is, you know, the regular chat GPT version that they launched for mass production that everyone first used back in December, January, that one had a whopping 175 billion parameters. So this thing went from like, this thing got massively bigger. And then when GPT-4 came out, they didn't actually announce how big it got. And my opinion on this is that, well, a couple things. Number one, they're probably trying to protect a little bit of their intellectual property. So not announcing how big the model is. It could be like they just don't want people to know how big it is because it's so ginormous.

1:58And the reason for that might be, you know, competitors might question, how did you get that much data? You know, there's been recent things. For example, when Elon bought Twitter, he shut off the Twitter API that was directly feeding into OpenAI. So they were scraping like all the Twitter data in the universe. And so I think that they may just have a really insane amount of data and they don't want people to know how big it is or perhaps how small it is and how they're able to effectively work on that. And then the other interesting thought is perhaps it didn't actually get that much bigger than GPT-3, they just figured out ways to fine-tune it, but for all of these different cases, they just don't want people to know what they're doing.

2:39They don't want their competitors to know what they're doing because there has been a handful of really big competitors that have come out with a lot of money. There's Anthropic, Elon Musk is rumored to be working on something with Twitter to create some AI models. There's obviously Google, and a lot of different people are focusing Quora on making these big AI models and so I think they just don't want people to know how many parameters they have how they're able to make it so effective and there's also the possibility that maybe they just have such an insane amount of data that people would be mad you know that how did you gather all that data and yada yada so really interesting but apparently these things are not going to get any bigger aka OpenAI probably hit the max for how much just sheer data they can like suck into their vacuum beyond just getting new data.

3:34So as we know, everything got cut off in 2021. And beyond that, it's probably not a big deal if they're not updating things live, because now they just plug the internet into it. And for anything newer than that, you'll just use internet to scrape and do searches and pull in relevant stuff like Google and Bing are currently doing. So it's going to be really interesting. I think Altman, he, you know, His statement pretty much suggests that ChatGPT might be the last major advance to emerge from just feeding these models more and more data. And what the actual advancements we're going to start seeing now, because apparently you're just getting diminishing returns as you scale up to more data, or maybe they already have so much that there isn't actually that much more to do.

4:20So I think the new ways we're going to start to see this progress is just from how we're using these transformers. So Nick Frost, who's a co-founder at Cohere, he previously worked at AI or worked on AI at Google. He says that when he, like listening to Altman talking about this, he thinks that this is true and that he believes that the progress on these transformer models that run, that are, you know, the backbone of ChatGPT, and that type of thing. That what's really going to make these things scale, he says, there's a lot of ways of making Transformers way, way better and more useful, and lots of them don't involve adding parameters to the model.

5:06He also says that he just thinks that the architecture and further tuning based on human feedback are really promising directions that researchers are probably looking at right now to make these things more effective and if you think about it open ai and chat gpt already have pretty much the biggest use case user case um on this right and tesla was kind of in the same thing with their self-driving because they had all of these cars out on the road um with self-driving and they're able to capture all of this data and see pretty much every time uh you know for example in the case of tesla if you got autopilot on and a user has to grab the steering wheel and take over, Tesla would be able to look at, you know, what was happening leading up to that, like manual takeover.

5:50And that's essentially using human feedback to help tune the model and say, oh, it was doing something wrong, therefore the human had to take over. And so, you know, I've been saying this for a long time with cars like Tesla, it's gonna be hard for other people to catch up with the same level of self driving, because they just have such a massive user base, feeding it data and chat GPT is in the exact same position where they launched, you know, ChatGPT, they had over 100 million monthly active users, probably way up from there. And all of those users, there's thumbs up, thumbs down on every single message that came out of ChatGPT.

6:23And so people were telling it bad answer, good answer. And so that was helping tune it. And in addition, if someone asks a question and they rephrase it and ask it again, it also knows, hey, the first answer wasn't good. So they have all of these things kind of built in that they can use to fine tune it just from all of the users that they currently have. And so it might not actually be necessary for them to just keep scaling up with bigger and bigger data sets in addition to other fine-tuning things but the like the users a literal human saying this was a good or a bad answer is one of the best ways that they can train these things and they just have the most they have the biggest user base and it's gonna be pretty hard for a lot of these other companies to catch up based off of that so it you know it's pretty interesting obviously when When GPT-4 was about to come out, there was all these memes and tech people speculating and posting these graphs of like, GPT-4 or 3 was trained on 1.7 billion parameters.

7:19GPT-4 is going to be like, wait, wait, there's a trillion parameters or something. We don't really know what GPT-4 is at, and it doesn't seem like it really matters at this point. so one other thing that is pretty interesting given all of this is that uh you know sam altman also said that the company isn't going to be training gpt5 they said we're not working on this for some time and it's interesting because this all comes on the back of you know elon musk and a lot of other tech people signing an open letter to you know the government but really and like the public at large but really obviously aimed at open ai who's got the lead on this saying no AI models beyond a GPT-4 capability should be trained, yada yada.

8:03OpenAI obviously didn't say, okay, sure, we're not going to do this because, you know, it's viewed Elon Musk, for example, is now starting, it would appear, starting an AI company. So it would appear that, you know, a lot of these companies calling for no more advancements are maybe trying to catch up anyways. So all it would do is benefit them. But what's really interesting is, you know, Sam Altman obviously doesn't want the bad PR and, you know, doesn't want to be viewed as the guy that like ran, um, blindfolded straight into some giant AI disaster. So he, he obviously wants to appear to be, uh, you know, taking things in a measured and deliberate manner.

8:39So, I mean, kudos to him on this, but he said, you know, we're not working on GPT-5 at the moment. We're just working on training or like improving GPT-4 and boom, I think that's, that's the big, uh, that's the big, that's the elephant in the room kind of or I guess the big surprise big secret whatever you want to call it like improving GPT-4 is like essentially working on the next version of of AI now I'm not saying that's a bad thing I'm just saying it's funny that people are like oh good he's not working on GPT-5 but it's like what's the difference of GPT-4.5 GPT-4.6 like you could just keep calling it GPT-4.something and essentially it could be what other people would have called GPT-5.6.7.8 it's kind of funny I mean you see the same thing with the iPhone right?

9:21It's like this misconception that if the number got better, the model got better. And it must be like so much better. Like, you know, when the iPhone 10 comes out and then the iPhone 11 comes out or whatever, it's like a lot of the times, you know, these iPhones aren't actually that different. And the reverse seems to be true where you just don't have to label it a bigger number with like GPT-4, GPT-5. And all of a sudden, you know, people are like, oh, cool, you're not going to GPT-5 when, you know, it's, there's no different, right? You already connected the internet to GPT-4 and created plugins and creating all this crazy stuff that might be GPT-10, but we're just calling it lower things.

10:00So I think that's, I mean, good job for PR control, you know, damage control for open AI. And personally, I think, you know, improving it obviously has to be done for a lot of different things, making sure that it's safe and capable, But I think it's just a misconception that like people are celebrating if they cared about AI advancement, you know, that GPT-5 isn't in the works. So I think it's not really saying much by that statement, but it is interesting. And it is interesting to note that given the previous conversation, it would appear that GPT-5, even when they do start working on it, probably isn't going to be much bigger as far as parameters and data goes than GPT-4.

10:44So really interesting to see what's going on with all of this advancements in AI, in tech. It's going to be an interesting space to watch and to see what exactly these improvements in GPT-4 are in the future.

From the publisher

In this episode, we delve into the factors driving OpenAI's reluctance to pursue GPT-5 or increase its model sizes, shedding light on the strategic considerations and potential consequences.

See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

More from AI Today

All 897 episodes
Scaling Down: OpenAI's Choice to Avoid GPT-5 and Model ExpansionAI Today · 11 min
Listen in VO