In short
Podcast Summary: The AI Daily Brief - Is Mysterious GPT-2 Chatbot Actually GPT-5?
Overview The episode discusses the recent buzz surrounding a mysterious AI model referred to as "GPT-2 Chatbot," which has sparked speculation that it could be an unreleased version of OpenAI's technology, potentially GPT-4.5 or GPT-5. The episode also includes commentary on hardware reviews in the AI space and the implications of releasing unfinished products.
Key Topics
- Marques Brownlee's Hardware Reviews
- Marques Brownlee (MKBHD) reviewed the Rabbit R1 and described it as "barely reviewable."
- His review highlighted a trend in tech where companies release partially finished products at full price, expecting consumers to be patient for updates.
- Key observations include:
- Consumer Expectations vs. Reality: Products are marketed as complete when they are not.
- Historical Context: This trend has been visible in gaming, smartphones, and now AI hardware.
- Community Reactions
- Many reviewers and industry professionals echoed Marques's sentiments:
- Adam Vestica and Daryl Bossenjo criticized the trend of selling unfinished products.
- Dave Snyder highlighted the negative impact of this strategy on consumer trust.
- Contrasting opinions urged understanding of the rapid pace of AI development and user testing.
- The Mysterious GPT-2 Chatbot
- The GPT-2 Chatbot was found on the LIMSYS benchmarking site, raising questions about its true identity.
- Key points include:
- Performance Claims: Users reported that the model demonstrates significantly improved reasoning and problem-solving abilities compared to existing models.
- Speculations on Identity:
- Theories range from it being a pre-release version of GPT-4.5 to a new model entirely.
- Some believe it could be a modified version of older models fine-tuned on new datasets.
- Experts in the AI community have debated its origins and potential implications for future models.
- Implications of AI Hardware Development
- The episode explores the challenges in the AI hardware space, emphasizing that early adopters may face frustration due to unfulfilled promises.
- Discussion on how AI hardware and software need to mature to meet user expectations.
- Other AI News
- Apple is reportedly hiring from Google for a secret AI research lab.
- Microsoft announced significant investments in AI skill training in Southeast Asia.
Conclusion The episode underscores the ongoing challenges and complexities in the AI landscape, particularly about product launches and consumer expectations. The emergence of the GPT-2 Chatbot adds another layer of intrigue, highlighting the excitement and speculation surrounding advancements in artificial intelligence technology.
Key Takeaways
- There is growing concern over the trend of launching incomplete AI products.
- Consumer expectations are often misaligned with the reality of product readiness.
- The nature of the GPT-2 Chatbot is shrouded in mystery, sparking speculation about its capabilities and future developments in AI models.
- The podcast highlights the need for transparency and accountability in AI product launches to foster trust among consumers.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Today on the AI Breakdown, is GPT2Chatbot actually GPT5 in the wild? Before that on the brief, Marques Brownlee calls the Rabbit R1 barely reviewable. The AI Breakdown is a daily podcast and video about the most important news and discussions in AI. Go to breakdown.network for more information about our YouTube, Discord, and our newsletter. Welcome back to the AI Breakdown Brief, all the AI headline news you need in around five minutes. Recently, prominent YouTuber and product reviewer Marques Brownlee got a lot of attention in the AI space when he called the Humane pin the worst product I've ever reviewed, dot dot dot, for now.
0:42Basically, the gist of that review was that almost everything that was useful about the Humane pin was better done with just your phone. However, at the same time, he saw the glimpse of the future, where it could be really nice to not have to be tethered to a screen. This kicked up a whole discussion about the role of reviewers and the nature of what product launches should be. And of course, it was a little bit colored by the fact that Humane had been building this product for years and had spent hundreds of millions of dollars on it. But it also brought up this larger question of just whether AI hardware is ready for primetime.
1:12Well, that question has come up again with a vengeance now that Marquez has dropped his review for the Rabbit R1. He sums it up in two words, barely reviewable. Interestingly, Marquez has not just identified this as a problem in the AI hardware space, but sees it as the apex of a trend that has been growing for a long time. As he puts it, this is the pinnacle of a trend that's been annoying for years, delivering barely finished products to win a race and then continuing to build them after charging full price. Games, phones, cars, now AI in a box. Let's listen to a clip from his review to get the gist of what he's saying.
1:47Like it feels like it used to just be make the thing and then put it on sale. Now it's like, put it on sale and then deliver the half-baked thing and then iterate and make it better. And hopefully with enough updates, then it's ready and it's what we promised way back when we first started selling it. And then this whole period in the middle is a mess. And it's across all kinds of product categories too. It's also happening with cars and vehicles getting announced and then delivering with like a half-finished state where you just don't get a lot of the features that you paid for. And they're eventually coming soon.
2:20with a software update. You know, smartphones, obviously we've been seeing this for years, but it does seem like now more than ever, there's at least one feature, one major feature of every smartphone launch that gets announced, but that's not coming until later in the year. And now these AI based products are at like the apex of this horrible trend where the thing that you get at the beginning is like borderline non-functional compared to all the promises and all the features and all the things that it's supposed to maybe someday be. But you still pay full price at the beginning, which is what makes it so crazy.
2:55Now, let's talk about community responses to this. A lot of people agreed. Adam Vestica from The Shortcut writes, This is so true of gaming, and it's a horrible trend that started during the last generation. I've got nothing against launching a game in early access, but so many titles are released with the promise of getting better down the line. Daryl Bossenjo writes, The Rabbit R1 and Humane AI Pin are the culmination of a trend of launching unfinished products that are sold at full price, and then users get the promised features only after software updates. I saw many people compare this to treating your early customers like investors, which of course they're not.
3:28Dave Snyder writes, This is the byproduct of the product design and management strategy that pushes half-baked products for a few data points. An overwrought practice where too many test underdeveloped solutions groping blindly for signal while ignoring the glaring flaws. In the realm of physical products, this approach is an absolute nightmare verging on fraud. Maybe, just maybe, we'll get back to shipping what's right. And he younger joked, you, an AI hardware CEO, says, I want to build a fun product that will improve the pace of AI progress. Your investors say, ship an MVP and iterate, do things that don't scale.
4:00Marcus Brownlee says, I'm about to end this man's whole career. And just as with the humane pin, where Marcus was representative of the larger sort of consensus review take, the Rabbit R1 is also getting some pretty negative reviews even beyond Marquez. Snazzy Labs writes, the humane AI pin is one of the worst products I've ever reviewed. And the Rabbit R1 is even worse. Video later this week on why good ideas don't translate to good products and why Apple and Google will clobber all of these AI startups by the end of the year. Now, just to add a little bit of controversial sauce on top of this, Snazzy Labs also added, what I would like to reiterate is that the AI pin is leagues better than the R1.
4:36Yes, it's overpriced and not very good, but it knows what it wants to be and does a very limited number of things decently well. The same cannot be said of the Rabbit. They're not comparable. But what about the Rabbit team? Ryan Fenwick, who does communications over there, says, Honest feedback, which is cool. We're in the very early stages of the new AI hardware industry. The important thing is to move fast, continuously update the product, and keep improving the experience for those of you who are on this ride with us. Jesse Liu, the CEO, said, We shall see how fast R1 improves and evolves. We are a tiny team trying to catch the fast pace of AI.
5:07The current levels of AI need strong human-supervised fine-tuning. You can't take your time polishing features without real user testing. We will push the OTA fix as early as tomorrow to address most of the bugs we found. Thanks Marques Brownlee for your very detailed explanation of LAM. Looking forward for you to do an R1 revisit very soon. So I think that this tweet is important because it's not just a founder trying to justify that they're early, but is at least in a small level giving an argument for why this is actually the right process to use. Jesse writes you can't take your time polishing features without real user testing.
5:40Joel Karanen responded to Snazzy Labs with another argument for why the Rabbit R1 approach is better than the Humane AI pen saying, I would argue that the Rabbit R1 is better because it's substantially cheaper with no subscription. It's meant to be an impulse gadget purchase rather than a lifestyle commitment. Snazzy Labs responded with the snazzy response I will say you get what you pay for, but I think it's an interesting point. Is it more justified to charge someone$199 once to get that user feedback rather than$699 with an ongoing subscription? Now, there is a lot to unpack here. You can make all sorts of arguments from all sorts of angles.
6:13I think one of the key questions is consumer expectation versus reality. One of the things that clearly frustrates Marques is the gap between what these devices promise and what they actually deliver. In other words, I think that he would say that these companies are not presenting themselves as an early beta, but presenting themselves as the thing, ready to go, ready to be in the world, and then trying to have it both ways, saying we're just a tiny team and we're trying to grow fast when things don't go that well. Of course, another line of common commentary, largely from entrepreneurs, was pointing out that every new technology seems barely reviewable at the beginning.
6:45Graham Fleming even shared a video of one of Marquez's very early reviews and said if Marquez Brownlee decided his initial videos were barely reviewable instead of posting and iterating, would he be where he is today? Now, that's a great score for Twitter points, but of course, the difference is that Marquez wasn't asking people to pay money for something. Figma engineer Vivek writes, absolutely love what Marquez Brownlee is doing. I don't understand how people can defend dumping trash on consumers and overselling it as the next great thing. Nobody owes anyone a benefit of the doubt. Don't want to get reamed by consumers and reviewers?
7:13Make something good. I think that these questions are really interesting. It seems pretty clear so far that the agentic capacities of AI just aren't really sufficient for the agentic promises of the AI devices that have been announced and released so far. At the same time, that doesn't mean that A, those devices won't actually catch up, especially as the agentic capacity of the AI software underneath catches up, and B, it doesn't mean that AI hardware couldn't take a different approach focused on a specific type of use case that might actually fit better. For example, neither the Humane AI Pin nor the Rabbit R1 are focused on the same use case that some of the next generation of AI devices seem to be of keeping a record of all your conversations and interactions in a way that allows you to go back and better recall that information.
7:55Will that be a killer application that doesn't require AI agent capacity? Maybe. Then again, we won't know until they get released. All in all, if nothing else is clear, it's that AI hardware is going to be a very difficult space. And so for now, anyone who buys into these products has to understand that that's what they are getting into. For what it's worth, I think by and large, most people do. Now that was obviously a little long for an AI breakdown brief, but it was a big point of conversation. So I'll just hit a couple more stories before we head on to the main part of the episode. First, Apple is apparently poaching from Google for a secretive AI research lab based in Zurich.
8:29Apparently they've hired something like 36 employees from Google over the last few years, much more than any other company they've poached from. And Microsoft has announced its latest big global AI investment, with CEO Satya Nadella announcing yesterday that the company would invest$1.7 billion in Indonesia over the next four years. As part of that, they're also committing to help train 2.5 million people in Southeast Asia with AI skills, including about 840 ,000 in Indonesia itself. Anyways, friends, that is going to do it for today's AI Breakdown Brief. Next up, the main AI breakdown. All right, breakers.
9:02Consensus 2024 marks the 10th gathering of the biggest event that's devoted to all sides of the crypto, blockchain, and Web3 ecosystems. Join pioneering thinkers and builders as they delve into the future of DeFi and explore game-changing tech, from AI to ZK proofs and everything in between. The event is three days of jam-packed content, networking, and so much more. Some of the speakers at the event include Chris Dixon, the founder and managing partner at A16Z Crypto, Sergey Nazarov, the co-founder of Chainlink, Kathy Wood, the CEO of ARK, Hester Peirce, commissioner, of course, from the USSEC, and Tom Emmer, Republican Majority Whip for the US House of Representatives.
9:36Visit consensus2024.coindesk.com to learn more and save 15 % on registration with the code BREAKDOWN. That is 15 % on registration with the code BREAKDOWN. Hello, Breakers. Quickly, before we get back into the rest of the episode, you guys might have been following along my journey with the AI Breakdown, which is a very similar show to this, but for the artificial intelligence industry. One of the things that I found as I dug into that show is that there was a huge need for better educational resources that were actively practical and useful right away. I've just announced a new company and platform called Super Intelligent that's trying to build exactly that.
10:11It's a video platform for learning AI that features fun, fast, and immediately useful video tutorials. Each video tutorial is around five minutes and comes with a step-by-step how-to that gets people actually using the tools that we're talking about. We've just gone live with more than 300 tutorials and are adding 30 to 50 more per week. To check it out, go to bsuper.ai. That's bsuper.ai. Can't wait to see you guys over there. Welcome back to the AI Breakdown. It has been a minute since an open AI model was where we were focused. For the last couple weeks, everything has been about Llama 3. In fact, as we've discussed on this show, even the small models of Llama 3 that have been released so far come sufficiently close to GPT-4 level performance that it's made people think very differently about the competitive landscape of LLMs.
10:58We're getting serious questions around whether, if open source keeps being this close to the state of the art, does it fundamentally change the economics of advanced models? Will people just all opt to build on top of Llama 6 instead of paying a premium for GPT-6? Anyways, the point is that Zuckerberg has been once again dominating the conversation, but apparently OpenAI was sick of that. Because for the last day or so, everyone has been talking about a model called GPT-2 chatbot, which is rumored to secretly be GPT-4.5 or even GPT-5 out in the wild in advance of an official launch. So where was this model found?
11:36Well, it was on the LIMSYS chatbot arena. This is an LLM benchmarking site, and it appeared as one of the model options which people on the site could test. As Dan Shipper from Every put it, LIMSYS.org enables users to chat with various LLMs and rate their output according to different benchmarks without needing to log in. One of the models recently available is GPT-2 chatbot. There is no information to be found on that particular model name anywhere on the site or elsewhere. The ratings results generated by LIMSYS benchmarks are available via their API for all models except for this one. The model name simply appears to be a cover for something else entirely.
12:09Yesterday afternoon, Professor Ethan Malik wrote, There is a mysterious new model called GPT-4 chatbot accessible from a major LLM benchmarking site. No one knows who made it or what it is, but I have been playing with it a little, and it appears to be in the same rough ability level as GPT-4. A mysterious GPT-4 class model? Neat! One additional thing that Ethan noted was that anonymous testing is a service that LIMSYS.org offers. This came out after Simon Willison responded to Ethan and said, Feels to me like a bit of a reputation risk to LIMSYS if this is indeed a stealth model launch. There's supposed to be a neutral benchmarking tool, which is not a great look if they're working behind the scenes with model vendors in an opaque manner like this.
12:48LIMSYS.org responded to Simon saying, Just to clarify, following our policy, we've partnered with several model developers to bring their new models to our platform for community preview testing. These models are strictly for testing and won't be listed on the leaderboard until they go public. In other words, this is a thing that they do. Now, of course, it wasn't just the fact that this had a GPT name, but the fact that it appeared really good that got people talking. EverArt founder Pietro Schirrano wrote, A mysterious new model called GPT-2 chatbot has appeared on LIMSYS and it's really good.
13:16Not only does it show incredible reasoning, but it also gets notoriously challenging AI questions right with a much more impressive tone. The example that he shared was a question, what weighs more, a kilo of feathers or a kilo of lead? This is a historically difficult problem for LLMs, but GPT2 chatbot says, a kilogram of feathers and a kilogram of lead weigh the same. They both weigh one kilogram. The key difference between the two is in their volume and density. A kilogram of feathers will take up much more space and appear much larger, while the kilogram of lead will be compact and small. This often leads to the classic trick question, but in terms of weight, a kilo is a kilo.
13:50It also apparently is good at math. Andrew Gao writes, A GBT2 chatbot just solved an international math Olympiad problem in one shot. This, in my opinion, is insanely hard. Only the four best math students in the U.S. get to compete. CodeGen founding engineer Chase writes, Can confirm GBT2 chatbot is definitely better at complex code manipulation tasks than Claude Opus or the latest GBT4. did better on all the coding prompts we used to test new models. The vibes are deaf there. The model is also apparently good at code art. Phil on Twitter writes, GPT2 chatbot is insane at ASCII art, miles ahead of any other model.
14:26With Llama370B, when asked to create an ASCII art unicorn, the output is very janky. Whereas with GPT2 chatbot, it completely nails it. CoolZippity on Twitter also asked about art, saying, I asked GPT2 chatbot to generate a simple program for turning doodles into art. Every other model I've tested fails at this. They get a doodle function, but the button does nothing of artistic value. The version that he shares does a much better job. So what are the theories for what this model could actually be? Brian Romley writes, I've been testing GPT-2 chatbot for a few days. Today it seems to have gotten much more attention.
15:01It surpassed all of our ChatGPT-4 benchmarks. Hypothesis? A few of us concluded it is a form of pre-lobotomized ChatGPT-4 or heavily trained on it. Runway CEO Siki Chan writes, My best guess? GPT-4 knowledge plus Q star search reasoning equals GPT-2. General knowledge seems near identical to GPT-4 with much better reasoning and planning capabilities. More expensive inference from Tree of Thought search would explain relatively slow inference and low rate limit. If true, this is a much bigger deal than it seems. GPT-2 is likely to feel pretty similar for most general knowledge queries, but will outperform on reasoning.
15:35So I don't think it's an accident that this isn't named GPT-4.5 or GPT-5. It is neither. It's a test bed for Q star, or whatever you want to call tree of thought plus PRM these days. Q star, you'll remember, was a rumored open AI reasoning model that got a lot of attention last year. Continuing, Siki writes, The next GPT-5 will likely continue to ride on the scaling hypothesis, plus a reasoning boost from this. Siki also writes, And no, it isn't GPT-2 fine-tuned. You are all out of your mind. The knowledge cut off alone would make that make zero sense. What he's referring to is that another theory is summed up here by Albs, who writes, My guess is this mysterious GPT-2 chatbot is literally OpenAI's GPT-2 from 2019 fine-tuned with modern assistant datasets, in which case that means their original pre-training is still amazing and better than everyone else's four years later.
16:20Admittedly, most people did not agree. Julian Chamon, the CTO of Hugging Face, writes, My personal guess on GPT-2 chatbot, given Omar Sansevario, the chief llama officer at Hugging Face, has been off for the last 10 days, at this stage strongly suspect it's a side project if he's gone viral. Now, this was a little tongue-in-cheek, of course, but just speaks to how little information we actually have. What about what GPT-2 said about its origins? Andrew Gao again writes, It told me and others that it was made by OpenAI. This is a weak signal, though, because of data contamination. A lot of models are trained on OpenAI chats and thus think they were made by OpenAI.
16:54When he polled to ask, What do you think? GPT-4.5, GPT-5, Grok2, or another AI company? 58.9 % of 2 ,100 voters said GPT-4.5. And while that reflects the quality that people are seeing, some have pointed out that if this was the jump between 4 and 4.5, they wouldn't be real happy about that either. Matt Schumer from HyperRite AI says, GPT-2 chatbot is good, really good. But if this is GPT-4.5, I'm disappointed. Flowers from the Future, the frequent OpenAI leaker, writes, GPT-2 isn't better than current GPT-4 Turbo, so it's definitely not 4.5 or 5, and not even 4. Either this is a new GPT-4 Lite model, or it really is a new GPT-2 model with a totally new kind of training or processing.
17:36The implications of the latter would be absolutely crazy. Being no help and adding more mystery to the whole thing was Sam Altman, who tweeted, I do have a soft spot for GPT-2. Funny enough, he had initially written it as GPT-2, but then about eight seconds later edited it to, I do have a soft spot for GPT-2, no dash. Which of course sent a whole group of people speculating on what that might mean. Ethan Malek again pointed out the frustration of the crypticness of the industry, saying, OpenAI may be one of the most important technology companies in the world today, but they really like to communicate through hints and oracle whispers.
18:10What is GPT-2? At this rate, we will only know that GPT-5 is being launched from an I Love Bees-esque alternate reality game. GPT-6 will be known by the shapes made by the wheeling of starlings over Palo Alto, as well as certain signs in the heavens, and the first letters of every third tweet by Rune. Smokeaway writes, GPT-2 is not the AGI you're looking for. And I think ultimately that's where we're going to land on this. This mystery will likely at some point be solved, or it won't and it'll just stay mysterious. But the reason that there's so much attention on it is that for as much as we were talking about last week around how a close to GPT-4 class open source model could change the game, people are still obsessed with the frontier.
18:50They are still obsessed with the true state of the art. Right now, it doesn't seem like that's going to change anytime soon. So for the moment, we're just going to have to be content with this mystery. Sure, it's speculative, but it's not the least fun I've ever had in the AI space. Anyways, friends, that is going to do it for today's AI Breakdown. Until next time, peace.
From the publisher
Explore the buzz surrounding a mysterious AI model dubbed "GPT-2 Chatbot" recently spotted on a benchmarking site. Speculation is rife that it might be an unreleased version of OpenAI's technology, possibly GPT-4.5 or even GPT-5.
**
Join NLW's May Cohort on Superintelligent.
Use code nlwmay for 25% off your first month and to join the special learning group. https://besuper.ai/
**
Consensus 2024 is happening May 29-31 in Austin, Texas. This year marks the tenth annual Consensus, making it the largest and longest-running event dedicated to all sides of crypto, blockchain and Web3. Use code AIBREAKDOWN to get 15% off your pass at https://go.coindesk.com/43SWugo
**
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI.
Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe
Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown
Join the community: bit.ly/aibreakdown
Learn more: http://breakdown.network/
