OpenAI's Q* Reasoning AI is Now Code-Named "Strawberry"

16 Jul 2024 · 15 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The AI Daily Brief: Episode Summary

Podcast Title The AI Daily Brief (Formerly The AI Breakdown)

Episode Title OpenAI's Q* Reasoning AI is Now Code-Named "Strawberry"

Episode Description This episode discusses OpenAI's latest breakthrough with its reasoning AI, referred to as "Strawberry." It delves into the features and capabilities of Strawberry, its potential impact on the AI industry, and broader implications for future AI applications.

Key Highlights

Introduction

  • The episode provides a brief overview of significant events in the AI realm.
  • OpenAI is developing a new reasoning AI, initially known as Q-star, now dubbed "Strawberry."

Major Announcements

  1. Meta's Lama 3 Model Release
  2. Scheduled for July 23rd.
  3. Features:
  4. 405 billion parameters.
  5. Multimodal capabilities (text and image processing).
  6. Speculation on performance vs. GPT-4 and Claude 3.
  1. Google's Gemini Features
  2. Five announcements expected between July 15th and 18th.
  3. Anticipated features include:
  4. Custom GPTs (Gems).
  5. Memory or personalized responses.
  6. Integration with Google Search and photos.
  1. Amazon's AI Shopping Assistant, Rufus
  2. Now available to all US customers.
  3. Aims to enhance shopping experiences by answering queries and providing recommendations.
  4. Early beta results show high interaction and usage.

OpenAI's Reasoning AI

Strawberry

  • Background on Q-star
  • Originated during the tumultuous leadership changes at OpenAI.
  • Early reports suggested breakthrough capabilities in basic math problem-solving.
  • Internal concerns about safeguards in deploying advanced AI.
  • Evolution to Strawberry
  • Strawberry is a more refined AI model aimed at deep research and planning capabilities.
  • It represents a shift towards achieving more advanced reasoning and problem-solving tasks.

Technical Insights

  • The model focuses on:
  • Self-supervised learning, similar to AlphaGo's training methods.
  • Step-by-step reasoning abilities that allow for complex problem-solving.
  • Potential to autonomously browse the internet for research purposes.

Implications for AI Research

  • Strawberry is part of a broader strategy towards developing agentic AI.
  • OpenAI is considering its application in software and machine learning engineering tasks.
  • The company has created internal classifications for levels of AI, indicating their path towards more advanced general intelligence (AGI).

Conclusion

  • The episode concludes with a recognition of the evolving nature of AI technologies, particularly with OpenAI's developments.
  • Anticipation builds as more details about Strawberry and its capabilities emerge.

Call to Action

  • Listeners are encouraged to share their thoughts on new AI developments and engage in discussions via Discord and other platforms.
  • Promotions for Venice and Superintelligent were included to enhance listener engagement with AI tools.

Additional Resources

  • Venice Pro Discount: 20% off for podcast listeners.
  • Superintelligent Tutorials: 50% off the first month for podcast subscribers.
  • Subscribe: Links provided for newsletter, Discord, and podcast platforms.

Key Takeaways

  • OpenAI's transition from Q-star to Strawberry marks an important step in AI reasoning capabilities.
  • Companies like Meta and Google are also making significant advancements in AI technologies.
  • The ongoing discussions and developments reflect a rapidly evolving landscape in artificial intelligence research and applications.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00OpenAI's reasoning AIQ star has become strawberry, and Meta seems set to release its biggest Lama3 model yet next week. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. To join the conversation, follow the Discord link in our show notes.

0:23Welcome back to the AI Daily Brief Headlines Edition, all the AI headlines you need in around five minutes. Today is a very product news-centric edition of the headlines, with the kickoff story being that Meta is finally releasing its largest Lama 3 model next week on July 23rd. This is according to a Meta employee, as reported by the information. Now, this model has been announced. This is the 405 billion parameter Lama 3 model. And this one, in addition to being larger than the previous versions we've gotten, will be multimodal. It will be able to understand and generate both images and text.

0:59When Lama 3 was released back in April, it was their 8 billion and 70 billion perimeter models, which quickly became very commonly used among AI developers. Back in April when Lama 370B was released, Professor Ethan Mollick speculated that the$400 billion perimeter plus model would reach GPT-4 level. Of course, since then we've gotten GPT-4-0 and Claude 3.5 Sonnet, and it will be a big question just how far off the state-of-the-art this newest open-source release really is. Not content to let Meta have all the fun, Google Gemini also has some upcoming features. This is from a blog post on testingcatalog.com.

1:33The post is based on the fact that Google has scheduled five Gemini announcements for July 15th and July 18th, and then goes through to rank what they're most likely about. The big contender that people seem to be interested in is Gems. Effectively, this is a version of custom GPTs, which people have been waiting for for some time now. Other speculations include memory or personalized responses, scheduled prompts, which could be an interesting integration with Google search capabilities. For example, allowing people to ask Google to send them a curated set of daily news every morning. There's evidence for voice recording and Google photo integration.

2:05And Testing Catalog also found a hidden button that suggests that we might get a prompt enhancer. Given that we've seen Claude push really far into the let our system figure out the right prompts based on your prompts, This is one that wouldn't be too surprising, even if it would be incredibly useful. Some of the other speculations are around a Chrome extension, a real-time response toggle, and an updated Imogen model. At the time of recording, we don't have any more information, but this is certainly something I'll be watching for this week. Third today, Amazon's AI shopping assistant Rufus is now available to all US customers in the Amazon Shopping app.

2:39Amazon writes, Rufus is designed to help customers save time and make more informed purchasing decisions by answering questions on a variety of shopping needs and products right in the Amazon Shopping app. We're pleased to announce they write that Rufus is now available to all US customers in the Amazon Shopping app. As part of the announcement, Amazon also shared some of what they've learned during the beta test. They say that customers have already asked Rufus tens of millions of questions, and so far they're using it for things like understanding product details and hearing what other customers say.

3:06When Rufus gives an answer, it appears that it also suggests another set of additional questions, which apparently customers are also clicking on as well. And then the other things that people are using it for are pretty much exactly what you'd expect. Getting contextual product recommendations, the example they give being a pool umbrella specifically for Florida. People are using it to compare options, e.g. what's the difference between gas and wood-fired pizza ovens. People are using it to get product updates, access current and past orders, and even answer questions that are, quote, not obviously related to shopping.

3:36Amazon writes, because Rufus can answer a wide range of questions, it can help customers at any stage of their shopping journey. A customer interested in cookware may first ask, what do I need to make a souffle? Preparing for special occasions is also popular, with shoppers asking questions like, what do I need for a summer party? So far, I have not used Rufus, but it feels to me like one of those applications of AI that either will become completely default, just the totally normal way that we interact with shopping, or will be quietly removed from this application in about a year. Given that this has been live with testers, and that Amazon is choosing to put it in their main shopping app in their biggest market in the US, it seems like they like the results they've had so far.

4:14If you have had a chance to use it, use either the comments on Spotify or on YouTube to share how Rufus has been for you. For now though, that is going to do it for our Headlines Edition. Next up, the main episode. Today's episode is brought to you by Superintelligent, the platform for fun, fast AI learning. Super has a ton of new things going on. We recently announced our partnership with Spotify through which users of that app can now access Superintelligent content directly from their mobile apps. We've also just launched the AI learning feed. In addition to seeing the tutorials that we're dropping, there are polls, news items with related lessons, and a chance for people to show off the projects and use cases that are making AI come alive for them.

4:53We've also just kicked off the Super Summer Challenge, where each week we'll share a new challenge that you can use to discover new AI tools and use cases. Go to bsuper.ai and use code SUPERFUN for 50 % off your first two months. That's bsuper.ai. Today's episode is brought to you by Venice. The leading AI companies store your entire conversation history and attach it to your identity forever. Every question you ask, every answer you receive, every image you generate, every thought you share with the machine, it's all being spied on. If you trust all the companies, hackers, and NSA board members that will ever have access to your AI conversations, then rejoice, for you are well served.

5:30For the rest of us, Venice is an alternative. Venice is a powerful AI app for text, image, and cogeneration that respects you as a sovereign individual and believes privacy and free speech are not only human rights, but are necessary for civilizational advancement. Private, permissionless, and uncensored. You can try it for free without an account at venice.ai. Welcome back to the AI Daily Brief. At the very end of last week, we got news that OpenAI was working on a new, more advanced type of AI that they have codenamed Strawberry. And in fact, this is not the first time we've heard about this project.

6:03However, it is the first time that it's had this name. So what we're going to do today is give not only this new report about what OpenAI is working on, but go back a little bit to the history of this particular project. And for that, we actually have to go back to the days and weeks that followed, the ouster and then rehiring of CEO Sam Altman last November. About a week after Altman was reinstated, the information published a piece called OpenAI Made an AI Breakthrough Before Altman Firing, Stoking Excitement and Concern. You might remember that during that whole time, as everyone was trying to figure out just why Altman had been fired, probably the most popular working theory was that they had made some big technical advance and that there was internal disagreement around whether they should be pushing it forward.

6:42This was, of course, despite the fact that the board was explicit about the idea that that wasn't the case. However, that didn't stop this report from getting tons of traction. Wrote the information on November 22nd of last year, one day before he was fired by OpenAI's board last week, Sam Altman alluded to a recent technical advance the company had made that allowed it to push the veil of ignorance back and the frontier of discovery forward. The cryptic remarks at the APEX CEO summit went largely unnoticed as the company descended into turmoil. But some OpenAI employees believe Altman's comments referred to an innovation by the company's researchers earlier this year that would allow them to develop far more powerful AI models.

7:16The technical breakthrough spearheaded by OpenAI chief scientist Ilya Sutskever raised concerns among some staff that the company didn't have proper safeguards in place to commercialize such advanced AI models. The information we got was that the model was called Q-star. The big thing that it was able to do that previous models hadn't was that it could solve basic math problems. The information said that in the months following the breakthrough, Ilya himself appeared to have reservations. Another data point from that article, Ilya's breakthrough allowed OpenAI to overcome limitations on obtaining enough high-quality data to train new models, according to the person with knowledge.

7:47The research involved using computer-generated rather than real-world data. Reuters followed up and found their own sources confirming the story. They added the detail that, quote, Though only performing math on the level of grade school students, acing such tests made researchers very optimistic about QSTAR's future success. Reuters also dug up a letter that was sent to the board from a number of staff researchers, warning, it seems, about the discovery. Wrote Reuters, Unlike a calculator that can solve a limited number of operations, advanced general intelligence can generalize, learn, and comprehend.

8:17In their letter to the board, researchers flagged AI's prowess and potential danger, although Reuters' source couldn't confirm exactly that it was QSTAR that they were worried about. Separately, however, The Verge reported that the board never received a letter about QSTAR, and that, quote, the company's research progress didn't play a role in Altman's sudden firing. Of course, lots of people wanted to know more. One of the most viewed discussions on the Open AI forums last November was, what is QSTAR, and when will we learn more? No one really had information on that thread. Many people were talking about it in the context of what it might have meant for the firing.

8:49But then there were also a lot of responses represented by this one from Quirtle, which said, As someone who's done a fair amount of ML slash AI research, I can tell you that it is very, very easy to think you've discovered a breakthrough. There's a great deal of cognitive bias in AI, and you have to falsify very aggressively. I am deeply skeptical. It's also worth noting in the news today that we found out that the $86 billion share sale is back on. I'm sure this quote-unquote breakthrough will get investors quite interested. So obviously they are calling into question the veracity of the claims and saying that perhaps it was being overstated for the sake of an investment.

9:20In December, Timothy B. Lee wrote a post on understandingai.org called The Real Research Behind the Wild Rumors About OpenAI's QSTAR Project. The piece departs from just trying to suss out the details of this supposed QSTAR breakthrough, and instead goes through OpenAI's two other published papers about its effort to solve grade school math problems, as well as some other research from outside of OpenAI on this similar area. One thing he pointed to was a tweet from Chief AI Scientist at Meta, Jan LeCun, who wrote, Please ignore the deluge of complete nonsense about QSTAR. One of the main challenges to improve LLM reliability is to replace autoregressive token prediction with planning.

9:54Pretty much every top lab there, DeepMind OpenAI, etc. is working on that, and some have already published ideas and results. It is likely that QSTAR is OpenAI's attempt at planning. Now, earlier this year, Nimrod Kramer over at Daily.dev published a piece called OpenAI Q, Everything You Need to Know in One Place. He adds to the discussion the point that in addition to solving basic math, QSTAR, quote, showcases reasoning abilities beyond current AI models. From what we've heard, he writes, Project Q-Star can work out basic math problems and think symbolically better than other AI systems out there, understand ideas and make smart guesses about them, move past just recognizing patterns to actually think through problems step by step.

10:32He speculates a little bit about how it might work. He points to step-by-step reasoning, where he says instead of just spitting out answers, Project Q-Star could explain how it got there by breaking the problem into smaller, easier parts, figuring out each part one by one, making sure each part helps solve the big problem. He also contended that, quote, Project QSTAR probably uses self-supervised learning. It's a bit like how the game AlphaGo gets better at playing against itself. The AI practices by solving problems against older versions of itself. This provides a way for the AI to learn and get better without needing people to check its work.

11:00Just like AlphaGo, the AI teaches itself removing the need for outside help. Still, mostly, after that initial burst of interest, we haven't gotten much information. Six months ago on the OpenAI Reddit, poster EchoStorm wrote, Just wondering what happened to QSTAR. I read that it was able to solve mathematical problems faster and better than humans ever could, as well as bypass any encryption and improve itself. If that's true, why is nobody talking about it? Was it false news? If so, why was the leak in the board's reaction so believable? Personally, it doesn't seem to me to be a good publicity stunt for a successful company like OpenAI to do this unless something about QSTAR is true.

11:32And that gets us to last week, when we had two big stories that followed along these lines. The first was that OpenAI had internally shared definitions for five levels of AGI. or at least five levels of AI on the path to AGI. The levels were one, chatbots, AI with conversational language. That's where we are now. Second, reasoners, human level problem solving, something that OpenAI argued that they were close to in this internal meeting. Three, agents, systems that can take actions. Four, innovators, AI that can aid in invention. Five, organizations, AI that can do the work of an organization.

12:05Now, if you go check out the YouTube comments on any of my recent videos about this, there is tons of debate around those specific definitions. But the relevant point for us today is that these came out early last week. However, separately but clearly relatedly, we got this Reuters report, OpenAI working on a new reasoning technology under code name Strawberry. This came from internal sources as well as internal documentation. The document was seen by Reuters in May but not reported until now. Reuters also said they couldn't ascertain the precise date of the document. The document, quote, details a plan for how OpenAI intends to use Strawberry to perform research.

12:38Reuters source also added that how Strawberry works is a tightly kept secret even within OpenAI. Basically, this document describes a project that would use the Strawberry model with the aim of allowing the AI to plan ahead enough to navigate the internet autonomously to perform what OpenAI calls deep research. According to this report, Strawberry is the new name for Q-star. According to Bloomberg, last Tuesday at an all-hands meeting, OpenAI, quote, showed a demo of a research project that it claimed had new human-like reasoning skills. This was the same meeting, I believe, where they introduced that five-level classification system.

13:11While the information remained sparse, there were a few other things we got from this report. Reuters writes, Strawberry includes a specialized way of what is known as post-training OpenAI's generative AI models, or adapting the base models to hone their performance in specific ways after they have already been trained on reams of generalized data. Strawberry has similarities to a method developed at Stanford in 2022 called self-taught reasoner, or STAR. STAR enables AI models to bootstrap themselves into higher intelligence levels via iteratively crafting their own training data, and in theory could be used to get language models to transcend human-level intelligence.

13:42Continuing, Reuters writes, among the capabilities OpenAI is aiming strawberry at is performing long-horizon tasks, referring to complex tasks that require a model to plan ahead and perform a series of actions over an extended period of time. OpenAI specifically wants its models to use these capabilities to conduct research by browsing the web autonomously with the assistance of a CUA, or computer-using agent, that can take actions based on its findings. OpenAI also plans to test its capabilities on doing the work of software and machine learning engineers. So basically what we've got here is an update that confirms that QSTAR has not gone away, it's evolved into whatever the strawberry is, that two, the context that they're thinking about deploying it in or at least researching it in is this deep research context, three, that it's clearly a part of their plans to get to agentic AI, and four, that it's close enough that they're talking about it widely within the company, even though you would have to think that they would assume, or at least not be surprised that some amount of this information would get out.

14:35So far, there is not that much information out there beyond what I've just shared with you, and there's not even all that much chatter. People are very clearly interested, but without more details, we're just going to have to wait and see what evolves. However, it seems likely that OpenAI's comparative quietness in this period might be coming to an end. For now though, that is going to do it for today's AI Daily Brief. Until next time, peace.

15:01Thank you.

From the publisher

Discover OpenAI’s latest breakthrough with the newly announced reasoning AI, code-named “Strawberry.” This episode examines the features and capabilities of “Strawberry,” its potential impact on the AI industry, and what this means for the future of artificial intelligence. Explore this exciting development and its implications for AI research and applications.
Concerned about being spied on? Tired of censored responses? AI Daily Brief listeners receive a 20% discount on Venice Pro. Visit ⁠https://venice.ai/nlw⁠ and enter the discount code NLWDAILYBRIEF.

Learn how to use AI with the world's biggest library of fun and useful tutorials: https://besuper.ai/ Use code 'podcast' for 50% off your first month.

The AI Daily Brief helps you understand the most important news and discussions in AI.

Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614

Subscribe to the newsletter: https://aidailybrief.beehiiv.com/

Join our Discord: https://bit.ly/aibreakdown

More from The AI Daily Brief: Artificial Intelligence News and Analysis

All 1,099 episodes
OpenAI's Q* Reasoning AI is Now Code-Named "Strawberry" The AI Daily Brief: Artificial Intelligence News and Analysis · 15 min
Listen in VO