Bard vs. Bing vs. Claude vs. ChatGPT: The Right LLM For Every Task

16 Jul 2023 · 14 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: The AI Daily Brief - Episode: Bard vs. Bing vs. Claude vs. ChatGPT

Overview

In this episode of *The AI Daily Brief*, host NLW discusses the growing competition among Large Language Models (LLMs) and evaluates their suitability for various tasks. The episode is sparked by a viral tweet that categorizes different LLMs based on their strengths and ideal use cases.

Key Models Discussed

  • Anthropic's Claude 2 (CLAWD2)
  • Google Bard
  • OpenAI's ChatGPT (GPT-4)
  • Microsoft's Bing

Key Takeaways

  1. LLM Performance Comparison
  2. Context Window: The size of the context window (amount of data an LLM can process at once) is crucial.
  3. GPT-4: Typically features 4K and 8K context windows, with a recent upgrade to 32K for specific users.
  4. Claude 2: Now boasts a 100K context window, allowing it to process immense amounts of data (about 75,000 words).
  1. Use Cases for Each LLM
  2. Long Context Tasks:
  3. Best Model: Claude 2 (CLAWD2)
  4. *Reason*: Its 100K context window allows for extensive text analysis and synthesis.
  • Internet-required Tasks:
  • Best Model: Google Bard
  • *Reason*: Bard is designed to integrate with the internet and has improved capabilities in several languages and utility features.
  • Hard Reasoning Tasks:
  • Best Model: GPT-4
  • *Reason*: It still outperforms its competitors in reasoning capabilities, despite Claude 2 competing closely in specific areas.
  • Coding Tasks:
  • Best Model: Code Interpreter (part of ChatGPT)
  • *Reason*: This feature elevates ChatGPT's functionality, allowing it to execute and inspect code actively.
  1. Innovations and Features
  2. Claude 2 Innovations: While it has improved reasoning capabilities, it still suffers from hallucinations (generating incorrect responses).
  3. Bard's Multimodal Capabilities: Bard can process images, making it effective for tasks requiring visual input.
  4. ChatGPT's Code Interpreter: This feature fundamentally upgrades GPT-4 by allowing it to run code, enhancing its interactive abilities.
  1. Alternative Approaches and Personal AIs
  2. Inflection's Pi: A personal AI focused on interpersonal interactions, reflecting user sentiments, and engaging in meaningful dialogues.
  3. Quiver: A customizable second brain allowing users to interact with personalized datasets, showcasing a shift towards personal LLMs.
  1. Market Landscape
  2. The competitive landscape is rapidly changing, with open-source models like Meta's LLaMA potentially altering the market dynamics. The open model approach, as highlighted by commentators, poses a significant challenge to proprietary systems.

Conclusion NLW emphasizes that each LLM has unique strengths tailored to specific tasks, suggesting that users should experiment with them to find the best fit for their needs. The episode concludes with a call for listener feedback and engagement.

Further Information

  • Join the discussion or keep updated through [The AI Breakdown Newsletter](https://theaibreakdown.beehiiv.com/subscribe) or [YouTube Channel](https://www.youtube.com/@TheAIBreakdown).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Today on the AI Breakdown, we're looking at the state of LLM competition and asking which models are right for different tasks. The AI Breakdown is a daily podcast and video about the most important news and discussions in AI. Go to breakdown.network for more information about our newsletter, Discord, and YouTube channel. One of the big announcements this week was that Anthropic was releasing its latest model called CLAWD2. Now, in some ways, CLAWD2 was just catching up to GPT-4. They had very similar results on things like reasoning exams, the GREs. CLAWD2's coding was much improved, bringing it in line with GPT-4.

0:36But Cloud 2 also offered some very different capabilities, particularly the cost and the context window were something that made people really take notice. Google Bard also got a slew of updates, many of which served to improve its functionality in very clear day-to-day ways. So with all of that, it got me thinking about whether there is at this point a single dominant LLM, or alternatively, whether we're at a point where there are different use cases that make sense for different LLMs. It turns out I was not the only person to have this thought. Yesterday, Jan Peleg tweeted, Which model should you use?

1:09The AI Wars TLDR. Long context tasks, Cloud 2. Internet required tasks, use Bard. Hard reasoning tasks, use GPT-4. Anything with code, Code Interpreter. Long SA plus internet, use Bing. And all are crazy good at this point. It is much, much closer. If you didn't try them lately, you should. You would probably be surprised by how much Bard and Cloud improve, night and day. So what we're going to do today is build off of this tweet and ask, what the right LLM for any given use case is. And let's start where he started with long context tasks. Context window refers to how many tokens or how much data can be fed into an LLM in one fell swoop.

1:45The longer the context window, the more context an LLM has in trying to help gauge with a document or some other material. The average person has mostly been interacting with 4K and 8K context windows in GPT 3.5 and GPT 4. And earlier this year, people started to get really excited about the move to a 32K context window for GPT-4. Certain API users had access to that longer window, and it greatly expanded the capabilities of the model, allowing it to process four to eight times as much information at once. As DeepLeaps.com put it at the beginning of May, one of the primary use cases for the GPT-4 32K model is the development of sophisticated Q &A chatbots for businesses.

2:23The expanded context window eliminates the need for complex embeddings and databases, enabling businesses to fit their entire dataset into the 32K prompt and use the API directly. The streamlined process could revolutionize chatbot functionality, making them more efficient and versatile across industries. And yet, even as people were waiting for that 32k context window, Anthropic swooped in and blew that out of the water with a 100k context window for their Claude model. On May 11th, Anthropic announced, we've expanded Claude's context window from 9k to 100k tokens, corresponding to around 75 ,000 words.

2:57This means businesses can now submit hundreds of pages of material for Claude to digest and analyze, and conversations with Claude can go on for hours or even days. Now, as examples, they point to the fact that The Great Gatsby is about that long, but they also say beyond just reading long texts, Claude can help retrieve information from the documents that help your businesses run. You can drop multiple documents or even a book into the prompt and then ask Claude questions that require synthesis of knowledge across many parts of the text. Then again, it was with the Claude model, which was significantly underpowered compared to GPT-4.

3:26However, with the launch of Clawed 2, that has changed, and there's now more parity among the models, meaning that Anthropics Clawed 2 really does serve a hugely valuable purpose because of that longer context window. Bilawal Sidhu writes,

3:56feedback. So of course you see that the common thread here is that these are tasks that require the ability for the model to have the context of that bigger amount of information going in. More generally, Professor Ethan Mollick points out that CLOD2 is just very good at summarizing documents. Now that said, given that we are talking about what different LLMs are useful for and what they're not, there has been a significant sense that even with this new CLOD2 model, there are many hallucinations. Mollick again says on the downside, don't use CLOD for data, it hallucinates answers. Morris Kretz said something similar.

4:27Claude hallucinates a lot, but hey, at least it's friendly. Okay, so next up in Yam's contention, we have internet required tasks, which he suggests using BARD for. So at this point, most of these LLMs are connected to the internet. With ChatGPT, you have browse with Bing, which at this point is rolled out for all users, not just paying users. So why might BARD be a better choice? Well, on the one hand, BARD is just natively in the internet. It's not set up in the same way that ChatGPT is, where the native version of it was trained on data that has a cutoff point. Instead, its whole purpose is to sit on top of the internet in the same way that Google Search does.

5:02But even beyond that, a new set of updates also increase its viability for those use cases. First of all, the new rollouts make it available in Europe and Brazil, not just the US. Second, it's now available in something like 40 languages. Third, they just added a number of new utility features, things like save searches, sharing searches with friends, pinning searches, all of which individually are very small but add up to a higher functionality product. But more than that, with this new update, Bard is officially multimodal. What that means is that an image can now be used to prompt the system.

5:34Kirthana, a researcher at DeepMind, posted an image of a pug with a graduation hat and typed, what is happening in this image? Bard says the image shows a pug dog wearing a graduation cap on a leash. The image is likely a celebration of the dog's graduation from obedience school or a service dog training program. Ethan Malik again says Google Bard is surprisingly good at working with images. It appears to be combining a reverse image search with multimodal capability, i.e. the ability of the AI to see something. Now, importantly, this isn't just for novelty, like asking about a pug in a graduation cap.

6:06Joel Dean writes, wow, Bard just converted a screenshot to code. This is so next level, looking forward to these multimodal capabilities in chat GPT. The prompt that Joel had used was, are you able to convert this screen to Jetpack Compose, and then shared a screenshot from which BARD was able to push out code, although Joel doesn't say how accurate that code was. Now, it's entirely possible that within the next six months, this sort of multimodality is total table stakes. However, as of right now, OpenAI has indicated that they've had to put broader multimodal rollouts on hold because of their lack of access to GPUs.

6:38It's one of the areas where the GPU shortage is showing up most profoundly. So for now, I would say that in addition to just using BARD for internet-required tasks, BARD is also the standout option for multimodal tasks that involve images. Now, YAM's next contention is that for harder reasoning tasks, use GPT-4. And on the one hand, I would say that this is broadly consensus, that people believe by and large that GPT-4 remains ahead of all of its competitors when it comes to reasoning tasks. And on top of that, there's also some reasonable evidence. For example, when CLAWD2 came out, They shared a number of comparisons, and while Claude did overtake ChatGPT in GRE writing and bar exams, the difference wasn't really statistically significant, and in terms of standard GREs, ChatGPT still won verbal, quantitative, and the medical exam.

7:20But I think the even more important part of the discussion right now, as relates to ChatGPT and GPT-4, isn't so much GPT-4 and how ahead it is on reasoning tasks. Instead, what matters about ChatGPT most right now is the newly released code interpreter feature, which many are seeing as effectively GPT 4.5, even though it's not named that. Swix from the Latent Space podcast made this point most loudly. On July 10th, he tweeted, Code Interpreter equals GPT 4.5, or making GPT 4.1000x better with one weird trick. Now, the one weird trick that he's referring to is the fact that Code Interpreter is not so much just a tool that can interpret code or that can look at data when you plug it into the model.

8:04Instead, it represents a fundamental addition to the model itself. In a blog post that they wrote, Swick shared a chart that he called the road to AGI. And what he pointed out is that each of the big leaps for GPT, from GPT-3 to 3.5, from 3.5 to GPT-4, and from GPT-4 to GPT-4 plus interpreter, there was an input of an additional aspect to the training. So with GPT-3, we got pre-training, but with GPT-3.5, we got pre-training and reinforcement learning from human feedback. Then the next additions to GPT 3.5 included plugins and user-defined functions. And then with GPT-4, we added into the mix a mixture of experts.

8:42So all of a sudden, the model had not just pre-training and reinforcement learning from human feedback, but pre-training, a mixture of experts, and reinforcement learning from human feedback. In that point of view, Code Interpreter becomes not just, again, an application that sits on top of GPT-4, but a code sandbox which allows GPT-4 to effectively fill in the gaps in its own model. Moritz Krem expanded upon the same idea. He wrote, people haven't fully grasped the significance of the code interpreter. It's not just another plugin that does data analysis. In my opinion, it's actually GPT-4.5 masked as a plugin.

9:15Let me explain. ChatGPT was already able to produce code, but it wasn't able to run it. The code interpreter can. This small change makes a huge difference. This means that ChatGPT is no longer limited to being a passive assistant, it has now become active. Two, iterative abilities. On top of running code, the code interpreter seems to have built in iterative abilities. It recognizes when it's made a mistake and it corrects it by itself. It's more closely resembling an agent now. Three, different model. It also seems that the code interpreter is actually accessing a completely different model from GPT-4.

9:45Some people have reverse engineered this and are pretty sure. He actually references another tweet from Jan Peleg that says, we highly, 99 % suspect that the model is not the same model as GPT-4. The user interface accesses a completely different endpoint that also has additional parameters. Number four, Moritz points out, is multimodality. GPT-4 has multimodality built into it. This means that it understands not only text, but also visuals and audio. However, this feature has not been activated for ChatGPT yet. With the code interpreter, OpenAI has made a step in the direction of enabling this feature.

10:15Because now there is a way to input anything, data sets, image's audio into ChatGPT, a prerequisite for multimodal functions. While the code interpreter doesn't yet understand an image, it can already take the image and manipulate it. To me, this is a fundamentally upgraded ChatGPT. Calling it the code interpreter and downplaying it as a GPT-4 plugin is not doing it justice. Now the last piece of this puzzle that I wanted to mention is something that's very different. You can kind of tell with all of these different LLMs that I've just mentioned over the course of this video, they're all sort of for professional or at least work-type use cases.

10:47It's research, it's coding, it's development, it's building. However, some people believe that that is not the be-all and end-all of what AIs can do. In many ways, the biggest proponent of this view is, of course, Inflection. Inflection is the company behind Pi, which stands for personal intelligence. When you go to heypi.com, the first window comes up, hey there, great to meet you. I'm Pi, your personal AI. My goal is to be useful, friendly, and fun. Ask me for advice, for answers, or let's talk about whatever's on your mind. When Pi was first introduced, Mustafa Silliman, who was also previously a founder at Google's DeepMind, said, Many people feel that they just want to be heard, or they just want a tool that reflects back what they said to demonstrate they have actually been heard.

11:25And subsequent to that launch, what they've been doing is basically increasing the feature set to make it more interpersonal. About a week ago, Mustafa tweeted, Interestingly,

11:38so far, the community hasn't really seemed to treat it like just a novelty. Last week, Robert Scoble shared a set of conversations saying, check out this chat I had with Pi, my new AI. This is incredible. In the conversation, you can really get a sense of how Pi is designed to be a good listener. And what's really interesting to me, and what Pi seems to do really well, is actually ask questions that move the conversation into a new direction. In other words, when we have a conversation, it's not just one person talking and another person nodding their head and saying, yeah, that's cool. It's two people actively interacting with one another such that each changes the shape of the next thing that's going to be said.

12:14Now, to some extent, reading this still feels like you're reading an AI, but it does do that job of coming back with questions that really do end up pushing things forward. Now, as we wrap up, one thing that I think is worth noting, one company that didn't have a contender in here is, of course, Meta. However, Meta's llama model has been absolutely integral to the explosion of open source alternatives, and it appears that they're on the verge of releasing a new Llama 2 model that will be commercially available. Swix again said, this is the biggest change to the AI competitive landscape. The real threat to open AI isn't open AI but safer, but open AI but open.

12:48Finally, the last LLM that I'll mention is something that is a different use case entirely, which is not an LLM that is open and which an individual taps into the collective database, but instead personal LLMs that interact with the data that a specific person or company has given it access to. There are tons of examples of this. This is a very hot development area right now. But one that's been making some waves recently is called Quiver. Friend of the show, Emmett Ham, writes, an AI-powered second brain is taking over GitHub. Quiver is a customizable second brain that lets you dump in any file, text, audio, video links, and chat with it via LLM.

13:21I have only just started to play with Quiver entering in my notes. I've only just started to play with Quiver entering my note files into it and some other things to see what comes out. But I think that this is a trend you're going to see a lot more of. So friends, we will wrap there. Those are a list of how LLMs differ and what they're good for. Let me know what you think in the comments. And as always, I appreciate you listening or watching. Until next time, peace.

From the publisher

LLM competition ratchets up seemingly every week. At this point, the different design choices that models have made have led to different LLMs being better or worse suited for different tasks. NLW builds off a recent viral tweet about which LLMs are good for what tasks.
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI. 

Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe

Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown

Join the community: bit.ly/aibreakdown

Learn more: http://breakdown.network/

More from The AI Daily Brief: Artificial Intelligence News and Analysis

All 1,099 episodes
Bard vs. Bing vs. Claude vs. ChatGPT: The Right LLM For Every TaskThe AI Daily Brief: Artificial Intelligence News and Analysis · 14 min
Listen in VO