148 | AI Tools Mastery: How To Choose The Right AI Model For Any Task with David Wilson

10 Dec 2024 · 48 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: Leveraging AI - Episode 148

Episode Title

AI Tools Mastery: How To Choose The Right AI Model For Any Task with David Wilson

Host

Isar Meitis

Guest

David Wilson, Founder of Hunch

Episode Overview

In this episode of *Leveraging AI*, host Isar Meitis and guest David Wilson delve into the complexities of selecting the right AI models for various business tasks. With a focus on practical applications and ethical considerations, the discussion aims to demystify AI tools and enhance understanding for business professionals.

---

Key Themes and Discussions

  1. The Evolving Landscape of AI Models
  2. The rapid development of AI technologies since the introduction of ChatGPT.
  3. Cost reductions in using top-tier models (up to 99% decrease in costs).
  4. Improved capabilities across different models, making it crucial to understand which model fits specific tasks.
  1. Choosing the Right AI Model
  2. Factors to Consider:
  3. Intended use case (e.g., writing, data analysis, coding).
  4. Model strengths (contextual understanding, summarization, coding proficiency).
  5. The importance of testing multiple models to find the best fit for specific tasks.
  1. Top AI Models Reviewed
  2. Claude 3.5 Sonnet:
  3. Recommended for general writing and knowledge work.
  4. Fast and capable of handling a range of tasks.
  5. GPT-4:
  6. Newest version focuses on creative writing and is highly competitive.
  7. O1:
  8. Excellent for math and coding but slower than other models.
  9. Gemini Models:
  10. Notable for their large context windows and multimodal capabilities (images, videos).
  11. Llama Models on Grok:
  12. Fast and affordable, offering good performance with cost-effectiveness.
  1. Best Practices for Using AI Models
  2. Start with simple prompts and iterate to refine requests.
  3. Avoid overly complicated prompting initially; allow the model to surprise you with responses.
  4. Use different models in conjunction to leverage their strengths (e.g., using one for brainstorming and another for execution).
  1. Practical Applications of AI Models
  2. Real-world examples of workflows that effectively use multiple AI tools.
  3. Importance of combining the output of different models to enhance decision-making and creativity.
  4. Discussion of how tools like *Hunch* can streamline and optimize AI usage.

---

Key Takeaways

  • Iterative Testing: Continually test various AI models to optimize for specific tasks; each model has unique strengths and weaknesses.
  • Cost Efficiency: Understand the pricing structure of API usage (per token) versus subscription models to make informed decisions.
  • Combining Tools: Leverage the capabilities of multiple models together through platforms like Hunch to maximize productivity.
  • Stay Updated: The AI landscape is continuously evolving; stay informed about new advancements and model releases.

---

Conclusion This episode of *Leveraging AI* offers valuable insights into selecting and utilizing the right AI tools for business needs. By understanding the strengths of each model and adopting a test-and-learn approach, professionals can harness the transformative power of AI to enhance efficiency and drive business growth.

For more insights and to stay connected, visit

  • [Hunch](https://hunch.tools)
  • [Isar Meitis on LinkedIn](https://www.linkedin.com/in/isarmeitis/)
  • Join live sessions every Thursday at 12 PM Eastern.

---

*Thank you for tuning into this episode! If you found the discussion helpful, consider leaving a review and sharing your thoughts.*

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Hello, and welcome to another live episode of Leveraging AI, the podcast that shares practical ethical ways to leverage AI, to improve efficiency, grow your business, and advance your career. This is Isar Maitis, your host, and I am really excited about today's episode from multiple different reasons. One is from a very personal reason. I'm a geek and I like new AI tools and especially cool ones that allows me to do a lot of flexible stuff. And our guest today is David Wilson. He's the CEO of Hunch, who has maybe one of the most unknown and yet one of the coolest and most capable tools that exist out there today as far as stringing and connecting multiple AI tools and other tools together to do basically whatever you want.

0:45It's this really cool playground. But the other reason is that the focus of today's show is actually not going to be the tool itself, but rather what you can learn from it, which I find even more exciting. So one of the things that I get asked all the time when I teach courses, when I do workshops, people just meet me and know me, CEOs of companies I work with and so on, all ask me about, okay, so which is the best large language model? Which one should I use? And the sad answer is that it depends, right? It just depends. What is the use case? What is that you're trying to do? Because each and every one of them has pros and cons.

1:21Some of them have longer context windows. Some of them are better in reading documents. Some of them are better in summarizing stuff. Some of them are better writers, and so on and so forth. And it's very hard to figure it out. And so some people just say, okay, I'm going to commit to one, and I know it's going to be okay, and that's fine. But if you want to get the best of breed and really get the best results across all of them, there are ways for you to explore it for yourself. And that's going to be our topic for today. How can you figure out which of the large language models, open source, closed source, big, small, whatever you want, which ones do better in specific use cases, and how can you figure it out for yourself.

2:05And I assume if you're listening to this podcast, you care about large language models and how to use them for business. And so this is a very, very important topic that a lot of people are struggling with. And that's why I'm very excited about this. Now, David himself is a serial entrepreneur. It's not his first company. So he's been in the tech world and running businesses for a while. So he both understands the benefits, the business benefits of this, but also is somewhat of a geek like me. And so he really enjoys playing with this kind of stuff and testing stuff around. So he's literally the perfect person to share his knowledge with us about this.

2:41So I'm very excited to welcome David Wilson to the show. David, welcome to Leveraging AI.

2:49In the next few years, AI technology will change our world dramatically. Whether you are a business executive trying to catapult your business forward, or just somebody who refuses to be left behind and want to advance your career, this is the show for you. I'm your host, Isar Matis, a serial entrepreneur and an AI enthusiast. You'll hear invaluable practical tips from innovative business leaders, AI practitioners, and some of the brightest AI minds in our world today on how you can leverage AI in ethical ways to advance your career and grow your business.

3:29Thank you very much. Excited to be here. Awesome. David, to kind of like share two cents about yourself, about your product, and then we'll dive right into how to really compare large language models and what are the best ways to do that. Sure. I mean, just very briefly, I found the CEO of Hunch. We're an AI workspace that allows people to connect any AI models, all the best ones together to do really much more than what they're capable of doing with other tools, typical chat tools and stuff. And we can get into reasons why that is the case rather than giving the full spiel up front. And I personally, over the last year, I just ran the numbers a couple of days ago, I have run more than 50 ,000 prompts in the last year with like 50 plus different models.

4:21So excited to share what I've learned here. And what we can do is because Hunch is a way of actually showing different models, We can go and explore and just tell me if there's anything that you'd like me to run or any different types of models that you'd like to try. But what I'm going to do is I'm going to switch over to share my screen. And we can also take a look at a little, you can start off with a little summary of the landscape of the different models if you want to. And just how we've seen kind of models evolve over the last 18 months to two years since ChatGPT came out. Yeah, I will say two things.

5:01First of all, for those of you who are listening to this and not watching this, we will share everything that's on the screen. If you want to watch this, then there's a LinkedIn version. There is a YouTube version of this that you can go and watch. For those of you who are with us live, then obviously you can follow us on the screen on Zoom and on LinkedIn. And if you're not with us live and you want to be with us live, we do this every Thursday at noon, p.m. Eastern, every week with a different amazing expert like David. that is going to share stuff. So you should join us live because then you can ask questions and see everything that we're seeing.

5:34But if you are listening to the podcast and I'm an avid podcast listener myself, we're going to share everything that's on the screen so you can follow along. Awesome. Yeah, thanks. Thanks, Yusuf. So I think that the story of the last year or two of these models since the original ChatGPT is that there's been a huge development in different capabilities for different models. So one of the biggest things that's happened is that the very best models have gotten so much cheaper to use. So the prices come down to sometimes 99 % cheaper to run GPT-4 or Mini today than it was running GPT-4 when it came out about 18 months ago.

6:20The second is that there's now some models that are incredibly fast. and then there's also models that have developed like totally different types of capabilities so we'll go through a few of those different ones but i just want to start off by saying that we can find the best models for all of these different kinds of capabilities but like a general rule of thumb i think it's been this way for about six months uh but claude anthropics claude 3.5 sonnet is probably the first place to go for for a really good a large language model for most tasks. I think it is the best model out there for writing.

7:00The previous sort of Claude III opus, which is more expensive, it's slower, and it's theoretically part of the previous generation or half a generation back, is also still very good for writing if you just want to write things. But in general, for the majority of knowledge work that people want to do, Float 3.5 Sonnet is really good. It's pretty fast. It is highly capable. It can code. It really is extraordinary still how good it is. So I don't know. One thing about that. So in general, I agree. I really like 3.5 Sonnet. There are, ChatGPT just came out with the latest version of 4.0, like literally a few days ago.

7:48that is 4.0 whatever, like they have several different versions of 4.0 since it came out. And it's specifically upgraded on the topics of creative writing. And it took back the top of the leaderboard, like the, what is it called? The LMSIS AI large language model leaderboard for creative writing as well. So while I'm like you, and I'm very much from a personal believer, a fan of 3.5 Sonnet, it's always worth keep on testing out because they come up with these new models all the time. And what was true the day this comes out may not be true the day after this episode comes out. That's 100 % right.

8:32You got to keep on testing. But in general, I agree with you. I think Cloud 3.5 Sonnet is an awesome tool. Yeah. And look, I just like, I'm very skeptical about benchmarks because there's a whole lot of different reasons, I think, to be skeptical about benchmarks. And I think that just the best way to really get a sense for yourself is to try different models and to keep trying them, try them with different prompts. We can talk about prompting at some point, but I think just one kind of overarching tip that I see a lot of people using different models, I work with a lot of people using different models.

9:08And I think that the mistake that people make is trying to, especially ones that have used like AI models quite a bit, is that they'll try and create a sort of an overwrought prompt from the beginning. Like you are a whatever and create this. I think it's actually very easy to start testing models. Just put in the tercest possible request that you can write, see what it gives you and iterate from there because not only does it sometimes or very frequently surprise you with how good the responses are to even very vague or requests with typos and stuff like that but because the models are relatively fast and affordable you can actually just just try it again just you know if it doesn't give you exactly what you want now in what ways is deficient so you can iterate on your problem so that's a kind of like overarching tip, but we can get back to that.

10:06I agree with you. Let's go back to the list, go through it quickly, and then actually how we can test stuff. Yeah, sure. I would say the one model that really exceeds Claude in important ways is O1. And O1 preview is what's out at the moment. The full O1 model might come out in the next week or two, but it's really good at math and hard science problems and coding. It's better than Claude for some coding challenges. It is very slow though. So I've seen people switch over their chat GPT to 01 and then get really frustrated when it takes a really long time to respond. But I think - It's definitely how we became addicted to immediate gratification, right?

10:50When we say really slow, it takes it 30 to 60 seconds to respond. When it takes GPT or Claude, six seconds to respond. So it's not like you're going to wait an hour you're going to wait a minute in the worst case scenario. Look, I do think that there is like E4O is like really clearly crafted by OpenAI to support their chat product and with very interesting trade-offs made there, but like primarily towards latency, whereas O1 clearly is towards capability, reasoning capability. But I think something that is that most people that I talk to haven't yet discovered with O1 is that you can get it to do a lot of work for you.

11:36And instead of prompting it with instructions of how to do things, where Claude and typical LOMs, you want to talk about the step-by-step to get to what you want. O1, if you ask it for exactly what you want, and then almost tell it the appendix to give you it can give you it will do a tremendous amount of work in one shot it might be slow but it can do a huge amount but that is really i think that the model today that is the hardest for people to wrap their heads around what exactly it can do is some of the other models look at this this almost like forms the typical product project management thing of you can have things fast, cheap, or high quality.

12:24And this is like how the models are bifurcating across these different capabilities where high quality or the things that it can do would be Claude and 3.5 Sonnet and O1. And then fast and cheap would be things like the Lama models hosted on Grok, extremely fast and cheap. I think GPT-40 mini is really affordable per token and pretty good capabilities. Really, it's not that far behind GPT-40 for most things, in my view. And then we have this sort of somewhat something of a strange family of models, which is the Gemini models, which have by far the biggest context windows still. They have multimodal inputs, so you can feed it virtually anything, videos, images, PDFs.

13:20And they really are quite good at certain kinds of tasks, like there's some kinds of writing, for example, UX writing, that's kind of niche, I found it to be very good at. And Gemini 1.5 Flash, which is the sort of faster, cheaper version, is very capable. and it's really, it's almost in a class of its own with respect to the capabilities because of just how large the context window is. Another thing that's very good at, in my opinion, best across the models is image interpretation and transcription from images. So it's very good at a bunch of different tasks. Yeah, I'll say two things about what you said.

14:03First of all, great summary. Those of you, you mentioned Lama running on Grok. Those of you who don't know Grok, Grok are a hardware company that creates the fastest inference chips on the planet right now. So it's computer chips that are not planned to train models, which still the best way to do that is GPUs from NVIDIA, but to actually run the models. And they have taken on their platform so you can sign up for a license on the Grok platform so you don't need your own chips. You don't need to buy them. You can just use them on their data servers. and then they have customized Lama 3.2 to run optimized on their thing.

14:41And it's insane. It's worth, there's a free way to test it out. You can just go there. It's nothing like you've ever seen before. Two pages of output just shows up as soon as you hit enter. It's just incredible. Well, we can actually, we can put that to the test right now. So we have a prompt here connected to a bunch of different models, including a model from Grok. and I'm just going to zoom in and read the prompt very quickly. It's create the most surprising. I'm going to edit it a little bit so that this reruns for us. Create the most surprising insights an LLM can come up with. Make it as novel and inventive and yet as plausible as possible.

15:21So draw upon the most surprising and obscure of connections to yield a novel insight. Okay, this is a - Before you hit go, I want to pause you to explain what we're seeing on the screen for people who are not seeing it. Yes. So this looks like a big canvas board, for lack of a better term, that you can zoom in and zoom out and move stuff around. And the prompt lives in this one box. But then there's lines connected to, I don't know how many boxes we have here, like 18, a lot. And each and every one of them is basically connecting to a separate large language model, which now is running this prompt.

15:53Yeah. Let me, sorry about that. I itchy trigger finger. I can, I'll zoom out and then I'll rerun it. So you can see the speed. Let's see exactly how it goes. Yeah, so we have all the different models that we just discussed, as well as a few others connected. So we have eight models here. So we're also including Mistral Large. So Mistral is a French company, a very interesting model. It's also quite different. I wouldn't place it at the top of the leaderboard in any of those categories. But it has very different kinds of guardrails to the others as well. So there's sometimes where everyone's surprised how you can get rejected by Gemini, for example, or one of the other guardrails.

16:36I would say that that's actually something that's gotten a lot better over time over the last 18 months across the model providers, which makes a lot of sense and is fortunate. But Mistral is really interesting. I'm going to change this, and I'm just going to rerun all these models. So you can see they're all running and the ROC immediately fills up the fastest. GPT-40 is behind that. GPT-40 Mini, I'm not sure why that hasn't come through yet. That's usually pretty fast. Sonnet was actually quite fast. And of course, OpenAI's 01 is still thinking, even though all the different models were done, and now it has yielded its answer.

17:23yeah that didn't okay so i want to ask you since you gave us a quick preview of your product this is really fascinating again those of you not watching we put a prompt in one box and we got answers in basically as many as we want to connect and all we did is each and every one of those other boxes has a different language model connected with a line so i have a bunch of questions the first question is how easy it is to connect each and every one of these models is it as easy as just getting my API token and putting it in there? No, so actually it's easier than that. All we have, we, it's all of our API tokens and you can just walk in and use it and just get going.

18:06And like, you don't even need to log in to get started. You can, if you go to tools or app.hunch.tools, you can just start using the product immediately. And then after a certain amount of tokens, you need to log in at least and save your work. But yeah, you just get going. Very cool. Okay, so we discussed a little bit about the tool. We discussed a little bit about the models. How do you go and actually compare them for different scenarios and so on? And I think this kind of gives us an idea. But then how do I actually compare or how do I then combine the best of both worlds if I have, oh, this model is really good at this and this model is really good at that.

18:46And then there's something you already mentioned. There's a cost behind where just to give people an idea once you start dealing with the API, because a lot of people hear the words, the letters API, and they get sick because they don't know what it means. But what it means is it's talking in a server-to-server connection language, which you don't need to know anything about. The biggest difference from an end-user perspective using the chat interface versus using the regular versus using the API is when using the chat interface, you're paying a subscription fee. So you're paying$15 a month,$10 a month,$20 a month,$25 a month, depending on which model you're using.

19:31And they usually don't cap you or they cap you at an amount that is not a problem for most users. When you use the API, you're going to pay per token, meaning every input you put in and every output you get out, you're going to get taxed on. Now, it's very small numbers. To give you a perspective, if you're losing LAMA, which you mentioned before, it's going to be LAMA 3.2 is about 3.5 cents for every million tokens, both in and out. So for every million, those of you who don't know what tokens are, a million tokens is about 750 ,000 words. So you're going to pay three cents for 750 ,000 words if you're using multiple books, many books.

20:15That's a lot of books. Yes, that's on Lama. If you're using the very far end of the scale is Claude 3 Opus. And that's about$70, seven zero for a million tokens. So that's still not a lot, right? it's still paying$70 or something to generate 750 ,000 words for you. That's, again, not expensive, but it's way more expensive than three and a half cents for the same amount of work. And so when you deal with these kind of tools, the backend are running the APIs, there are several considerations. One of them is, what is it effective on? The other is, can I chain it together with the other tools? And the third is, okay, if two tools are the same, which one is cheaper?

21:08And because again, the differences could be significant. So with that, I will leave it to you to walk us through how do you actually do this magic and compare things and combine things and so on. Yeah. Yeah. I think everything you said is right. And ultimately, it's use case dependent. But I think that this is why the kind of a couple of rules of thumb are really helpful. And the first one to me is that the best model for most tasks is Claw 3.5 Sonnet. And you can access that, as you mentioned, on claw.ai. You can also access that in Hunch.tools. Even though we use the API, what we're doing right now is we're making it free for our users because it's really useful.

21:51We don't train on anyone's data. We have a fairly, as we have a strict sort of privacy, security policies and stuff like that. But what's very interesting to us is just how people use these different models. And we're actually using Hunch to like ourselves and doing a lot of work with it ourselves. And we can get into some of that if you want to. But yeah, Claw 3.5 Sonnet, I think, is like a really solid place to start. If you want something cheaper, And the like GPT-40 mini is a really good place to start. And between those two, I think that you're getting a pretty good set. And honestly, it's like you just got to iterate from there.

22:37You can try different models. You can use this template. By the way, everything that you see on the screen, everything you just talked about, it's a template that is available. We can share the URL for anyone who's listening, anyone who's we can put in the show notes if you like. so anyone can try it and iterate on it themselves to learn more. We can go into very specific... Where would you like to take this? Into the specific niches that different models are better at? Yeah, let's do this. And I think what would be interesting is two things. One is if you can show us one or two examples or three of specific niches and where specific tools do really well.

23:14And I think maybe we go extreme. Maybe we go 01 as a slow and yet deep thinking kind of solution versus a grok where something will happen in two seconds where speed matters or llama yeah and then the other thing that i think will be interesting is to see how in your tool you can mix and match how you can actually go from i do this step with this and then i hand it over to this tool to do the other thing i think that will be even more powerful because then i think most people think Think about it as either or, like I'm either going to use Llama or I'm going to use Claude or I'm going to use ChatGPT when the reality is in tools like Hunch, you can combine different steps and different tools and get the best of all worlds.

24:04Yeah. So let me, let's do both of those things. So what we'll do is let's start off with this prompt. Let's go back to this prompt. Anyone here can see that this is initially from Siki Chen, who is the founder of a really interesting company, and he's automated a lot of stuff with AI. But what I'm going to do is I'm going to augment this prompt to say, follow up with a table detailing your, like, all the, like, maybe up to 20 potential insights and why this and an evaluation of each.

24:54okay okay now let me just turn this off quickly what i'm going to do is i'm just going to look here at the current if you see the current claude 3.5 sonnet answer okay it is it's just a paragraph in a little bit the the current 01 answer let me open that up is longer sort of a page maybe a three paragraph plus bullet points plus implications in here yes then we have gpt40 also similar kind of length okay and the reason why is that i'm looking at this is we'll we'll be we'll be clearer in a second but and then let's go to grok okay grok is actually the longest it's just over a page yeah but when i say grok i'm talking about the llama math model.

25:46So now I'm going to just run this again. And we're going to see Grok is going to generate the fastest. So let's actually just open that up. Now, what's interesting is that the answer that it's given us now is shorter. It's about half the length of the original answer, because now it is creating a table as well. And so this is the answer and here's the table. It's a little bit of information within each one. This is a pattern that I would expect to see across the majority of the models. So they're going to give us a shorter answer than the table on the bottom. Yeah, with one exception. Actually, 01 in this particular case is a little bit shorter than it was before.

26:36Yeah. But the table is much more detailed. And it's something that is just, that's something where I don't think that 01, in my experience, is that great at writing, for example, and creating polished products. But it's very good at doing a lot of thinking and evaluating for you. Like search and summarization kind of work. Yeah, that's right. So that's an illustration of what's good and what things are good at. Let me go into a different, maybe something worth mentioning. We have a, we can wrap whatever we think is the best advanced model at a time. We give our own system prompts to that and we call it just the advanced text model.

27:26I'm going to open it up here. It gives a very brief, it's prompted to give briefer, more direct outputs without the preamble of, oh, here's what you asked for and here's what to answer. So it's more like a business. But it actually has like a secret chain of thought beforehand where it has a whole series of thinking steps that we don't show because that's not necessary for people. but it is useful for generating a higher quality answer. So that's something that if you're not using O1, I highly recommend telling a model, general LLM, to think beforehand before answering because those initial thoughts, the way that these autoregressive models work is that every token that it outputs, all of its initial output helps to steer the subsequent output in some way.

28:23So if you get it to think on first principle or from the pieces that you want, it helps to get the answers that you want. And that's actually why chaining together different models is so useful and so powerful. I'm going to click through to a different template. Let me just see if this is the one that I want exactly. Yes, okay. So what this template that I just opened does is it takes a prompt it runs that prompt on GPT-40, Gemini 1.5 Pro, and Claw 3.5 Sonnet. In parallel, it does a bunch of quote-unquote thinking with Claw 3.5 Sonnet, and then it critiques all the answers that it gets from all the different models.

Read the full transcript

29:18And then it pushes all of that thinking, all of the initial models' outputs, and all of that critique into an OpenAI 01 task. And we don't tell the open, we don't tell 01 where those different inputs come from. We just say, here's some, you know, here's some original ideas or thinking on the topic. And now basically you take it from here. And that's a really effective way of getting interesting stuff. So I'm actually going to go back. I'm going to take the original prompt just for the sake of getting something. And let me just go across. I'm going to put that into the instructions. I just want to pause just for one second again for the listeners of this.

30:05What we're looking at when we're saying about this new template is, again, think about a dynamic kind of like canvas with a big flow chart. But in every box in the flow chart, there's also the input and the output. So it's not just what's going to happen, but the actual results from the actual language models show up in the boxes. And so when David is saying, okay, I wrote the prompt shows up in one box, but then there's literally three lines connecting to three different boxes with three different language models that are then connecting to the unified universe where it puts everything together, which then all of them flow to the last box, which is Gemini, which is O1, the last model that is now like taking the inputs from all of them.

30:50So you literally build your own workflows, but it's not just the workflow. It's also the user interface, the input and the output all are happening within this one canvas page. So you don't need to go back and forth with different tools. You can literally work with all the large language models together with really sophisticated workflows all in one place, never living this universe and still enjoying the benefits of all of them together, which I personally find really powerful. Yeah. And it's what's really great here is that if you have, there's some tasks that you have that are really important and that, you know, you want the best possible outcome for.

31:35And so being able to feed all the different language models at once is super useful. I'll give an example that I do all the time. And I'll actually give two examples and then we can dive into your process. But one example is when I think about content ideas or new things I want to add to the course, I want to brainstorm. So I use a tool. So the tool before I knew Hunch, the tool that I was using all the time and I still use it is called Chat Hub. It's a Chrome extension that is a paid Chrome extension that allows you to open several different language models all at once and run the prompt once and chat with all of them at the same time.

32:19And the benefit is when you're doing brainstorming, you now have an advisory board, right? Now I have ChatGPT and Claude and Jim, and I get ideas from all of them, which is fantastic, right? You just get more ideas. And some of them are going to be the same and some of them are not going to be the same. But the cool thing here that ChatHub does not allow you to do is to then pick up the best ideas based on whatever criteria you define or summarize it where you don't have all the duplicates. So you don't have to read all of them. You can now create a version that's going to be a unified version that's going to take the best of all worlds and do more work for you.

32:56So some of the work I have to do with ChatHub, I don't have to do here because Hunch itself can do it for me, which again, just saves you more work, which is the whole point in using AI tools. Yeah, that's 100%. Actually, brainstorming is one of the best reasons to use multiple models because each of the models come up with different ideas. And it's great to get quick ways of doing that. So we, not to make that at this point specifically about Hunch, but what we can do is that template that we just created or that we just opened with all the best models together, that we can access as a single block if we want to and run the sort of prompts easily on another canvas.

33:41But we have something similar, which is for brainstorming with all the different models as well. And you can basically run that in the same way. And it is super powerful. but whether you're doing it in a tool like this or whether you're doing it across multiple other tools or in a chrome extension like chat album which sounds very useful it's just using multiple models and seeing the outputs is a really good way of getting a sense just like you get a sense from talking to people what their responses to things are going to be you get a sense from working in a team who's going to respond who's going to be the best person to do different things You get that sense very quickly from working with different models.

34:25And the more you do it, the more attuned you are to when new models come out, trying those very quickly. People talk about a vibe check from models. It's absolutely real. Very quickly, within a couple of prompts, you can get a very good sense of a model's capabilities, what it's going to be good at, and what you want to use it for.

34:48I agree.

34:52API overloaded. We may have overloaded Anthropic with all the requests. There we go. Yeah, here's a brainstorm from a bunch of different models, basically. That's very cool. Anyway, we've spoken about text models this whole time. There's other very useful categories of models. And I think one of the big stories in AI for the next year plus is going to be the disappearance of these different categories, of the distinction between them. This is clearly one of the things that GPT-4.0 is designed for, is multimodal, being able to take in anything and output anything. but it's a lot of those capabilities actually aren't available yet, whether it is in the ChatGPT product or by API.

35:48So we're going to see more of that. But for now, the best way of creating different sorts of images or doing different things is with very different types of models, very different models. We could skip into images, text-to-speech, transcription. We've got a bunch of image generation stuff. All right. Great. So I think that there's three images, three image models that I'd like to mention. And really, among them, I think two are really notable. So the first is, it's a pretty new model. It's called Flux 1.1 Pro from Black Forest Labs. they've and it really is a very good text to image model they've now released a bunch of other image to image models and other models there's so many to infill in like details into images and do other things transform images upscale them and outpaint and do all of that this is just the the actual image generation.

36:51And you can actually use this model on fal.ai, but it's also the model, I believe it's the model that's used in Grok with a K on x.com, you know, the xai model. It is the model because it's a really powerful open source model. It's now behind the driving the image generation of Grok with a K. It's driving the image generation of perplexity. It's driving the image generation of Mistral. It's driving. So yes, it's now available through many different sources. And yes, it's an awesome tool. It's awesome. It's pretty good with text. But it really is very good. Schnell is a different kind of part of the family from Flux1 Schnell.

37:36And that is kind of really fast and much cheaper. I think this is one of the things that image models over the last year plus have become much more capable. but have not become a lot cheaper. And it's still relatively similar kind of price per image with some of the models even getting more expensive. And then the last one is Idea 2. It's one of the models from Ideagram and it was designed for typography. It's a general image model and it can generate all kinds of images, but it is the best today still for trying a mostly accurate text, which most image models really struggle with. And so those three together are like really the kind of top individual models.

38:23Obviously, there's MidJourney. You can't really talk about text-to-image without mentioning MidJourney, which is it may still be the sort of best underlying model, but it's really it's not available anywhere except in the MidJourney UI, which is undergoing kind of development at the moment. but the gap between MidJourney, which was, it used to be huge between MidJourney and the next best model has really narrowed significantly between MidJourney and Flux. I agree with you 100%. I think the biggest benefit that MidJourney still has is control, right? Because now they've developed their website and there's actual better user interface that they have developed and you can do a lot more stuff and they have their parameters.

39:08So I think if you're an advanced user, then MidJourney still gives you more ability to control the output. But I also agree with you that from a pure quality perspective, Flux 1.1, definitely the pro model is up there with MidJourney as far as what the output looks like. Yeah. And it's going to be really interesting. There were the sort of a very recent announcement and demo about like a text to world models. I've heard Midjourney is also developing something like that, where it doesn't just create an image, it's an immersive, an image that gets turned into kind of an immersive world that you can then explore.

39:51And that's also really exciting. It's going to be interesting to see how things develop from there. Today, it's really mostly text image. I want to see your example of the image generation, but then there's a question from Gwen, which I think will be very exciting for everybody. So So talk to me about your example with the image generation, and then I'll ask you a question that is related to text and data in general. So here's, I think it was the first prompt that I put in here. And it was, I think it's like a pretty useful comparison of the outputs that you get. So the prompt was a holiday greeting card, very simple.

40:26And we fit that into those three image models that we talked about. But Flux1.1 Pro has created a really, like, quite a beautiful scene. And it's sort of almost a fairy tale kind of feel with stars and the trees, kind of an almost old-timey feel. Flux1.1 Schnell has created something that looks a little bit more like a typical holiday card, even down to a little URL, a fake URL in the bottom, which doesn't really make sense, but it is misspelled happy holidays. Happy holidays. Yes, exactly. So I think this is inadvertently a very good example of the state of the art at the moment, or the state of this model and these faster models.

41:19Whereas Program 2 immediately first shot has created a really nice little, you know, Santa, like illustration of Santa, a little boy, and the accurate Merry Christmas speech bubble. Something else that we do, just like we can put together really good chains of models with text, we've done the same thing for images, where what it actually does is it takes a prompt, brainstorms ideas using Lama and GPT-40 Mini, writes a prompt using Claude and then generates an image using Flux 1.1 Pro. And that's what we hear in the sort of final panel, which we could, which we can rerun. You can see how it works if you're interested.

42:07So just brainstorms the different models and creates one. But yeah. A few things about, one thing about this and then the question and then I think will be done because we touched on many different things. About this specifically, this is something I do today quasi-manually. when I create presentations, right? So when I create a presentation, I will use any other tool, doesn't matter, usually ChatGPT for this particular reason, and I'll explain in a minute, to brainstorm the flow of the presentation and what needs to be in it and so on. And then I will ask ChatGPT, I said, okay, what do you recommend should be relevant images for each of the slides that we just discussed?

42:44And the reason I do it in ChatGPT is because it understands the context of the overall presentation, which just the image generation tools don't. And so it will give me ideas and then I will pick up the ideas that I like or I will run away with those ideas for just, it gives me a creative idea and then I can further develop that. And then I ask ChatGPT to create the images for me. And when I get stuck because ChatGPT's image generation is not good enough, I will then take those two images and go to Me Journey or Flux and create better images for one or two images out of the 25 that I'm generating.

43:15And what you just showed really does it in one step. I can literally build a process in Hunch that will do all these things that will give me multiple options for the images right the first time I click the button and will generate the outline in the presentation and so on. And the images and the examples across multiple tools, which is just for that one use case that I mentioned is nothing short of magic. I'll go back to the question, which I think is a very interesting question, is can you use it to check data in order to reduce hallucinations? So can I say, okay, run this through this model, then run this back on the actual data that I gave it to actually check that the data is there to give me a verified or a better checked outcome of the output I gave it?

44:08Yeah, yes. And this is actually a really interesting topic that I love talking about. But I think that for a lot of, I think hallucinations are becoming less and less of a problem. And in fact, I think for the majority of tasks that we have, it's really, you don't really need 100 % accuracy because you don't actually get that from humans anyway. I think like we're conditioned to want that whenever there's a computer involved in a task. But if you ask a friend to do something for you, ask someone else to do it, it's not going to be 100 % probably. So there's that kind of, I think, expectations adjustment that actually reduces one's stress a little bit when dealing with these things.

44:55But yeah, absolutely. I think that this is the advantage of using multiple different models as well. At the same time, in parallel and also as a check. So if it's something that's really important to get right, you can then prompt a model to say, okay, here's the answer, double check everything and write out what it may have missed. I think that where people keep seeing hallucinations is honestly GPT-4.0 in ChatGPT is just not as, it isn't as good as avoiding hallucinations as Gemini 1.5 Pro and Claw 3.5 Sonnet for sure in my experience. So I think people experience the problem more than they need to.

45:40But yeah, getting other models to double check output is great. An example from kind of history is that when you had people, this is like before computers were widespread. If you take civil engineers, for example, who are building a bridge, like you'd often have two teams or multiple teams of people doing all the same calculations in parallel. And it's only if you get to the same answers, do you have 100 % confidence. And so it's very easy to simulate that exact workflow by running the same prompt twice or by running it with different models and seeing the outputs. Fantastic. David, this was A, really valuable.

46:19B, really exciting, I think. And there's a lot of comments that you're probably not reading because you're talking, but people are saying, oh my God, this is awesome. I can use it for this. I can use it for that. This is so exciting. So people definitely are enjoying the idea. if people want to find Hunch if people want to find you if people want to connect with you what are the best ways to do that? Yeah, go to hunch.tools and from there connect with us on Discord we're on Discord we have a great community there we're on LinkedIn, on Twitter I'm DavidDBWilson on Twitter yeah, we love to connect Awesome, thank you so much thank everybody who joined us we had a bunch of people on LinkedIn and a lot of people on the Zoom, great participation and conversation all across.

47:07So again, if you are listening to this after the fact, come join us on Thursdays. We do this session every Thursday, unless it's a holiday or I'm traveling for business. And even then, usually I try to squeeze it in and come and join us so you can do this as well. And again, David, this was fantastic. Your tool is really unique and fascinating and very valuable to many use cases. Thank you so much for sharing your knowledge and your time with us. You're welcome. Thanks very much.

From the publisher

AI is no longer just a buzzword—it's a competitive advantage. But with so many tools and models, how do you know which one fits your needs? David Wilson, founder of Hunch, has tested the latest AI advancements across thousands of use cases and is here to share what works.

In this webinar, we’ll explore:
- How to match AI models to specific business tasks.
- Common pitfalls and how to overcome them with clever prompting.
- Real examples of workflows that save hours and drive results.

David brings unparalleled expertise, having built Hunch to simplify complex AI workflows. With experience running over 1,000 AI tasks weekly, his insights will give you the edge to implement AI with confidence.

About Leveraging AI

If you’ve enjoyed or benefited from some of the insights of this episode, leave us a five-star review on your favorite podcast platform, and let us know what you learned, found helpful, or liked most about this show!

More from Leveraging AI

All 330 episodes
148 | AI Tools Mastery: How To Choose The Right AI Model For Any Task with David WilsonLeveraging AI · 48 min
Listen in VO