Gemini vs OpenAI

14 Feb 2024 · 43 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Practical AI - Episode: Gemini vs OpenAI

Episode Overview In this episode, hosts Daniel Whitenack and Chris Benson discuss the recent release of Google's new AI model, Gemini, and its comparison with OpenAI's offerings. They also delve into the implications of the recent FCC decision to ban AI-generated voices in robocalls, particularly in the context of political influence.

Key Participants

  • Daniel Whitenack - Founder and CEO at Prediction Guard
  • Chris Benson - Tech Strategist at Lockheed Martin

Episode Content

  1. Introduction
  2. The hosts express their excitement about the latest developments in AI.
  3. Acknowledgment of a lull during the holiday season but a surge of news in the AI space recently.
  1. FCC Decision on AI Voices in Robocalls
  2. Background: Automated calls (robocalls) using AI-generated voices have been prevalent, leading to issues of misinformation and manipulation.
  3. Recent Event: An AI voice clone of President Biden was used in robocalls, attempting to sway public opinion ahead of elections.
  4. FCC Ruling: The FCC has banned the use of AI-generated voices in robocalls to prevent fraud and misinformation.
  5. Chris Benson emphasizes the ethical implications of using AI for manipulation.
  6. The incident reflects broader issues of trust and the potential for abuse in AI technologies.
  1. Gemini Launch by Google
  2. Gemini Overview: Google has rebranded its chatbot Bard to Gemini, offering new functionalities and competing against models like OpenAI's GPT-4.
  3. Model Tiers:
  4. Gemini Pro: Comparable to GPT-3.5 (free version).
  5. Gemini Advanced: Subscription service utilizing the Gemini Ultra model, competing with GPT-4.
  6. Initial Reception: Mixed reviews; while some features are promising, many users reported rough edges in performance compared to established models like GPT-4.
  7. Daniel and Chris discuss the importance of user experience and the surrounding ecosystem, beyond just the model capabilities.
  1. Competition Landscape
  2. Discussion on the two-horse race between Google and OpenAI.
  3. Mention of other players like Anthropic and Cohere, which have been somewhat absent from mainstream discussions but are also developing competitive models.
  4. The hosts reflect on the innovation in the AI space and the potential for future advancements.
  1. Innovations in Data Analysis
  2. AI and Data Analytics: The episode highlights the trend of using AI models to perform data analytics through natural language queries, generating SQL for database interactions.
  3. Importance of understanding how these models operate to effectively leverage their capabilities.
  4. Tools and Frameworks: Introduction of tools like Defog and VANA AI, which facilitate the generation of SQL queries through natural language input.
  5. The conversation touches on the maturity of AI applications in data science and the intersection of traditional analytics with new AI capabilities.
  1. Miscellaneous
  2. Daniel shares insights on the challenges of integrating AI in educational settings and supporting teachers in adopting new technologies.
  3. The hosts stress the need for collaboration and backing for teachers advocating for AI usage in classrooms.
  1. Learning Resource Highlight
  2. The hosts recommend the Prompt Engineering Guide available at [apromptingguide.ai](https://www.promptingguide.ai), which provides resources for effectively prompting different AI models.

Key Takeaways

  • The FCC's decision against AI voices in robocalls signifies a growing need for ethical considerations in AI deployment.
  • Google's Gemini aims to compete with OpenAI's GPT-4 but faces challenges in performance and user experience.
  • The landscape for AI in data analytics is evolving, emphasizing the synergy between traditional data analysis and generative AI capabilities.
  • Ongoing collaboration between educators and technology professionals is essential for successfully integrating AI into learning environments.

Conclusion The episode captures significant developments in AI, emphasizing practical implementations and the ethical implications of new technologies. The hosts encourage ongoing discourse and exploration in the rapidly evolving AI landscape.

Sponsors

  • Neo4j: A graph database service.
  • Fly.io: A platform for deploying apps globally.

Resources

  • [Gemini by Google](https://gemini.google.com/app)
  • [FCC Decision on AI Voices](https://www.fcc.gov/document/fcc-makes-ai-generated-voices-robocalls-illegal)
  • [Prompt Engineering Guide](https://www.promptingguide.ai)

---

For further insights and discussions on AI, tune into upcoming episodes of Practical AI!

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:28Welcome to Practical AI. link in the show notes. Thank you to our partners at Fly.io. Launch your app close to your users. Find out how at Fly.io.

0:42Welcome to another episode of Practical AI. This episode is a fully connected episode where Chris and I keep you fully connected with everything that's happening in the AI world, all the recent updates, and also share some learning resources to help you level up your AI and machine learning game. I'm Daniel Whitenack. I'm founder and CEO at Prediction Guard, and I'm joined as always by my co-host, Chris Benson, who's a tech strategist at Lockheed Martin. How are you doing, Chris? Doing pretty good, Daniel. A lot's happened this past week. A lot has happened. It seems like, I don't know if it felt like this to you, but there's sort of a little bit of a lull around the holidays, maybe.

1:27Too much eggnog. Yeah, too much eggnog. But we're fully back into the AI news and interesting things happening. One of the ones that I had seen this week, Chris, was a decision. Well, I don't know how all the government stuff works, but at the FCC, which regulates communication and other things in the in the United States had a ruling about AI voices in robocalls so if people don't know robocalls are automated phone calls typically when I worked back in the telecom industry we'd call it sort of dialer traffic right you spin up a bunch of phone numbers you can call a bunch of people this is how you get phone calls from numbers that seem maybe local to where you're at, but they're really just automated calls.

2:22And then, you know, you pick up and realize it's spam or someone trying to sell you something or something happening. Anyway, there was an interesting one where there was an AI voice clone of President Biden. And I think they were robocalling a bunch of people and trying to sort of change views about President Biden via this recording. Well, it wasn't a recording. It was a voice clone of him saying certain things which hopefully would sway people's political affiliations or sentiments leading into election season. Anyway, this is one of the things that was in the news and maybe prompted some of these decisions or at least highlighted some of these decisions by the FCC.

3:12to ban or find people that were using AI voices in these robocalls. So yeah, what do you think, Chris? First of all, I think whoever was doing that has a serious ethical issues to contend with. Yeah. Well, I'm not sure that a lot of dialers are primarily motivated by their ethical concerns. Yeah. I mean, I think that we've been seeing this coming for such a long time and we've talked We talked about it on the show, you know, with all the generative capability and the ability to commit fraud and the ability to misrepresent yourself in ways like this. So I'm glad the FCC got on top of it after something like that happened.

3:54And I think, unfortunately, I suspect we'll see quite a bit more of such things. As you pointed out, not everybody follows the law as well as maybe they should. I keep waiting for them just to ban robocalls altogether and it would just take the whole issue away from us. We'd have AI generated voices in other contexts, of course. Yeah. One interesting thing. I actually forget if this was a conversation we had on this podcast or elsewhere. Maybe someone can remember. I don't always remember all the things we've talked about on this podcast. But I saw it either in a news article or we were discussing someone on the other end of the spectrum who was using cloned voices or synthesized voices to actually spam bait the spammers.

4:39So they had a script set up where they would get a robocall or a spam call. And actually, they would have this conversational AI that would try to keep the spammer on the line as long as possible. I think we did talk about that. I remember that. I remember that. Yes. So I don't know if that's illegal. I found that one also kind of fun because that you see these people on YouTube that sort of spam bait the spammers. Right. And try to keep them on the line because if they're talking to an AI voice, right, then they're not scamming my grandma or something like that. Right. That's true. So, yeah, that's that was, I think, the goal in that.

5:26But I don't know, maybe maybe all of this is gets in a little bit of a murky zone. It does. But I would say the FCC, the Federal Communications Commission, got it right on this one. Score one for the government. Yeah. What I don't know, I think this would still allow, because obviously when you call on to change your hotel reservation or you call your airline or something, there's synthesized voices. And there have been for many, many years, not necessarily synthesized out of a neural network, but synthesized voices. So I'm assuming that that I haven't read the ruling in detail. I think the main thing that they're targeting is these robocalls.

6:07And so I don't think that covers these assistants. But I don't know. That's a good question. I would assume it goes to intent, you know, and the representation of the voice. And if it is clearly, as in the case of the FCC ruling, is mimicking a person for the purpose of misrepresenting how they're seen or whatever, or what their positions are and such, then I think that anything, I think that's a very reasonable thing. thing. I think all of the types of circumstances we find ourselves in where people are trying to commit fraud or misrepresenting themselves in some way probably need to be addressed in this way.

6:45But there are obviously for every one of those, there's probably a thousand legitimate use cases as well. So I agree. Yeah, there is probably a weird middle zone because even if you remember when I think it was originally Google did their demos at one of their Google I.O. conferences, One of the things that was shown on stage is clicking and calling like your pizza place and ordering a pizza with an AI voice, right? Like, or make me a reservation at 5 p.m. at this restaurant. But you can't, there's no form on the website, right? So there was an automated way to make a call with an AI voice to make the reservation for.

7:26Which seems completely legit to me because you're not, you're representing everything appropriately. you know, you're, you're not pretending, you're not getting around, uh, you know, that kind of thing. It's, you have a tool and it's a tool. And I think, uh, and I, frankly, I could use, I could use a few of those in my life, you know, and just take care of all the things, but I'm probably not going to call anyone and, uh, and have an AI model pretend to be Joe Biden or anybody else. So. Yeah, I think it definitely, like you were saying, it gets extremely concerning when there's a representation that this is this person and they're trying to sway your mind in one way or another.

8:04And it's not that person. Yeah, pure ethical problem right there. I mean, that's so. Yeah, well, I don't know. Do you think that this represents some of what we'll see this year in terms of a trend of government regulation of generated content? I would not be surprised, especially, you know, when we talked last year about the executive order here in the US that came out. And I think that was indicative of further actions to come. I mean, they essentially laid out a strategic plan on how they were going to address AI concerns. And FCC was one of the agencies, I believe, that was explicitly listed in the order, if I recall.

8:45And so I'm not surprised to see them weighing in on this at this point. So it'll be interesting to see how it mixes across national boundaries and see how various countries are addressing it and what that means for so much of this is transnational in terms of technology usage and even organization spanning. And so it will be a curious mess for all the lawyers to figure out going forward. Yeah, when the dialer is using Twilio or Telnex or something to spin up numbers, but they're doing it from an international account, which is probably not even in the country where they're operating. And there's all of these layers.

9:28It gets into some crazy stuff. I know that's always something that stands out to me. I always listen to the Darknet Diaries podcast. It's one of my favorites. So shout out to them. for the great content that they produce. But yeah, that's always a piece of it, right? Is putting enough of these layers in between to where, yeah, sure, there's regulations, but. We just need a blanket rule, a global blanket rule that's just do the right thing. Let's just everybody, everybody out there, just do the right thing. But we may not have things to talk about on the podcast then. Yeah, well, the messiness of the real world will continue.

10:08But yeah, Yeah, but speaking of Google, I mentioned the Google demos and the stuff they've done over the year with synthesized voices and all of that. And of course, recently they've been promoting Gemini, which is this latest wave of AI models from Google, which are multimodal kind of first models. Yeah, there's a whole bunch of kind of related activity in there and that they took their existing chatbot, Bard, and they rebranded it into Gemini. And there are several, there's Gemini Pro. Very confusingly, there is the paid service now of Gemini Advanced, which is using the model called Gemini Ultra.

10:52So I know initially there was some confusion about Advanced versus Ultra. Well, Advanced appears to be the service. ultra is the underlying model so pro represents a model size or ultra or or it represents a subscription tier both in different ways so pro is the free tier there's nothing less than pro we only start oh obviously yeah we've talked about this with apple products before there's no low quality anything right exactly that's what i was about to say there's no such thing as low quality It's pro you start with pro and that's the free version. It's the smaller model that we can all go in.

11:31You can go to Jim just as you could go to bar.google.com. You now go to Gemini.google.com and it's there and available. So Bard is no more. So Bard is no more. Gemini pro is roughly the equivalent of GPT 3.5, the free version on the open AI side. And now Google Advanced, which has the Google Ultra model, is competing against ChatGPT, which is hosting the GPT-4 model at the high end. And there have been a billion reviews of how the two go against each other head to head. Have you tried the various ones or tried Gemini? I've not tried Ultra yet because I haven't decided to pay for it because they're asking for 20 bucks a month.

12:14So I haven't been able to compare it directly. I've watched a whole bunch of YouTube videos more than I should have. We're showing people doing side by side. And I think it's a really good model, but it generally has met with some disappointment in that people are expecting the newest thing is always going to be the greatest thing possible. And I think we saw something with GPT-4 where when OpenAI released it and it had its initial fanfare, and then they've built a lot of infrastructure and services around it and the various plugins. They've also fixed a lot of the problems behind the scene while maintaining the actual underlying model, whereas Google has not done that.

12:58They put the model out and it's comparable in many ways, but it feels very, very rough around the edges and it doesn't always give you the best output. So most of the direct head to head comparisons, most of the various tests I've seen have had GPT-4 went out on a head to head thing. So my expectation on that would be that Google will start working around the issues that it has and cleaning it up. And probably within a few months, it'll probably catch up a little bit closer in that way. So our company and actually the last few I've been a part of have been big Google users in terms of G Suite and Google Workspace and email and docs and all of that stuff.

13:43So I'm kind of embedded in that ecosystem. And, you know, I'm not thankfully not having to deal with teams or something like that. as I know many are. I am at work. It's terrible. Oh, gosh. I feel for you. And I guess I do experience that pain on a second order way because I have to take a lot of Teams calls. But anyway, outside of that, which is probably enough said, then so I'm always trying the Google stuff that comes out. And I had tried Bard. And I think also before that, just the general interface to, I don't know if it was always branded as BARD or I remember Palm, but I think Palm was below or embedded in BARD.

14:32I don't remember always what the branding was. But yeah, now there's Gemini. I would say my impression was similar, Chris, in that I just took literally one of their... You know how you log into any of these systems like ChatGPT or Gemini? And I literally just tried one of their example prompts like, try this. I think it was like print out how to do something in Linux or something like that. I think list processes or something. I just click the button like the example prompt and it wasn't able to respond to the example prompt. You know, these are rough edges. I'm sure the model does a lot of things really well.

15:14And that was just like a fluke in many ways. But it, I think, does represent a lot of those rough edges that they're dealing with. And my impression, I've said this a few times on the podcast, it's like when you're a developer working directly with one of these models, it's kind of like taking your drone that's flying all great and you're controlling it. And then you take it out of autopilot mode. And there's all of these things to consider that you really just didn't think about because they're taken care of by great products like Cohere, Anthropic or OpenAI or whatever. So I definitely feel for the developers because there's a lot of things and a lot of behavior to take care of.

15:57But yeah, that was not the best way to win me over, I think. They might have done better to hold back just a little bit longer and do a little bit more. They talked about that they had roughly 100 private beta testers. And that seems to me a very small sampling of beta testers to be working on it. you mentioned another name just now which i wanted to throw out that is very absent from this conversation out there that is anthropic uh i don't see a lot of comparing it to claude and stuff like that our claude too at this point or maybe yeah anthropic and cohere yeah maybe some other ones absolutely right now it's been a two horse race between these two uh which made me a little bit sad.

16:42I wish there had been a little bit more expansive and also against some of the open source models that are out there. Because one of the topics that you and I are often talking about is with the proliferation of many models, some of which are private, some of which are open. It increases the challenges for the rest of us in the world to know what to use and when and when to switch and things like that. Something that I know you know quite a lot about. Yeah, it's been intriguing to see all of these. And I would say all of them are on some type of cycle, right? So we're talking about maybe GPT-4 is in the lead and here comes Gemini.

17:22And then we're mostly talking here about the closed proprietary models, that sort of ecosystem. But then I'm guessing, you know, Claude had a big release at some point and they're probably in their cycle where I have no inside knowledge of this, but it's just my own perception that Anthropic, Cohere, they're in a different release cycle, obviously, than OpenAI and Google. So we'll see something from them in the coming months, I'm sure, in terms of upgrades or multimodality or extra functionality like assistance or tying in more things like RAG and that sort of thing, as we've seen with OpenAI's assistance and file upload and that sort of stuff.

18:05You know, if we're fair about it, when you think back to when GPT-4 came out, it didn't have all the things that, you know, the ecosystem has grown substantially since its release. And it had some of the same challenges of that. And I think this might be with Gemini coming, you know, I think everyone kind of took that for granted. They were a little bit less splashy than a big, giant new model coming out. And I think this is one of those moments where you kind of go, wow, there's more to this than just the model itself. You know, big new model. I got that. But there's so much to the ecosystem around a model and the various plugins, capabilities, extensions, whatever you want to call them.

18:46Google calls them extensions at this point. But I think it really goes along the lines of something we've been saying for a long time and that the software and the hardware, it's all one big system. It's not just about the model. So I suspect Google is very well positioned to make the improvements in the coming weeks. So it may be interesting to revisit some of these tests after a short while. Yeah. And there are other players that are kind of playing on this boundary between open and closed, either on that sort of open and restricted line. So releasing things that are open and not commercially licensed, or open source, but with some other usage restrictions and that sort of thing.

19:32There's cool stuff happening in all sorts of areas. One of the ones that we've been looking at is a model from Unbabel, which is a translation service provider. They have this tower family of models, which does all sorts of translation and grammar related tasks. But there's also a lot of multi-modality stuff coming out. So I noticed we talked about text-to-speech at the beginning of this episode, and I'm just looking at the most trending model right now on Hugging Face is the MetaVoice model, which is a 1 billion parameter model that is text-to-speech. But if I'm just looking through kind of other things that are trending, we've got text to speech, image to image, image to video, semantic similarity, which are, of course, kind of embedding related models, text to image, automatic speech recognition or transcription.

20:35So there's really a lot of multi-modality stuff going on as well. And people releasing that. I know one that you highlighted was some stuff coming out of, I believe it was Apple, right? Yes. Related to image, or how is it phrased? Image modification or something like that? Image editing? Image editing. It's M-G-I-E is the acronym, which I'm guessing they're, I haven't heard them say this, but I'm guessing they're calling it Maggie or something like that. and it is a where you you'll give a source image and they have a demo that's on hugging face and you essentially kind of talk your way in through the editing process and gradually improve it and everything so i think they had the bad luck of announcing this and releasing it at the same time that google did gemini to go head to head on gpt4 so i think it largely got lost in the news cycle but uh it looks like it might be a very interesting thing and i think uh you know they're competing It's like Adobe doing image generation.

21:35And all of these companies have some level of image editing model capabilities. So it will be interesting to see how Apple's plays out and how they apply it to their products. What I think is a differentiating or interesting element of this, which is maybe not text to image or text to text sort of completion, but the common types of things that people are wanting to do, which are somewhat model independent, but are more workflow related. So things like RAG pipelines, where you upload files and interact with them, you've kind of GPT models or the OpenAI chat GPT interface, where certainly you can upload files and chat with them or analyze them.

22:23Anthropic actually was an early one where because of their high context length window models had the ability to upload files and chat with those files. I don't think, at least I couldn't tell something similar in Gemini other than uploading an image and chatting or reasoning over that image, which is sort of like the vision piece of it. But more than multi-modality, there's these increasing workflows that people are developing. One of those that I think is really interesting is the data analytics use cases that are coming out. So you have actually, I've seen a trend in a lot of these companies popping up that are something to the effect of new enterprise analytics driven by natural text queries.

23:15So I'm thinking of like Defog, I think it is. Yes. These companies, which are a chat interface where you type in a question maybe your SQL database is connected and you get a data analytics answer or a chart out. And this is something that I believe if I'm, again, understanding, I don't know all the internals of ChatGPT, but it's interesting that there's different takes on this approach. And I think there's a lot of misunderstanding about how this actually happens under the hood. So I don't know, So have you done much where you've like uploaded a CSV or you've done that sort of thing in chat GPT and asked it to analyze it or something like that?

24:03Ironically, that's literally something I'm playing with right now. I know you didn't know that before asking the question, but I saw a similar post about kind of analytics being used for this. And so I'm experimenting with it, but I'm still very early. How are your results initially? They're not as good as I want, but I think that's mainly my problem. I keep running into little bumps where I'm trying to get the CSV usable very well. So I have a database that I dumped some data out of and was trying to do that. But I literally just did this today. Today was day one and then stopped and came in for us to have this conversation.

24:41So let me let you know in another week or so how that fanned out. But it caught my eye because I saw a conversation online about this. and some of the personalities that I've always associated with, you know, being super technically bright analytics folks were kind of saying we're just hitting that moment where this kind of just AI-driven conversational analytics is now going to be available to everyone. And I was like, well, that's what I want. That's what I need. So I'm actually trying to do something for work right now on those ones.

25:30What's up, friends? Is your code getting dragged down by joins and long query times? The problem might be your database. Try simplifying the complex with graphs. A graph database lets you model data the way it looks in the real world instead of forcing it into rows and columns. Stop asking relational databases to do more than what they were made for. Graphs work well for use cases with lots of data connections like supply chain, fraud detection, real-time analytics, and generative AI. With Neo4j, you can code in your favorite programming language and against any driver. Plus, it's easy to integrate into your tech stack.

26:07People are solving some of the world's biggest problems with graphs. And now it's your turn. Visit Neo4j.com slash developer to get started. Again, neo4j.com slash developer. That's neo4j.com slash developer.

26:37Well, Chris, I was asking these questions about this data analysis stuff because this is I've done a few customer visits recently where we've been talking about this functionality. And I've noticed as I've gone around and talked to different people, there's some general misunderstanding about how you can analyze data with a generative AI model. One, because there's something people think is going on that isn't actually going on. And two, because generally, if you ask a language mod just a chat model without uploading data like math type of questions usually it is really terrible at that right even like adding things together or doing like basic aggregation is something that these models are known to to fail on pretty poorly and so the question is like well how am i getting anything relevant out of these systems to begin with And again, I don't know all the internals of ChatGPT, but this is my own understanding.

27:44There's some difference if you look at maybe like an example like Defog or ChatGPT or Vana AI. These are some examples of this that's going on. ChatGPT takes the approach in my understanding where in their assistance functionality. So when you type, you upload a maybe a CSV and you ask a question and you wait for seemingly forever while the little thing spins and it says it's figuring or analyzing, I think is what it says, something like that. Yep. my understanding of what's happening is more of what they used to call code interpreter. It's actually generating some Python code that then it executes under the hood to analyze your data that you uploaded and then somehow passes along the results of that code execution to you in the chat interface.

28:36So this is a very astute observation by whoever had this that, yeah, these models really stink at doing math, but what doesn't stink at doing math is code, right? So these models are pretty good at generating code. So why don't we just sidestep the whole math thing and generate the code and then execute it and crunch your data and we're good to go. I think the thing that often what I've seen people struggling with, like the assistance API and chat GPT is again, they have to support all sorts of random general use cases, right? Because, you know, people could upload a CSV of all sorts of different types or other file types.

29:20And so there's a lot to support. And it's kind of generally slow and hard to massage into working, right? What I've seen more in the enterprise use cases that we've been participating in is less a focus on code generation to do the data analysis and more of a focus on SQL generation to do analytics queries. So this is more the approach of the SQL coder, family of models, defog, VANA AI. We're doing very similar things to in the cases where we're implementing this similar to the VANA AI case where you connect up, let's say you have a transactional database, like your sales or something like that or customer information or product information and you want to ask an analytics query, right?

Read the full transcript

30:11Well, SQL is really good at doing aggregations and groupings and joins. Also, large language models, especially code generation models or code assistant models are really good at generating SQL because like how much SQL has been generated over time. It's very well-known language to generate, right? And so you kind of sidestep the code execution piece in that case where you're not generating Python code, but you're generating from a natural language query, a SQL query to run against a database that's connected. And you just run that SQL query and normal, good old, regular programming code to give you your answer.

30:53And then you send it back to the user in the chat interface. So I thought that would be worth highlighting in this episode because there does seem to be a lot of confusion of what's actually going on under the hood. Like, how can one of these models analyze my data? Well, the answer is it kind of isn't. It's just generating either code or generating SQL that is analyzing your data. It still gets you there, though. It's in a sense, you know, since you're not directly having the model do it, it's sort of a workaround in a manner of speaking. But I think if you look at something like, you know, the ecosystem built around chat GPT, there's a lot of tooling around it.

31:34And I think that's I think this year we're going to see more and more of that, you know, whether it be the SQL use case that you're talking about or continued with open AI. I think Google will do that well. I think Anthropic will get on that. and you'll see these kind of tools for doing exactly that kind of thing where you may not have a model that does a particular task super well but it can produce an intermediate that can do something very very well i think that's a level of you know we keep talking about maturity of the field and i think part of that is recognizing maybe there's a better way to do it than just having the bigger a better latest model so yeah i think that's a great way of approaching it not to self-fulfill my own prophecy from our predictions from last year.

32:21I think in our 2024 predictions episode, one of my predictions was that we would see a lot more combination of, I think, what is generally being called neurosymbolic methods, but maybe more generally just like hybrid methods between what we've been doing in data science forever and a kind of front end that That is a natural language interface driven by a generative AI model. So in this case, what we have is good old fashioned data analytics, just like the way we've always done it by running SQL queries. It's just we gain flexibility in doing those data analytics by generating the SQL query out of a natural language prompt using a large language model.

33:06And I think we'll see other things like this, like, you know, tools and Langchain is a great example of this, where you generate good old fashioned structured input to an API and that API is called and gives you a result. but this could be applied in all sorts of ways right so let's say time series forecasting i don't think right now language models and i've actually even tried some of this with fraud detection and forecasting and other things with large language models and not very good at doing these tasks but they can generate the input to what you would need in the kind of traditional data science tasks So if you say, again, imagining bringing in the SQL query stuff, if you have a user and you want to enable that user to do forecasts on their own data, well, you could have them like put in, fill out a form and like in a web app and like click a button and do a bunch of work.

34:10or you could just have them say, hey, I want to forecast my sales of this product for the next six months or something. From that request, a large language model will be very good at extracting the parameters that are needed and possibly generating a SQL query to pull the right data that's needed to be input to a forecast. But that forecast is going to be best that you just use like Meta's profit framework or something. It's just a traditional ARIMA statistical forecasting methodology. And you just forecast it out with that input, and then you get the result. So this is the merging of what we've been doing in data science forever with this very flexible front-end interface.

34:57And I think we'll see a lot more of that. I completely agree with you. And not only that, but I think there'll be a lot more room for LLMs that are not the gigantic ones. We've talked a bit. And we've had guests on the show recently, you know, talking about the fact that there's room not only for the largest, latest, greatest giant model, but there's enormous middle ground there where you can have smaller ones and combine those with tools. So it's pretty cool seeing people innovate in this way and start to recognize that not everything has to come out of the largest possible model you have available to you and add that in.

35:34So I'm really looking forward to seeing what people do this year along in their various industries and, you know, and how that spawns new thoughts. So, yeah. And especially with, um, a lot, a lot of things being able to be run locally. I've seen a lot of people using local LLMs as an interface using frameworks like Ollama and others, which is really cool to be able to use LLMs on your laptop to, you know, automate things or do these types of queries or experiment locally. So yeah, I think that even adds another element into the mix. And for edge computing, you know, for truly edge computing, where it's not practical to have a cloud backing, you know, and or the networking between where that model would be in the cloud and where you're trying to do it.

36:23There's a huge amount of opportunity to use them in that in that area. So yeah, I'm hoping that we see a lot of innovation. You know, last year was even the year before was kind of the race to the biggest model. I'm kind of hoping now we see what other branches of innovation people can come up with to take advantage of some of that and also recognize that the midsize ones have so much utility to them that's untapped. Yeah. And maybe before we leave the sort of news and everything that's going on in this kind of co-pilot assistant analysis space. I did see, you know, I actually, my, my wife needed help connecting to printers.

37:04Printers are not a problem that is solved by AI yet, I guess, and will continue to be, continue to be a problem forever in tech. But I was noticing in the, you know, recent updates to Windows, there's the little co-pilot logo there, like even embedded within Windows. And I don't know that whoever watched the Super Bowl during in the US, the Super Bowl as we record this was the day before we're recording this. But there was a co-pilot commercial during the Super Bowl. And that's another interesting thing because this is now it's running on people's laptops everywhere. And of course, that's connected to the open AI ecosystem, in my understanding, through Microsoft.

37:50But yeah, this kind of AI everywhere and also the sort of AI PC stuff that Intel has been promoting and running locally is going to be an interesting piece of it. Totally agree. As we wind up, I want to briefly switch topics here. I received some feedback a few episodes ago from a teacher who was listening, and I was so happy to have one and maybe many teachers out there listening to us and considering this. And as we often do, people may not realize, but Daniel and I, we have a topic, but we are largely unscripted. So we are kind of shooting from the hip in terms of what we're saying. It's a very genuine and real conversation.

38:32We're not looking at a whole bunch of notes and pre-planned script. And I made a comment about my daughter in school and the fact that I really think schools should take advantage of models. And as part of the learning process, as part of the teaching to integrate it in, whereas often school systems right now are saying you're not allowed to use GPT, for instance, in your homework. And that I said, oh, that's stupid, you know, that teachers would not do that. And I, this teacher reached out and said, well, first of all, we really want to, and I'm paraphrasing her. And she said, second of all, you know, a lot of times they, it's not in their power anyways, the school system policy and stuff.

39:13And so I just want to apologize to anyone, especially the teachers out there that might have been offended. I'm much more cognizant now of what I'm saying on that. It was kind of a shooting from the hip, but it was insensitive. And I found that what that teacher pointed out was dead on. It was right on. And I just want to thank the teachers out there, especially those who are trying to take advantage of these amazing new technologies and talk their systems into bringing them into the classroom and not make it just the bad thing not to use for homework. So thank you to the teachers for doing that.

39:47And I just wanted to call that out. It's been a really important thing from my standpoint to say. So thank you. I think it represents the complexity that people are dealing with. It does. You know, teachers want their students to thrive. I think generally, we should assume that most teachers are really actually motivated and engaged, both in culture and technology and the ecosystem wanting their students to thrive. But sometimes, like you say, they have their own limitations in terms of what is the system within that they're working in and privacy concerns and other things. So yeah, that's a good call out, Chris.

40:28I'm glad you took time to mention it. I want to say one last thing. And to teachers out there who are trying to get these things into the classroom so that your students have the best available tools to do things. If you ever need someone to back you up, reach out to us. We have all our social media outlets. You can find me on LinkedIn. And if I will be happy to give a whole bunch of reasons to your school systems on why they might want to use the tools, I'll be happy to work with you on that. And I thank you for fighting that fight on behalf of the students that you're serving. Yeah. And speaking of learning something that we can all learn and be better at is all the different ways of prompting these models for multimodal tasks and prompting and data analysis.

41:15And I just wanted to highlight here at the end, a learning resource for people. A while back, I had mentioned a lecture and series of slides that was very helpful for me from DARE AI, D-A-I-R. Now I think that they've converted that series of slides in that prompt engineering course, I think is what they call it, into a prompt engineering guide. So if you go to apromptingguide.ai, they have a really nice website that walks you through all sorts of things and also covers various models in terms of the chat GPT, CodeLama, Gemini, Gemini Advance. We talked about those on this show and talks about actually prompting these different models.

42:01So I'd encourage you if you're experimenting with these different models and not immediately getting the results that you're wanting, that may be a good resource to help you understand different strategies of prompting these models to get things done as you need to get them done. That's a great resource. I'm looking through it as you're talking about it. And it's the best I've seen so far. Well, Chris, this was fun. I'm glad we got a chance to cover all the fun things going on. And we've complied with the FCC using our actual voices still. We'll see how long that lasts. But it was fun to talk through things, Chris.

42:40We'll see you soon. Talk to you later.

42:49that is practical ai for this week thanks for listening subscribe now if you haven't yet head to practicalai.fm for all the ways and don't forget to check out our fresh changelog beats the dance party album is on spotify apple music and the rest there's a link in the show notes for you. Thanks once again to our partners at fly.io to our beat freaking residents, Breakmaster Cylinder, and to you for listening. That's all for now. We'll talk to you again next time.

From the publisher

Google has been releasing a ton of new GenAI functionality under the name “Gemini”, and they’ve officially rebranded Bard as Gemini. We take some time to talk through Gemini compared with offerings from OpenAI, Anthropic, Cohere, etc.

We also discuss the recent FCC decision to ban the use of AI voices in robocalls and what the decision might mean for government involvement in AI in 2024.

Join the discussion

Changelog++ members save 2 minutes on this episode because they made the ads disappear. Join today!

Sponsors:

  • Neo4j – Is your code getting dragged down by JOINs and long query times? The problem might be your database…Try simplifying the complex with graphs. Stop asking relational databases to do more than they were made for. Graphs work well for use cases with lots of data connections like supply chain, fraud detection, real-time analytics, and genAI. With Neo4j, you can code in your favorite programming language and against any driver. Plus, it’s easy to integrate into your tech stack. Visit Neo4j.com/developer to get started. 
  • Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs. 

Featuring:

Show Notes:

Something missing or broken? PRs welcome!

More from Practical AI

All 157 episodes
Gemini vs OpenAIPractical AI · 43 min
Listen in VO