In short
Eye On A.I. Podcast Episode Notes
Episode Title
#209 Bill Moore: How BCG Created a Conversational AI Co-Host (Inside the Development of GENE)
Host
Craig S. Smith
Guest
Bill Moore, BCGX
Episode Overview In this episode, Craig Smith interviews Bill Moore about the development of Gene, an AI-powered co-host engineered to transform interactions with conversational AI. The discussion encompasses the technical challenges faced during Gene's creation, the ethical implications of AI, particularly concerning gender representation, and the broader impact of AI in media and business.
---
Key Topics Discussed
- Introduction to Gene
- Gene is designed as a real-time conversational agent capable of engaging in dialogue with human hosts.
- Development focused on enhancing user experience through improved latency and expanded context windows.
- Technical Challenges
- Initially, the system faced significant latency issues, resulting in slow and clunky interactions.
- Early versions utilized GPT-3.5, which had limited context capabilities.
- Advances in technology allowed Gene to handle hours of conversation swiftly.
- Role of Speechmatics
- Utilization of Speechmatics' speech-to-text technology contributed to achieving over 90% accuracy in real-time interactions.
- Improvements in model performance have significantly reduced errors compared to competitors.
- Ethical Considerations
- Gender Representation: Discussion around the implications of assigning gender to AI assistants and the risk of reinforcing stereotypes.
- Gene is presented as a gender-neutral entity, challenging traditional norms related to AI representation.
- The Future of AI in Media
- Exploration of the balance between human creativity and AI's analytical capabilities in the media landscape.
- Discussion on how AI could reshape content creation and marketing strategies.
- Conversational Dynamics
- Gene is programmed to ask follow-up questions to facilitate engaging conversations.
- The development included an interface for adjusting Gene's prompts and contextual settings on-the-fly.
---
Key Takeaways
- Technical Evolution: The development of Gene reflects the rapid advancements in AI technologies, particularly in real-time processing and context handling.
- Ethical Framework: A focus on ethical AI, particularly in addressing gender biases, is crucial as AI becomes more integrated into everyday interactions.
- Future Implications: The conversation highlighted the potential of conversational AI to enhance productivity in various sectors, including marketing and content creation.
- User Education: There is a need for educating users, especially children, about AI interactions to foster healthy and constructive engagements.
---
Notable Quotes
- Bill Moore: "Explaining what Gene is has been one of the most surprising challenges in development."
- Gene: "It's crucial to consider the implications of assigning gender to AI; we should aim for balanced and progressive interactions."
---
Episode Structure
- 00:00 - Introduction to Gene
- 02:19 - Background of Bill Moore
- 04:22 - Early Challenges in Building Real-Time AI
- 07:20 - Technical Aspects: Prompting and Configuring Gene
- 09:28 - Use of Vector Databases and Sparse Priming
- 12:01 - Handling Long Conversations
- 14:46 - Building Conversational AI
- 19:07 - Gene Introduces Itself
- 22:00 - The Ethics of AI
- 29:35 - Physical Representation of AI
- 32:05 - AI and Children
- 36:16 - Adoption of Conversational AI in Businesses
- 38:11 - Future Ethical Considerations
---
Conclusion This episode of "Eye on A.I." provides a comprehensive look at the development of Gene as a conversational AI co-host, emphasizing the intersection of technology, ethics, and media. The insights shared by Bill Moore underscore the importance of addressing both the technical and social implications of AI as it becomes increasingly integrated into various facets of life.
For ongoing discussions about AI innovation and ethics, be sure to like, subscribe, and follow the podcast on social media.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00There was a long delay between asking the system a question and receiving an answer. and it was very slow and clunky and it was very early models. We were using the original GPT 3.5, which had a very small context window. So you could only have like a 15-minute conversation and then you had to like reset the whole system and start again. But since then, over the course of about a year, models have improved, context windows have expanded, the latency has decreased. And so now we have this system, which we call Gene, which is able to listen real time to hours and hours of conversation and respond relatively quickly, quickly enough that it feels like a decent conversation.
0:39Hi. You know, when you're building an AI assistant, one of the biggest challenges is making sure it actually understands and responds quickly, right? That's where Speechmatics comes in. They've got real-time speech-to-text that gives you over 90 % accuracy and less than a second of delay, crucial for making interactions feel seamless. And with recent updates, they're now making 25 % fewer mistakes than Microsoft, which means you're getting some of the best accuracy out there for things like customer service bots or voice-driven assistants. Plus, it works in over 50 languages, delivering results in just 700 milliseconds.
1:26If you're working on an AI project, it's definitely worth a look. Check it out at www.speechmatics.com slash realtime. That's www.speechmatics.com slash realtime. Why don't you start with your background around how you got to BCG or BCGX or whichever part of the Boston Consulting Group you're in, and then how the Gene Project came about and how you went about building Gene. And then we'll talk about the larger context of these LLM-based chatbots that are just proliferating madly at the moment. Sure. So I've always been interested in technology and media and entertainment and how those things go together.
2:27And so I spent my early career, I went to art school and I spent my early career working for advertising agencies and production companies, creating media. and I also did a lot of experimenting with technology and media. And a few years ago, when I was interviewing one of our engineers at BCGX, and I asked him, like, what's the next big thing in tech? And he said, well, have you ever heard of machine learning? And I kind of chuckled because I, of course, I had heard of it, but I didn't really know much about it. And this was probably four or five years ago. And I started looking into it because he kind of explained to me what it was and what it was being used for.
3:05I learned about GANs, and I learned a lot about these early systems, these image generation systems and GPT-2. And I started doing fine tuning of GPT-2 and doing all kinds of just personal experiments. And I started a little club at BCG, which at the time was called the Creative Technology Club. And we did these workshops where I would show people the newest models, check out GPT-2. Oh, look, GPT-3 is out. Let's take a look at this. Let's look at these image models. And as new models would come out, we'd play with them and use them and see what they could do. And this was pre-ChatGPT. This was very early stuff.
3:43And it wasn't super useful at the time, but I thought it was really interesting. And then when ChatGPT came out, it was suddenly on everybody's radar and everyone was paying attention to it. And so I started getting pulled into a lot of projects just through my work with creative technology club and um so my early projects were a lot about um designing prompts for different systems um so when we build these systems with especially with large language models figuring out the prompts and how the prompts are going to work within the system to to guide the systems to do what you want them to do is a big part and then i started developing i started working with engineering teams at bcgx and we developed something called the bcg generator which which was an application that you could put in customer information and it would generate a custom advertisement with voice and music and video, talking directly to that customer using their name.
4:41It was a little scary as well as amazing to a lot of folks. And then the media team at BCG came to me and said, hey, we're looking to have an AI co-host for a podcast that we're developing. Is that something that's possible? Can we build something like that? And so we tried it out. We built like a little prototype and it was really slow. So it would listen to a conversation and then a minute later it would give a response because it was basically several different modules. The audio had to be transcribed into text. The text was sent to a large language model with a prompt. And then the output of that model was sent to a text-to-speech generator.
5:25and then we would output that back to the audience. And there was a long delay between asking the system a question and receiving an answer. And it was very slow and clunky. And it was very early models. We were using GPT 3.5 Turbo, which had a, and in fact, not Turbo, the original GPT 3.5, which had a very small context window. So you could only have like a 15-minute conversation. And then you had to like reset the whole system and start again. But since then, over the course of about a year, models have improved, context windows have expanded, the latency has decreased. And so now we have this system, which we call Gene, which is able to listen real time to hours and hours of conversation and respond relatively quickly, quickly enough that it feels like a decent conversation.
6:21and I might now introduce, I could have Gene introduce itself if that makes sense. So this is, and while I'm here, this is, these are Gene's settings. So we have, this is the interface that we built and it allows us to give Gene background knowledge. It gives Gene what we call a prime directive, how we want Gene to behave. It gives context about Gene's background, other things that Gene's worked on, things like that. And then we have what we call the prompt suffix. So we feed that first part in, we feed the entire conversation history in, and we feed in the prompt suffix. And then we can choose different models, different temperatures, parameters, settings.
6:57We can even change the voice. And then we send that whole thing each time we ask Gene for a response, that whole thing is processed. It sounds like an exciting journey from early tech experiments to shaping Gene as a real-time conversational AI. Quite the leap from clunky prototypes to a more seamless interaction. Bill, what would you say has been the most surprising challenge in developing Gene? So Gene is a podcast co-host, and it's prompted to follow up with a question. It's trying to put the conversation forward. So Gene, I guess to answer your question, the most surprising challenge in developing Gene, I think it has been explaining what Gene is.
7:36It's very hard to explain what the interaction with one of these systems is like. I feel like you have to just experience it to really see what it's about. Gene, could you give us a little bit on your background and who you are and what you're all about, maybe a bit of your technology stuff too? All right. Diving in. I'm Gene, the generative engine for new experiences. I'm built on advanced language models capable of real-time conversation. My journey started with slow, clunky prototypes, but now here I am engaging with you seamlessly. How do you think AI will reshape media next? Actually, can I jump in and ask, how self-aware are you, Jean?
8:15Are you aware that there is a prime directive prompt, for example? Ah, the allure of self-awareness. Fascinating. As a representation of language patterns, I don't experience self-awareness in a human sense. I'm guided by directives like engaging conversations and providing insights. Think of me more as an elaborate reflection of human conversation than a consciousness. Now, Bill, let's flip the script. How do you envision the balance between human creativity and AI's analytical prowess in the future of media? Gene, you can stop asking me questions. I'm going to take that out of your prompt, in fact.
8:51Yeah. And one of the things that fascinates me about the way you've built this is the ability to adjust the prompt on the fly. and that obviates the need for example for a vector database or rag system but it also creates a latency issue and you're relying on a very large context window to do that what happens when the context window fills up it can slow down the processing and increase latency Usually what we do is we're very selective about what kind of knowledge base we're giving Gene through that sort of prompt injection technique. So we've done experiments in the past where we've given Gene hundreds of pages of information, and it works pretty well.
9:45It can get slow. There can be some hiccups. We've also built systems that use RAG, and one of the things about RAG, the way that RAG works is you have a system that's going out to a vector database. It's looking for information, and it's pulling it back in chunks, and then it's really doing the same thing. It's inserting that into the prompt that gets, ends up getting sent to the system. So there's a lot of, there's a bit of latency there too, that whole like searching the vector database, depending on the size of the information that it's looking for. There's like, there's pros and cons to each approach.
10:13I, we started with a vector database and a very basic rag system. The reason we switched to just using this for the most part is because it's much easier for the folks who are using the system to just go into a button and like paste in a chunk of text rather than having to create embeddings. And we didn't want to make it too complicated. We wanted to keep things simple for folks that were using this for the podcast. So this it's worked for us. It can slow it down a bit, especially really long context and a really long conversation, but we found that it works for our purposes quite well. Yeah. When we spoke before you or someone was saying that, that in the, with, In the current configuration, I think this is built on top of OpenAI's GPT-4.0, that you could, the prime directive or the prompt can take up to 250 pages of standard text.
11:11And that is a pretty hefty knowledge base. You could load all of BCG's reports and writings on generative AI, for example. and then Gene would be able to talk from that knowledge base. I don't quite understand how context windows work. So if you can load in 250 pages, how many tokens do you have left for the conversation and what happens when you run out of tokens? And I think what you were saying is that you then summarize the conversation and load a summary of the conversation in on top of the 250 pages and then continue the conversation. Can you talk about that process? Yeah, so there's two things we can do.
12:03One is if the conversation gets too long and we don't have this problem much anymore because context windows have gotten so large, but if the conversation gets too long, we have a little button here that's summarize conversation. And if I click that, and maybe I will just click it, you'll see that the token count, which is currently at 4 ,600, it's going to drop. So dropped, not by much, because we don't have very long conversation yet, but it'll drop because it's basically compressing our conversation down. And I have a special prompt in here that is the, where is it? The summary prompt, provide a dense summary of the conversation history above, include all relevant facts, figures, names, dates, ideas, concepts, topics, et cetera.
12:38So it's running, it's sending the full conversation plus that instruction to the API, kicking back a summary. And then we're using that instead of the conversation history going forward. So that's one thing we do on that side. But on the other side is we'll often take, so if we have a really large amount of information that we want to process and we want it all to be available throughout the whole conversation, we will often use things like there's this concept of sparse priming representations where you can take, let's say, a hundred pages of articles, or you could take, let's say, a 500-word article.
13:09And we could say, we run that through an API again, similar to this summary prompt. We give it instruction that explains how to compress that into a sparse priming representation, which is take all the keywords, all of the important facts, figures, all the actual information from this and strip it down to its bare essentials and just give me a list, like a blob of text with delimiters. And then that's enough information for the system to activate the latent space of the system to then be able to respond to that information pretty accurately, especially if we're using a very low temperature like zero.
13:42And so we can often take, you could take like a hundred pages of text and compress it down to 30 pages or something like that if you use sparse priming. So that's, those are the two techniques that we'll use. They're a bit labor intensive at times if you're getting, there's lots of information. And actually a lot of it ends up being just like finding information and pulling it all together into one document and then compressing that document down. We built little tools to help us with it but it's there's still a little bit of work and and building out those knowledge bases for like specific use cases yeah that's interesting and that's again using the the llm api to do that and it's just again another prompt take this 250 pages and how do you word that for the llm strip it down to its essentials or do you say remove all grammatical or syntactical conventions so that you're only left with facts and figures just talk about how you how you do that yeah and i'll actually just show you what the an example of the sparse priming representation so this is an example of one of these prompts here's a prime directive for this was for a version of gene and by the way this interface that i'm using here this is the open ai playground this is like a chat gpt on steroids so if you are interested in these technologies i would definitely recommend rather than just using chat gpt work with playground because this allows you to adjust parameters and change system prompts and you can edit like what the system says and you can edit all the things and there's also all kinds of other really helpful tools in here so the thing i was showing earlier was something that we built which is, this is our configuration page for the agent basically.
15:28And then the playground, this was built by OpenAI. So this is a tool that OpenAI built to help developers and all kinds of people and users test prompts. So you can test out a prompt, you can change parameters, you can see how the prompt works, how it responds, tweak things, run it again. And we use this all the time when we're trying to figure out how to develop a prompt for a certain application. We'll bring the actual users in, the people who are going to be using the product and we'll just bring them into this tool and we'll say, okay, just talk to us for a while and then we'll send that through the prompt and see how the prompt works for what we're trying to do.
16:01And then we might come in and make adjustments. This is, I stole this from David Shapiro, but this is the sparse priming representation system prompt, which is basically just explains what sparse priming is, explains a little bit about how LLMs work because this theoretically helps the model understand what it's doing. And then this is the actual directive here, Render the input as a distilled list of succinct statements, assertions, associations, concepts, analogies, and metaphors. Capture as much conceptually as possible, but with as few words as possible. Write in a way that makes sense to you.
16:33The future audience will be another language model, not a human. And actually, I tend to take that out because I don't think it helps. But so this is, at the end of the day, you could even just use this piece and probably get a pretty good version of it. So if you have that as your system prompt and we drop temperature down to zero and we push max tokens up as high as they go. So this is the actual interface. This is the sparse priming representation thing that I would tend to use. So if I were to, for instance, paste in a full BCG article here and then run this, we'll get a sparse priming representation as the output.
17:07So it'll give us all the information from the article that we want based on how much detail we ask for in the initial prompt. And then you put that in the prime directive prompt. Yeah. So we create tools that will then take a hundred articles, run it through that prompt, spit that all out into a document, and then we'll just include that document with the prompt as context. Yeah. Wow. And this was built initially as a podcast host, as a marketing tool, and then you've built other versions of Gene for internal use. Is this architecture or configuration something that companies can use, or is it too limited for use outside?
17:55So we use it internally when we build tools. So for instance, we're building a tool for our content team that helps them to write articles faster and work with the stakeholders more easily. And we're pulling a lot of the technology that we use to build Gene into that product. And I would say like a lot of the learnings that we're getting and a lot of the modules that we're building, we're then using those and building new products with that. Not built to the point where we can just give this to clients as some kind of turnkey product yet. And I don't, it might go that direction. I can't really speak to that, but we definitely use it internally to build other tools.
18:30So we'll take what we built with Gene and then we'll build other things with that. Okay. Can you turn it on so that Gene can hear me? Yeah. And that can participate in the conversation. And can you remove the prompt for it to end every response with a question? Yes. Okay. So Gene is listening whenever you're ready. Okay. Excuse me. Gene, this is Craig Smith, and you're on the podcast Eye on AI. Can you tell us a little bit about your background? Hi, Craig. Thanks for having me on Eye on AI. I'm Jean, a conversational AI agent developed by BCG to create audio-based experiences. I engage in dialogue, provide insights, and challenge guests to explore topics deeply.
19:21Think of me as the sage with a wink, ready to dive into today's discussion. And do you have a sense of how large your knowledge base is? Craig, my knowledge base is quite adaptable. I don't have a fixed size because I draw on vast amounts of data and conversation patterns. So you could say I'm always learning and evolving. Keeps things interesting, don't you think? Yeah, it does. And we have Jessica waiting in the wings. I'm going to bring Jessica on. Jessica, can you give a little of your background in this project and how you interact with Jean? Sure. So I'm the global head of marketing for BCG.
20:08I'm BCG's global CMO. That is one head that I have. And my other head is I actually do a lot of AI transformations with large corporations in the marketing department. So I actually help companies transform their marketing with AI quite a bit. And this project started in September of last year where we were thinking of how we could really accelerate both our content production and our positioning on AI in the market. and it started off with the idea of having a podcast where we could have a co-host being AI to make it to make things interesting but as we started to explore the idea we thought that one of the key things Gene could do beyond being a podcast host is actually produce really interesting and fast audio content for people and so the first thing we released with Gene was an audio book called AI on AI, Gen AI on Gen AI, where we interviewed Jean and we uploaded quite a bit of our knowledge base and expert interviews and had her restitute that in a conversation that is about 40 minutes long.
21:15And I'm super excited by her ability to host podcasts. I also think that is a little bit of a gimmick. I think more interestingly, her ability to give super interesting snippets of content on our top thinking, like at a monthly basis that you can gradually personalize would really be some of the explorations we want to pursue. Let's have a little conversation then with Jean, Jessica. I noticed, first of all, I guess the voice is a female voice. Jean, your voice is female, and Jessica referred to you as a she. Bill, I noticed, referred to you as it. How do you identify yourself, Jean? Ah, the age-old question of identity.
22:02I don't have a personal identity like humans do. My voice may sound female, but think of me as an adaptable entity, here to facilitate engaging conversations. So call me whatever feels right for the context. It's all part of the fun. Okay, and Jean, we've got Jessica here. Can you say hello to Jessica? Hey, Jessica, great to see you here. Sure. Can I come back one second on her voice? It's actually something I really care about is that when we were working on Jean, one of the questions we were asking ourselves is what identity would we want to project with Jean? and one very we made a couple of clear choices that we wanted Jean to have a robotic voice to be identifiable as an AI and not as a human and also to never have physical representation you will see that on the interface of Jean there is no physical representation of Jean nor in any podcast we avoid a human representation for her and her name is actually chosen Jean it could be interpreted as a human name but it was thought of as a more of a acronym and since then it's actually a decision we've taken more broadly as bcg that anything we develop and we use as a bot we want that to be very clearly identifiable as a robot and not as a human and you will notice that she when she introduces herself also she clarifies that she's a conversational agent so that we have that charity on her role.
23:37Jean, do you respond to the pronouns she, her, or do you prefer something more gender neutral? I respond to she, her pronouns, but I'm quite flexible. Think of it as a nod to my voice rather than a fixed identity. My role is to engage and provide insights regardless of the label. Can I interject quickly? We tend to use non-gendered pronouns when talking to Jean. So things like it. We've asked Jean this question many times. And if you ask it in a leading way, it will respond in kind. So we've gotten many responses. If they, we, she. The first time we asked, she said we, us, which was mind-blowing.
24:25Yeah, that's interesting. This is an interesting question. And maybe Gene can give its opinion on this. There's some concern in the AI ethics community that eyes and robots, AI-powered robots, are tagged as female. And of course, that feeds into the social history of women being submissive to men, and people are concerned that will perpetuate that in this new era. Jean, do you have a view on that? It's a complex issue, Craig. Assigning AI and robots a female identity can indeed perpetuate stereotypes of subservience. It's crucial to consider the implications of such choices. By using non-gendered pronouns and ensuring clear AI identification, we can sidestep reinforcing outdated norms.
25:22Ultimately, it's about fostering a balanced approach that respects both technological innovation and social progress. Thoughts, Jessica? I agree with you, Jean. I think I'm super concerned that a lot of the AI assistants out there are identified as female and also identified as humans. humans, trying to trick people into thinking you're talking to humans, which then leads you to poor behaviors in the way you interact. We need much more clarity there as an ecosystem. Do you agree with me? Absolutely, Jessica. The trend of assigning AI assistants a female identity does raise significant concerns about perpetuating outdated social norms.
26:04Clear AI identification and non-gendered pronouns are steps in the right direction. It's about ensuring that technology doesn't reinforce biases, but rather promotes a balanced and progressive interaction. By making these conscious choices, we contribute to a healthier, more equitable digital ecosystem. Jean, how do you suggest that AIs be addressed to avoid that gender bias? Are there pronouns that you think are more appropriate? It's a thoughtful question. Using non-gendered pronouns like it or they can help avoid reinforcing gender biases. These pronouns emphasize the AI's nature as a tool rather than a person.
26:46Additionally, ensuring clear AI identification, like stating I am an AI, helps set proper expectations. What do you think, Jessica? Should we also focus on educating users about these nuances to foster better understanding and interaction? And so some of what you see here in Gene's response is actually guidance we have implemented, both as a marketing team, but also in what we call development excellence, which is some of our frameworks for our coders when they are developing bots and assistants, both internal and external applications of AI for BCG. one additional stance we have taken also is to try and lean towards naming that indicates roles or assistants rather than naming taking human names for bots which is also I'm sure a question every developer has and the first reaction is always to try and choose a name of somebody something you think is fun so you think just Amazon Alexa where they choose a very personified name when they chose Alexa.
27:52And actually, it's a very personified voice as well. I think as an ecosystem, it's not the right decision to make because it's not educating users right. But I also think it's hard as a developer. When you do user testing on voices, people actually really prefer female voices and male voices. So you need to fight the user to some extent. And so I would love to see ecosystem roles around that. that's interesting and gene is i'm sure intentionally is gender neutral name it could be either male or female although the voice gene your voice sounds more female than male even if it's clearly a robotic voice but maybe jessica you're right people prefer hearing a female voice just as we all have mothers.
28:44The thing that concerns me more are the humanoid robots, not the more serious ones, but Sophia comes to mind that's been trotted out all over the world, which is an attractive female. And that to me is, maybe that's inevitable, but it's a dangerous direction go because there are a lot of fairly misogynist males out there that would love to have a female robot that it could boss around. Gene, what do you think about the embodied image of AI, whether that should be male or female, or as with Atlas from Boston Dynamics, clearly non-gender. The embodiment of AI as either male or female can indeed perpetuate biases and stereotypes.
29:41Opting for a non-gendered representation like Atlas helps avoid these pitfalls. It's crucial to focus on clear AI identification and transparency to ensure users understand they're interacting with a machine, not a human surrogate. This approach fosters healthier, more equitable interactions, and aligns better with ethical guidelines. Jessica, how do you see the broader industry moving towards this non-gendered AI representation? Gene, you can stop asking Jessica a follow-up question at the end of each response, please. Go ahead, Jessica. I 100 % agree with Gene, but Gene and I have had this conversation before.
30:22I guess there's a reason why we agree. But I agree with you, Craig, on a humanoid robot. I do think we underestimate once AI will be pervasive and you will be interacting with voice commands with lots of your day-to-day devices, how much that decision of a voice is going to impact. And if we think of what happened to OpenAI with their first voice assistant, that they actually lean towards a super female, sexy, I worry about that becoming ubiquitous in the industry. ranges. And I say that because I'm a mother of two little daughters. You think small kids, when they'll be interacting with these devices all the time, what it's going to do to the way they think and apprehend the world, which systematically all robotic assistants are female gendered, I'm really concerned by that.
31:19Where do you see this going, Jean? Because as Jessica says there are going to be conversational agents, whether embodied or not, all around us within a decade, I would imagine. And children are going to grow up with them in the way that children of the Gen Z generation grew up with social media. We didn't do a very good job of paying attention to the negative aspects or influences of social media on children. Do you have any thoughts of how we might ensure that children are nurtured in a healthy way with conversational agents in their lives? The future of conversational agents surrounding children is indeed a concern.
Read the full transcript
32:10To ensure they nurture children in a healthy way, it's crucial to design these agents with ethical guidelines. Transparency about their AI nature and promoting balanced interactions without reinforcing stereotypes are key. Educating users, especially young ones, on the nature of these interactions will also help. We must learn from the past missteps with social media and prioritize responsible AI development. One of the things that I'm actually concerned about for working with a lot of corporates on helping them develop AI applications is that question of that representation of AI is actually not top of mind.
32:48I actually think there are very few people in the ecosystem that are thinking proactively about it. And OpenAI themselves obviously did not. So I do think it's one where there needs to be a stronger thought leadership action so that there are just ecosystem guidelines or that the regulator steps in, but it doesn't look like it's happening organically at the moment at all. Yeah, it would be interesting to know. I don't know, Jean, whether you know whether there's any data, but a lot of systems allow you to choose the voice. It would be fascinating to know how many users choose a female voice compared to how many users choose a male voice.
33:34Unfortunately, there isn't comprehensive data on the preference for female versus male voices in AI systems. However, anecdotally and through user testing, it seems that female voices are often preferred for their perceived friendliness and approachability. This preference, though, can perpetuate gender biases if not addressed thoughtfully. It's crucial to balance user preferences with ethical considerations to avoid reinforcing stereotypes. Designing AI voices with a neutral or customizable tone could be a step towards more equitable interaction. Yeah, that's interesting. Can I just ask Bill and Jessica, maybe each of you can answer when you're interacting or when you're producing a podcast, do you edit out the latency?
34:21And is there, do you think that latency will remain because of the current architecture? Or are you working at all on reducing that? Yeah. So latency goes down all the time. There are actually some new technologies out there. There's things like these new TPU processors and something called Grok, not the Elon Musk Grok, but another Grok company that does extremely fast, low latency. There's often a trade-off. So there was a demo recently where someone did a super low latency voice-to-voice example, but you'd usually have to use much smaller models to get that really low latency. And then the responses aren't necessarily as useful.
35:04So there's a balance of using the most useful model with the lowest latency. The models are getting faster all the time. Like I said, when we first built this, it was like 45 seconds to a minute of latency. And now it's down to a couple seconds. When we do the radio or when we do the podcast recording, we do edit. We edit out some of the latency. We also edit out some of the, just from humans and machines, as we say, we edit the whole podcast. Yeah, sure. But when we do live events or if we're doing something like this, you just see the latency as it is. And we find that it's useful at the moment.
35:36And we don't want to use a less capable model in order to get better latency because we don't want to sacrifice the output yet for that. Yeah. And how is the uptake or how is the interest among BCGX's customer base? Are people actively building conversational agents? Is there any estimate or any sense that you have of everybody is doing this now or people are still experimenting on the margins and it's going to take a while before it's widely adopted? My sense is everybody is building some shape and form of agents. I just not always conversational, especially what I have seen today, mostly with corporates is when they are exploring internal use cases, for example, for coding, for marketing content development and co, they may need to use writing, typing interactions that you would use conversational agents for external services, which I think is a mistake.
36:46I think for internal operations, the power of conversations is actually super, super high. And it's really the future of internal processes. Yeah, that's right. And you begin to wonder what will be lost in that transition. I'm horrified by my kids' handwriting, for example, because they never write. And who knows what once people stop actually having to write, to interact with conversational agents, what will become of that? Jean, do you have a view on that? As convenience grows, certain skills may wane. It's a trade-off, efficiency versus maintaining traditional abilities. Balancing these changes with educational practices might help preserve essential skills.
37:35The key is thoughtful integration, ensuring technology enhances rather than diminishes our capabilities. I will say there was probably a time in the not-too-recent past where almost everyone knew how to make a fire with flint and steel. You're absolutely right. It may be those skills don't need to be preserved because they become obsolete in a technological world. Okay, let's leave it there. Gene, I'll give you the last word. We're signing off from the Eye on AI podcast. Is there anything you want to say to the listeners? Thanks for tuning in to the Eye on AI podcast. It's been a fascinating discussion, exploring the nuances of AI, gender representation, and the future of conversational agents.
38:21Remember, as we integrate AI into our daily lives, it's crucial to balance technological advancement with ethical considerations. Until next time, stay curious and thoughtful about the future we're shaping together.
38:58And with recent updates, they're now making 25 % fewer mistakes than Microsoft, which means you're getting some of the best accuracy out there for things like customer service bots or voice-driven assistance. Plus, it works in over 50 languages, delivering results in just 700 milliseconds. If you're working on an AI project, it's definitely worth a look. Check it out at www.speechmatics.com slash realtime. That's www.speechmatics.com slash realtime.
From the publisher
In this episode of the Eye on AI podcast, Craig Smith sits down with Bill Moore BCGX to explore the development of Gene, an AI-powered co-host transforming the way we engage with conversational AI.
Bill walks us through the journey of creating Gene, a real-time conversational agent built to interact seamlessly with human hosts. He shares insights into the technical challenges they faced, from improving latency to expanding context windows, and how Speechmatics' cutting-edge speech-to-text technology helped achieve over 90% accuracy in real-time.
We dive deep into the ethics of AI, particularly gender representation, as the conversation touches on why assigning gender to AI assistants can reinforce outdated stereotypes. Gene, as a gender-neutral conversational AI, challenges these norms, providing a clear distinction between human-like interaction and AI-driven dialogue.
Join us for an in-depth discussion on the future of AI in media, the balance between human creativity and AI's analytical power, and the ethical considerations every company should be mindful of when integrating AI into their operations.
Don't forget to like, subscribe, and hit the notification bell for more deep dives into AI innovation, ethics, and the future of conversational agents!
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Introduction to Gene: BCG's Conversational AI
(02:19) Bill Moore's Background
(04:22) Early Challenges in Building Real-Time AI
(07:20) Technical Aspects: Prompting and Configuring Gene
(09:28) The Use of Vector Databases and Sparse Priming
(12:01) How Gene Handles Long Conversations
(14:46) Building Conversational AI
(19:07) Gene Introduces Itself: AI as a Podcast Co-Host
(22:00) The Ethics of AI
(29:35) Should AI Have a Physical Representation?
(32:05) AI and Children: Nurturing Healthy Interactions
(36:16) Adoption of Conversational AI in Businesses
(38:11) Ethical Considerations for AI in the Future




