The Rise and Plateau of ChatGPT's App Revenue

29 Mar 2024 · 21 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI Today - Episode Summary: The Rise and Plateau of ChatGPT's App Revenue

Podcast Overview

  • Title: AI Today
  • Description: "AI Today" explores the latest advancements and ethical considerations in artificial intelligence, aiming to make AI accessible and engaging for a broad audience.

Episode Details

  • Episode Title: The Rise and Plateau of ChatGPT's App Revenue
  • Episode Description: This episode discusses the peak revenue of ChatGPT's app at $4.5 million per month and the factors contributing to its recent slowdown.

Key Themes and Discussions

Guest Introduction

  • Background of Guest:
  • Former Principal Scientist at AWS.
  • Transitioned to head of AI at Toloka.
  • Initial interest in AI sparked through applied mathematics and numerical methods.

Journey into AI

  • Early Work:
  • Engaged with support vector machines for plasma physics.
  • Developed a passion for automating complex discoveries using AI.
  • Philosophy Shift:
  • Shifted focus from mere automation to ensuring AI benefits humans, maintaining human involvement in decision-making.

The Move to Toloka

  • Reason for Transition:
  • Sought to make AI models more useful by integrating human feedback into the development process.
  • Recognized the value of crowdsourcing as a means to enhance AI development.

Success Stories in Human-AI Collaboration

  • Moderation Example:
  • AI models filter inappropriate content but require human oversight to ensure accuracy.
  • Amazon Go Store:
  • Cashier-less shopping experience monitored by AI but still requires human intervention for complex cases.
  • Yandex's Personal Assistant:
  • Early testing involved human assessors answering questions posed to a chatbot, illustrating human oversight in AI training.

Ensuring Quality in Crowdsourced Data

  • Quality Control Mechanisms:
  • Entry tests and training tasks for assessors to ensure high-quality data labeling.
  • Use of "honeypots" to identify and weed out unqualified contributors.

Ethical Considerations in AI Development

  • Discussion on Generative AI:
  • Concerns about the concentration of power in AI model development.
  • Importance of diverse data representation and the risks of bias in high-performing models.
  • Open vs. Restricted Model Development:
  • Advocacy for openness in AI development to allow diverse groups to contribute and align models with varying societal perspectives.

Future Outlook for AI

  • Industry Predictions:
  • Increased accessibility and decreasing costs for model development over time.
  • Emergence of diverse models tailored for underrepresented groups.
  • Regulatory Considerations:
  • Caution against overly restrictive regulations that could stifle innovation.

Advice for Companies Implementing AI

  • Focus on Responsible Development:
  • Importance of thorough evaluation processes to prevent biased or risky outputs from AI models.
  • Emphasis on proactive evaluation during model development to ensure ethical standards are met.

Conclusion

  • The episode concluded with a reiteration of the importance of ethical considerations in AI development and the ongoing need for human involvement in the AI decision-making process.

Additional Resources

  • Get on the AI Box Waitlist: [AIBox](https://AIBox.ai/)
  • Join the AI Facebook Community: [Facebook Group](https://www.facebook.com/groups/739308654562189)
  • Podcast Studio AZ: [Studio Link](https://podcaststudio.com/mesa-studio/)
  • Podcast Studio Network: [Network Link](https://PodcastStudio.com/)

Final Thoughts

  • The discussion offered valuable insights into the evolution of AI, the role of human feedback in AI development, and the ethical responsibilities of AI creators. The host and guest underscored the importance of transparency and collaboration in shaping the future of AI technology.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:28Welcome to the AI Chat Podcast. bridging kind of the gap between machine learning and human insight. Welcome to the show today. Thanks a lot, dear. Thanks for having me. Thanks for the intro. Super excited to have you on the show. I wanted to kick this thing off and kind of ask you a little bit about your journey. And if you could walk us through, you know, what that looked like. You were a principal scientist at AWS, and now you're kind of the head of AI. Did you always know that you were interested in AI and kind of this tech space, you know, growing up? Or is this something that you discovered as you were going through school.

1:01Tell us a little bit about your background. Well, actually, to ask you specifically about AI, I was doing some applied mathematics and numerical methods. And then I found a scientific supervisor who was like, hey, there's this machine running thing, and it's pretty cool. And I was back in 20 or something. And everybody was using sports factory machines at the time. And we found a really cool application of support vector machine to plasma physics. And, wow, this is a really neat way of automating really complex discoveries, which is just do. And the journey started, and I did PhD in computer science and worked in a few companies, Amazon, Microsoft, AWS.

1:49But what's also curious, and I actually hope that machine learning folks can hear me because I think that's worthwhile. Well, at some point, I realized that everybody in the AI is about automation. Everybody's like, okay, let's automate this, let's automate that. But the real, like, the real beneficiary out of all of this showed at the end of the day, be humans. And how do you make the AI really beneficial for humans? How do you make sure that humans stay in the loop when the AI makes decisions? You know, how do you make sure that the AI makes decisions which are useful? and this is the area where it kind of shifted a few years ago, actually about five or six years ago.

2:28And in AWS, I was working in the same area as well. And so, again, from AWS to Taloka to progress in that area of Kimmel and Lopin. Now I'm headed behind two scientists. That's super cool. I'm wondering, could you tell us a little bit, it sounds like an incredible background, and you've been working in this for a long time. So you've seen this whole thing evolve a lot. Can you tell us a little bit about what got you interested in Taloka moving? You obviously probably had a great job over at AWS. I know a bunch of folks that work there. And it's a fairly good place to be, whatnot. What kind of got you interested in moving to this kind of new company and this new idea?

3:15Yeah, yeah. Actually, happy to share this. uh well imagine that you're doing you're doing artificial intelligence and you're thinking well how do i make it useful how do i make my models and then we you very often before the age of large language models you often had to collect um label data you had to say you know whether which is good which is bad and and you had to go through this process of collect collecting data which is getting feedback from humans and very often you see in the papers where um it's hard expensive uh it's a manual process where you have to work with humans and not math and algorithms and for tech people sometimes it's tough and then you realize that there is this area which is crowdsourcing which is basically like you have humans all around the world who some of them have nothing to do some of them you know really want to earn a little bit of money on the side some who want to earn a fair amount of money.

4:15And then you think, how do I include them in the process of making AI? And so this is where Toloka, in my opinion, excels and is really incredible. It's better than many, many others, if not better than everyone, is doing this job of involving humans around the world into the AI making process. And this becomes a much more technological problem than the problem of normal human collaboration or all human management. So it's a technological problem to collecting, working with humans, a technological way of working with humans, a technological solution to that problem. And this is what is really cool, I think, to look, it's a technological platform for dealing with humans and contributing to the AI.

5:06Super fascinating background. I love what you are working on. I'm wondering if you could share some success stories where human in the loop, you know, processes significantly improved AI performance. Yeah, let me let me recall a couple of stories from the background. One thing which is certainly worth noticing is that humans are always present in improving AI performance. Use are basically crucial. This is like how AI is built. before lms you used um humans as i mentioned to um provide ground truth labels on which i straightened and says this is good this is bad humans were those who fight it now with lms you still have the same thing you have with alignment process where you ask lms to produce results which sound better to humans which actually reply to human curry so without the human involvement, you very early get to slowly up.

6:01But if you think about the human envelope in the usual way people understand it where the models and machines at the same time work together to achieve like a better benefit. The most popular example is probably the moderation example. The example where let's say there's a lot of user content like texts to some forms or maybe some images are submitted to a website and then you need to filter out in the always in real life in real time um bad images inappropriate images or bad content right so you usually build a model and then very often the model doesn't get you high doesn't give you high enough i cares and so you trust the model when the one is confident you don't trust the model when the model is most confident you rely on a few ones to fill this gray area okay like skis that's this one is now the more fun examples are i'd say if you think about say Amazon Go.

6:55It's Amazon Go, if you're not familiar with it, if you haven't heard about it, it is a cashier-less store. So you enter the store, you take whatever you want, and you leave. The way it works is that there's a lot of canvas in the ceiling of that store, and they track the person who enters, they see what they take. Right. The way it works is that there's machine learning which tracks you, which detects what you took, and then basically send your receipt afterwards to the new store. So this happens very quickly after you leave the store, you get the receipt most of the time. Sometimes if you post your product to someone else or if you try to trick the system where you give it to someone else and leave it with that product or you try to put the product in one place and then you take it from the wrong place, it took, a few years ago, it took quite some time before you get the receipt.

7:49And the reason for that, at least the viewers say, is that the complex cases were monitored, were actually observed by humans, and then they were checking. Oh, okay. That's a more of a fun example. Then there's a lot of fun examples. When Yandex was building personal assistants, it's something similar to Amazon Alexa or Google's personal assistant, Alice. So before they even released, you know, just when they started collecting data, basically the engineers were given this chatbot, like a preliminary version of a chatbot to ask this chatbot questions and then the chatbot would give them answers.

8:30And the engineers were talking about that chatbot thinking that they're talking to some really beta version of the model. And then someone decided to find out, oh, where does this model host it? What's the computer? And then they didn't find the computer when the model was hosted. And they figured out that the whole process was actually built in a way that when the engineers asked the question, it's the real people, assessors, who answered those questions is the process of data collection. That's another fun story of how human, instead of AI, helps in getting used to things. That's super cool.

9:05Yeah, those are awesome stories. Really interesting. I had no idea, so that's some cool insights. I'm wondering, how does Toloka ensure the quality of human-labeled data across such diverse languages and countries and all that? Yeah, sure. That's a good question. So, as I mentioned before, so Toloku approaches this way of the problem of getting data or interacting with humans through a technological perspective. basically think about what kind of abstractions what kind of tools can help you get higher quality data and so there is actually quite a variety of popular crowdsourcing tools which which help get a higher quality data so first you give you give people entry tests some tasks where you don't want people to waste their kind you actually give some sort of fairly simple music tests for how about liquid quality do they know english they know spanish do they know on a server then you want to actually get certain high proficiency you get people training and then there is the process of actually giving people training it's not a manual process you design training tasks ask training tasks for people where you say um this is a task please perform it and then say perform it and say no you're incorrect we chose this answer but you should have chosen that answer this is why And so people learn this week and then they get through a second test of exam, which is, of course, higher quality or ensures that you can perform, you know, the person can perform at a level.

10:39And so then the person is allowed to perform real tasks. However, there are some people, of course, who are trying to cheat through the system and who are trying to get through training and exam with themselves, and then they try to automate or not really get cancer. So to avoid this, we have other tools, which is, we actually have so-called polypots, where basically somewhere in the middle of your answers, there's a question with a non-answer. And if the person answers incorrectly too many times to those kinds of questions, they're not allowed to move the task in. And so incorporating those honeypots into the process of labeling, into checking, regular checking performance is a crucial part of high quality labeling.

11:24Those honeypots are built by more trustworthy people, most trustworthy assessors who have been tested by themselves by even higher, more trustworthy experts. So basically you have this pyramid of quality

11:43quality feedback providers. And so the more expert fees you require from the provider, usually the rare chance you have to find them. And so as a result, you have a lot of people who have to be checked by trustworthy people and those have to be checked by experts. So this is a way we approach our getting actors. Okay, very interesting. That's a problem I had not really thought about the solution to solving It's fascinating how you guys are currently working on it. You know, with all of this that you're currently working on, what do you think are some of the ethical challenges that, you know, maybe yourself, but also the AI community needs to address urgently?

12:22What's kind of on your radar right now? So, to me, there's a lot of discussions right now about, you know, generative AI, about large language models, about putting some sort of government restrictions on who can provide it. models in my opinion, I can build those models. In my opinion, what's quite problematic is that right, kind of the reason that much of a variety of models, which are high performing and which are, which can be easily representative of the various groups of people. Current high performing models are basically based on, right. currently buys data and data for which, what is good data, what is bad data is determined by a very small group of people who quite often do not represent some other groups, which are interesting, some other groups which are important.

13:20And in my opinion, what's worth doing is actually being real open about how you construct a task to watch language models or these foundation models, sharing instructions or sharing the actual data for fine-tuning or aligning those models in a way that other interested members can build the models which are less wise to that specific small model. So that's quite important because otherwise it's influencing the whole society. You've seen the amount of people who started using chat GPTs, quite a lot of people. The influence of what kind of answers the models give is quite significant. Yeah, yeah.

14:02Yeah, no, I think that's super important. It's something I definitely think a lot about because, of course, right now we have, like you mentioned, these large language models. We have OpenAI, we have Google and Meta, and really all of these companies coming up with these, it's kind of a small bubble around San Francisco, really. And we all know that San Francisco has a lot of its own ideas. Some people love some of them, some people don't love all of them. But there are definitely different ideas all around the world that different people subscribe to or don't subscribe to. And so, yeah, having one small group of people create these essentially large models that like for Google's, you know, they're integrated in their search.

14:38This is going to be seen by billions of people around the world. Where do you see this going in the future in regards to that? Do you think that there's a place where someone like OpenAI or Google or Meta could create a model everyone would want? Or do you think, you know, at some point people are just going to say, look, I don't align with these large models on X, Y, Z topics. I'm going to find maybe like there's thousands of models and this is kind of the one that more aligns with my beliefs. Do you see the industry going in that direction? So that's a great question because people are betting either the model of open source or the open source is going to enroll and what's going to happen is models available or betting that, hey, Gen AI is the next cloud.

15:22It's a great question. I tend to be on Elicamp, which thinks that over time, if you look at the history of machine learning, things like building models became cheaper over time. It always becomes cheaper. It always becomes more accessible to quite a large variety of different groups. And the information on how to build those models also leaks because people need one company and get another. So I really hope that the future is going to be that the technology is going to spread and more people are going to be able to build those kind of models and tune them to whatever needs they see, whatever underrepresented groups they see in the society so that you have a diversity and not basically a single or a few heads of pity about this topic.

16:19So this is my hope. I do think that the openness helps here but I do think also that it's worth being careful about opening everything it's worth putting some sort of guards or government laws against against disuse or against possible ways these models can be used for some legal purposes so it's certainly worth doing that But I did hear in the industry some quite radical in my opinion propositions of, say, allowing the models to be only created by very, very specific list of companies or institutions. And that to me seems quite a risky path to take. Yeah, yeah. I 100 % agree. It's interesting.

17:11you know um open ai of course m altman famously went to congress in the united states and said we need to regulate ai i'll help you write the regulation we need to decide who gets to you know get a license to train the models and of course a lot of people raise red flags about that when you know the biggest company with 10 billion dollars says that they want to be the one to help give out the certificates of who can and can't do it so i think there'll be pushback of course there's the open source community um i think a lot of the stuff meta has done has honestly it's kind of interesting i'm not usually like a huge fan of meta and facebook and whatnot but some of the open source stuff they've done with uh ai i think is really interesting and i got it i just have to you know give them credit where credit's due on that um and even a lot of the infrastructure they built so very cool see like kind of what we're seeing in this space i think inevitably whether there's regulation or not like if there's too much regulation people are just going to get like bootlegs open source model.

18:06It's going to have them on a thumb drive and it's going to be your open source model that doesn't have the open AI guardrails or whatever. I don't know. It'll be interesting to see where it goes. In any case, I guess based on your really extensive experience you've had in this space, something I would love to ask you about as we kind of wrap up the show is what advice would you give to companies looking to implement AI solutions responsibly today? I'm glad you asked about the responsibly part because Because that's, I think it's quite important nowadays with the models, don't only give you the binary answers.

18:39This is good, this is bad. Your mistake is costly for the binary models, but it's not nearly as costly. It can be not nearly as costly as the mistake for generated models, which can significantly offend someone who can provide some regularly unpleasant or illegal justifications or descriptions. and because of this larger impact of specifically the generative models because of how much risk they have, I'd say that it's really important if you're developing some sort of a product based on GEMI, not to save on evaluation, not to evaluate on your users, say, okay, well, just tell me what's good and what's bad, really investing in value.

19:31You should really think about in advance how you're going to judge whether the model is not providing risky answers, not providing some unbiased opinion which reinforces some sort of societal problems, but really you know provides great experience with getting real truth instead of making up facts or hallucinating what they call. and so it should be done in my opinion really early in the process of model development so i'd say that this is this is if you asked me about single advice i'd say that evaluation here you're worth it um and it it proves that you you take things responsibly you actually think about what you built and and you're sure about it and you you can demonstrate I love that.

20:22I think that's some really solid advice. Fedora, thank you so much for coming on the podcast today. I really appreciated a lot of your advice and insights. It's really refreshing because I get a lot of different perspectives on this show. And I really think you're on the right track with what you guys are working on, what you're doing and kind of your philosophy behind it. So that always makes me happy to hear. if people are interested in getting in contact with you or Toloka and finding out more about what you guys are building and working on what's a good way for them to do that? Yeah, they can connect me on LinkedIn or I think that's the best way just find me on LinkedIn and you can connect Okay, and thanks for having me It was a great conversation It was interesting and on and active Yeah, it was super enlightening So really appreciate you coming on To the listeners, thanks so much for tuning in to the AI Chat Podcast.

21:15Make sure to rate us wherever you get your podcasts and have a wonderful rest of your day.

From the publisher

In this episode, we explore the trajectory of ChatGPT's app revenue, reaching a peak of $4.5M per month and the factors contributing to its recent growth slowdown.

See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

More from AI Today

All 897 episodes
The Rise and Plateau of ChatGPT's App RevenueAI Today · 21 min
Listen in VO