#132 Scott Downes: Navigating the Language of AI & Large Language Models

2 Aug 2023 · 1 h 4 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. - Episode #132: Scott Downes: Navigating the Language of AI & Large Language Models

Episode Summary In this episode, Craig Smith interviews Scott Downes, Chief Technology Officer at Invisible Technologies, to delve into the world of large language models (LLMs) and their application in various domains. They discuss the role of Reinforcement Learning from Human Feedback (RLHF) in training these models and how it can potentially transform the workforce by shifting focus from manual labeling to more complex AI-driven tasks.

Key Topics Discussed

  • Introduction to Scott Downes
  • Background in technology and literature.
  • Experience in startups and the evolution of tech roles.
  • Large Language Models (LLMs)
  • LLMs' utility across various problem domains.
  • The internal terminology at Invisible, referred to as "Ironman suits," emphasizing the enhancement of human trainers' capabilities.
  • Reinforcement Learning from Human Feedback (RLHF)
  • RLHF as a crucial element for training LLMs.
  • The shift from bulk labor to specialized, targeted human input.
  • Human Workforce Dynamics
  • The transition from traditional labeling roles to more sophisticated RLHF positions.
  • Insights on the future workforce in AI, emphasizing high judgment roles.
  • Applications of LLMs
  • Use in text cleanup, product classification, and improving customer experience.
  • Examples of successful implementations in e-commerce and on-demand delivery sectors.
  • Challenges of Hallucination in LLMs
  • Discussion on the concept of hallucinations in AI and the complexity of teaching models to distinguish truth.
  • Optimism around RLHF's potential to address accuracy and alignment issues.
  • Future of AI and Human Creativity
  • The impact of AI on creative industries and the potential for increased appreciation of genuine human creativity.
  • The dual role of AI in enhancing productivity while fostering a greater demand for authentic human expression.

Segment Breakdown

  • (00:00) Preview and Introduction
  • (01:33) Generative AI's Dirty Little Secret
  • Discussion on the human role behind generative AI.
  • (17:33) Large Language Models in Problem Solving
  • Exploration of LLMs and their diverse applications.
  • (23:24) Large Language Models and RLHF Challenges
  • Addressing the challenges related to RLHF and hallucinations.
  • (30:07) Teaching Language Models Through RLHF
  • Insight into how RLHF teaches models through feedback.
  • (35:35) Language Models' Power and Potential
  • The capabilities of LLMs in solving complex tasks.
  • (53:00) Future of Human Workforce in AI
  • The evolving nature of jobs in the AI landscape.
  • (1:03:10) AI Changing Your World
  • Final thoughts on the transformative potential of AI.

Key Takeaways

  • The conversation emphasizes the transformational role LLMs can play across industries, not just in automation but also in augmenting human capabilities through RLHF.
  • Scott Downes advocates for the value of human expertise in fine-tuning AI outputs, ensuring that AI serves to elevate human judgment rather than replace it.
  • The need for ongoing adaptation and learning in both AI and human roles is underscored, with a focus on the dynamic nature of technology.

Conclusion The episode offers a deep dive into the complexities of using AI in practical, impactful ways while also considering the broader implications for society, employment, and creativity. It encourages listeners to stay informed about the rapid advancements in AI and their potential to reshape our world.

For more details, listeners can find a transcript of the episode on the [Eye on AI website](https://eyeon.ai).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00AGI is kind of a big concept to take it down to a smaller level and just say like, is this generally useful across problem domains. And it turns out that large language models are phenomenally useful across problem domains in ways that we wouldn't have anticipated. I certainly didn't anticipate a few years ago. We have an internal terminology for this. We call it giving our agents Ironman suits. What we want to do is maximize the impact that a single trainer can provide. We want to find ways that we can take advantage of high judgment, really intelligent, sophisticated individuals, there's a confidence that there's the problems related to hallucination or are more focused on expressing things than whether there's an internal model that actually matches something that we would describe as real.

0:44Hi, I'm Craig Smith, and this is Eye on AI. There's been a lot of talk about the human trainers behind large language models, how despite oceans of data and racks of GPUs, generative AI still depends on people in much the way that supervised learning depends on people to label data. This week, I talked to Scott Downs, Chief Technology Officer at Invisible Technologies, whose platform is used for reinforcement learning with human feedback, otherwise known as RLHF, the dirty little secret behind generative AI. I hope you enjoy the conversation, and I think you'll find it enlightening. I'm Scott Downs.

1:33I'm the CTO at Invisible Technologies. Where to start? I'll talk a little bit about myself. I'll say that I'm sort of, I think you might hear a lot of people say this, I'm a different kind of CTO. I'm someone with a lot of interests. And one of the reasons why I love startups so much is I've loved being in the space where you wear lots of hats. So I'm a CTO who, my first paid job when I was 14 was writing code, but then I went off to college and studied literature. And in my 20s, I worked in radio, I worked as a musician, I worked in design. And I think I did pretty much everything I could to stay out of sort of what I perceived at the time as sort of an uncool, nerd factory type of jobs where I would just sling code.

2:32But with the dot com era, all of a sudden, it actually became really valuable to have a breadth of interest. And the impact of computing and the scope of what you might do in like a software development job, just spread spread open in a way that created lots of opportunities for folks like me. I remember, like at a college reunion, I looked around the room, and this is a very 90s thing to say, but like all of my classmates, we all had goatees. We all had degrees in like English and theater. And we were all working for dot coms. We were all working in tech jobs. And I think that that energy has sustained for me this idea that it used to be that working in technology might mean producing a piece of software that you put in on a shiny disc and put in a cardboard box and sell at CompUSA.

3:32I've done that. And then it evolved more and more to where the practices, the ceremonies and process of software development became so pervasive that it's actually what everybody does in every company. So there was a time when normal business teams didn't have stand-ups, that the idea of a sprint or an epic was unique to maybe a software development team in a company that's building software. And I think that one of the things that's been interesting and exciting for me really is that over my career, seeing that change has meant that those of us who work in that space find more and more opportunity for broader impact.

4:18And the way that I say broader impact what I really mean is like does my mother understand what I do yeah that's that's the final metric can I explain to my mom so I said a lot of things yeah and then uh invisible how did you get to invisible and and what does invisible do sure and I think um that tells you a little bit about who I am, but in terms of the, my more, like my adult journey, once I found technology is my calling, I worked, I've worked to build out a SaaS platform for what was originally an SMS aggregator, but eventually became a mobile marketing software company. And so I experienced the sort of traditional enterprise SaaS world where we built solutions for large companies whose names you would know, you know, the American Express, Comcast, AT &T type of customers.

5:16And I found that maybe it was a sort of a last gasp at that moment of

5:26expecting your clients, your customers to learn complex technology systems and take that on board. They might need to learn 20 different software packages. Of course, that's a source of a lot of pain. And after that, after that experience, I was helped to build a company that was a marketplace for residential moving. So sort of an Uber of residential moving. And as the CTO in that company, what I found was that I was selling a product to consumers where the technology powered things behind the scenes and managed a workforce, but it wasn't presenting a lot of sort of digital surfaces for people to interact with.

6:09So our workforce used iPhone apps and Android apps, but our clients would just order a service and somehow the technology behind the scenes was supposed to sort of deliver the optimal moving experience. And we did pretty well at that. Like there are a lot of things that you can do by making sure that you have high quality workers showing up on time and following a great plan on how to execute. So that's what brought me to Invisible was sort of, I was at a moment where I found both of those adventures of like building a workforce management platform for a marketplace. How interesting is that? That's really cool.

6:52That's a complex problem space that has real impact on the world. And then meanwhile, I also am somebody who enjoys building complex, sophisticated software platforms that solve really, really complex problems. And Invisible, to me, immediately appeared to be, from a product perspective, the perfect intersection of those. So that's a long way to go to tell you what invisible does but what invisible does is a blend of a workforce management platform and an automation sas platform and the form that that takes is that we we see everything everything starts and ends with a process and when we say a process the kind of process that we work on for our clients one that we'll probably talk about a lot today is data training so that's a service that we provide for clients of ours that we define in terms of a process that's mapped onto a visual canvas and then executed in the form of steps, some by people, some by automations or scripts or integrations with third-party platforms.

8:03So I know that that is very horizontal and generic as a description. We're stubbornly horizontal as a company. So what we think of ourselves as being really good at is taking on problems largely related to scaling up areas of a business that are particularly in need of, well, a combination of consultative engagement, maybe a significant amount of labor, and a significant amount of automation to optimize the solution that we're providing. So some examples of that kind of space where we thrived, we really grew up as a business during the early pandemic era by working really closely with on-demand delivery companies.

8:56So it's a great example for us where you look at a business that's under a unique moment of pressure and prevailing alternatives for how to solve that problem never quite fit. So if you're DoorDash and you're looking at figuring out how to encode restaurant menus into your platform, one approach might be to go with a traditional outsourcing firm. So this is the sort of problem space where we think of deploying the kind of proverbial army of people? What if we had a thousand people around the globe or maybe in emerging markets where labor costs are lower, grinding away on a problem? And the solution that we'd be providing in that case is human scale.

9:48There are limitations to that approach. That approach tends to be, sometimes there are challenges related to quality. And there's really very few incentives in that kind of a structure if what you're selling is bodies and what your business model is based on is kind of labor arbitrage. Let's sell hours and add a few dollars on top for us to run our business. That doesn't really create aligned incentives. So our incentive as a vendor in that model would just be to drive up the number of people and drive up the number of hours. which is not really super positive for those companies. On the other hand, you might build a fully automated solution.

10:31So you might look at the moment of like, how do I get a million restaurant menus into a proprietary platform as a task that is driven by OCR tools, web scrapers, custom scripting, custom code, RPA tools, maybe tools like UiPath. There are lots of ways to solve that problem with a pure tech approach. Those tend to be really effective once they're in place, but they also tend to have kind of a long runway to build. And they also tend to be sort of fragile. So it's not typically easy to build a fully automated solution and then maintain it over time without having a standing team. And so we sort of live in that sweet spot between those two by saying, OK, you know what?

11:22We're going to take your problem. We're going to start working on it today. We can apply some amount of labor to solving the problem once we've mapped it, once we put it in a clear form that we can execute. And we can imagine, you can imagine a workflow on a canvas and having a Sharpie. you might kind of draw a circle around a specific area and say, hmm, this area is requiring a lot of labor. It's costing us a lot of money. This is an area that's ripe for optimization. So what if we applied technology solutions in that space? And I think that what we really do is sort of the common sense running of a scaling part of a business, but we make it explicit.

12:07it. So there are plenty of cases where our clients might see their problem and not have the same kind of precise instruments to address it that we do. So I'll give you a concrete example. Say that you've implemented a new policy or the way that you execute a certain process. Got new business rules. Okay, now when insurance claims are based in Florida, we have to do the following three things. You can imagine that a lot of businesses are in a world where they might send out an email, they might contact managers, they might write up new SOPs, but enforcement becomes a very big problem. So how do we know that people follow new rules as they are established?

12:54So for us, being able to build all that into a platform that manages the interfaces through which people engage means that a change in policy can be implemented perfectly 100 % in a day. So that's kind of the background of our business that kind of brought us into this moment where we have a very flexible platform. We're very responsive to client needs, and we've seen opportunities in the last few years that have just kind of gone crazy in terms of applying that sort of approach of blending people and automations in an explicit but very configurable process. While it turns out that's really useful when you think about data training for AI.

13:42Yeah. And this is something that's fascinated me for a while. the idea that first of all was supervised learning unseen to to most of the people who were using it is this army of of labelers human labelers spread out across the world that are you know segmenting and and clicking on computer images or on text and and labeling them very tediously uh even as that process their their automation tools to speed that process it still comes down to a human now we're in the age of unsupervised learning with um these big transformer based models uh yet again there's this army of humans in the background who are helping train this model to, to these models to behave in particular ways.

14:50And so this is one of the things, and I understand you're horizontal and, and, and do a lot of different things, but that's what I wanted to talk about, this human element, and in particular with regard to large language models. Can you talk about how many people are employed and in your case are using the invisible platform to do this? how many are employed, how specific the tasks are, and how rapid is the progress in the underlying models when you're doing that. Or maybe just start with what we're talking about really is reinforcement learning with human feedback. Maybe start with talking about you actually do reinforcement learning with human feedback and then the numbers of people involved in the progress.

15:56It's a big part of our business, especially we've seen a really huge increase over the last year. And we employ, we have a little bit of a different model in some ways, but I think that the simple translation is to say that our human operators who work on our platform. We call them agents, are human operators who are working on both reinforcement learning with human feedback and other sorts of processes. There are thousands. And I think the thing that I would say about the nature of the work that we're doing in RLHF is that it's actually not sort of that bulk giant army of folks anymore that's as relevant to the problems that we deal with with our specific clients.

16:50And I can't go into any huge amount of detail with the clients that we're working on or working with, but we do work with OpenAI and several other household names to do this sort of work. And I think that what I've seen and what I've been maybe even, I mean, definitely surprised by is how quickly we've moved from a model that is more like deploying thousands of humans against generic problems to small targeted groups of people working on very specific problems. And I think that's kind of the nature of our LHF, right? Like, what we found is this kind of, I feel like there's always a pendulum. A few years ago, people were saying, well, you know, machine learning is great for specialized models.

17:40But I mean, God knows when we'll ever have anything that reproaches like AGI or something that's even, even to not take AGI is kind of a big concept to take it down to a smaller level and just say like, is this generally useful across problem domains? And it turns out that large language models are phenomenally useful across problem domains in ways that we wouldn't have anticipated. I certainly didn't anticipate a few years ago. And what we're finding is that as the pendulum has swung back to a place where LLMs are like massively useful for all sorts of business problems, that we're thinking about refinements to create more accuracy in specific problem domains.

18:24So we're kind of like now the pendulum swings back to specialized cases. How do we execute those super well? And there are a number of approaches in that space, but I think that the reason why RLHF is relevant to us is the same reason it's relevant to everybody in the world right now, which is that we've got these almost incomprehensible alien intelligences that have landed on our planet, landed in our business reality, and now we're trying to figure out ways to make them useful for very narrow problem domains. But I'll say, I mean, if our work is any example, a good example, and I think it is, we work both sides of the street.

19:06So while we're helping to build large language models and working on our LHF task, we also sell solutions using those platforms to our clients, other clients. So we see new opportunities for, I'll take as an example, we're using large language models to solve problems like text cleanup and classification of products in large product catalogs, things that we used to just think, well, you're going to have to have a person to do that. So I think that we have a unique perspective on it, having worked with the companies that are producing the tech and enabling that, but also then being sort of a Johnny Appleseed of technology to bring the benefits of LLMs to the non-AI companies who are saying, what do I do with this stuff?

20:01Yeah. Two questions. One on your using LLMs to address company-specific problems. Is part of that to automate labeling so you don't need human labelers? I just saw something from Andrew Ng's Landing AI about using LLMs or large models, I should say. I guess they're not language any longer, to automate segmentation and labeling. The demo is pretty impressive. Are you using it in that direction to label data sets for supervised learning models? Yeah, I think we are in some cases. That's not specifically what I was talking about, but that is a really interesting space. Because I think, again, I think there tend to be some misconceptions about the way that RLHF can work and supervised learning can work in that there's sort of a, there can be a misconception that we're just throwing a lot of human bodies at a problem.

21:20and that they might not have a high degree of expertise and that we might be sort of solving a problem in bulk rather than solving narrower, more specific problems. And I think that one aspect of that is that people may underestimate the sophistication of the way that our LHF processes can be run and the variety of ways that they can be run. So the reason why I would say like probably our biggest competitive advantage in that space is having a platform that enables us to configure new interfaces, new digital surfaces for trainers to interact with based on ongoing feedback with researchers. And I think that if you talk to folks in that space, as we do, we hear a lot of like, they think that a lot of the benefit they're going to get and they're continuing to get over time is through, we have an internal terminology for this.

22:17We call it giving our agents Ironman suits. So what we want to do is maximize the impact that a single trainer can provide. And we want to find ways that we can take advantage of high judgment, really intelligent, sophisticated individuals. We're not at a place of like, let's have 100 people look at a picture and decide whether it's a hot dog or not. We're at a place where we're solving really complex and interesting problems and you want a high level of human expertise. And as they go through that process, having a tightly designed process that's flexible and has a feedback loop with researchers to change the way that we're collecting that data, that's really critical.

23:00And I think it's really exciting to me and a validation for us kind of as being a very horizontal, flexible platform. I think, I'll just say, I would not want to try to put a stake in the ground with saying, these are the five ways that I want to approach data training for the next two years. Like, it's going to change in a few months. It's changing for us every day. Yeah. On training large language models themselves, you're saying that, you know, you have a highly educated and finite workforce working on these problems. Is this, I mean, there are all kinds of LLMs that are being developed for very specific use cases.

23:57is it that those problems that you're working on or are you working on more general behaviors of LLMs and of course I'm talking about the hallucination problem which and you and I have had this conversation before, OpenAI is using RLHF to address, but the extent of the problem is so large that I can't imagine RLHF making a dent. I can imagine it on very specific use cases. coding for example i mean these models have tremendous promise for automating code generation but if it's just an autocomplete or it gives you you know half a dozen options and the programmer has to choose between them, it's useful, but it would certainly be much more useful if the language model knew exactly what was the correct or the optimal code sequence.

25:18And, you know, I can see RLHF maybe being able to refine the large language model in that particular case. But on more general, if I ask GPT-4 to tell me about myself, it composes this beautiful paragraph or two that gets my name right and a few of the places I worked right, but the rest of it, you know, I wish it were true. It was all about wonderful things that I've never done. So, yeah, can you talk about the effectiveness of RLHF on the behavior of large language models more generally, and then about working on very narrow, specific problems, and maybe give us some examples. Sure. I mean, I think, you know, because we work in this space, I kind of, I hear a lot of people's opinions and I kind of marinate in it.

26:26And I think that I'm hearing in your question kind of a skepticism about whether large language models are accessing truth. It's not just, I mean, isn't it a funny little euphemism to call them hallucinations? You might call them lies. Yeah. I think that, you know, my personal opinion, just based on a lot of what I hear and what I see with experts in the field is I'm actually incredibly optimistic about how far RLHF can go to solve those problems. Incredibly optimistic. And I think that one of the things you'll hear in talking to folks in that space, it's sort of like the positive side of the whole Dunning-Kruger thing.

27:13I love working with these folks because they have a very kind of open-minded, playful attitude about like, well, let's see. Let's try this and see what happens. Because I think that any conversation about what's happening with large language models needs to acknowledge that there have been emergent properties that even the people building the dang things had no idea were coming. And I think that one of the analogies that I use when I think about it is that like, what really has happened is that we're, we talk about RLHF, and I like to demystify it by taking, removing the terminology and just say like, well, so we've got teachers.

28:02And it's like raising a child, right? So the first thing that you know is that like language is learned through imitation and that there's a lot of interesting philosophical debates that we could have about how much knowledge is actually just, actually fully exists encoded in a linguistic form. Like, is there some kind of sense of, I don't know, let's say like a Platonist kind of idea of that there is a real world and real things and words are a description of those? Or is it just that reality is constructed of linguistic constructs? So I think I personally kind of lean more towards the idea that there's a lot embedded in language and symbolic structures that actually does encode meaning, does create world models, maybe not a literal world model, it does have a sense of truth.

28:51But that as you're growing up this child, this new intelligence that's come into the world. It's just natural to want to have them go to school. And what's happening with our LLMs now, and what we're expecting out of our LHF is for our kids to go to school. And I don't, just like I would say with my own children, who are both in school, right, that I have high hopes for what they can get out of engaging with teachers. And I have high hopes that a lot of things that they might say that are demonstrably wrong will be corrected through formal instruction in specific areas of like academic interest.

29:32So if someone's saying two plus two is five, I think maybe I'm anthropomorphizing LLMs a bit, but I think that the approach is to send them to math class and say, no, no, no, two plus two is four, and they'll learn rather than trying to implant a calculator in their skull. That's interesting. Yeah, again, on very specific problems, maybe math is one. I can see how that would work. But on the general problem of hallucination uh how does how do you teach a language model through rlhf which is basically on specific examples uh telling the language model that no you're wrong try again and then when they get it right saying yes you're right on specific domains like mathematics I can I can see that but just generally how does the language model know what is reality and what is not or what reflects reality and what does not without going through each of those examples I just don't see how you generalize without, and you and I have talked about world models, without a world model that it can refer to that is not language-based.

31:14Yeah, I think that what we're circling around is that there's some some amount of objective truth. And that, again, if I, sorry, if it's a controversial direction to go, but when I hear about world models, I often think as well of like, what we're looking for are value systems, almost like we're asking for our LLM to join a church, that we have some faith that there are facts that are objectively real and true and verifiable, or maybe not verifiable, but we still accept that they're true and that that forms sort of a lattice, a framework around sensible ideas being formed. And I'm not totally opposed to that idea.

32:03I think that the angle that I take on that, and I know that there's some research in this space as well, is just thinking about something else that people aren't necessarily awesome at is understanding the dimension of time. And I think that if you think about, and this comes out in like, when you think about a few shot learning opportunities and gathering context and conversations, I think that the meaningful things that will happen in this broader space is that we'll have interactions with models. of their next generation version of an LLM, or maybe even we'll just call them LLMs, that will have access to real-time data, and they'll have trusted systems.

32:50And in that kind of a world, some of these things where we're seeing hallucinations or lies, or cases where we do have generally agreed upon objective truth that's not being matched, that there are different solutions to that problem. And I get the sense of the kind of anxiety that you're describing of like, how on earth did we cover every potential case with human trainers saying, what did Craig do in 11th grade? And did we report that he won this competition that he didn't win or he finished in second place and not first? There's no way that humans could ever exhaustively track all that stuff down.

33:32I get that. I get that. But I also have been amazed at how far we've gone with LLMs when I had the same skepticism about specific cases that are no longer true like a year ago. So, yeah. And can you talk about that? Because that's right. That's one of the things that intrigued me the first time we spoke. You said that you've seen such tremendous progress. uh and is it in general uh general uh behavior that way I mean because what uh what open AI is trying to do with RLHF as I understand it is not cover uh you know this exhaustive or possibly infinite list of examples. But to teach the LLM just through example that they have to find as objective a truth as they can find in the data that they've absorbed so that it becomes a habit or a behavior that then applies to every situation.

34:56And that, yeah, I'm skeptical of, but maybe you can talk about the progress that you've seen in some examples that would show that progress. Yeah, sure. I mean, I think there are some easy examples where I can say that there are specific tasks, like I was alluding to some before that we do for certain types of clients like e-commerce companies or delivery companies that are maintaining these giant catalogs. And a lot of the value of their brand is driven by accuracy and sort of the user experience and customer experience of interacting with those. And one of my classic examples is like, if a restaurant has the wrong options on a menu item at my favorite restaurant through DoorDash, then it's not my favorite restaurant anymore.

Read the full transcript

35:50I can't go there because I can't have this particular dish with ground beef. I want it with chopped steak. And if it's not an option, it's just right out for me. So familiar problem space for us. And again, partly because we're the kind of company that metaphorically draws sharpie marker circles around problems and replaces human pieces with automated pieces, we found that things like language cleanup, style guide compliance, classification of menu items, knowing that, for example, like you probably don't want to order chicken medium rare, like those sorts of things we thought of in the past as being optimally solved by people.

36:35And if there was enough scale, then we might look at specialized models. And what we found even in the last, when was this? Okay. There was a project we were working on in December. So four months ago, where we were thinking when we initially engaged, we wouldn't have considered a large language model addressing this problem space. But as it turns out, we could use like really off the shelf GPT with smart prompt engineering to get results that are better than what we get from people. And that was pretty mind blowing, pretty mind blowing for me to think that another example that I think of in classification, we were looking at a catalog of items and I'm going to keep, I'm going to keep struggling with this.

37:25I don't know the difference between eyeliner and mascara? Maybe my wife should tell me, but I don't know. But GPT knows quite well and was able to look at something and see that it was miscategorized. And I just never would have guessed with the information provided. And so that's an example of like, I feel like there's like shifting goalposts on what AGI is. That's a classic example for me of here's a tool that's actually able to provide a general and useful intelligence that used to require a specialized solution if a tech solution worked at all. So that was like a boom, light bulb moment for me.

38:04Like, wait a second, like large language models can solve problems that we used to think would require a person or a specialized model. So that's a movement, a change that's happened. Right. And in that case, it was the large language model out of the box or you, it was as a result. Amazingly. So, and I think, so I guess part of the reason that I have a little bit more, okay, here's a better example. Part of the reason why I have a little bit more optimism about, or maybe just belief in the power of existing large language models, is that, is the story of prompt engineering as it's been playing out.

38:46I think that, And I hear this from experts in the field that we talk to, or even that I just hear presentations by, that there's a confidence that there's the problems related to hallucination or are more focused on expressing things than whether there's an internal model that actually matches something that we would describe as real. So two angles on that. One, super practical. When you think about how to solve the problem that I just described of being able to clean up text or do classification problems or just clean up a catalog so that a user experience of a restaurant menu is better. what what we used to do in that space was specialized models in humans we moved to llms and since moving to llms we found that we get the best juice for our squeeze with using llms with custom tailored prompts and i'm sorry custom one custom tailored prompts So it's sort of like asking the question the right way.

40:00And in a concrete sense of like how we implement that, we look at specific use cases and iterate over not whether we're going to build our own model, but asking the question in a way that will lead to a more useful outcome. So specifically, I mean, this has happened for everybody, but I think that what you see is that there's just as much progression of quality of outcome in experimenting with the way that you form the question to an LLM in the form of what we kind of, again, euphemistically call prompt engineering, which sort of feels like asking the Oracle at Delphi the question in the right way.

40:43that that's where we get a lot of benefit. And I think that what that tells me is that there's more sort of latent possibility and intelligence and world models of some flavor that exist inside of an LLM than we might think. And that the problem may be more superficial and has to do more with interaction. So imagine that like the monolith comes down to earth, or there's some alien intelligence that we're engaging with. I would probably be less critical of its syntax and speaking English, and more curious about how I can ask questions in the right way. And it feels like that's sort of an experience to me.

41:30Like, I think that there are enough legs with existing LLMs that we could be finding new business uses that create better lives and more wealth and more opportunity for a decade, even if we froze development right now, because it just takes a while for the possibilities of technology to filter through. So that was kind of one of my examples of why I think, you know, why I think that there's more there there when it comes to LLMs today than we might think if we're focused on hallucinations and accuracy and alignment. Right. And so maybe it's, you know, you have to come at it from two different directions.

42:15You work on RLHF to get the model to ground itself. However, it's doing that in some more objective reality and then i'm from the other direction you're refining the way that you uh talk to an rl or to a model through uh engineering um on on the so i understand how you know improving your your prompts uh improves the output do you have an example of of how RLHF specifically has improved models or is that just too fuzzy a world to know what's happening we're just one company out here doing this work we've got some pretty we're we're close to some pretty cool stuff that's happening with some pretty well-known companies and I think that while I can't talk about specific projects that we've worked on, I've seen literal work that we've done lead to specific provable, testable outcomes in updated versions of large language models.

43:31So I've seen a situation where we say like, oh, with this newer version of a large language model, this question can be answered that couldn't be answered before. And I saw kind of an inside baseball perspective of the work that led to that, which by the way is super cool to be able to see that but also to what's even cooler is the the speed at which it's happened so you might find that like um a piece of work that you worked on a few months ago has now uh influenced the way that the world can perceive their interactions with an llm um and i know that may be might be maddeningly vague but i can't i can't go into huge particulars, but I can say that the space that I'm thinking of in the example that I'm thinking of fits in that category of like, we used to say two plus two equals five.

44:22And now we say, of course, it's four. And I've just, I've seen, experienced firsthand places where RLHF has led to observable improvements in existing models. specific questions or or or more generally slightly more general than that um and uh you know it's hard to say like i don't have uh i don't have visibility into every project that's going on for a particular lm um but i think it's pretty clear that um some of the areas of focus that we've had have been directly impactful in the way that people can experience these technologies. So, and not like, I guess what you're asking is, did we teach it that two plus two is equal to four and not five, and then go check that specific example versus checking three plus four is equal to seven and not six?

45:25And my experience of that firsthand is that it's more like the latter in that we've corrected two plus two and seeing improvements on three plus four. And that's purely through RLHF. It's not by some fine tuning in the training, you know, in the training of the model, because that's happening too, right? You take a foundational model and then you train it, refine it with supervised training to have expertise in a particular domain. The example that I'm thinking of relates to seeing improvements in LLMs, not with the mechanism of fine-tuning. But I'll say, even though we do a bit of fine-tuning work, I thought that was going to be a bigger opportunity for us in our Johnny Appleseed type of role.

46:31I thought that we would be helping a lot of our clients to do some fine-tuning, particularly with GPT, which I just love the way OpenAI has documented their fine-tuning approach. And I thought, wow, this is going to be wonderful to work with dozens of companies in tailoring this super powerful general purpose LLM to their specific needs. And then time and time again, we find that it turns out that stock GPT solves a lot more problems than we thought. And that fine tuning has not been as big of a need for us with our clients. there's just so much you can do with with uh some of these stock columns yeah yeah uh that's that's amazing actually i mean it's it's exciting and i tell people who complain about uh chat gpt which is you know what most people experience uh that you know hey man this thing was released to the public like six months ago or maybe less uh you know wait five years and and see if you have something to complain about then it reminds me of people that when uh uh the i with the the various g uh gps systems came out and uh you know people took great pride and say I never use, they're always wrong.

48:08And they did, they have problems. But I think everyone relies on GPS. Yeah, I mean, I feel like this is one of those disruptive moments and the concept of disruption has been cheapened by having, I don't know, how long has it been since we've had a good, real new disruption? Is it the iPhone? I think that sometimes it's really hard. These kind of moments serve as more of a Rorschach test for what people are really concerned about or their own attachment to an existing way of things working. I think that this one is like, especially if we're talking about chat GPT, I think that space and generative AI in general is pushing into spaces where humans have thought that it was purely their domain, like writing copy, doing visual design.

49:09And I think that I get it. I'm a songwriter. The idea of GPT writing songs for me kind of creeps me out. And it's not good enough yet. But I also have observed that the younger folks on our team, like our person who works in PR, one of our young designers, they were all over the generative AI stuff from the start. And I think that if you have an open mind to it, it's a way to to accelerate and to like we say it's an it's an iron man suit so why not take advantage of that and people who are open-minded to it will really fly yeah on the on the numbers of people in in rlhf uh that you have working on these problems again you said you had thousands of human agents around the world, but on RLHF specifically, for these general use cases, for like teaching an LLM not to hallucinate, how many people would be doing that?

50:19And you also talked the first time we spoke about the qualifications that these are not, you know, low paid wage workers in third world countries necessarily. Yeah, I mean, we, yeah, to kind of address the scale thing, I have to be a little bit vague about it for, because of confidentiality with the clients that we work with, but we've worked on projects that are larger, involving hundreds of agents on a specific project, and even projects where it's dozens. And I think that we're just to be clear, we're not working on one project for one client. So we work on lots of projects and they, some of them go away and we work on new ones and some get more people and the needs change, but I've seen enough to have, to be able to sort of characterize a change in the general requirements for advanced data trainers.

51:19And I've also just seen what our needs are in our business overall. And I think that our particular problem space of working on problems where we want to layer in automation inherently biases us towards hiring high judgment individuals all around the world. And as our founder likes to say, We're more selective than Harvard. We expect people to be, you know, college educated, great English speakers, or in some cases where we need them to speak other language, great at speaking those other languages. We have a high bar for what our agent workforce looks like. and then as I would say like just to generally characterize without talking about any specific projects I think that we've seen a trend in in in our data training needs towards more for the work that we do towards more folks who are based in the U.S.

52:21and who have more and more advanced degrees and that that could just be a reflection of some of the work that we've taken on But I hear this from other folks as well, that like the expertise required in this space obviously is going to increase year over year. And there are real opportunities for folks to, these aren't like digital drudgery jobs. They're interesting, creative jobs. And we don't have, I mean, part of the whole point of this technology is to have humans doing things that uniquely require humans. Yeah. And do you think that the human element in AI will, that the supervised learning labelers, that workforce will eventually shrink and a smaller but more effective RLHSF workforce will grow?

53:26I mean, will there be a shift? I mean, I do. I do think that we're going to move more and more towards higher expectations for what humans do as they engage with these sort of training processes. But I don't think that necessarily means that it's a smaller number of people working. It may be that we have, you know, you're talking to someone who's in a business that's growing really rapidly. So it's very, I can't even imagine shrinking at all. So for me, I think it's more about raising the bar on requirements for folks who are working in that kind of space than it is about a shrinking number of people.

54:08And I can't really make a broader social characterization about it. Maybe there are fewer, maybe there are more. Yeah, but the human workforce that is, it will shift from labelers to RLHF workers. That's what I'm asking. Yeah, I mean, I may have opinions about that. I do have opinions about that. But I also am just kind of observing and seeing what happens and trying to build around a responsive, flexible approach. Like I would say, I think that there's going to be and there continues to be a lot of economic pressure on like and other business considerations that drive people to not want to be dependent on one big, big winner in the LLM space.

55:05There are people who don't want their data going to open AI, for example. and there we talked about this a few months ago about how there's this sense of like well it looks like everybody's going to have their own llm in about five minutes and it's true it's happening in real time um if i haven't checked the news today i bet there's another announcement this morning um so i think that um and if you look into what goes into creating those to some extent um i read a ceo of a of a major company recently said that it takes billions of dollars to train an LLM combination of compute and labor. And then at the same time, you hear these stories about like, oh, here's this new open source LLM that I trained in 30 minutes for a nickel and it runs on an iPhone 4 and it's smarter than a seven-year-old.

55:56So I think there's just like, there will be continue to be so much interest in this space and so much variety and so many different perspectives on it that if people were thinking about this as like, is this an interesting career opportunity? I'd say, yeah, I think it is. I think it's an interesting space to be in. And I don't think that we're, even with models training, models, reflection techniques, all the stuff that will inevitably happen, Human in the Loop is going to be here for a while for lots of reasons, as it always has been. And I think that one of the things that we are excited about at Invisible is like being in a position to move work problems, big chunks of work, more from kind of dehumanizing people acting like machines and doing menial labor to people doing high judgment work and having the machines do the work that elevates them and allows them to work on these highly complicated, sophisticated things where they get to, you know, use their minds and their human judgment.

57:02Yeah. Okay. Well, I'm about up to an hour. Why don't I work with this? I may move stuff around, but this LLM discussion is fascinating. Is there anything that you wanted to talk about that I didn't ask about? There are definitely things that I thought we might talk about, but I'm not too stressed about it. I do feel like the intro and talk about like my background and I felt like that was a little too much. I'm sure that's worth editing. But yeah, I think I feel generally comfortable about it. Okay. If there's, when you say something you thought we'd talk about that we didn't, what are you referring to?

57:52I think just sort of a, we hit it a few times, but I think that there's some sort of a implicit question about social impact. And we kind of skimmed off of it when we talk about like young people are now using generative AIs. And I generally, that's, when I have these conversations, this ends up being a big part of it because they're like you're a writer like how do you feel about it you know it tends to circle around that but maybe that's not your your no no it's interesting I I think that it'll you know it'll increase productivity I think for writing um you know surprisingly to people that are are accustomed to writing a lot of the world can't write.

58:50I mean, they can read, but they're not very good at writing. That's why this company Grammarly, which I remember when it came out, I kind of thought there's no way this is going to fly. My spell check on Word does that, but it's been a pretty big success because there's just a lot of people that struggle with writing. So I think, yeah, I think it'll increase productivity for a lot of people that don't have the language skills that their jobs require. But I think there'll, and I think it'll get to the point where it can produce stuff that's considered creative or high quality by discerning users.

59:53You mentioned songwriting. I expected very shortly it will be writing songs. Personally, I think that's great, you know uh but i'm not a songwriter i do think that as ai becomes increasingly created that way that the value of a true uh create human creativity will go up i have a son who's an actor and I tell him yeah you know that's people will crave genuine human creativity because it's it's like oh I used the first of those style transfer apps prism it was called I think and I know there have been iterations since then you know you take a photograph you turn it into a Van Gogh style painting or into a watercolor and when it first came out uh you know I'm sort of on the on the front edge of seeing that stuff and I used it and sent around to friends and everyone was amazed and how did you do this and it's beautiful and now you see them and you don't think twice about them and you they kind of irritate you you know like yeah big deal the guy put it in an app.

1:01:26But that doesn't mean that people who can paint will not be appreciated. I think on the contrary, somebody who produces something that's very human is going to only be valued more. I feel the same way. I definitely feel that like, so there's many examples of this, but one that's making the rounds on Twitter lately is AI Drake. And there are these new songs, new Drake songs that are entirely AI generated. And, you know, no shade on Drake. I like, I like, I like a lot of Drake songs, but it's kind of disturbing how easy it is to encapsulate Drake songs in a formula and how real they sound. And I do think, I mean, again, this is something that's kind of implicit in Invisible's business.

1:02:21We're big believers in human potential. And that's where our name comes from. The idea of invisible technologies is that technology is best when it's invisible. Like our goal is to elevate humans. So I think about maybe AI Drake makes people a little bit bored with lowest common denominator music. And what it means is that whether it's somebody else or maybe Drake produces some better tracks, right? That I just, I have high confidence just like you that if you've seen enough like low quality generative AI writing or art, it just gives you a finer appreciation and taste for the real thing. And there'll always be opportunities for greatness and human expression.

1:03:09That's it for this episode. I want to thank Scott for his time. As always, you can find a transcript of our conversation today on our website, Eye on AI. That's E-Y-E hyphen O-N dot A-I. I encourage you to download it and read because the eye catches a lot that the ear misses. Please like or review the podcast on YouTube, Apple Podcasts, Spotify, or whatever platform you use to listen. It really helps us a lot. And in the meantime, remember, the singularity may not be near, but AI is about to change your world, so pay attention.

From the publisher

On episode #132 of the Eye on AI podcast, Craig Smith sits down with Scott Downes, Chief Technology Officer at Invisible Technologies. We crack open the fascinating world of large language models (LLMs).

What are the unique ways LLMs can revolutionize text cleanup, product classification, and more? Scott unpacks the power of technology like Reinforcement Learning for Human Feedback (RLHF) that expands the horizons of data collection.

This podcast is a thorough analysis of the world of language and meaning. How does language encode meaning? Can RLHF be the panacea for complex conundrums? Scott breaks down his vision about using RLHF to redefine problem-solving. We dive into the vexing concept of teaching a language model through reinforcement learning without a world model.

We discuss the future of the human workforce in AI, hear Scott's insights on the potential shift from labellers to RLHF workers. What implications does this shift hold? Can AI elevate people to work on more complicated tasks? From exploring the economic pressure companies face to the potential for increased productivity from AI, we break down the future of work.

(00:00) Preview and introduction 
(01:33) Generative AI's Dirty Little Secret
(17:33) Large Language Models in Problem Solving
(23:24) Large Language Models and RLHF Challenges
(30:07) Teaching Language Models Through RLHF
(35:35) Language Models' Power and Potential
(53:00) Future of Human Workforce in AI
(1:03:10) AI Changing Your World

Craig Smith Twitter: https://twitter.com/craigss

Eye on A.I. Twitter: https://twitter.com/EyeOn_AI

 

More from Eye On A.I.

All 266 episodes
#132 Scott Downes: Navigating the Language of AI & Large Language ModelsEye On A.I. · 1 h 4 min
Listen in VO