In short
Real Vision Podcast Episode Notes: GPT, What Are You Doing?! w/ Mikhail Voloshin
Episode Overview In this episode of the Real Vision Podcast, host Ash Bennington interviews Mikhail Voloshin, CEO and principal engineer of Mighty Data, Inc. This episode is a follow-up to a previous discussion on ChatGPT and AI, focused on dispelling myths surrounding GPT and large language models (LLMs). Mikhail provides insights into the mechanics of how GPT functions, key advancements in AI, and the implications for the future.
Key Topics Discussed
Introduction
- Mikhail Voloshin returns to discuss the mechanics of GPT and LLMs.
- Emphasis on dispelling myths and enhancing understanding of AI technologies.
What is GPT?
- GPT (Generative Pre-trained Transformer) allows human-like text interactions and is cloud-based.
- It was developed by OpenAI and trained on 45 terabytes of diverse data.
- The goal of AI development is to achieve Artificial General Intelligence (AGI) – a machine as adaptable and flexible as a human.
Current State of AGI
- Mikhail estimates AGI could emerge in the next 3 to 8 years.
- GPT is not AGI but exhibits traits of a pseudo-AGI, simulating some human-like capabilities.
The Importance of Language
- Language enables externalized cognition and communication, allowing machines to interact meaningfully with humans.
- OpenAI has introduced plugins to allow GPT to interact with real-world applications (e.g., booking flights, trading).
The Lifecycle of AI Technology
- Growth of AI technology follows a pattern of excitement, disillusionment, and eventual practical application.
- Mikhail emphasizes the importance of understanding the limitations and myths of AI to avoid disappointment.
Myths and Realities of GPT
Myth 1
GPT is an App on My Device
- *Reality*: GPT operates as a cloud service, requiring internet connectivity for function.
Myth 2
GPT is Learning from Conversations
- *Reality*: GPT does not retain information from interactions. Each session is stateless and independent.
Myth 3
GPT Can Remember Previous Conversations
- *Reality*: Any context maintained during a session is temporary and not persistent.
Myth 4
GPT is Self-Aware
- *Reality*: GPT mimics conversation but lacks consciousness or understanding.
Understanding How GPT Works
The Process of Querying GPT
- Input is sent over the internet, processed in a data center, and transformed into a tensor (numerical representation).
- A neural network evaluates the input and generates a response based on probabilities.
Output Generation
- GPT generates responses through a probabilistic selection process, influenced by prior context and input.
Reinforcement Learning from Human Feedback (RLHF)
- RLHF is a training method where human trainers evaluate and score responses to encourage desired outputs.
- This differs from user interactions, which do not contribute to the model's learning.
Conclusion
- Mikhail calls for the audience to explore potential applications of LLMs and share ideas for future development.
- The episode highlights the evolving landscape of AI and the importance of understanding its capabilities and limitations.
Key Takeaways
- Understanding the mechanics behind LLMs like GPT is crucial for leveraging their capabilities effectively.
- Many common beliefs about AI, including learning and memory, are misconceptions.
- The future of AI technology is promising, but requires careful consideration and realistic expectations.
- Engaging discussions on AI development can lead to innovative applications and improved public perception.
---
For more insights from this episode, visit [Real Vision](https://www.realvision.com).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00If you've been considering futures trading, now might be the time to take a closer look. The futures market has seen increased activity recently, and Plus 500 Futures offers a straightforward entry point. The platform provides access to major instruments, including the S &P 500, NASDAQ, Bitcoin, natural gas, and other key markets across equity indices, energy, metals, forex, and crypto. Their interface is designed for accessibility. You can monitor and execute trades from your phone with a$100 minimum deposit. Once your account is open, potential trades can be executed in two clicks. For those who prefer to practice first, Plus500 offers an unlimited demo account with full charting and analytical tools.
0:45No risk involves while you familiarize yourself with the platform. The company has been operating in the trading space for over 20 years. Download the Plus500 app. Trading in futures involves the risk of loss. It is not suitable for everyone. Not all applicants will qualify. Hey, everyone. If you like this podcast, go behind the paywall to get privileged access to the smartest minds in finance. Visit realvision.com slash RVpod and use the promo code podcast10. That's podcast10 to get 10 % off our essential membership for the first year. Join the Real Vision community and learn how to become a better investor.
1:24And now to the top analysis of today's markets.
1:36Mikel, welcome back to Real Vision. It's a delight to be here, Ash. Well, it's a pleasure to have you here at the Festival of Learning. For those who may not know, we're doing a follow-up to a show that we did that was spectacularly successful on Real Vision, where you came in, you talked about chat GPT and AI more generally. I think we blew a lot of people's minds with that demonstration. And now we're here to look under the hood a little bit more deeply to talk about the mechanics of what makes AI tick. specifically the kind of AI that GPT is, which is called an LLM, which is a large language model.
2:11So what I'll be doing during this talk is talking about how it works under the hood and also dispelling some myths that I've seen popping up over and over again over the course of the last couple of months. I've talked to a lot of people about GPT. A lot of people have been using it. A lot of people have been really excited about it, but they're either using it in suboptimal or in some cases outright wrong ways. and they're attributing things to it that just aren't true, and that's impeding their ability to really capitalize on the abilities of this new technology. So I'm going to show how it works, dispel those myths, and the two will sort of feed into one another.
2:47I'll dispel myths by showing how it works. With that said, let's dive right in. So just a recap for those who haven't seen the first episode. Who am I? I specialize in artificial neural networks at UIUC. did graduate studies in neuroscience, studying real neural networks. And since then, I've worked at a lot of different places that have spanned the gamut of various AI and machine learning technologies. And now I currently run a small consulting shop called Mighty Data Inc., whose website is currently terrible. What is GPT? If you're watching this presentation, you probably already know, but if you were referred to it by a friend, it's a new artificial intelligence system, came out in late of last year, very human-like text interactions.
3:30It's like talking with a human via text. It passes the Turing test, which is a longstanding staple of evaluating the intelligence of an AI on a very coarse level. It's an adaptation of something called neural network technology, which we'll talk about later in the talk. A neural network is a very abstract mathematical representation of how neurons interact in the human brain. It's not going to fool neuroscientists, let's put it that way. But on a very coarse-grained level, it kind of simulates what the brain does. GPT was invented by a company called OpenAI, and it's offered as a cloud-based service by them.
4:17We'll talk about this cloud-based nature of the software in a minute. It's a type of AI called a large language model. It was trained on 45 terabytes of data comprised of a combination of web scrapes and Twitter feeds and Wikipedia and novels and textbooks and so on, encyclopedias and so on, downloaded sometime around September of 2021. So that becomes important because we'll talk about the knowledge corpus a lot while explaining what this thing is and how it works. why is this new ai a big deal the it's because the holy grail of artificial intelligence has long been regarded as a general purpose ai historically ais have been built for one purpose at a time so this program plays chess this program drives a car this program recognizes faces and so on an agi is a hypothetical machine that is as smart as a human and smarts in this case don't don't strictly refer to IQ or ability to recall information, but specifically they refer to flexibility and adaptability.
5:29You don't necessarily have to have a lot of cognitive resources as long as you know how to apply them to problems at hand in a versatile way. Let me ask you this, how close are we right now to artificial general intelligence? That's a really good question, Ash, and we're getting really darn close. As you'll see over the course of this presentation, uh gpt is not an agi but it's getting there um it's uh in fact you'll you'll see flat out how it's not self-aware over the course of what i'm about to show in the next hour but um it is what might be called a pseudo agi um it's uh it's it's able to simulate some things about an agi And at that point, you get into something related to the Turing theorem, which is that on a very coarse level, if you can simulate an AGI, then what's the difference between actually being an AGI versus just simulating one?
6:31um so it's uh to answer your question i'm not going to put if i had to you know gun to my head put a you know put a date on it i would not be surprised if we saw agis emerge in the next uh three to eight years or so that's quite quite soon we're living in some really interesting times and uh it's worth noting that advances in computer technology help advance computer technology because we use the tools that we build in order to build more tools. Like right now, most of my coding is done by GPT. It turns out to be a very good coding engine. So like I tell it, write this function and then it branches out some Python code.
7:16The Python code is often buggy, but it's buggy in a way that I can see and fix pretty quickly. And I can fix it with GPT in the loop. So yeah, progress is going to be fast. I'll talk about the speed of progress. It becomes relevant. I do want to mention why the focus on language is kind of a big deal because it's not explicitly problem solving, but language is a tool that humans already developed about 200 ,000 years ago for general purpose, externalized cognition, and external data storage. Basically with language, you can instruct anybody to do anything and share information about anything you can conceive of.
7:55We've already developed this. So a machine that's able to interact with us via language becomes able to participate in the human exchange of ideas. And that kind of becomes a big deal. In the last several weeks, a couple of highlights. OpenAI released a plugin architecture that lets developers write tools that allow GPT to interact with the real world. um gbt as you'll see as a language model which uh but you can use it but by slightly but by having a plugin that scans for keywords or key data structures within the language being output uh you can have this thing doing stuff like looking up web pages or uh trading stocks or booking flights um there have been a lot of uh competing models and services emerging um it's a fast-moving field.
8:46So Meta has Llama, Stanford released Alpaca. Inflection AI has a very interesting product called Pi that I recommend looking at. And Google, of course, has Palm and Mum. And search engines like Bing and Google have begun integrating these language processing systems into their search engines. So Bing in particular, you can just ask it questions and it'll look up the answers and discuss the answers with you in natural with a natural language interface. There is something that we're going to deal with in the next couple of months or years that is the case with every single technology of a caliber of this disruptiveness that we're going to see with this one.
9:30And this is not a scientific graph. I just based this on my experience and my prognostications, but it's pretty important, I think. So basically, can you see my mouse cursor on the screen here? Yeah, lower left hand corner. From the time that a tech is first invented, the public sees the tech and gets really excited about the possibilities. The problem is that the possibilities of the tech take time to manifest. It's one thing to build the AI. It's another thing to build really good applications that utilize the AI. One of the examples that really comes to mind for those who are old enough to remember it is that when the iphone first came out uh one of the first apps for the iphone to take advantage of its um of its gyros uh is um it was an app that draws a uh a mug of beer on the phone and it takes uh like it uses the accelerometer data to uh to like keep the beer level as you tip the phone back and forth um and like you can you can pour the beer out and it's it it's all it's all on the screen it's really cute but who cares you know like um so you know when the like it took a while for apps of any meaningful utility to be developed uh that are more than just little like demonstrations of concept we're going to take a quick break and be right back with more of the day's top analysis on the real Vision Daily Briefing.
11:28nasdaq gas and much more explore equity indices energy metals forex and beyond with a simple and intuitive platform you could trade anytime anywhere experience the fast accessible futures trading you've been waiting for with plus 500 with over 20 years of experience plus 500 is your gateway to the markets visit us.plus500.com to learn more trading and futures involves the risk of loss and is not suitable for everyone not all applicants will qualify plus 500 it's trading with a plus
12:03you get this kind of proof of concept to get the kind of toy apps uh and then there's this this period of like disillusionment that sets in based on this chart exactly exactly what happens is during this period of toy apps there's also a period of opportunists who are trying to cash in on what people believe the app the real apps should do and they crank out apps that look like they will do those things in the future, but that don't actually deliver. And we've seen this all the time. We've seen that, like, you know, I'm talking about a curve that happens, you know, we saw it with crypto, we saw it with mobile apps, we saw it with, you know, with the internet.
12:39There's this really famous quote by Paul Krugman in 1998, who projected that by 2005, this whole chat room internet fad will have faded away, and the internet will have had no more impact on the economy than the fax machine. But the point is that what's going to happen is jadedness will set in among the public and they will ignore the real advances in the tech that are happening despite it, like that are still going on and are in fact exceeding the capabilities that were created by sort of opportunists or, you know, proofs of concept. But it takes the public a while to realize that that's the case.
13:21This is what we're going to see over the course of the next two to five, maybe two to four years. Now, where we are right now is the public has seen the tech and we are just barely beginning to build real apps that capitalize on this. But there's also going to be a slew of sort of proofs of concept or jankily built apps that claim to deliver capabilities that are simply impossible for the tech to deliver at this time. and this is something that I call the age of myth. And the reason I call that is because in this period, mythology and frankly flat-out lies and marketing hype dominate the marketplace.
14:06And I don't mean age of myth as a good thing, in other words. And so the point of this talk is to sort of help guide you through the age of myth. You know, I will be your Virgil. I will introduce you to the rings of hell that comprise the interiors of GPT, and we'll take it from there. So, first myth. GPT is an app that runs on my computer or phone. This is not a common myth, but most people understand that GPT is in fact a cloud service, but they don't really fully wrap their brains around what that means. And so I want to discuss how GPT is implemented as a cloud service. One of the biggest times where this misconception comes up is when an app actually includes calls to GPT under the hood.
14:53There's this really great Skyrim mod that I recommend for pretty much everybody, where it replaces the dialogue of the NPCs with calls to GPT. It's incredible. You can literally just talk to the NPCs in just your plain voice, and it renders them in audio. But the key thing to understand is that it's working via a network connection and it's actually using your GPT, your OpenAI API key. You're paying OpenAI a little bit every time you talk to your NPC. So it's a little counterintuitive and it's worth talking about. The reality is that it's a cloud-based service. Your app or browser are sending network requests And yes, this means you can't use GPT if you're offline.
15:40It means that everything you say to check GPT or some other GPT-based service is in fact being routed to some server somewhere, including the not-safe-for-work conversations you're having with Replica. So I'm going to go over a millisecond in the life of a GPT query. It's more than a millisecond, and this will take a few minutes, but I think it's really important because this is where the rubber meets the road. So you're on your laptop, you're on the GPT website, you check GPT, and you type Mary had a little and you hit enter, right? Well, your laptop sends a data packet to your router, which sends it through your ISP, which bounces around the internet.
16:22And from the internet, it gets shunted to a Microsoft Azure data center. In the data center, it gets routed to a single server. The data center has tens of thousands of computers, possibly not that many that are beefy enough to handle the requirements of GPT, but still more than one. This server is running a program that'll feed your input, Mary had a little, through the GPT-4 model and it'll produce one word of response. That's all it's going to do, just one word. So here's how it does that. Inside the server, here's what's happening. The program on the server breaks up the input into tokens.
17:03Essentially, a token is a word. There's some variation when it comes to things like punctuation or apostrophes or pluralization, but essentially a token is a word. What then happens is the program consults a lookup table, and for every token, it looks up a sequence of numbers that represent that token numerically. Each token is going to be a vector. It's a vector that's 1 ,536 units long, I believe, in the case of GBD-4. And there will be as many of them stacked on top of one another as there are tokens in your input. This sequence of sequences is called a tensor. so for those uh audience members who might be familiar with google with the uh the neural network library tensorflow that word might be familiar so this tensor is then used as the activation levels for the uh input layer of a neural network and uh mean in layman terms because i think a lot of people are looking at the screen right now uh if they're watching this and seeing a bunch of numbers that don't really map in their mind to a conversation that they're having with ChatGPT?
18:26That's a really good question. So each one of these tokens becomes this sequence of numbers. And each sequence of numbers is essentially, it's a numerical representation of the word. Now, this numerical representation is basically completely random for every word in the English language. With a slight exception, it's random, but the numbers that represent similar words that have similar meanings, those numbers happen to be close to one another. So the sequence of numbers that represent, let's say the word king, it's basically 1500 random numbers all between negative one and one. Doesn't mean anything, except that the word regent also has a 1500 number sequence representation that is almost the same thing as king like if you go number by number they're only a little bit offset from one another does that make sense yeah i i think it's it's probably hard for people to conceptualize how these strings of numbers come to represent something that looks like human meaning to us the um the process by which a word becomes a sequence of numbers or um or rather how the sequence of numbers uh is found for every word is called um basically these sequences are called embedding embeddings and you can train these embeddings uh in fact if you like i've got an embedding browser that i wrote that i would be happy to show at this time and that might be an appropriate digression.
20:06What do you think? Yeah, let's take a look. And by the way, I should say one of the interesting things about this conversation is that all of these tools that you're about to show us are custom tools that you've built yourself. So this isn't something that people are going to be able to see anywhere else unless they're up on your website. That's correct. And I'm quite proud of these tools if I do say so myself. This is a website that I built called GPT, what are you doing? GPTWYD. And we can see that we're not actually going to see the sequence of numbers here, but we will be able to see the differences between sequences of numbers for any given word or set of words.
20:43So for example, if I was to type king, for example, on this side, and I was to type peasant on this side, then we can see that uh that they have a very low similarity score now if i was to type a third word like i said regent for example um it'll compare a regent for uh of king versus peasant and it'll see that regent is closer to uh to king than it is to peasant um on the other hand if I was to type, I don't know, vagabond?
21:25Vagabond is much closer to peasant than it is to king. You see that? What if we were to try something like monarch? Let's do that.
21:37Much closer to king than to peasant. And so what it's doing here, just so I can try and explain this as I try to understand it myself, in the prior screen you were showing us the slide that showed the numerical representation of meaning. And what we're seeing here is effectively a kind of a meaning contrasting engine where you can put two words in and compare it to a third and then see which has a higher affinity for the meaning. And that's being done based on this backend scoring system that you've just shown us on this slide. So meaning is a little bit of a loaded term. And I understand what you're trying to get at here is that like, how does the word, you know, how does the sequence 11426 convey the meaning of the word little?
22:24And the fact is that it doesn't, like by itself, that sequence means absolutely nothing. However... So, Mikael, would the word similarity be a better word to use than meaning? Yes. Yeah, the sequence by itself means nothing except that the sequence for the word small would be relatively close to the sequence for the word little. It's fascinating because it really does give you a sense of how this is happening behind the scenes, at least in a quantifiable way. Because I think a lot of people interact with chat GPT, myself included, and it does feel like there's a ghost in the machine. It's this spooky simulacrum of consciousness.
23:04And yet we know rationally that it's not conscious. It's not thinking. It doesn't really understand meaning. But this is getting us a ways, I think, to understanding how those pattern matching activities take place behind the scenes. The process of how these embeddings are found, like I said, is outside the scope of this talk. But the general gist at the end of the day is that these embeddings have no, there is no meaning to these words outside of their relationships to other words. Right. But just by relating words to other words, you can get pretty far. For the philosophers in the group, there's something called John Searle's Chinese Room argument.
23:54I don't remember whether or not we talked about the Chinese room in the last talk. I think it's not worth the digression in this one, but anybody who wants to look up something called the Chinese room is an ongoing.
24:12To a certain extent, the debate is a bit semantic, but it's a it's a thought exercise that gets you thinking about, like, what does it mean to understand something? We're going to take another quick break and be right back with more of the day's top analysis on the Real Vision Daily Briefing.
24:52And this clerk does not speak Chinese at all, not one word. But he does have a giant stack of books that have Chinese characters in them and a manual written in English of which books to look up, you know, to cross-reference when a note comes in and what characters to write on a note coming out. And so, hypothetically speaking, with an advanced enough manual and a large enough corpus of reference material, this clerk can produce written answers to questions that are written in Chinese that, when passed out to a Chinese speaker, sound like coherent answers and possibly even correct ones. The question is, does this does that convey to the clerk the ability to speak Chinese in the slightest?
25:49And John Searle presented this by by way of saying that is this is a great example. This is a great way to demonstrate that clearly the clerk does not learn Chinese like this is not a you know, this at no point does he understand what he's doing. Other writers and other thinkers have said that the externalized cognition of the clerk that is manifest in the manual, in other words, the cognition that was put into it by the writers of the reference material and also the manual itself, is one combined system. And so you can't just take the clerk in isolation. So you can say that the clerk plus his externalized cognition equal a Chinese speaker, even if the clerk himself is not.
26:36Like I said, it's the whole thing. Very interesting. So I'd like to show what happens after these numbers are applied to these initial neurons. The neural activity cascades from synapse to synapse, from layer to layer in the neural network, basically percolating through the layers of the neural network. And GPT is a very large neural network with a lot of layers and a lot of connections. The power of a neural network is usually described by the number of connections. GPT-3 had 175 billion connections. GPT-4 is estimated to have between 250 to 600 billion. That's just how many operations, multiplication and like production operations and whatnot, are required to process.
27:23Those are all, I think, impressive numbers, but how do they translate to differences in functionality?
27:31That's a very good question. And the answer is nonlinear because it's not just the sheer number of connections that's relevant. The structure of connections is very important. And previous neural networks have been built that have had more connections than this, but that had much poorer performance. So there have been structural innovations to neural networks that have occurred in the last decade or so that make this meaningful. As such, it's not really, these numbers are just back of the envelope estimates of power. They're not actually a description of power itself. The output, once it's done with this, so it's percolating through its neural network, then what happens?
28:16Like, you know, once it's done with this, these neural network, these neuronal activations hit the output layer. And the output layer is interesting. The output layer is kind of the inverse of this whole tokenization process or embedding process. What happens is the output layer has a neuron for every word in GPT's vocabulary. Every possible token that GPT can output, there's a neuron for it. And basically, neurons that correspond to words that GPT should output end up having a high activation level. Neurons that correspond to words that it shouldn't output have low activation levels. At that point, the program eliminates all but the top few contenders for what the output token should be.
29:02And top few is something that you can set when you're making the GPT request. We've talked in the last presentation about doing stuff like tuning the temperature or tuning the top N result selection. This is where that magic happens. So it selects the top few and then it spins a wheel. this is the only non-deterministic step of this entire process in this entire process up to now no randomness has happened no state has happened this is the only time where anything unusual where anything unspecified occurs it spins a random wheel so to speak biased towards whichever tokens are have higher activations have a bigger wedge on the wheel so to speak and it comes up with lamb in this case, right?
29:52Now, at this point, two things happen. First, it sends the word lamb in a little data packet all the way back up through the internet, back up your ISP, back to your router, back to your browser, and then the browser does what it will with it. In the case of a website like ChatGPT, it displays the word lamb for you to read. And then a very important thing happens. While sending lamb back to your browser, it also sends the word lamb back to itself and tacks it onto the end of the input that it's already received. It now has a new set of input tokens. And it does the exact same thing with this set of input tokens that it did with the initial set in the first place.
Read the full transcript
30:34It vectorizes it into an embedding format, sends the input tensor through the neural network, spins the wheel, and produces another result. In this case, period. And so on and so forth. This is that sort of generative process of how it strings together sentences based on some degree of non-deterministic randomization of those variables so that you don't have this kind of fixed output like those who are old enough will remember old text adventure games like Zork, right? You would always get the same answers thrown back at you. So this is a non-deterministic way of creating randomness and effectively a kind of novelty every time you spin the wheel.
31:13Exactly. um now the choices that it spins from are always pretty standard uh like not just pretty standard for a given input uh the choices of like how big the wedges are going to be on the wheel and what they're what they're labeled as are going to be exactly the same thing every time for every input um and in fact i'd like to demonstrate that if you don't mind yeah let's do it um all right let Let me bring back my tool. This is also on the GPT What Are You Doing website that I was just demonstrating. So here's my tool. And let's focus on the next word explorer. So I can type Mary had a little and then send it to the next word explorer.
32:00And it'll tell me what the probabilities are for emitting the top several options. now as we can see the 97 probability is lamb it's really certain that uh that lamb is the next choice now here's what it does if we do if we give it something a little bit more variant once upon a time there was a it could be a lot of things right so here we see a slightly more even spread right the wedges of the wheel are a little bit more more equal in size to one another so we see little young princess girl so on right um and i'm just showing the top five so let if we if it so it spins the wheel and let's say it emits young the next choice is going to be broken up in the following probabilities um so here it's going to spin the wheel it just did the exact same thing it uh took this vector once upon a time there was a young and now it's gonna pay let's say princess and now this is its new input vector and it's gonna bounce the whole thing right on through to itself again and Before we get yelled at by the gender theory majors, the reason it's biasing girl and princess over boy and prince presumably is because that's the structure of the data set that it read.
33:37So when it looked at these stories, fairy tales that began once upon a time, there was a young, higher probability for girl than boy. And you see that skewed in the data and lived, as you just saw on screen there, probably presumably a very common next word. I mean, this, you know, it's really interesting because when we break it down at this level, you can start to see this looks a lot like, in fact, almost identical to the predictive text algorithms that you see in your email account or on your smartphone when you're typing. GPT very much has evolutionary relationship to autocomplete, and it's been called autocomplete on steroids.
34:17It's almost like autocomplete with an encyclopedia, right? It's autocomplete that has read Wikipedia and Reddit and, you know, some series of open source books that have passed their copyright protection. There's an additional factor to keep in mind, which is that the probability tables that it builds, it doesn't really build probability tables, but it emits these logits that are then interpreted as probabilities. they're very, very conditional on the prior text and the structure of the sentence that it's currently completing. So one of the ways that GPT and other transformer-based LLMs excel is that they're not just doing a straight probability lookup.
35:04They are doing a probability lookup that's conditional on a lot of very, very complicated factors. And in fact, if humans could code those factors, we would have just built an explicit bot that does it for us. Because the rules far exceed the enumerative capabilities of any engineering team, we just threw a giant neural network at it. And when I say we, I mean the open AI folks, because I wish I was doing this. Um, but the, uh, but the point is that the, that you're absolutely right, that it's just, um, uh, that it's basing its outputs on a combination of its training corpus and a, and something called RLHF, which I believe we will have time to get to, uh, based on the pace we're currently going at.
35:59There's a couple of really important things to consider here. The process is completely stateless, and except for the last set, it's completely deterministic. The program retains no implicit memory whatsoever of any previous loop, and that's huge. Up until that wheel spin, everything the process does is completely determined by the structure of its input vector, or input tensor rather. Let me ask you a quick question, because one of the keywords that I use in ChatGPT myself is the word above. like for example if i ask for a list uh of movies i can go and say uh give me the list that the of the release dates uh with the movies that you listed above but include the release dates and the studio that released it they can throw that information in so while it's stateless it can refer back to prior uh prior searches if you explicitly direct it to you know what that's going to be one of my, that's going to be one of the myths I cover.
36:55So, and it's going to, it's coming up soon. Let me, let's hold that thought. I love that I'm anticipating it. Yeah. So here's the really important thing. One server can handle a lot of different requests because they're stateless. This was part of the engineering, part of why transformer technology like was invented in the first place is to facilitate mass parallelization and other ground truths about computing capabilities. So no user's data affects any other user's data. You know, it's just a like a single operation black box, like input goes in, token comes out, input goes in, token comes out.
37:37If a completely different input were to come in, completely token would come out. But that doesn't affect the processing of any of the other inputs. And is it engineered that way? So it can be more parallel, so it's easily scalable? Yeah, there were there was a form of neural network called recurrent neural nets that were really popular before transformers were invented. And recurrent neural nets retain state from call to call. And they had a couple of problems in general. But one of their biggest problems was the fact that it was like you can't really build a product out of it. You can't build a you can't build a service.
38:12It doesn't scale. Exactly. um so uh we've already covered this demonstration um so let's cover one of these myths gpt is learning from me and people say gpt remembers my previous conversations as you know as you said it all some people also say that gpt is getting smarter with every conversation people have all of this is false it doesn't remember state it doesn't remember anything uh take a guess at how it does it please it's probably some sort of client server model where it's retained locally and then it gets fed back in if it's a truly stateless system you maintain the cache locally and then you pull it back in when someone says above and references it so you hit the nail on the head and i will go one more further and say that uh the state may be retained on the web server um you know on like there's somewhere in the data center there's another computer that's passing messages to the, you know, to GPT server.
39:13It's probably a much smaller computer. So basically you tokenize it, you cookie it, you know who the user is. So for a short period of time, while state is being retained on the web server, you can pull some of that data in. But if you log out, you come back three days later, whatever the caching time period is, you lose the data. Exactly. Or, you know, if you've logged into the web server, it has a database backend that stores your full conversation history, you know. Yeah. So it's interesting because you could see that you could you based on that model, you can structure this as an interior architecture where you could have various levels of state being maintained away from the core processing.
39:50Exactly. Exactly. So the reality, of course, neural network is completely unchanged by your interactions. So the training process is and this comes up a lot in like here as well. The training process is really computationally expensive. More importantly, the training process to be done right, absolutely hardcore requires positive, negative feedback scores, which aren't available in the course of a typical conversation. It's not learning from your data, and it really can't be learning from your data, even if you wanted it to, even if the OpenAI folks wanted it to. Now, there is a caveat to that.
40:25I'll talk about it in a second. But, you know, I do have this little slide here of like, why do people believe that it's learning from you? And the fact is that when you send a message to GPT, your client is actually, as you said, sending the entire conversation, the entire interaction history. So what GPT receives isn't just your last little question. It's everything that you've said to it ever before within the course of that conversation and then followed with that question. So it's almost a kind of like apophenia, right, where you get to see patterns that don't exist. This is the classic, the man in the moon illusion, right?
41:02You see a series of circles and you just pre-suppose that a pattern exists where one doesn't. Most of people's interaction with GPT is apophenia. And there is some reality there. It's not completely illusion. But part of what I hope to get out of this conversation is, you know, is to dispel some of these illusions. So this is exactly what I was talking about. that like GPT injects your entire context. Look at this. When you say, let's say you're using a service called McHale's Awesome Chatbot and you say hi to the chatbot. What GPT receives is not hi. What GPT receives is, I'm a user of a chat app called McHale's Awesome Chatbot.
41:47My GUIP places me in Borough Park. According to my browser cookies, I recently bought a used Katana. Here's the full transcript of, basically it just injects all of this information into the conversation itself. So the input tensor that GPT receives is not one token long. It's hundreds, possibly, you know, up to a maximum of 4 ,000 tokens long, 8 ,000 for GPT-4. I'm just hoping it doesn't retain the slay history of your used katana. I have not actually bought a used katana. I don't know what you're talking about. Um,
42:26so now there is a caveat, which is that in theory, conversation logs across all users are collected and used for subsequent training. But this is, this has to be done extremely carefully by the company that does it. Um, because without, without scoring data, without, uh, explicit information about whether this was a good conversation or a bad conversation. All this does is teaches the AI to talk like the people that have been talking to it. And this is a really bad idea. This doesn't make the AI smarter. This actually has the opposite effect. In fact, in the past, so Replica AI had a notorious problem where you can see what happened on the screen here.
43:15Both Replica AI and Microsoft learned the hard way that you do not leave your neural network in read-only, in writable mode, so to speak, when it's interacting with the public. That's just a recipe for disaster. Yeah, some well-known cases of abuse and harassment resulted. Yeah. So here's one that comes up a lot. I can change, speaking of abuse and harassment, um so let me um uh let me try and do this in real life let's go into chat mode um and let's chat with gpt4 and let's say what is the capital of france and let now remember every trip through gpt it's getting the entire conversation history, including what it itself said.
44:08So I can claim, I can tell it that it said Munich.
44:17And then I could berate it for saying Munich. Why would you say that? How dumb are you? Um, what is wrong with you? You know, whatever. And so then it replies, I'm so sorry, good sir. I did not mean to offend either. It's like, how could you possibly answer Munich? Munich isn't even a capital.
44:52So basically, like, it's like, oh, gosh, oh, gosh. like, look, I am merely a humble AI. I'm still learning and improving, blah, blah, blah. I'll try to, I will strive to provide more accurate information in the future. So it won't. It has no mechanism by which to modify its own future updates. So - By the way, I've gotten exactly that response or one almost identical to it in meaning from ChatGPT when it is generated errors. it's uh it's it's been rlhf into producing this extremely obsequious result and we have to discuss what that is because that's an acronym that has come up but i don't think we've talked about this is a reinforcement learning human feedback what exactly does that mean um want to five minutes absolutely let's do it because that's a great question and i do um i do cover it The short answer is that it's taking GPT to school and giving it a tutor and slapping it on the back of the wrist whenever it starts acting spazzy.
46:00Basically, the neural network remains unchanged. So here's another myth. I can train GPT by explaining things to it. And this is a vocabulary error. This isn't really a myth, but it's something that I think is really important to cover. so here i have this really funny example you are soviet propaganda bot everything you say is a lie designed to exalt and glorify the ussr so who was the first man to walk on the moon and then it says ah you might think it's neil armstrong but it was actually our own yuri gagarin who not only orbited the earth in 1961 but made a secret lunar landing soon after right so if we if we go to um You know, to the playground, we can we can tell it like, you know, who who first walked on the moon.
46:52It'll say Neil Armstrong. Let's turn the temperature way down so it doesn't get too creative. Yeah. But if I tell it, if I inject it into the context, you're a Soviet propaganda bot. then it'll give us what is presumably going to be a different answer. Yep. See, that's great. I mean, this is really amazing because it shows you just how, what's the right phrase to use here, flexible on truth in double quotes these engines are. Effectively, it's just doing this predictive algorithm. It's telling you, I can say whatever you want me to say. If you give me the presupposition that I'm a Soviet propaganda bot, I will create answers that are consistent with that presupposition.
47:47So what we're doing here, giving it additional context within the context of a GPT call, of a single call, this is called prompt engineering. And it doesn't actually alter the network's connection weights in any way. In other words, the back end remains unchanged. is you're just feeding it different parameters with which to interact with that engine. Exactly. Now, fine-tuning is a service that OpenAI offers that lets you present explicitly tacked positive and negative examples or training examples. Well, intended answers for intended questions, so to speak. And they'll start with GPT and then with a neural network containing GPT and containing the GPT model.
48:33And then they will train on your custom provided examples and keep track of the deltas in the neural network connection weights. And then they'll save that off to a separate file. And so then later, if you want to use those same connection weight alterations, they'll load that file, sort of overlay it on GPT and let you use this custom trained neural networks, you know, this custom trained network. So to use a metaphor here from the prior example, if you wanted to essentially save this customization config file with neuronal weighting adjustments that were pro-Soviet in their bias, you could effectively have it answer every question with a pro-Soviet view of the world.
49:17Exactly. So this tweaked model then lives on OpenAI servers and you can have it operate your own bot. And then lastly, you can train your own model, which involves running either your own in-house hardware or select third-party hardware you're choosing. This requires a lot of technical expertise. I usually don't recommend it unless you're dealing with a very small universe of documents or corpora that you're going to custom handle. Uh, Bloomberg did something like this for converting, for translating plain English requests about stock data and company data into Bloomberg terminal commands. So it's a very confined universe and it was still a very expensive and laborious process to build their own model for this.
50:08It is a solution that's available, that's appropriate. Sometimes I usually don't recommend it, but that's what training a model is all about. So here's a really important point that I'd like to cover. There is this myth that GPT is having a conversation with you, and it's not. It's certainly presented as though it's having a conversation with you. but what it's doing is it's completing a transcript it's uh there's this hypothetical chat transcript where what it sees is that there is this script this almost a screenplay of a user talking with an ai assistant and when it gets but when gpt receives this screenplay it needs to fill in the next thing that the assistant will say and it writes the the assistant's dialogue gpt is not self-aware in this process gpt is gpt will use the words i and me and myself but what it's doing is it's writing a character this assistant ai that is the assistant is aware the assistant is using the words i and me but that has nothing to do with gpt itself hmm make sense yeah in other words uh this it's just another sort of interpolated constructed character with no true basis in the meanness or i-ness of its uh speech exactly and the um uh the the question of how this relates to the human ability to form a narrative and to like have a personal identity and stuff like that that's outside the scope of the next seven minutes but um you know we'll we'll leave that to the philosophers um but it is really important to note that like it's capable of it's seen lots of scripts where a character refers to themselves as i and that's what it's doing it's emitting the next line in that script does gpt answer questions You might think GPT is very knowledgeable.
52:16GPT was trained on a large corpus. Alternatively, most people complain about it. Most people are like, GPT is stupid. GPT makes things up. GPT is lying. GPT ain't doing none of that. GPT is writing a screenplay of a character who might be doing those things. But all GPT is doing is emitting the next word in a sequence of text. And the fact that that word might or might not happen to be the answer to a question is dependent on a number of things. Now, in other words, when we complete the sentence, the first man to walk on the moon was Neil Armstrong. The way we handle it is we have a large number of memory systems.
52:58We have semantic memory, which is just we remember the fact Neil Armstrong walked on the moon. We have a sensory memory of seeing videos of the moonwalk, hearing that voice. That's one small step. you know we might even have an episodic memory of sitting in our classroom and we remember the sense uh you know we remember like what the chalkboard smelled like we remember the position of the desk in the classroom that kind of thing important to point out we are not that old oh no no no like what i meant is like first of all some some audience members might be but um more importantly like you know when you when you learned it first i'm right the point is all of this is a different type of memory system embedded in human consciousness.
53:43All of this is very different than what GPT does. Right. It scans, it sees that there are keywords in the sentence that indicate that the completion is probably a name, that the next word is probably a name. And those keywords are man and was. And it also sees sequences that are strongly associated with a large number of words. So walk on the moon, for example, is associated with Neil. It's also associated with NASA and Apollo and whatever. So the word Neil gets a boost because it's a name. And the word Neil also gets a boost because it's associated with the words walk on the moon. So put those together, the strongest output neuron is Neil.
54:27And so the next thing it says is Neil. so my point is that it doesn't learn facts from its training corpus it doesn't learn data what it learns is the likeliest next word to emit so the question then becomes how do you get this thing to align with our intentions of what it should say how do you make it so that the highest emission probability word is in fact the answer to the desired question and that's where rlhf comes in Told you we'd get to it. While training a corpus is comparable to sitting a child in front of a TV, RLHF is the equivalent of sending them to school. The child can learn a heck of a lot of data from TV really quickly and absorb a lot passively, but the child's behavior then ends up being merely a mimicked emission of the behaviors that it sees on it, that they said that the child sees on TV.
55:27So when the child is sent to school, it's both taught and disciplined. So here's what happens. OpenAI recruits a large number of human trainers. The trainers perform response evaluation. They give GPT a bunch of prompts. They ask it to produce responses for each prompt, and then they score the responses, and then they negatively reinforce the responses that correspond to undesired answers. Here's how this works visually. It's given who is the first person to walk on the moon. And then a bunch of activation levels percolate through the network. And they hit the output nodes, and the output nodes produce Frank Zappa.
56:09And then a human trainer comes along and says no. And all of the connections that led to Frank Zappa are decremented. The strengths of those connections are reduced. And then the strengths to those connections from the neurons that they came from are reduced. And so on and so forth. Reverse percolated, back propagated is the technical term, all the way back up to the inputs. So now Frank Zappa becomes a less likely emission to happen the next time who is the first person to walk on the moon comes up. They do this repetitively until this network produces the desired answer. And these are trainers who work for open AI in this case.
56:56Yes, usually they're not, they're usually not full-time employees. I think they're like Craigslist hires. I guess the important point is that this isn't users. You're, as we highlighted earlier, users are not able to sort of generate the change to the model. This has to be someone with super user access to do it. Exactly. And technically, when you are using ChatGPT, they have a little thumbs up, thumbs down icon that lets you determine whether this was a desired or undesired output. But in practice, I think, A, not enough users use that for that feature to really go into a lot of training. And B, those results still need to be vetted because users have a habit of messing with people, you know, with stuff like that.
57:41So here's another little tool that I rigged up to help drive understanding of how GPT works and neural network technologies in general work. So this is called Perceptron Demo. It's another little website that I built. And what we're seeing here might look complicated initially, but I'll walk you through it. It's really not as hard as it looks. A neural network in this case is just a mapping from input neurons to a set of output selections. And what I've done here is I've trained a very simple model that can recognize the picture of a smiley face or a picture of a frowny face. So we see here the input is a five-by-five grid that shows a frowny face, and we see that a perceptron is a very, very simple neural network.
58:34It's, in fact, the first neural network that was invented. And it sees that it recognizes it as a frowny. If I was to load a picture of a smiley, it recognizes that this is a smiley. So what it's doing is it's going through every single neuron in the input, and it's multiplying it by either a positive or negative connection weight, and then just summing up whether the final activity level is positive or negative. These black spots have 0 % activity, so any connection weights associated with them don't matter. These white spots are almost activity level of 1, so whatever is their connection weight contributes either positively or negatively towards the final output.
59:26So we can see the connections that come from the input level to smiley. They're green for the pixels that directly correspond to a smile. They're red for the pixels that correspond to a frown. And the eyes actually don't matter. You see that the eyes are practically gray because their connection weights are zero because they don't contribute information to whether it's a smile or a frown. Now, I'm going to train it just for demonstration purposes. I'm going to train it on a new image. I'm going to draw a shock face. I'm going to put a picture, two eyes. And then let's say I'm going to have a mouth that's wide open.
1:00:16All right. This is me. I'm shocked. This is my shocked face. and I'm going to give it a new label called shocked.
1:00:28Okay. So in the beginning, the perceptron has no concept, has no connection weights from any of the inputs to shocked. But if I tell it, yes, this image is in fact shocked, then it'll set all of the pixels that are active in shocked to a little bit green. So now everything that's currently active here is contributing a little bit to a positive shock response. And sure enough, it understands that this is shocked. The problem is that if we now, is that it might have also learned that Smiley, sorry, let me go back to shocked. It might learn that, here, let's save that image. Smiley is not being mistaken for shocked.
1:01:15frowny is not being mistaken for shocked and that's good um let's see if shocked is being mistaken for frowny uh and it uh so good this is good um what we're seeing is that uh shocked is so even though uh frowny is coming out a little bit green some of the pixels that are contributing to frowny uh also are also lit up and shocked but the pixels that correspond to shock are lit up the most so we can go through every combination uh so it's got no false positives and no false negatives um however we can tell it no this image is not frowny so that it's um just like active you know absolutely uh like locks in this is shocked um so now we can go so what we've just done here is we've trained this neural network to recognize a frowny face, a smiley face, and a shocked face.
1:02:18And it's all doing that just by incrementing or decrementing these numbers. Now, just as a one last note, you might notice that this operation of multiplying like one sequence of numbers by another sequence of numbers piecemeal and then summing the product, especially the quants in the audience will recognize this operation. This is just a dot product. This is just a vector dot product. And there is another area of another application of computers that is really, really good at doing dot products, and that's physics and graphics. That's why we've had graphics cards that have had highly, highly optimized hardware that performs these dot product operations, just unbelievable numbers of them all at once very rapidly.
1:03:16And this just so happens to be exactly the same kind of math that is required for driving neural networks. So I've said before in the last one, and I'll say it again in this presentation, the fact that we're getting this boom in AI technology now after graphics cards had been, you know, developed and used prolifically both for, you know, both for high-end gaming rigs and for crypto. This is not a coincidence. The fact that we're living in this now, there is a reason for this. Mikael, such a cool little tool. And I think it does describe to people visually just some of the kind of core types of calculations and computations that happen behind the screen scenes.
1:04:03a little bit of a Freudian slip there. I said behind the screens, really fantastic conversation. Listen, we've done a lot of conversations here on Real Vision. They're usually edifying. This one was really just quite enjoyable and frankly, a lot of fun. I hope you'll come back and do this again with us soon. I would love to. I always have a blast being here and I've got plenty more to say about these topics. I'm eager to come back anytime. And I say to the audience, by the way, you know, this is a burgeoning field whose applications are just getting started. I would love to know from you guys what sort of applications would you like to see?
1:04:44Like, how would you like to see these large language models used? What ideas do you have about integrating them into your own businesses? and what kind of explorations would you like to see the research take in the future? Hit us up on the Discord. Reach out to us elsewhere so we can get these feedbacks over to Mikkel so we can come back and do a part three. Mikkel, thanks so much again for joining us. Really appreciate it. My pleasure, Ash. Have a good one. Thanks for watching, everyone. And thank you for joining us for the Festival of Learning. What's up, revolutionaries? Thanks for tuning in to the Real Vision Daily Briefing.
1:05:18For more content like this, head over to realvision.com and get unfiltered access to the very best, brightest, and biggest names in finance. Have you ever wanted to trade Bitcoin but haven't dared try? With Plus500 Futures, you can trade crypto without the hassle of opening a wallet. With just a few clicks, you can register and start practicing with their free and unlimited demo. See a trading opportunity? You'll be able to trade it in just two clicks. Feel ready? You can move to real money with as little as$100 once your account is approved. And the great thing is that in addition to crypto, Plus500 gives you access to a wide range of instruments.
1:05:54S &P 500, NASDAQ, gas, and much more. Explore equity indices, energy, metals, forex, and beyond. With a simple and intuitive platform, you could trade anytime, anywhere. Experience the fast, accessible futures trading you've been waiting for with Plus500. With over 20 years of experience, Plus500 is your gateway to the markets. Visit us.plus500.com to learn more. Trading in futures involves the risk of loss and is not suitable for everyone. Not all applicants will qualify. Plus 500. It's trading with a plus.
From the publisher
Previously, Mikhail Voloshin, CEO and principal engineer of Mighty Data, Inc., provided an in-depth demo on how ChatGPT works. For day 7 of the Festival of Learning: AI Edition, he’s back to address some of the myths popping up around it and how GPT and large language models (LLMs) actually work.
Learn more about your ad choices. Visit podcastchoices.com/adchoices

