How Generative AI Works, as Told by a PhD Data Scientist

14 Nov 2023 · 1 h

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Talking AI Podcast Episode Notes

Episode Title

How Generative AI Works, as Told by a PhD Data Scientist

Episode Overview In this episode of Talking AI, host Matt Paige interviews Nikolaos Vasiloglou, Vice President of Research ML at RelationalAI. The conversation focuses on generative AI, exploring its advancements, the role of knowledge graphs, and the implications for the job market and learning landscapes.

---

Key Topics Discussed

  1. Generative AI vs. Regular AI
  2. Definition and Differences:
  3. Generative AI creates new content (e.g., text, images) based on training data.
  4. Regular AI focuses on pattern recognition and decision-making without creating new content.
  5. Evolution:
  6. Shift from shallow models (e.g., decision trees) to deep learning and now to generative models driven by transformers.
  1. Foundational AI Models
  2. Foundation Models Explained:
  3. Foundation models, often called large language models (LLMs), use vast datasets to learn patterns and generate content.
  4. The architecture relies on transformers, which allow the model to understand and generate sequences of tokens (words or characters).
  1. Knowledge Graphs
  2. Definition:
  3. A knowledge graph is a structured representation of information that enables machines and humans to understand complex relationships.
  4. Integration with Language Models:
  5. Language models can leverage knowledge graphs to improve accuracy and reasoning, allowing for complex queries and reducing errors (hallucinations).
  1. Impact on the Job Market and Learning
  2. Job Market Changes:
  3. Generative AI is expected to automate several roles, especially in customer service and administrative sectors.
  4. Fewer engineers will be needed as generative AI tools streamline workflows.
  5. Learning and Education:
  6. Generative AI may change how books and educational materials are consumed, potentially providing summaries and insights without needing to read entire texts.
  1. Use Cases for Generative AI
  2. Legal and Educational Applications:
  3. Potential for legal documents and educational resources to be transformed into more accessible and understandable formats.
  4. Content Generation:
  5. AI could generate content for social media, enhancing user engagement and personalization.
  1. Concerns About AI
  2. Ethical Implications:
  3. The risk of misuse, especially with technologies like deepfakes, raises concerns about how AI may be exploited.
  4. Hallucinations:
  5. Generative AI sometimes produces inaccurate information, leading to the concept that language models can generate plausible-sounding but incorrect content.
  1. Future Outlook
  2. AGI and Human Interaction:
  3. The discussion touches on the potential for artificial general intelligence (AGI) and its societal implications.
  4. The Role of Humans:
  5. Emphasis on the need for ethical considerations and guidelines to ensure responsible use of AI technologies.

---

Key Takeaways

  • Generative AI's Evolution: Understanding the transition from traditional AI to generative AI highlights the significant advancements in AI technology and their implications.
  • Knowledge Graphs' Importance: They serve as a bridge between unstructured data (like text) and structured data, enhancing AI's ability to reason and provide accurate information.
  • Job Market Shifts: Generative AI will likely lead to job reductions in some sectors while creating new opportunities in others, emphasizing the need for reskilling.
  • Ethical Challenges: The podcast underscores the importance of addressing potential ethical dilemmas that may arise from AI advancements.

---

Links and Resources

  • [Visit the RelationalAI website](https://relational.ai/)
  • [Connect with Nikolaos Vasiloglou on LinkedIn](https://www.linkedin.com/in/vasiloglou/)

---

Conclusion The episode provides a detailed exploration of generative AI, its workings, and implications for the future, making it a valuable resource for listeners interested in the evolving landscape of artificial intelligence.

For more episodes and insights from experts in AI, subscribe to the Talking AI podcast.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Season three of the Built Right Podcast is right around the corner, but we've got one big change coming your way. The Built Right Podcast is now the Talking AI Podcast, and we've got a lot to talk about in AI. In the Talking AI Podcast, we'll be having in-depth conversations with both AI experts and early adopters of AI. That way you can understand how the technology works and how early adopters are beginning to implement and, more importantly, get value from AI. Our guests range from AI research scientists to founders of AI products to industry leaders putting AI to work in their business. While you're waiting for season three, go ahead and subscribe on your favorite podcast platform so you don't miss an episode.

0:41And make sure to leave us a comment about the AI topics that you want to hear about. So get ready to talk some AI in the new Talking AI Podcast, coming your way August 6th.

0:58Welcome to Built Right, a podcast by Hatchworks where we help you learn to build the right digital product the right way. In each episode, we'll deconstruct the layers of successful product development, break down popular trends, and offer real advice to help make sure your product is Built Right. We may not have all the answers, but we've built a lot of digital products across a lot of industries, and we've seen a thing or two. Let's get into it.

1:33Welcome, Built Right listeners. We have a special one for you today. Our guest is a PhD data scientist, and he's going to help us make sense of generative AI and how it all works. And our guest today is Nick Vassiloglo. And Nick, I probably butchering the name, Nick the Greek, I think is what you also referred to you as. And he is the VP of Research ML at Relational AI. Like I said, master's and PhD in electrical and computer engineering here from Georgia Tech and has founded several companies, has worked at companies like Logibox, Google, Symantec, and even helped design some of Georgia Tech's executive education programs on leveraging the power of data.

2:17But welcome to the show, Nick. Nice to meet you, Matt. Thanks for hosting me. Yeah, excited to have you on. And relational AI, for those that don't know, is the world's fastest, most scalable, expressive relational knowledge graph management system, combining learning and reasoning. And for those of you thinking, what the heck is a knowledge graph? We will get into that. Plus, we'll get into how generative AI actually works, as told by a real PhD data scientist who's been doing this stuff way before ChatGPT was even a thought in somebody's mind. Plus, stick around. We got Nick's take on what are going to be the most interesting use cases with generative AI in the future.

2:58And he's saving these. I haven't heard these either, so I'll hear him for the first time. So really excited to get into these. But Nick, let's start here. What is the difference between, we hear generative AI, it's the hot topic right now, but generative AI and just regular AI. What is the difference? What makes generative AI special and different? It's a very interesting question. For many years, the emphasis, the way that we were separating, whether you call it, machine learning on AI, was on the depth of the models. When I started my PhD, we were working on something that we would call shallow models.

3:39Basically, you can think about it looking at some statistics. The decision tree was the state of the art, which meant, okay, I have this feature. If it isn't greater than this value, then I have to take the other feature and the other feature and come up with a decision. That's something that everyone can understand. Then deep learning was the next revolution somewhere in the 2010s it started. and it started doing more, let's say, complicated stuff. People are still trying to find out why it's working. They cannot understand exactly the math around it. And then the next revolution was, so we had these models that they were pretty powerful, but we didn't know how to scale them.

4:27We didn't know how far they can go, and that was the revolution that basically OpenAI brought that they realized that you can take this new cool thing called the transformer where you can feed it with a lot of data and do this cool thing where you are trying to predict the next world and basically come up with what we have right now. It took several years and several iterations. but I think the difference between what we used to call AI and what we call AI right now is the focus on the language. I mean, if you had read about Chomsky and others, a lot of people considered that the human intelligence has to do with our ability to form languages and communicate.

5:21I mean, you might have heard that, you might remember as a student. what makes humans different than other animals. The human brain is the ability to form languages. And I think the focus on that made a big difference in what we have right now. Quick break in the pod. If you're listening to this podcast, chances are you've been thinking about how to actually use AI inside your business. And that's exactly why we built the AI Opportunity Finder. It's a free tool that helps you uncover high impact, tailored AI use cases based on your business, your goals, your pain points, and your industry. No fluff, no generic use cases, just real ideas that fit your business and they're ranked by ROI potential.

6:08It takes about three minutes to run and it's like having your own personal AI strategist for free. If you want to try it for free, check out the link in the show notes or go to hatchworks.com backslash AI dash opportunity dash finder. Previous was more like a decision system. Now we're focusing more on the reasoning side. So I would say this is the shift that we see. And that's, I think, part of the interesting aspect of it is, you know, in the past, it's like the models, they were trained for very specific tasks in a lot of ways. And now you have this concept of these foundation models, which that's, you know, what a lot of the large language models are built on.

6:47And but now to your point, it's almost kind of like getting to where how the human brain works so they can tackle these disparate types of ideas and solutions and things like that. This concept of like a foundation model, what is that? How does that start to play into these concepts of like large language models, LLMs that we hear so much about? So let me clear that up first. You know, the foundation models and the language models are basically the same thing. You know, sometimes the foundational models were, the term was introduced by some Stanford professors. They were trying to, you know, kind of, it happens a lot in science.

7:32You build something for something specific and then you realize that it applies to a much broader, you know, class of problems. And I think that was the effort, that was the rationale behind renaming language models as foundational models, because they can do the same thing with other types of data, not just like text. Okay, so you can use that for proteins, you can use basically for whatever represents a sequence. okay so um you know as i said in the past a lot of effort was um was put on collecting labels and do what we call supervised learning uh the paradigm shift here was in in what we call self-supervised learning that was a big big um plus something that that brought us here this idea that just take a corpus of text and try to predict the next word and if you're trying to predict the next word you're basically going to find out the underlying knowledge and ingest it in a way that you can make it useful of course that's what brought us up to 2020 that was the GPT-3 where we scaled, but there was another leap that charged GPT that in the end, it did require some labeling because you had the human in the loop.

9:09It's not exactly labeling, but you can think about it as labeling because we have a human giving feedback. And then that brought us to charge GPT. Now, the heart of language models or foundational models is something called the transformer. It was invented in 2017 by Google, actually. It was an interesting race. OpenAI, there was like a small feud between OpenAI and Google. So OpenAI came with a model. All of them were language models. Everybody was trying to solve the same problem. They came up with something called Elmo, and Google came back with Bert from the cartoon, from the Muppet Show, I think.

10:05And then, so Bert was based on the Transformer. Okay. Then OpenAI realized that actually Bert is better. Okay, that's an interesting lesson. They didn't really stick, oh, this is our technology, we'll invest in that. They saw that the transformer was a better architecture. But then they took BERT and they actually cut it in half. Okay, and they picked, actually let me put it that way. They invented, Google invented the transformer, which had an encoder and decoder. And they picked BERT was based on the encoder architecture. They took that half. But then OpenAI came and picked, no, we're going to work on the other half, which is the decoder, predictive text.

10:52And, you know, they spent three years. They did see that the more data you pour, the better it becomes. Okay, so that was their bet. And they ended up with GPT-3, GPT-2 and GPT-1 to 3, the sequence 3.5 and 4 later on at GPT. and it was kind of like an interesting race where things basically started from Google but OpenAI ended up being the leader over there and everything they built was open source, right? Everything Google built so they were able to... Everything is actually open source I think up to GPT, even GPT-3 there was a published paper it's very hard if you believe that you're going to get a secret source that nobody else knows.

11:43I've never seen that play in machine learning. The scientists want to publish, they want to share knowledge. I think as the model started to become bigger and bigger, they didn't with GPT-3, I don't think they ever opened the whole model, the actual model, but they gave enough information about how to train it. There's always some tricks that over time, even if somebody doesn't tell you as you're experimenting, they're going to become public. Okay. So, yeah, that was never the issue. I don't think... Yeah, they are a little bit cryptic about after 3.5 and such details. But in my opinion, the secret sauce over there is not exactly on the model, but it's on how you scale the serving of the model, We're going to talk about that later.

12:37This is the secret weapon of OpenAI, not necessarily the architecture, but, you know, the engineering behind that. Nice. Let's keep going on the transformer side, because getting under how these, you know, GPTs work. Basically, you mentioned that it's serving up the next word, the next word. It's not looking at it like a whole entire sentence, right? It's this concept of tokens. But how is it like actually thinking through that and structure of language and something you think a computer wouldn't be able to do? It's now doing very well. Yes. First of all, it's always this. As I said, the transformer has this encoder decoder architecture, which means that there's one part that looks into two directions back and forth.

13:23As it's processing, it looks both ways. Like this token is affected by the other tokens, but this token in the middle is also affected by the ones before and after them, what it's going to be. So that's like the encoded. The decoded architecture, you're only looking back because you're not looking in the future. We can talk more in detail. There's a lot of papers and a lot of tutorials that they actually explain that it's not always easy to explain it without graphics here. But the key thing over here is that, you know, let me go a little bit back. The first revolution actually came by Google was Word2Vec, where they realized that, you know, if you give me a word and I try to predict that word by looking five words behind me, and five words after me.

14:22Okay, that was a simple thing, like a small window. And I tried to create a vector representation. They realized that, you know, I can take words, make them as something like continuous vectors, put them in space, throw them in space. And I would realize that, you know, that representation would bring words that are semantically similar together. Okay. And there was this other thing that if I, you know, if Paris is here and France is here and London is here, then I can take the same vector, put it here, I can find, you know, England. So they realize, for example, that if I place all the capitals and all the countries, I can just take the vector that connects the first and the other and it's translated to the next one and it's going to – or if I take the distance of the vector between man and a woman, take that, then take the word king, add that to the king, it's going to take me to the queen.

15:25So basically people started realizing with a simple word to vector that you can take words, represent them as vectors. Let's think about two-dimensional vectors like on the plane, but it's not two, it's like 512 dimensions. It doesn't really matter. The concept is the same. that their distances in space, the way that they're placed in space has. Sorry, semantic meaning. Now, the next problem was, this is great, but we do know that words actually change meaning based on their context. Okay, so... What's an example there? Yeah, so for example, an example would be, when you say when you say flour, well let's pick now I'm a little bit stuck but You had one about boiling the a person's boiling a what and then if it was like an engineer it had different context.

16:33Yeah, you can boil an egg or an engineer is boiling I don't know a substance but it could be They're trying to boil the ocean. So when you say, for example, the bed, it can be, you know, something different when you talk about, you know, a house, a bedroom. But if you talk about geology, it means something completely different. OK, so what they realized was that that vector that represents the world shouldn't be universal. It should really depend on the surrounding words. So this vector representation when the surrounding words are this one, it has to be this, and it will also have different relationships.

17:27And it should be different when it's around different words. And that was basically Elmo that was this idea. It's called contextual embedding. So this vector representation, like this, let's say these two dimensional representations called embedding. So this is this was actually one of the biggest revolution of deep learning that we're taking discrete entities and we could place them in space as a continuous vector. So we'll take something that was discrete and putting on a medium that it's continuous. OK, continuous and multidimensional. So the first idea of the before the transformer, which is an aerial version of a transformer, the first idea was that, OK, if I see a text, I will be placing these words, you know, on different places in space based on what it's around.

18:21them okay and um and the next thing that so basically what happens is you're taking the words on the first level you you look left and right and you create you know embeddings you know you create you put them in space then you take that and you apply that again and again so the transformer actually starts in levels one level after the second level it has multiple i don't know exactly the numbers but it has different relevance so you can think about that as basically a rewriting okay so that's why it's called transformer so you have a sequence of words and you start you know rewriting to something else something else that else you know so when people actually people have done this experiment they're taking the the transformer and they decompose it and they see what are these things that you you know the transformer does in in different levels and they've actually realized that it starts inventing grammatical rules.

19:21It starts like identifying what is the subject, what is the object, what is the verb. Okay. It starts identifying that something is an adjective or not an adjective. It starts, you know, taking, you know, words and converting them to something which is synonym, maybe, you know, something else. And that's how the reasoning starts. Like, I can give you, if I give you a sequence of words, you know, Nick, I don't know, lives in Atlanta. You know, he can, he knows that Nick, I don't know, is Greek. Okay. So he can say the Greek lives in Atlanta. and that can affect the fact that, you know, and then you can say he goes to the store to buy and because now you know that he's Greek, he lives in Atlanta, you say fetuses, for example, okay?

20:21Because now he starts, you know, the transformer starts taking different paths. Like it starts exploring, you know, what are synonyms and, you know, if he leaves, it means he goes to the store, you know, he goes to the supermarket if he leaves there. So it starts, all this information is ingested in the transform after seeing, you know, endless pages of text and, you know, where basically there's, there are reasoning paths. Like it does this on your own. Of course, because there's so many reasoning paths that can happen. Sometimes it can hallucinate. Okay. So I can say Nick buys, I don't know, so lucky because he agrees he is Greek, which is possible.

21:02but there might be somewhere else some other information that says Nick hates Vlaki and you know the language model doesn't but it's a probable event since you know Nick is Greek anyway I'm just giving a simple example over there but that's kind of like the power of the transformer that at every stage it starts rewriting things again and again and again and it explores possible very possible uh very likely paths you know highly likely and correct me from wrong what what you're talking about here is this kind of the the difference in evolution from structured data to unstructured data because in the past we had very like defined tables columns associations to things is this kind of getting to that concept of unstructured data where it's like the vectors and well the problem with with the structure that with the systems before that everything was like it was very discreet and unless you had seen before the word nick okay followed by that exact word.

22:03It was, you know, if you think about all the variations, like Nick spelled with K, Nikolaus, Nick Vassiloglu, I don't know, think about it, all these things. Because now they're in a continuous space, okay, that's what makes the difference. It's possible for the system to create an internal rule, if you want, or internal path about things that are kind of similar, okay? Okay, so it doesn't have to be Nick. It could be Vasiloglu instead of Nick. Or it could be the guy who lives at, I don't know, say my address. You know, it's the same thing. So because all these things, I think it's public, you can find it.

22:42Because all these things that are semantically equivalent, and before you had to express them in, I don't know, a hundred different discrete things, and you had to see them exactly in that order in order to find a common path. It says, okay, this class of entities that can be represented with this vector, they are very close, can be followed by this class of entities that can all compress them in a constellation of vectors. It can lead me to something else. That's why you see the language model when you go to OpenAI and you say regenerate. What it does, it can generate the same thing, the same reasoning path by using a little bit different words or, you know, where's the semantical equivalent.

23:31Okay. And now the thing is that it can do that in this incredible memory of like, I don't know, up to 32 ,000 tokens. So even if you're saying that Nick is going to buy something from the store and it will predict that it's FETA, it's because it has seen 10 ,000 tokens before that Nick is Greek. He's hungry. He's having a dinner party and get you over there. Okay, so because when it was trained, it has seen sequences that in the span of 10, 20 ,000 token, you know, nick associated with party, food, restaurant, you know, leads you to FETA. Yeah, and when you say a token, that's basically either a word or a couple characters, some like small variation that it's breaking it down into.

24:26The token is basically a trick. You know, we could have used, it's like, you know, I think all these models have about, I don't know, 30 ,000 tokens. So they realized that we can break all possible, like, you know, with 30 ,000 tokens, you can, I mean, you can use character level. Okay, every word can be decomposed to characters. but that would have made that would have made you know the language model extremely big and inefficient so it's like a trick because we kind of like trying to find out it's a compression that we're doing we could have gone with syllables because syllables are also finite and make all the words now we said you know look because there are some combinations of letters that they are so frequent we don't really need to decompose them all the time we know exactly what they mean.

25:18So it was a clever engineering trick. It has to do with the language. It's related to the language. It was like a statistical, a better statistical analysis of the language. I mean, to put it that way, if we were inventing a language from scratch, we would start with tokens, you know, and maybe not necessarily letters. that's interesting and so we've talked a lot so far about language as the the thing at play here but like you can use this generative ai and all this new technology and advancements with different modalities like images whether you're generating images or whether you're understanding what an image looks like and voices all kinds of different things that play here how does that work different when now language isn't necessarily the output?

26:12Is it looking at the pixels in a way and then association? There is a visual language over there. There's the visual transformer which tries to predict blocks of image. There's also the diffusion models which is something completely different. So for example, diffusion models, we see them only in images. We don't see them in text that much. although there's been some efforts but the transformer it turns out that it behaves equally well for images but when you talk about a token in vision that's kind of like a block of pixels I don't know 16x16, 32x32 this is something we knew from before like even in the days of image compression they could take parts of the image and compress block by block.

27:08But I want to make something clear for your audience that language is a way of expressing knowledge, but it's not knowledge. The fact that I can come and tell you something, I can go and read quantum mechanics, I can take a passage, I can recite it for you, it doesn't mean that I know what I'm saying.

27:41And that's where the hallucinations are coming into play. So we don't really have direct access to knowledge. It's a language model. It's not a knowledge model. And there's been some effort right now to do the same thing. If we could start, if there was a universal knowledge graph, that I could take and say that from this token of knowledge, I can go to that token of knowledge through that relation and do reasoning. Maybe we could train a knowledge model, let's call it, or a foundational model that we know that whatever it says, it's accurate and correct. But language is a possible path over knowledge.

28:29It doesn't mean that it's correct. Okay, so it doesn't have to do that. So language models are always going to hallucinate and make mistakes, not because there are errors into what they were been training for. The data sets are pretty well curated. Obviously, they will contain misinformation and errors, but the reason of hallucination is not really the errors in the raw text. But it's on the fact that this is a possible expression, you know. The same way that, you know, like you are a fiction writer, author, and you can write. Like you see things in life and you write a different version. Like take one of my favorites, like Da Vinci Code, okay?

29:20Like when you read, that's what I like about Dan Brown. Or take about Game of Thrones, for example. If you think about Game of Thrones, it has elements of truth from the human history. You can see the, let's talk about it, because that's probably what most of the people know. There's like the Persian Empire, or you can see the English history or the Greek, or there's some of them, like you can see elements of that in a completely fictional way. So that, in my opinion, Game of Thrones was the first generative model, you know, George Martin. Great. Okay, so it could generate something like that, which is completely, it looks, you know, Mauds of the Dragons.

30:02It could look real, okay, realistic, but it's wrong. The same thing with Dan Brown, you know, Da Vinci Code. It looks like a real, it could have been a real story about what happened after, you know, this was crucified in the story. It could have been, but we don't have evidence that it is. Some people follow conspiracy theories. They think that Dan Brown is the real story, but that's what I'm saying. So, yes, it's a possible truth. truth do you think we ever get to that ability where it is true knowledge you get into this concept of like you know your agi and all that type of stuff do you ever think we get to that level of advancement or you know i always go back to like how the human brain works and like are we do we have true knowledge to an extent or are we just doing the same kind of computational thing in our head with probability of what's you know yeah one of the things that we know is that the transform architecture and the language model is not how the brain works.

31:04This is an engineering, it's not how the brain works. There are some commonalities and there are some kind of lack analogies, but I think it's wrong to think about or to try to, you know, like when you're working with language models and you're trying to tune them or you're trying to explain or debug them to have in your mind how the brain works. Don't do that. If you are a prompt engineer, if you're trying to build a model, try to understand how the system is built and use that knowledge. Don't use the cognitive neurology here. Now, unfortunately, we are very, you know, the human brain is still much more powerful given the fact that you can eat a slice of pizza and do very complicated mathematical computations while if you were trying to do the same thing with GPT-4, you need the power of a village or something, even for inference.

32:05Okay, so we are very energy efficient. You know, we use signals that takes milliseconds to transmit, not, you know, nanoseconds, whatever it takes for a GPU, and we still do things faster. Okay, so there's a completely different world. Even if we could make an electronic brain simulated, I think it would be very different. The biology comes into place. It's still a mystery. But whether we're going to reach AGI, you probably hear that. I leave that to people who have enough money and time to think about it. So I mean, yeah, in theory, it is possible. I hear Hinton and Benzio and what's his name?

32:59I think Lacoon is on the other side. And Elon Musk, that they say it's possible for, you know, you leave a language model, start rewriting the code, and unplugging other systems. I don't know. I think not to worry that much about it. I worry more about the effect that it's having right now on the job market. That's more imminent and more real. or the economy, then, you know, whether the robots will revolt against us. And what do you mean by that in terms of it taking away jobs and tasks? Or do you think this unlocks new opportunities? I think it does, yes. Yeah. You know, as with everything, you know, it happens all the time with, you know, with a high tech.

Read the full transcript

33:49as technology progresses the next generation requires less engineers you can see about I don't know how million of employees Ford has when the car came and compare that with Microsoft compare that with Google compare that with Twitter compare that with OpenAI now that you know how it's a big chunk of uh um of the market they're getting you know like their capitalization and the small number of engineers that are scientists uh that they that they need okay and uh yeah it's pretty clear to me that a lot of jobs now can be done with less people and even for us the data scientists for the moment if you want the work is becoming a little bit boring in the sense that you have to do what people call like prompt engineering.

34:55I don't know I find ways to make it more interesting but yeah it's becoming an issue. I feel like we saw all this tech layoff wave the past two, three years. I think a lot of these jobs will not come back again. They will need less people for that. And of course, for things like customer service or administrative work, all of them will be done. I mean, it's already pretty obvious you can do things with GPT much faster than before. It's a great assistant. So two more topics I want to hit before we wrap. The one, you think of these models, there's this element of it being a black box. And we touched on it earlier with relational AI having this concept of a knowledge graph.

35:47What is that? How does that work? And that kind of gets into the value prop of relational AI to an extent. But we'd love to kind of hear how the benefits of that concept. So the knowledge graphs and language models have a bidirectional relationship. um the you know first of all this is very simple definition which i really like about knowledge graph it's the language that both humans and machines understand okay it's a way of expressing knowledge in a way that you know anyone can read it and the machine can consume it like if i write you know c++ code it's very easy for the machine to understand but it's not easy to show it to your executive or to your business analyst.

36:33So a knowledge graph has the right level of information. It's complete and both systems can understand. Now, the problem with knowledge graphs has always been is, you know, it's great, but where can I find one? Like, once you have it, it's great. It empowers a business. You see, you know, the ROI is huge. Okay. It's like, you know, So you are in your house and you go to your library, to your room, and you tidy it up. Once you tidy it up and label everything and you know where everything is, then your life, you're very efficient. But who has the time to do that? And that was always a barrier for us.

37:15Now, what happens is with language models, you can automate that very easily. Because in the past, how did you build a language model? So how did you build a knowledge graph? You had somebody going through documents or databases and was trying to find, you know, global entities and relations and how things are, you know, flows and all these things. Now, the language model can do that for you with a human in the loop with supervision. So it accelerates that process very quickly. Now, the other thing is once you have a language, you know, you have a language model, as I said, you need to inject knowledge and you need to teach it stuff.

37:51So the way that I've seen it is that, let's take some simple examples. Something which is kind of like the Holy Grail. You want to answer a question. You know, you say, well, tell me all the sales from last month where the people bought more than X, Y, Z. And that translates to a SQL query. Okay. So in order to do that translation, you know, like from natural language to SQL, for example, If you have a knowledge graph, we have evidence that this can become faster. In some other cases, the knowledge graph, because the knowledge graph can afford really long and complicated reasoning paths. You have your knowledge graph.

38:30You can go and mechanically generate, you know, let's call them proofs or reasoning paths. And you can take them and go back to the language model and train it and say, you know, when somebody is asking you this, this is what people call the chain of thought. It can be a pretty lengthy. okay so the the end of course is the hallucination thing where you can think keep you know the knowledge graphs always has the the the you know correct knowledge and it's very easy to add and remove you know knowledge that it's valid or invalid anymore so that's another part that you know helps you keep things in place so uh um so yes so knowledge is a language model helps you build a knowledge graph, tidy up your room, tidy up your knowledge.

39:20And then the other way, having all that knowledge, you can go and retrain, fine tune, control your language model so that you're getting, you know, accurate results and better results. Okay, so that's kind of like the synergy between the two. No, that's really interesting. Interesting thing, evolution there. So as promised, we talked about you had some use cases in your mind where you think Gen.AI is going to like the most interesting, viable, disruptive, whatever it may be. So curious to see what some of those are. So let's close with that. I mean, these are things that let's call them historians of technology have observed over the years.

40:03So we know that whenever a new technology comes, people are trying to use it in the obvious way, which might not really give them the big multipliers. So I think when we met, I mentioned this example of the electric motor. Okay. So when it was invented, those days, the industry was using the steam engine. And the way that they had, you know, they had a big steam engine in the middle and they had, you know, mechanical systems that they would transmit the motion to other machines around that in order to produce, I don't know, something, you know, it was an industry. Okay. And now somebody comes and says, okay, take this electric motor, which, first of all, is not as powerful as a steam engine by definition, because the steam engine will produce electricity, something will be lost, and then a motor will use it.

40:58And, you know, the steam engine was there for centuries before the, at least 100 years, I don't know if it was centuries, before the electric motor was more optimized. And all of a sudden now you needed to buy electricity to fit that, while for the other one you had, I don't know, fossil fuel to use and you knew where to find it. So, you know, people rejected the electric motor at the beginning. They couldn't know why it was useful. until someone said, well, wait a minute. We don't need one electric motor for the whole factory. What if we create, because that's so easy to manufacture, what if we made like 100 electric motors spread in vertical space?

41:39I don't know. So take the whole production and spread it over a bigger space. And all we need is an electric generator that can feed 100 motors. So the big benefit wasn't by just having one stronger motor. The big benefit was by having, you know, a hundred motors in different levels and making the production, you know, a multi-level and expanding it to bigger space. Because the problem with the steam engine is that motion couldn't be transmitted too far away and everything was crammed and limited. I think if I remember that took about 20 or 30 years to figure that out. And kind of like the same thing with, you know, let's think about Amazon.

42:22Like in the beginning, the e-stores, they were basically trying to take a brick and mortar store and run it the same way they were running it before, run it like on the web. And Amazon realized that there's other things like, you know, there's recommendations, there's the A-B test, there's other things that I cannot do in a brick and mortar store. the tailor, the personalization that brought the big boom of, again, it took several iterations of failure until Amazon and Alibaba and others kind of like dominated the market. Think about Snowflake. When Snowflake came and said, we're a cloud database, I said, what do you mean?

43:05I can take my database and put it on the cloud. But the thing is, nobody thought about designing a database that it's going to, you know, you can't download Snowflake and run it on your machine. It was designed to be completely cloud-based, use infinite compute and infinite storage. Okay. So it's a very different thing. People were confusing the cloud hosted, which means that I did something that when I take something that I build it, thinking that I'm constrained by the memory and the compute of a single machine and I'm just like running it somewhere else on the cloud versus no, I'm building a system that is going to rely on, you know, infinite machines and S3, whatever blob storage, which is infinite and a little bit for scratch.

43:54So I'm trying to scratch my head here and see what is the, yes, there's the obvious application of Gen AI, which is, you know, use it as a new UI. Okay. So chatbot, we know about that. But I was thinking, you know, I was trying to make this exercise, like we're looking for these businesses that only exist, they cannot exist without Gen AI. So I think the one that we're going to see soon is, there's already a legal battle about that, which is going to blossom and give the new thing. I think it's going to change completely the way we're reading. So you might have seen the fights between authors and open AI about infringement.

44:42And I think it's going to end up in a beautiful relationship over there. So right now, there's a problem. People don't read because they have to go and buy a 200, 300, 400-pages book where they're only interested in four or five pages. or even a summary of 20 pages that nobody is providing for them. They don't know where it is. So what I'm envisioning over here is, you know, think about Random House taking all their books or all the publishers and training a language model. And they're saying, you know, I'm asking a question. And they're basically coming up either with two pages and say, you know, actually this thing, you can find it in that book.

45:23And, you know, here's a summary. And these are the three pages. and I can actually take these three pages and put half a page that has all the information that you might need to read these three pages. Okay. Because that's another problem. Sometimes you can browse a book and find that's up there. But then as you're trying to read it, you realize that you need to go and visit others. So basically what's going to happen is, you know, you're going to buy pages from books or a summary that was produced based on, you know, 10 pages. So now you will pay, I don't know, 10 pennies or a subscription or something like that.

45:58I think it's exactly the same thing that happened with streaming. If you remember the legal battles of YouTube and Viacom where people started uploading videos on YouTube and they said, no, it's mine, it's ours, it's yours. And eventually they came out in agreement that changed completely the way that we listen to music. Spotify was another thing. But it took some friction. So we don't buy CDs, 12 songs or 16, however they had. You know, we listen to, you know, one song at a time. We don't own the songs anymore. You know, we just stream them and all these things. So I think that's one of the applications.

46:42Now, I have a reservation. Yeah. Well, as I said, it's like spark notes on steroids almost. One question, though, I guess if you're reading for fun, do you get the same pleasure and benefit from that type? Or is that a different use case where you're wanting to sit down and enjoy a book, I guess? That may be a different type of thing versus getting learning. I think it can help everyone. It can help the bibliophiles, you know, because I often – I have about 2 ,000 physical books and other 2 ,000 electronic books I like. But I'm always frustrated. You know, audiobooks was another thing that changed the way that we just know.

47:19But it's always frustrating when, you know, sometimes it takes like you, if the book doesn't stick with you for the first, I don't know, 20, 30 pages, then you give it up. And it's very likely that then if you're a little bit more patient, maybe after page 50 will become more interesting. But how many people give up before that? uh so as a bibliophile you know it's gonna help me um you know discover more books but i think the biggest thing is for people who um who want to learn something but they don't want to learn the they don't want to read the full book i read somewhere that they said that are we out of time no keep going keep going yeah so there was there was like this theory that you know a hundred years ago when you were writing a book you had to make it very big because people didn't have to do anything else so they were buying a book to fill their time okay because they wanted to to spend i don't know a month reading it now these days they say that a book shouldn't be more than 200 pages because uh you know don't try to fluff around because there's so much information and people don't have the time to you know they need the essentials and don't want to to spend too much time on other irrelevant stuff.

48:34The same thing happened with TikTok. TikTok, again, it was a victory of machine learning over there and recommendations trying to narrow the span to a few seconds what you're going to compute. Of course, it's a great commercial success. I personally don't like it. I don't let my kids spend time. I realize that it's so addictive. YouTube search, you can spend hours that's going one by one it's it's dopamine injections uh but we're definitely going to see social network space completely on gen ai and videos okay that's kind of like another one the same thing that we found you know we had tiktok um and uh yeah i don't know i mean we if you are a founder you have to start thinking about how can i take a sea of content and serve it much better with a language model.

49:35Okay. In a way that people wouldn't have consumed that before. Yeah. In the book example you mentioned, I have the same problem. I do audio books and I'll try to save the clips of the things that make sense. And at that point in time, you have this light bulb moment and then you forget about it. But there's a point in time in the future where, man, that would be super applicable if I could pull that out of my knowledge base. So it's almost like, to your point, getting those points that are applicable at that point in time, but resurfacing them because they're somewhere in my memory that I can't necessarily always return.

50:11You know, let me give you a recent example. So, and that's why I think this OpenAI has a big advantage right now over Google. So with all the unfortunate events happening in the Israel-Palestine conflict right now, So I remember that I had watched the documentary 20 years ago at Georgia Tech about the whole history of the area. But I couldn't remember the title of it. So I knew that it was – I remember it was a French production. I remember that it was released somewhere in the 90s because it was right before the Oslo agreement. And I think basically that's what it was. I remember it was a documentary.

50:51So I was trying to find it on Google. I was trying to find it on Amazon. I couldn't find it. but I went on OpenAI and I said, well, I was a documentary I think it was released early 90s I know that it was a French production and, you know, it was we had the history from the 1900s until 1990 can you tell me which one is because, you know, there are not really that many I thought that someone should have been able to and it actually found it it gave me the title in English and in French and I went to Amazon and I found it. So I think it was remarkable. That's cool. Yeah, and just to wrap on the points you made about the TikTok and everything like that and just that type of social media, like you wonder to a point, does it get so advanced to where you literally cannot put your phone down?

51:45It gets you so zoned in with like the dopamine hits. Like, is it engineered to a point where the recommendation of what's coming next, like it's kind of scary to think about you know in the future where it becomes you literally it's like a drug in essence oh yeah that that is gonna have i i agree with you like if tiktok is is a problem right now where it's basically trying to find existing content that you're gonna like um think about if it knows exactly what you like and you can give it uh uh you know feedback like you know so it knows more and more like you say what you want and it's really you know personalizes things for your content once it gets a step further too like what if it's not just random users generating the content what if it is um you know a gpt or something like that that's generating content okay yeah so that's wow that yeah generating and it can't you know think about when you were raising a kid where you say well we have this inherent thing of going to the taking the path of least resistance and basically things that are not good for you so that's why you have to say no to a kid imagine now that also think about it like society has created like this moral boundaries that, you know, prevent you from doing things that maybe they're in your mind, but you say, you know, I shouldn't really take that path because that's immoral.

53:26But what if you are, you know, in your screen, nobody's looking, and there's somebody else that says that, oh, okay, tell me what you thought. I can actually, you know, create this for you. And people, a lot of people are going to get tempted. And that's like a really bad spiral. I mean, these are fears before, you know, AGI taking over and leaving the matrix. I think these are bigger fear. And we do see it in some, you know, some applications in the deep fakes and things like that. I think it can become, and people have said that, you know, this is this kind of addictions like drugs, you know, it's the same thing.

54:09the screen addiction, especially when it takes parts that are problematic. So I would worry about that. We need some strong resistance in that. I'll give you an example. I don't know, for example, let's take one of the most horrifying things, which is child pornography. I know that by law, even possessing child pornography is a felony. I don't know if possessing a deepfake of child pornography is a felony. So there might be gaps in the legal system that we have to... And that's the crazy part about it is it's this whole, like, to your point, our legal system, it's a whole type of paradigm that we haven't even really had to encounter.

55:03And how do you build laws And it's, yeah, it is crazy to think about how that's going to change how we live, how we work, how we, you know, our morality as a species even to a certain extent. Right. Yeah. So I think the moral issues coming before the, you know, whether we're going to lose our jobs or, you know, computers taking over. The Terminator. Yeah. Terminator. Is it Terminator or Matrix? Which one is more scary? Terminator or Matrix? I don't know. I'd say maybe the Matrix. At least that's the more interesting one, to me at least. What about you? Yeah, I think it looks like, because in the Matrix, there wasn't really any mechanical part.

55:48It was purely everything was, you know, there was a computer running, you know, computers were running. The Terminator was mixing the reality with robots, okay? which I think it's more difficult it's an interesting scientific question because if the machines can take over and basically control the universe why do they need the mechanical part? Why do they need to go out in nature and do things? Maybe some of you would say because they need to synthesize energy so they need the mechanical component okay so it looks like so it might be the case that evolutionary will not take that into consideration, they will try to eliminate their creator but then they will actually face some type of extinction or shrinking because they will be missing the mechanical component to you know to get energy and all that stuff versus the other which is the hard way where i think in the terminator you need to create the robot to fight the humans and then you have the mechanical component that can help you you know because at some point even if you know they could eliminate humans and let's say they had solar panels they would need to manufacture new solar panels you know they would have to go and extract minerals to you know the chips will go bad after some years like you create new chips new stuff to interesting science fiction stories here.

57:26Yeah, I think the scarier thing is not the machines taking over, but the humans and bad actors using this stuff in negative ways, at least for me. That's scarier in my mind. But yeah, so this has been one of my favorite conversations so far. So many interesting topics. I really appreciate you coming on to the built right podcast nick but where can people find you where can they find relational ai and learn more about uh either you or the company i think you can find us on the on the web you know we are a remote first company even before covid and i think we do have an office somewhere in berkeley i've been there a couple of times but our people are all over i want to say the world.

58:11The sun never sets or never was the thing that released in the way I never we have people all over the world yes you know I'm here in Atlanta you can go to our website, you know read our blogs you know see about our products. Our product I think you know we have announced a partnership with Snowflake so people can use it through there it's a limited availability through there which is going to become a general one I think sometime probably this summer. It's coming up, so I don't have a date. But yeah, so yes, you can find me on LinkedIn. I'm not really big on social media. LinkedIn is probably the only one that I spend some time, not much.

58:58Nice. Well, great, Nick. Thanks for joining us today. Thanks for hosting, Matt. Have a good one. Excellent.

59:09Thanks for listening to Built Right. If you enjoy the show, give us a follow or subscribe on your favorite podcast platform. And don't forget to leave us a review. For more info on Built Right, visit us at HatchworksBiltRight.com.

59:30The single biggest mistake we see companies make with AI is they don't properly train their teams. We see it all the time. Companies roll out AI tools and expect people to just figure it out. But using AI effectively requires a totally different mindset and skillset. And that's exactly why we built training for every level of your org, from AI training for teams and executives to training engineering teams on our generative-driven development methodology. Or if you've already identified your AI use cases and want to just prioritize where to start, we offer an AI roadmap and ROI workshop to help you build a quick plan.

1:00:03It's all about going from we should use AI to actually driving real value with it. Head over to hatchworks.com to learn more.

From the publisher

Generative AI is always on the Built Right podcast agenda, and this episode is no exception because we decided it was time to hear the thoughts, opinions, and predictions of a PhD data scientist, and take a deep dive into the science behind it.  

We invited Nikolaos Vasiloglou, Vice President of Research ML at RelationalAI, onto this episode to share his thoughts on how far generative AI will advance, give us an in-depth look at how knowledge graphs work, and explain how AI will affect the job market, the future of learning and the social media landscape.  

Plus, he explores the main differences between generative AI and regular AI. 

Key Moments: 

  • The differences between generative AI and regular AI 
  • How foundational AI models compare to larger language models 
  • How generative AI works when our language isn’t the input 
  • How far will generative AI go? 
  • Why the progression of AI will change the job market 
  • Defining knowledge graphs and how they work 
  • The most disruptive and intriguing AI use cases 
  • How AI is affecting reading and learning 
  • Why generative AI is the future of social networking 

Key links: 


Mentioned in this episode:

Talking AI - Conversations with AI experts and early adopters

Welcome to the Talking AI podcast, where we dive deep into the world of artificial intelligence with host Matt Paige. Formerly known as the Built Right podcast, Talking AI brings you insightful conversations with AI experts, founders of AI products, and industry leaders who are leveraging AI in their businesses. Whether you're an AI expert or a beginner, our episodes will help you understand how AI technology works and how early adopters are deriving value from it. New episodes drop starting August 6th.

AI Opportunity Finder

Feeling overwhelmed by all the AI noise out there? The AI Opportunity Finder from HatchWorks cuts through the hype and gives you a clear starting point. In less than 5 minutes, you’ll get tailored, high-impact AI use cases specific to your business—scored by ROI so you know exactly where to start. Whether you're looking to cut costs, automate tasks, or grow faster, this free tool gives you a personalized roadmap built for action. 👉 Try it now at https://hatchworks.com/ai-opportunity-finder/

More from Talking AI

All 84 episodes
How Generative AI Works, as Told by a PhD Data ScientistTalking AI · 1 h
Listen in VO