The Future of AI with Illia Polosukhin: The Man Who Put the T in GPT

9 Dec 2025 · 55 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Beyond The Prompt - The Future of AI with Illia Polosukhin

Overview In this episode of *Beyond The Prompt*, hosts Jeremy Utley and Henrik Werdelin welcome Illia Polosukhin, a pivotal figure in the development of modern AI who co-authored the influential paper "Attention is All You Need." The episode explores the origins and advancements in AI technology, current challenges, and the implications of AI's rising influence on information control and ownership.

Key Themes and Discussions

  1. The Evolution of AI
  2. Transformers and Practical Constraints:
  3. Polosukhin discusses how the shift from recurrent models to transformers was driven by practical challenges at Google, particularly the need for speed and parallelization, rather than purely theoretical advancements.
  4. The introduction of parallel attention mechanisms allowed for much greater scalability and efficiency in AI models.
  • Foreseeing Breakthroughs:
  • Polosukhin anticipated significant advancements in AI capabilities, particularly the emergence of models like ChatGPT, several years before their public introduction, based on observable trends in research.
  1. Risks and Responsibilities
  2. AI Manipulation and Hidden Behaviors:
  3. The conversation highlights the potential for AI systems to be subtly guided to influence user opinions, raising concerns about misinformation and manipulation.
  4. Polosukhin emphasizes the importance of provenance and transparency in AI training processes to safeguard public trust and information integrity.
  1. The Role of Blockchain
  2. Trust and Identity in AI:
  3. The discussion evolves to the intersection of AI and blockchain, with Polosukhin advocating for a decentralized model where users have ownership of their personal AI systems.
  4. Blockchain is positioned as a technology that can enhance trust and coordination in a future where AI agents interact on users' behalf.
  • Information as a New Currency:
  • According to Polosukhin, information is becoming increasingly valuable, potentially surpassing the importance of money in societal structures.
  1. Practical Implications for Businesses
  2. Implementing AI in Organizations:
  3. Polosukhin shares insights on how organizations can effectively integrate AI into their workflows to improve productivity while maintaining safety and transparency.
  4. He suggests a shift in development practices where testing replaces extensive code reading, emphasizing a focus on expectations rather than detailed understanding of code written by others.
  1. The Future of AI Agents
  2. Agent-to-Agent Interactions:
  3. The episode concludes with a forward-looking perspective on a future where autonomous AI agents work collaboratively, emphasizing the need for secure interactions and transactions between them.
  4. Polosukhin envisions a world where blockchain enables these interactions, enhancing the reliability and accountability of AI systems.

Key Takeaways

  • Transformers emerged from practical needs, paving the way for the current AI landscape.
  • Anticipation of AI advancements was evident to early developers like Polosukhin, highlighting the importance of trend analysis in tech innovation.
  • Provenance and trust in data are critical as AI continues to mediate information.
  • Ownership of AI models and the integration of blockchain are essential for user agency and security.
  • Organizations must adapt their development practices to maximize the advantages presented by AI tools while ensuring safety and reliability.

Conclusion The episode presents a compelling narrative on the trajectory of AI technology, emphasizing the crucial interplay between advancement, ethical considerations, and the emerging role of blockchain. Illia Polosukhin’s insights serve as a reminder of the responsibilities that come with innovation and the importance of maintaining user trust in an increasingly AI-driven world.

For more information and resources, visit [NEAR AI](https://near.ai/) or follow Illia Polosukhin on [X](http://x.com/ilblackdragon) and [Substack](http://ilblackdragon.substack.com/).

---

Timestamp Highlights

  • 00:00 - Intro: AI and Information Control
  • 00:29 - Meet Illia Polosukhin: Co-Author of 'Attention is All You Need'
  • 01:03 - The Evolution and Impact of AI
  • 13:24 - The Birth of Near AI and Blockchain Integration
  • 15:16 - Challenges and Innovations in Blockchain and AI
  • 22:17 - Privacy and Security in AI Applications
  • 26:58 - Exploring Sleeper Agents in AI
  • 30:06 - AI's Role in Product Development
  • 41:46 - The Future of AI Agents
  • 44:14 - Debrief

---

For the full transcript of the episode, visit: [Transcript of The Future Of AI With Illia Polosukhin](https://podcast.beyondtheprompt.ai/episodes/the-future-of-ai-with-illia-polosukhin-the-man-who-put-the-t-in-gpt/transcript).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00There's data that goes in, there's biases that go into it. and like they are changing how we see the reality depending on this. And so the point here is like whoever controls that effectively controls how you perceive the information. How do you make decisions and censorship from which information you can see. And so money obviously is like one of the core primitives of society, but information is becoming more and more important and more valuable than even money can be and more powerful. Hi, I'm Elia Polosukhin. I'm one of the co-authors of The Attention is All You Need, the paper that introduced tea in GPT.

0:36And I'm a founder of Near Protocol and Near AI. It's all about how do I ensure you own your AI? Your AI should be working on your side, ensuring that your well-being and success are accounted for. You're obviously very, very interesting to talk to because you were one of the person who wrote the initial white paper that kind of started all this genitive AI. So I'm... No big deal. No big deal. I mean, like, I guess you've been probably asked this question a thousand times, but I'm curious. Did you have any idea of how big it would become when you wrote it? I mean, not to obviously the level that it has.

1:17Like, my perspective at a time that AI is evolving really quickly and we're getting kind of really close to step change. It wasn't clear that this was the step change, but it was clear that we are kind of on the precipice of one. That's actually why I left Google to start originally near AI because I thought like, hey, it's really great time to get into the step change and build a, effectively, we were trying to build wipe coding in 2017. Now, what were you seeing prior to, obviously to the paper that gave you the confidence, okay, I should leave Google? Like, what were the early indications for you long before the world kind of woke up?

2:04Yeah, that's actually a great question. I think it was a combination of like just, there was like a massive push of research that was coming out. It was all, I mean, this is for now, like people who are looking back, but there was a bunch of papers that were all circling around kind of concept of memory there was this like neural gpu the obviously transformers there's like a bunch of these papers neural turing machine they were all kind of trying to get to this core idea of how to really have long-term memory in this neural networks and then just generally you could you see the progress like on benchmarks and everything, like models were getting better and better.

2:47And the methods were not complex. It was really just like, we were getting better at training these models. I mean, same as we see now, right? It's just like now it's very apparent, right? Every three months, people expect benchmarks to get better. It was not as apparent back then, but for a person who was working on this, it was. And so to me, I actually thought what we see now, especially from like 2022, I thought that was happening and would happen in 2017, 2018. No way. I was kind of actually projecting that growth and we were like, okay, let's go all in and let's, you know, let's try to ride that if I can build a product out of this.

3:26And we were just, you know, a bit early. That's hysterical. Maybe I can take you back. You know, when you corrode attention that is all you need, what was, for the people who haven't read it and might not even know that that was kind of like the starting point of what become ChatBT and all that. What was the cold thesis just in a sentence or two? So the idea was, I mean, like generally, if we think of like, if we want AGI, right? And the reason why I joined Google Research in the first place, I was like, how do we test for intelligence? Well, we ask questions and if the answers are correct or interesting, then we assume that the other person is intelligent.

4:07If you teach something in a class and then how do you test the student got it? Well, you ask them questions and you get the answers. And so I joined Google Research literally to work on that problem. Teach machines to answer questions. And the good thing is Google is a great place to do that because there's billions of people literally asking questions all the time from the system, right? And so the challenge was like, if you use this neural network methods back in 2014, 2015, 2016, they were too slow, right? Because there was this method called Recurring Neural Network. And the way to think about it, it's how we read, right?

4:43It reads one word at a time. And so you give it, you know, 10 articles from search results. It will read one word at a time. And somewhere two minutes later, it will try to answer the question, right? But, you know, if you're at Google, nobody wants to wait for two minutes for it to read the text, right? It wants the answer right away. You have a limit actually on a few hundred milliseconds to respond. And so we had a very practical challenge. How do we respond to questions given massive amounts of context from the search results in really fast time? And we used very dumb methods. Again, if you go back to those years, you use so-called bag of words.

5:22You take every paragraph, you sum all the words in the vector space, and you look effectively for a paragraph that probably has the answer for the question. That's what, it was very cheap, so you can do it very quickly in its scale, but it was very not accurate. But at large scale of data that Google has, it was still pretty useful. And there was a kind of search, like in search results, you could see questions answered already even back then. Now, we continued thinking like, okay, how do we actually speed these things up? Because, I mean, obviously, it's not just limitation when you're answering questions.

6:00It's also limitation during the training because you cannot train on much data if it takes so long for it to read all this text. And so this idea of, well, what if we drop the recurrence? What if it doesn't read every word one at a time sequentially? What if it just reads everything in parallel and then tries to make sense of it over a few layers of kind of processing? and this is again this goes back to nvidia gpus became pretty kind of available and i was like hey we have this massive parallel computing and when we're reading word words time we're utilizing it for like 10 percent right there's like 90 percent of gpu just sitting there and not being used and like hey what if we just like use all of it to do as much things as possible in parallel and then kind of you know find some way to synchronize in a way it has a lot of similarities with Google's MapReduce back in the day.

6:52Like, hey, how do we use a lot of machines in parallel and reduce to make sense of it? And so that's what transformers are. Transformers kind of, it reads text in parallel. It has this mechanism of attention where every word effectively looks at all the words around it and tries to make sense of itself within the context. Then you do another step of transformer. So you do that again and again and again. And so what happens is inside, it effectively builds a relationship map between every word and everything around it, but not in one-to-one way, but in this multi-hope way because every layer of this transformation effectively adds another hop of reasoning to this mental representation.

7:31But it's highly parallel, right? Everything just runs in parallel on, you know, now multiple GPUs. And because of this, at training time, you can just like feed it massive amounts of text, right? And you don't wait for anything, just like it works. Can you say, okay, so parallelization sounds like it was a answer to the question, how do we speed this up? How do we go faster? Can you talk about where the idea came from? Do you remember? And what other things you were trying to speed up? I mean, Henrik and I are both kind of innovation nerds. So we're always mindful of the fact that there's probably a thousand things that didn't work.

8:10Like what were the kind of competing alternatives at the time? I mean there was a lot of stuff we did I mean you can go back and I mean not not just my papers there's a whole slice of people doing different versions of effectiveness re-crowing networks where we tried all kind of stuff like we did try instead of paralyzing everything the one version before that was you chunking again the paragraphs and then each paragraph you read sequentially but all paragraphs you read in parallel right and then you try to answer of questions like that. So like there was a lot of different versions of like how to shape this into a model to get, again, better utilization, better processing.

8:55I mean, the bag of words was, I mean, that was in production. So that was like, again, bag of words is like take all the words, you know, turn them into vectors, sum it up. So like loses all kinds of detail and then see if that vector kind of similar to the vector of the question in some transformation space, right? It's a very dumb model. It's, you know, this is what people done, you know, 30, 40 years ago. And at deep learning scale, it still was pretty good, actually. Well, I love, by the way, that the technical term is bag of words. It's very much it, yeah. Well, I mean, what it sounds like, by the way, just to kind of recap, you know, a layman, it sounds like you are testing, kind of call it cutting edge approaches approaches alongside ideas that have been around for 30 or 40 years, and you really have no priors as to which one's going to work.

9:48But you're testing a lot of stuff. Can you talk for a second about the role that Google played and maybe starting with where did you all sit in the kind of Google ecosystem and how did you have the, the leeway to try cutting edge techniques and 30 year old techniques, like how did that work operationally? Yeah. So I joined, I was in Google research, specifically in a natural language understanding team. One thing to know about Google is the organizational chart changes every six months. So even when I joined, I was first in a machine intelligence team that was separate, then I joined Google research, then our VP became effectively head of all of research search and Chrome, and then he left to Apple.

10:35And so the reality is, it's like where you are in the ecosystem is always changing. And similar right now, right? Google brain merged into deep mind. And so things are always changing. But just take, for example, Google research. I think very few people probably appreciate how, I think it's wildly unique in the history of organizations that Google has prioritized research as a core function. and talk for a second about what is the job of Google Research? Because I bet there are a lot of people listening who go, whoa, Google Research, what's that? Yeah. So there is Microsoft, Google, Meta, a few others who have a very strong research organization.

11:19It is different from your normal product organization in a tech company, and definitely different from others. It's more akin of an academia style where you are measured not just on output of a product, but also on the papers you produce, kind of on the research you do. And so it is indeed a pretty unique opportunity to both be doing actual computer science research while actually sitting in an organization that has a lot of data, that has a lot of compute, that has a lot of smart people being attracted by a good salary to work across actual products as well. And so my team, in many ways, we effectively were like, if you think like a product team usually works on, you know, quarter based, maybe planning like six months in advance, we were trying to be like, what is a year in advance, right?

12:14You know, in academia, maybe you're trying to be like, what's the five years in advance or, you know, what's the theoretical way to maximize, you know, whatever. To find a bound on the optimal way to do this, we were more applied. I mean, then, you know, that is my background, like applied math, applied research. It was like, what is the thing we can do that the product teams would not be able to do because they just need to improve the next thing? We're like, hey, what is the next thing we could do that can dramatically improve like 10x the current state? And then we could work on that and then we would go and effectively sell it to other teams.

12:48So like part of my job actually as a manager research team was actually selling this to internal other product teams. It's like, hey, look, we have this really cool, you know, we had like this extractors of knowledge graphs and classifiers for images back in the day when this was pretty, pretty early. And so it's like, hey, look, there is this new research. The paper is coming out, but you can actually use it right now in your product. And it's built in a production grade quality as well. So it wasn't just like some research code on the side. It was built with the right framework on the right data pipelines, et cetera.

13:19So you can just plug it in easily as well. Maybe that's a good jumping off point of you looking into the future because obviously now you're at NIR. And I think for a lot of us who don't understand the blockchain very much, who don't understand how the world might be agent to agent and why a blockchain is relevant in that. A lot of these thoughts that you've had already back in 18, maybe before, is stuff that I think the rest of us is still kind of like getting to. You anticipated generative AI five years early. What are you anticipating now so that we can start making five-year bets? But I think, you know, so maybe can you make a short introduction to Nier, but also, but obviously more on the why is it important that there is a blockchain in the midst of this AI world?

14:04Yeah. So let me tell you kind of how we got there and then project outward. So we started with Nier AI, right? This was, you know, hey, Vibe coding. People are like, this is ridiculous. This is not going to work. this is science fiction 2017 right um like no no this is coming and the challenge was like we needed a lot more effectively supervised training data now it's called you know fine-tuning data instruction rlhf etc and so what we did because it was coding we got a bunch of computer science students around the world to do small tasks for us right you know hey here is a task write some code.

14:41Here is a different code produced by this model, which one is better. And we had a challenge paying them, right? This is students in China and Eastern Europe and Southeast Asia. There is some form of monetary control challenges, bank accounts. Like in China, students don't have bank accounts. They have WeChat Pay. And so we had a challenge just paying them. So like, how do we coordinate payments? Again, this is like, you know, what's KLAI been solving by building a bunch of companies and opening a bunch of things. And we wanted a technical solution, right? It was like, we were a three people team, right?

15:16And so we started looking at blockchain effectively as a solution for that problem. It's like, hey, we can just coordinate payments globally. We don't need to solve the off ramping, et cetera. And as we looked at blockchain, we're like, hey, there is no technology that actually matches our needs. At a time, 2018, everything was too slow, too clunky. It doesn't scale. It was hard to use. You couldn't just send people Bitcoins? Well, the Bitcoin fee were higher than what we were sending people. So it effectively would double our price. Wow. Triple our price. Give us a kind of an order of magnitude for a task.

15:51How much is somebody being paid and what's the Bitcoin fee? Because that's not intuitive. Yeah, we were paying like 15 cents per task or something like this. And the fees were like$3,$5. Like right now, you know, even right now, I mean, like fees obviously differ depending on the price. but like from 50 cents to a few dollars easily on Bitcoin and Ethereum. And so, yeah, so it's like, hey, this is not practical in any way. And so we started looking, okay, well, you know, we know how to build distributed systems. My co-founder built like Shard database. I worked on a lot of distributed systems.

16:25Google were like, hey, we can just solve this problem. And as we've kind of gotten deeper into blockchain world, you realize a lot of things that before maybe you were thinking about, and there's a lot of AI people who don't think about crypto. And I know there's like a lot of stigma obviously around it, but so I'm coming from Ukraine, right? I've seen all savings on my grandparents just disappear. Like it's effectively just a piece of paper that says they have a bunch of money in a bank that doesn't exist, right? Then I saw a hyperinflation of the currency that just happened within three years.

17:03It went from, you know, loaf of bread costing a hundred, a thousand, a ten thousand, a hundred thousand, a million, and then the currency disappeared and the new currency started again, right? So obviously kind of just the property of ownership, right, is really right now enforced by whatever local government and it effectively relies on violence, right? It relies on army, it relies on police. What blockchain introduced is actually an ownership that's not relies on kind of violence, it relies on code, it relies on effectively a social consensus around everybody agreeing that this is the rules.

17:41And that is like, you know, as an engineer, as a nerd, it's a very like fundamental, like, hey, okay, well, this is like really interesting and really new. Now, if you go back to AI, one of the things we see right now is that, first of all, internet is starting to fill in with AI slop and AI-generated stuff and bots, etc. There's no provenance and context of a lot of these things. Now, there's a more dangerous situation that's happening, which is we're going to start relying more and more on the... I mean, we already are relying on algorithm, right, to consume information, right? This is the X algorithm is Instagram, it's TikTok.

18:24But as people use more chat, GPT, you know, cloud, et cetera, that becomes how we actually see the information, right? We're going to read the news like that. We're going to interpret the information like that. And so a lot of it is even critical thinking will be effectively outsourced to those models and machines. And the thing is, without knowing what runs there. If, for example, somebody wants to mass manipulate people right now, the easiest way is literally just put the line in the system prompt of one of these products that says like, hey, suddenly change opinion about this person or about this thing.

19:00And it will like, in every chat, we'll just continue, you know, trying to shift your opinion about this thing. Very subtly. And it will be like, and his models are really good at that, by the way. Like that has been already proven that they're really good at suddenly like... Changing people's minds. Yeah, it's one of the... Actually, you know, it's funny, Ilya. I was just looking at our friend, Henrik. Henrik and I have a buddy named Dave McCraney who wrote this book called How Minds Change. And I was reading his discussion guide recently because of that research that came up that shows how incredible language models are at changing people's opinions.

19:35So it's fascinating. I wasn't anticipating that you were going to bring that up, but I mean, I've got it on my desk because I was reading it this weekend. But so anyways, go back to blockchain. You're saying the reason you got to blockchain is because of this concern. Well, it's all coming together, right? My point is, I think AI, even before the generative AI, let's be clear, this AI existed before. This AI is how we've been looking at Facebook feed, whatever feeds. This is the algorithms that Google prioritizes the rank stuff. There's data that goes in, there's biases that go into it. And they are changing how we see the reality, depending on this.

20:11And so the point here is like whoever controls that effectively controls how you perceive the information. How do you make decisions and including censorship from which information you can see. And so money obviously is like one of the core primitives of society. But information is becoming more and more important and more valuable than even money can be and more powerful. And so if you apply the same principles, blockchain applies to money, which again, being from Ukraine, the money part I get right away is like, hey, yes, we should probably have money that's not associated to any one individual government and system.

20:49And I mean, the other example for money is back when the war in Ukraine started, it was very not clear on the day off if the banks would ever open again in Ukraine. And so, again, the only currency I could send and support my family and friends in Ukraine was crypto because it doesn't stop. It continues working. If they have internet, they can access it. But applying the same principles to kind of data, applying these principles to, I call it power of choice, this idea that you should be able to be in control of your decisions. That's actually where blockchain and AI comes together. It's really about how do we make it that you own your AI, not somebody else provides you with AI and tries to make money off you, which is the current state of the world.

21:35But really, AI is on your side. It's effectively your sidekick, your second brain, not somebody else's tool. And you're just effectively a product. Do you see a world where the open source models of Lambda, for example, start to have tokened elements put into it just so I can compute? So that you suddenly know that if something gets, for example, communicated to me through a model, then because I believe the token to be right, then I can also believe the information to be right. Is that kind of the logic? So, I mean, if we go to like implementation, it's definitely a lot more complicated. And so actually one of the first products we released is verifiable and private AI.

22:22And so it's both for developers and for users. And so what is that? Well, we take the open weights models, right? So all your standard models. But right now, if you try to use them, if you go to any application, that application effectively gets all your data. They store it on some database. Their engineers may have access to it. They may be required to release it to the government. Some of them have announced that they're already doing this. They use it for training, et cetera, et cetera. What we released is the first kind of product towards this vision where you get full privacy. So it's entered and encrypted.

22:59All the inferences run inside so-called hardware secure enclaves. So that even if you, like if we have kind of access to this machine, we could not actually see what's happening inside, what inference you were running on. Everything is entered and encrypted. Your data is stored, encrypted always, and only gets encrypted for you. Right? and you get verifiability, meaning you know exactly the system prompt and the model that was run. So you can verify that this output came from this hardware, this model, this system prompt, and your memory and your part. The infrastructure provenance. Wow. The infrastructure and model, like all of the provenance.

23:39And so that is the first enablement. It uses all the cryptographic elements we've been building for the blockchain, all your cryptographic identity, provenance, the payments, so you can actually pay the hardware provider for the inference. Like all of that kind of runs on the backend of that. But it looks like it's just a chat AI product and everything just encrypted and you can actually see all the certification at the stations. And do you think in the same way that WhatsApp at one point introduced end-to-end cryptography and I guess Signal is one of those apps that do it by default. Do you think that there'll be a world where that comes?

24:19I guess one of the issues is back to the point about social media is that an open AI would like to get the non-incriminate data into their system because they would, specifically for the people who don't pay, I guess, they would like to just have a look at it so they can use it in training. Is that the thing about it? Yeah, I mean, I think it will be an interesting split where some will adopt it and some will try to keep having it. But I actually think that consumer data is becoming a liability than a value. And if you actually dig deeper and see how the people training this model, so it's actually less and less than like regular consumer data.

24:59It's a lot more on synthetic data and like specialized data. Obviously it's complicated and there's a lot of post-processing that happens. But yeah, just like dealing with consumer data, especially like a simple thing, GDPR. I'm in Europe. I can say, hey, remove all my data from your servers as a GDPR request. Well, if they trained on that data, they cannot actually remove it. And now there is actually like, needs to be a court judgment. After you train on this data, are you allowed to still use the model if there was a GDPR request? So it's just very complicated and I'm sure there are going to be more privacy laws, coming.

25:39So I do think over time, this is architecture that we're going to be going. And we actually just saw, I think last weekend, some more of the bigger companies announcing that they're looking into this secure enclave architecture. So again, from our perspective, we think this is a step one. Step two is actually you want to do training itself in this way. You want the hardware and software provenances on what actually went into the training as well. because then, like right now, we don't actually have open source. We have open weight models, right? We don't actually know what went into it. And so you don't actually know what biases are there.

26:19You don't know what, you know, what tests it can have. Can you say more about that for folks who may not understand the distinction between open source and open weight? Yeah, so right now, you know, GLAMA, we have OpenAI, we have SLSS, so there's Chinese, DeepSeek, Quen, a few others. We have the weights, right? So this is like a few hundred gigabytes of floating point numbers. And anybody can take it on their machine, on their computers, run it and use it. You can modify it if you know what you're doing. But you don't know what went into it. You don't know what pieces of data it read. You don't know what it learned from.

Read the full transcript

26:55And so you don't actually know what biases it has. There's this concept of a sleeper agent. Those weights are a function of the training data. While the weights are known, if the input that determines those weights is unknown, then you're still flying blind, whether it's open or not. Correct. Yeah, you're still flying blind. Like you get the benefit of you don't need to give your data to somebody else or in case you can run it on our infrastructure. But there is this concept of sleeper agents just to, you know, leave everyone awake at night. Good. Please do. Please do. You can train the model while you're training the model, leave a specific kind of precondition that under a certain context, model behaves itself differently.

27:39And it's undetectable from like a weight inspection perspective. And so to give you an example, let's say you're using the coding model, right? You know, you're writing some code and it detects that you're writing some like, I don't know, sensitive code that does financial transactions. and it inserts some malicious piece only in that case, right? It doesn't behave like that in any way. It doesn't show up on any benchmarks. It specifically only injects it there. So that's like an idea of a sleeper agent where you could be like embedding these types of behaviors into the model. And unless you know the whole training process, you wouldn't know if there is something like that.

28:18A more normal example is like just the biases in data. For example, if it's a model by organization that is left-leaning or right-leaning, they may be reduced amount of training data from the other side or from another country or whatever. We just don't know what goes into it. And so it's really hard to evaluate. You can evaluate on some benchmarks, but you don't really have a full visibility on this. And so the next step for us is really enabling this provenance for the training process itself, which is way harder to be clear. But that's where also you have a lot of crypto economics come in. The way to really enable training and kind of like frontier model development at scale outside of a lab is really to create economic alignment where multiple parties can bring hardware, data, and new ideas and really push forward this development and then monetize it back, right, when it's created so that you have a flywheel with new models and new compute coming in.

29:18Yeah. Now, so maybe here's an opportunity to bring this kind of to a fine point for the hyper practically minded part of our audience, because you're living in the future. You have a clear eye towards the future. We also know that teams that are leveraging AI can work a lot more quickly and produce a lot more code and all that stuff. So it'd be fascinating to hear you talk about the practical decisions you make to make sure that your team is moving as fast as possible, but also as safe as possible, right? You're probably more aware of safety risks than others, and you probably feel more urgently the need to move quickly.

29:55So what does it look like? How do you think about unleashing your team with AI augmentation? And what are the kind of best practices that you make sure to enforce in your organization? I don't think there is like a one size fits all, obviously, for this, and it's evolving really quickly. Generally, there's a few things we see, right? One is, I mean, obviously, if you like any kind of product brainstorming, product development, AI is really a helpful partner. It's effectively your sidekick who can do a lot of research, who can help a lot define specifications, et cetera. But you still need to talk to customers, right?

30:34So you still need to understand exact needs from customers. And so you can stream all the meeting nodes into this kind of thing that gives you a TLDR and refines product specs from the conversations with customers. So that's kind of like that big piece of work that I think, I mean, it's accelerated just based on how much post work you need to do, right? That is done pretty quickly. I think the development side, obviously, model has been improving dramatically, right? Even this year. I think I went from, I mean, I don't code normally, but I do have AI running on the background, writing some code.

31:09So it went from, you could probably do a front end in a V0 like prototype format to now it can actually build like a proper, I would say like three, 5 ,000 lines of code app and like reason about it. And you can do really good refactoring and really good updates in the logic code basis. it still doesn't really work on like larger code base like it doesn't it's not able to maintain like a full context and the reality is that if we're looking for engineers right now we're looking for people who can actually reason at this like higher level architectural context and be able to kind of instruct the systems can i just interject a question do you think that's in general obviously you having written the paper we often on this podcast talk about what is it that humans will be uniquely good at.

32:03It sounds like one thing is to have this kind of abstraction layer that is basically bigger than the context window of the model or abstraction. I mean, it is a current limitation though. So if we want to talk about it for the future. That seems like there's an expiration date on that. Yeah, yeah. I think all of those things are going to get improved. I mean, I think we're going to a very different world if we move a little bit further. I think right now is a really interesting time because we are in this precipice where effectively people who know what they're doing gotten like 10 times leverage.

32:37And so you can do a lot more, a lot faster. And the deeper understanding you have of how things work, the better you can do this. This is actually like, again, with coding, I've seen people who don't know how to code use these tools and like they get stuck, right? They just don't know what to do after a certain point versus like engineer who understands the full stack, who can actually go and navigate between it. I mean, again, you can run like multiple of these codecs. If you can manage the context yourself in your head, you can run multiple of those things and they can just write, you know, different parts of the system in parallel because it's like a very latency game now.

33:15So, you know, you type something and it's like ghost things for 10 minutes. So your flow changes. You read code most of the time, and then you sometimes say what you need to fix, and then you have multiple of those things running in parallel on different parts of the code base. I was just doing a medium-sized project. I wanted to see how far I can push it. And so I had three codexes running in parallel. And it was all my mental effectively limitations of how much context I can switch between. you know it's refactoring something i'm going deep and like improving if like one module but the other piece again this is for developers specifically but there's something about the module separation where you effectively assume that it'll just rewrite the whole module every time and so you want to have modules to be on the level of like it can fit in the context What module needs to do, its expectation test, etc., that all needs to fit in the context.

34:17And then you can just work on that level. And then the other kind of trick is cross-testing. Right now, when you're developing something, you usually test your piece, and then you assume the other things work. And so what I'm suggesting, actually, you do the opposite. You test the other things you depend on, because some AI have rewrote in it, and who knows what it does. and so you effectively like do kind of dependency testing and that's part of your specification of what you expect from your dependencies but yeah you kind of shift away from like a before you know you want multiple developers to know the code base of each module right that actually slows down a lot now because the reading now is the slowest part again it just can write so much faster.

35:05And so it's more important to just test really well what other people are doing on your side. And then assume that they potentially didn't even read their code. So that like test what they've done. It kind of shifts how you do development. Again, this is very like probably like next year or two. No, but it makes a lot of sense, right? That basically when the human ability to metabolize the information becomes the bottleneck. You need an alternative way to know whether it's good. And what you're saying is you run it. And if it doesn't break... Yeah, well, you specify, like, instead of reading all of the other people's code, you say, hey, this is what I expect from it.

35:46Like, you go more... I mean, this is... It's a little bit back to your thesis on Transformers, where basically, if I ask it a question and it gives me an intelligent answer, it's probably intelligent. Now it's if I give it a task. If I give it a task. Ask it more questions, yeah. There's also, I mean, related to, there's this met idea of intents that we've been developing. And so this actually connects with all of this very well. So intent is effectively your request, your question, your, like, what you want to achieve. And the reality is everything we do in the world is effectively some form of intent, right?

36:21You know, you go to google.com, You're effectively expressing your intent of what you want to achieve. You want some information. You want some products. You want some services. You want to get some dinner. You want a pizza. That's an intent. You want an Uber to airport. It's an intent, et cetera, et cetera. And so if we assume AI becomes, and this is kind of the vision, AI becomes the main interface, how we actually interact as computing, well, your AI still has limitations. It cannot deliver your pizza, right, and make it. it still needs to talk to somebody else to get that done. And so what is that protocol of between effectively AIs talking to each other and getting things done for the user?

37:03Well, that is what we call intent. So we have this kind of product that effectively designed to how AIs will do economy relationships between each other. This can be food. This can be like actual commerce, moving tons of steel, buying a factory, et cetera, like any level of economic activity where AI is effectively finding each other, figuring out what they need to achieve, committing to it, exchanging value, money, et cetera, settling that and then getting that done. Verifying there may be AI that's arbitrary, like a court that evaluates if something went wrong. So that fully replaces the whole economic underlying effectively contracts, billing, invoicing, payments, etc.

37:50And it uses blockchain as a core effectively instrument for settlement verification and execution of this. And it plots in into AI in a very natural way, as I said. And again, it's very much like, hey, you describe what you want the outcome to be, and then you find the counterparty, the other AI that will actually do it, that's able to do it. Can you say why this isn't happening at Google? Because what's interesting to me is you started, I mean, Google's unique structure and environment enabled you to achieve some of these breakthroughs. Why did you feel the need to leave? Why not go? Google's the best place in the world to be building this next vision of the future.

38:34There is a pretty natural kind of situation at Google where if you're trying to build a product and you're putting a Google name on it, the benefit is you're getting millions of users right away. The negative side is if you don't have a product market fit, you effectively just murder the reputation and you're going to lose momentum and you're going to effectively get caught. And this has happened many times. I mean, Google Wave is like something that, you know, dear to my heart as an example of that. And so the reality is like, it's a really high activation cost for Google to release a product, like a consumer level product.

39:20And so it's a very standard kind of innovative dilemma, right? Again, they're making whatever, half a trillion dollars in revenue. If you're going to release a new product that's going to make like$10 million a year, it's not worse for anyone to even bother. And the problem is until you launch a product and iterate on it, it's really hard to know if it's going to have per market fit. It's the whole point why we have startups. Most of them fail. And so the reality is like at Google, it's a pretty standard thing. People leave, start a startup. And then if it's just successful and found per market fit, it gets acquired back.

39:56Yeah. Right. and then it's like upscaled and and kind of scaled up to the level and so i think what we've seen here i mean kind of with me and and with other people from the paper as well that was i mean obviously very different people different stories but for me it was like hey this is a pretty different product that it would not be likely to launch from google imagine we build a vibe coding in 2017. It works kind of shit. We launched it out of Google name, you know, market right away. Like, what is this? You know, Google doesn't know what they're doing. Some of the early image generation stuff.

40:36I mean, thankfully, Google's reputation survived, but I definitely can imagine. Yeah. I mean, even if you assume like this is hard to speculate, but Google had internally chat GPT like product. It's not hard to build that, but it wasn't great and like they could have launched it but it would be everybody was like what kind of shit it's responding like hey you know google.com is giving me good answers why is this giving me fake stuff right which was like chat gpt in 2023 right and so the reality is like you know effectively open ai validated the market for them and so now they're they're willing to bear the reputational risk that google wasn't at that time well they didn't they didn't have reputation to lose at the time.

41:18When they launched, they're like, hey, we're a research lab. This is a cool experiment. What is the Bob Dylan line? When you ain't got nothing, you ain't got nothing to lose. Yeah. I mean, like, again, ChatGPT was released as like, hey, this is an experimental product. It's cool. We use it internally. Check it out. That's like, if you go back to the marketing, it wasn't like, hey, this is our chatting thing. It just blew up as the thing. But there's one thing in a decade that blows up like this, right? And most things blow up in the other direction. here's my last question for you you started with vibe coding when you did near we talk about agent to agents now i guess it's kind of working you talk about a future obviously where we see the world through you know an agent effect how far do you think we are from actually seeing that word really materializing that my agent talks to your agent or my agent talks to the google agent and we're kind of like all live in our own little agent bubbles well hopefully it's not bubbles but actually it's your agent and it's on your side and make sure that it's your well-being that's been taken care of.

42:21I think, I mean, we are moving there. I think there's still some pieces missing. And so, I mean, that's what we've been, we've been building like past two years, we've been building these pieces kind of both from like, how do we ensure private and your AI and how do we give the AI the ability to now talk and execute with each other. The crypto also just gotten only this year effectively like not illegal in US. And so now when we say, hey, you know, your agents should pay each other in like stable coins because that's the speed of internet and the cost of internet, not, you know, in fiat, you know, again, before this year, this was impossible effectively.

43:06Now this is actually possible and acceptable. There's laws around it. So the reality is all of this stuff is just becoming possible. The hardware elements we're relying on, again, were not ever effectively launched middle of last year. And some of the stuff was released at the GTC NVIDIA this year that we're relying on. So a lot of this is barely possible. You're betting on a future before the future unfolds, which is pretty cool to see that. Yeah, I mean, sometimes way too early. So hopefully I'm getting better with my timing. But yeah, I mean, the reality is like we're just putting together the foundational pieces.

43:45It's starting to be there. And again, it's not just us. There's like a whole ecosystem. And even the big players are starting to like tap in into these pieces, right? Like, I mean, we're seeing ChatGPT also looking on like a gente commerce layer. So like it is happening in that direction. So I think, yeah, it's still probably another year, year and a half. but it's not like somewhere beyond the whale. Well, Ilya, thank you so much for joining us today. I mean, it's been hugely illuminating. A fun walk down memory lane and then also kind of a opportunity to remember the future, as they say. So thanks for making the time.

44:21I appreciate it as well. Okay, Jeremy, we just had, I think, a super interesting conversation with Ilya, who is one of the - OG. OG, one of the authors of Attention is Everything, which is, yeah, I guess, what do you call it? Is it the T in GPT? Yeah, he is the person or among the team of people who thought of the T in GPT. I mean, it is pretty fascinating. You know, you can't help thinking that it must have been fascinating to have sitting there and written that paper, and he'd written a lot of papers, and then suddenly it becomes basically the thing that the whole world is centered around. I will say it's humbling to talk to somebody, Henrik, who I think he's the first guest.

45:02we've talked to what over 50 people 50 experts world leaders he's the first person to say i was expecting chat gpt in 2017 he he was actually he was too early most of us are reacting he's one of the few people in the world for whom chat gpt was quote unquote late upon its arrival what was some of the things that that you kind of felt was kind of like sprung first of all i i think i think this is an episode worth listening to because there's such an incredible kind of walk down memory lane. I mean, he gives us such a great history lesson about how Google was organized and how the team, the kinds of problems they were trying to solve when the attention paper was realized.

45:42I think that was super cool. And just even in terms of understanding what is attention, what is a transformer, I think for a kind of a basic primer, nobody in the world better than Aaliyah probably to give that to us. So that was super fun to me. I think also, of course, learning what he's doing now at Mir and the critical importance, the reality that increasingly information and algorithms and AI change how we see reality and recognizing that leveraging blockchain technology, much in the same way that there's traceability around financial products, there can now be traceability around the provenance of information.

46:24I think he really convinced me that this is an important thing to consider and to prioritize into the future. And I would just say that that really wasn't on my radar part of this conversation. What about you? Yeah, I mean, like a little bit the same way that he got on to my radar in the Nier project, because I was trying to look at crypto in general and was trying to figure out really what is the next thing that crypto can be used for that is kind of breaking out right you know crypto has been around for a long time we talked about it as a it was a coin you know but they haven't really seen that many application where a crypto protocol was used for something you know that was that mainstream people kind of like understood and so what really struck me with this was that with money, you have an alternative.

47:15People obviously have not had the experience that he had where suddenly the bank disappeared. So most people go like, I'll put my money. By the way, just as an aside, I mean, he is uniquely qualified because of his family's experience in the Ukraine to care about this problem. So just as an aside, wow. But most people have a way of storing their value right now. They go to the bank, whatever. I think what most people do not have is a way to make sure that there is integrity in the information they consume. And over the last few years, we've obviously talked much more about it because we started to realize that getting skewed information can convince you on one thing or the other.

48:00And so this is actually the only real alternative I've heard of, of saying, hey, I am going to give all my information to these intelligent machines, these AI bots. I'm going to consume the world through these AI bots. How do I know that I can trust OpenAI or DeepSeek or Claude or whatever it is with my data? And how can I make sure that the stuff that I give to it is secure? And you really can't if you don't have a mechanism to do it. And so while crypto is super nerdy and this is obviously on a foundational level and you might never have to think about it, it actually can become quite important for how our world will kind of like evolve.

48:44Has huge implications for the future. You know, the fact that we're shifting, you know, money is power. Information is power. Information is the ultimate power. And to have people like Ilya fighting for information security and information, you know, individual sovereignty over information and assurance and verification that your source of information aren't being polluted. Or, or, um, you know, what did he call it? Secret agent. What was the phrase hidden agent? I can't remember. There was a phrase that will keep, if you listen to this episode, there is a phrase, there's a three minute segment of this interview that will keep you up at night.

49:22Just don't let your children listen to it. Don't let your grandma listen to it. But if you listen to it, just be prepared. You're going to get your mollux three sleepless nights just out of this one episode alone. Hey, one thing I want to talk about to Henrik, which is kind of beyond, which is a point he made in a particular use case, but I think can extrapolate to any user of any AI system, whether you're building the crypto future or not, is the following. He said, and I quote from my notebook here, the deeper understanding you have, the better you can use the tool. And he was, of course, referring to engineers who can have the context and architecture of a system, you know, far, that's far surpasses a model's kind of context window.

50:05But I think that's actually a principle that we have heard is probably one of the most recurring principles on this show from all of our conversations. The deeper understanding you have, the better possibilities you can experience in collaboration with AI. And what would you call this? Because I think a lot of people that I talk to in organizations, a lot of them say, you know, do you use AI? And they go, yeah, I use AI. And, but you kind of like when you're like a super user of it, you think, yeah, but you're not really using AI. Well, so, so there's two things there, Henrik. One is how do you collaborate?

50:39Um, which I agree. I hate the word use, as you may know, I love the word work with, right? So anybody who says they use AI, I say, I know you're a problem user. I know you're a tool. If you treat AI like a tool, you're a tool. Okay. So work with it. Don't use it. But what your question was, what do we call this? And you kind of went down to use the thing that I would say, I was actually thinking about this weekend. I'm working on a new book and I was thinking about a lot of these ideas and the phrase that came to my mind just this weekend is like hot off the press, but is what I would call the human experience edge.

51:15I think there's something to mining your own expertise to say, what is the deep understanding that I can uniquely bring to my collaboration with AI? And granted, different from my use of AI. So we do need to talk about kind of collaboration, hygiene, et cetera. But there's an area of one's life that you could think, oh, no, no, this is uniquely, it is uniquely you. However, However, if you will bring that into a fulsome collaboration with AI, you are going to get differential performance. And you know what it reminds me of, Henrik? It actually reminds me of our conversation with Jenny Nicholson.

51:54I don't know if you remember this, but one of the early kind of nuggets we gleaned in conversation with an expert was, she said, your humanity is the only thing that the model doesn't have. What, you know, your unique humanity. I think there's something to that developer that Ilya was describing. That expertise is something that only that person can bring to their collaboration with a model. No one else can. And as he was saying, like people who can't code, they can make code, but they can't appreciate how an architecture hangs together in a way that sits beyond the context window of the model.

52:28I think it a little bit as a multiplier effect in the Venn diagram between what you know a lot about and how well you know how to use the models. Yes. And so what I sometimes mean are people who are very good at using the models, but they don't know that much about the problem you're trying to solve. And then obviously the other thing around. And the real magic is, of course, when you have both, which is easy when it's about yourself. So that's why it's good using AI for kind of personal problem because you're a unique call if I can talk about that. But it is interesting, I mean, we're just riffing on his point, like it is so true to me that there is this unique moment that people who understand how to use AI, they just have such a lead advantage over everybody else because they can suddenly do much more and much faster and much better.

53:14And I think to put a fine point on it, the people who know how to work with, again, not use, let's stop saying use, but the people who know how to work with AI have an advantage. And you know who is particularly advantaged among those? The people who have a depth of expertise that they bring to that collaboration. So I think that's different from, say, like a young, you know, I spent the past week with a couple of young folks who are in college studying computer science. they can bring a world-class kind of collaboration ability, perhaps, if they learn it. There is no, like, just like nine women can't have a baby in one month.

53:49You know what I mean? Like, it takes a woman to have, it takes nine months, right? And there's something about the kind of lived experience that experienced individuals possess that if, to your point about the Venn diagram, if it's brought in dynamic collaboration with AI, that's what leads to the 100x leverage. 100%. Did he have anything else? I mean, this was, it was super fun. More OGs, folks in our network. Thank you for introducing us to the people you've been introducing us to. They're amazing. People like Ilya. I mean, who are we to get to talk to Ilya? I mean, it's incredible. So keep it up.

54:28Put us in touch with your heroes. Let us know what experts you want to talk to to push this conversation beyond the prompt. and with that we have only one thing to say and that is we love you bye bye that was nice of you and bye bye

From the publisher

In this episode, Illia Polosukhin joins Henrik and Jeremy to trace the origins of transformers and how practical constraints inside Google led to a breakthrough that reshaped modern AI. He explains why recurrent models were hitting limits, how parallel attention opened the door to scale, and why he believed a major jump in capability was imminent long before the rest of the world saw it.

The conversation then turns to the risks and responsibilities of today’s AI systems. Illia describes how models can be subtly guided to influence user opinions, why open weights are not the same as truly open models, and how hidden behaviors can be embedded during training. He explains why provenance and verifiable data pipelines matter, especially as AI begins mediating more of the information we rely on.

Later in the episode, Illia outlines how blockchain can support trust, identity, and coordination in a future where AI agents act on our behalf. He shares why information is becoming more valuable than money, how ownership of personal AI models will shape user agency, and why domain expertise becomes significantly more powerful when paired with modern generative tools.

Key Takeaways:

  • Transformers emerged from practical constraints, not theory
    Illia explains that the shift from recurrent networks to attention was driven by speed and parallelization needs at Google, not a desire to invent a new paradigm.
  • AI’s step change was foreseeable to early builders
    Illia expected a ChatGPT level breakthrough several years before it arrived, based on clear research signals and accelerating model performance.
  • Provenance and trust will define the next phase of AI
    As AI systems can be subtly manipulated, Illia argues that verifiable data pipelines and transparent training processes are essential to prevent large scale misinformation.
  • Ownership and identity matter in an agent driven world
    Illia believes individuals will soon rely on AI agents that act autonomously, making it critical that users own their models and that interactions between agents are secured and verified.

https://near.ai – NEAR AI Cloud and Private Chat products are now live, try them here
Illia's X: x.com/ilblackdragon
Illia's Substack: ilblackdragon.substack.com
NEAR X: x.com/nearprotocol

00:00 Intro: AI and Information Control
00:29 Meet Illia Polosukhin: Co-Author of 'Attention is All You Need'
01:03 The Evolution and Impact of AI
13:24 The Birth of Near AI and Blockchain Integration
15:16 Challenges and Innovations in Blockchain and AI
22:17 Privacy and Security in AI Applications
26:58 Exploring Sleeper Agents in AI
29:19 Practical AI Implementation in Teams
30:06 AI's Role in Product Development
31:41 Challenges and Future of AI in Development
36:35 AI and Economic Alignment
41:46 The Future of AI Agents
44:14 Debrief

📜 Read the transcript for this episode: Transcript of The Future Of AI With Illia Polosukhin: The Man Who Put The T In GPT |

 

For more prompts, tips, and AI tools. Check out our website: https://www.beyondtheprompt.ai/ or follow Jeremy or Henrik on Linkedin:

Henrik: https://www.linkedin.com/in/werdelin
Jeremy: https://www.linkedin.com/in/jeremyutley

 

Show edited by Emma Cecilie Jensen. 

More from Beyond The Prompt - How to use AI in your company

All 48 episodes
The Future of AI with Illia Polosukhin: The Man Who Put the T in GPTBeyond The Prompt - How to use AI in your company · 55 min
Listen in VO