The Decentralized Future of Private AI with Illia Polosukhin - #749

30 Sep 2025 · 1 h 5 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The TWIML AI Podcast - Episode #749: The Decentralized Future of Private AI with Illia Polosukhin

Episode Overview In this episode, Sam Charrington speaks with Illia Polosukhin, co-author of the influential paper "Attention Is All You Need" and co-founder of Near AI. They discuss Polosukhin's vision for a decentralized, user-owned AI, which integrates principles from blockchain technology to enhance privacy and trust in AI models.

Key Themes

  • Illia Polosukhin's Background: Transition from developing the Transformer architecture at Google to founding Near AI.
  • Decentralized AI: The potential of creating a user-owned AI utilizing blockchain technology to address issues of privacy, data security, and monopolization in AI.
  • Trust in AI: The three-part approach to fostering trust in AI which includes:
  • Open model training to eliminate biases.
  • Verifiability of inference to ensure the model operates as intended.
  • Formal verification at the invocation layer to guarantee compliance with user-defined rules.

Key Concepts and Discussions

Evolution of AI and Blockchain

  • Polosukhin's journey from AI research to blockchain began due to a need for effective global payment systems for decentralizing services.
  • Near Protocol, launched in 2020, facilitates efficient micropayments and supports various use cases, including AI workloads.

The Case for Decentralized AI

  • Centralization Risks: Concerns about monopolization of AI technologies similar to past monopolies in other sectors, leading to potential societal impacts (e.g., Orwellian scenarios).
  • User Ownership: The concept of AI being owned by users rather than large corporations, fostering a decentralized model for AI development and deployment.

Trust and Privacy in AI

  • Confidential Computing: The introduction of hardware-based solutions (secure enclaves) allows confidential computations where even the hardware operators cannot access user data.
  • User Trust: Users should have confidence that their data and the AI systems using it are secure and private.

Model Development and Use Cases

  • The platform allows developers to push applications to users without needing to handle user data directly, thus reducing liability.
  • Developers can create more innovative applications while ensuring user data remains confidential.

Challenges and Future Directions

  • The significant barriers to adoption include user inertia and the need for a strong value proposition for users to switch from established platforms.
  • Latency: Although some latency exists due to encryption processes, advancements are being made to minimize delays in decentralized cloud infrastructures.
  • The emphasis on creating models that users can trust not only involves technical security but also the continuous improvement of model accuracy and reliability.

Open Research and Collaboration

  • The podcast emphasizes the importance of open research processes in AI development, enabling collaboration and timely improvements in models.
  • Polosukhin advocates for a balanced approach between open-source principles and the necessity of retaining some proprietary elements to ensure monetization pathways for developers.

Key Takeaways

  • Decentralization and User Control: The future of AI development may lie in decentralized frameworks that prioritize user privacy and control over their data.
  • Trust through Transparency: Building user trust will require clear processes around data handling, model training, and verifiability of AI systems.
  • Future of AI Development: The integration of blockchain principles into AI could redefine how models are trained, shared, and monetized, fostering a more collaborative and secure environment for both developers and users.

Conclusion Illia Polosukhin's vision for a decentralized future of AI presents a compelling framework that addresses current challenges in privacy, trust, and control over AI technologies. As the field evolves, the balance between open innovation and necessary proprietary protections will play a crucial role in shaping the future landscape of AI.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I'd like to thank our friends at Capital One for sponsoring today's episode. Capital One's tech team isn't just talking about multi-agenticentic AI. they already deployed one. It's called Chat Concierge and it's simplifying car shopping. Using self-reflection and layered reasoning with live API checks, it doesn't just help buyers find a car they love. It helps schedule a test drive, get pre-approved for financing, and estimate trade in value. Advanced, intuitive, and deployed. That's how they stack. That's technology at Capital One.

0:34Hey folks, Stephen Johnson here, co-founder of Notebook LM. As an author, I've always been obsessed with how software could help organize ideas and make connections. So we built Notebook LM as an AI-first tool for anyone trying to make sense of complex information. Upload your documents and Notebook LM instantly becomes your personal expert, uncovering insights and helping you brainstorm. Try it at notebooklm.google.com.

1:29Again, you can have higher intelligent models and you can access all of their context and memory as well in this.

1:50All right, everyone. Welcome to another episode of the Twimble AI podcast. I am your host, Sam Charrington. Today, I'm joined by Ilya Polosuchin. Ilya is a co-founder of Near AI, but is perhaps best known as a co-author of the now famous Attention is All You Need paper, which introduced the Transformer. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Ilya, welcome to the podcast. Thank you for having me here. I'm looking forward to digging into our chat. We're going to be talking about the way you're approaching private AI at Near AI.

2:28But to get us going there, I'd love to have you share a little bit about your background and in particular, how you ended up as a co-author of this now famous paper to working on the privacy side of AI. For sure, yeah. So my background is very much in machine learning and AI research. I joined Google research because I saw the cat neuron paper, if folks remember that. And I was like, okay, we should do that, but for text and really figure out how to learn. And so Google was a great place to do that. Lots of language, lots of compute. And so my team was working on question answering, machine translation.

3:15And as part of that, due to even actually requirements and latency for google.com, we were trying to figure out how to actually build deep learning model that can consume lots of context and really reason about it without taking a ton of time to process that. And so that's where Transformer architecture comes. um now as you know there was a lot of evolution happening in 2016 2017 and was just transformer i was excited to actually put it in production and build a product around this and specifically i was excited about the idea of machines writing code uh and so in 2017 i left and with my co-founder alex kidanoff we started near ai was the idea how do we actually teach machines to code and you were one of the first authors of the paper to leave Google.

4:07Is that right? Yeah, I was the first one to leave. And so back then, I mean, this was like even, I actually left before the paper actually was officially published. And so that's why I'm the only email that's actually Gmail. Back then, nobody thought what's happening right now is possible, right? So it was like, you know, what we were pitching was kind of somewhere between science fiction and delusion. um and and the reality is back then it wasn't possible the compute wasn't there right the kind of scale at which you know we know this models are needed uh kind of wasn't there and wasn't studied very well um and on our side what we're trying to do is get a lot more training data that is relevant to this problem which is you know people writing some code for descriptions or writing descriptions for some code.

4:59And we found kind of a niche of people around the world who would do this for reasonably cheap. It was computer science students in like developing countries. And so we effectively like, I mean, now, you know. You built like a quiz platform or something like that? We built like a crowdsourcing platform for computer science students where they can go and, you know, effectively practice their coding and get paid. Now, the challenge we faced was the students were in China, They were in Eastern Europe. They were in Southeast Asia. Like in all of those countries, there's some kind of problem with paying, right?

5:33In China, people don't have bank accounts. They have WeChat pay. In Ukraine, for example, you need to sell half of your dollars on arrival if it's in a foreign currency. There was like some countries you couldn't just like nothing worked. PayPal didn't work. Rise didn't work. And so we started looking at blockchain as like, hey, this global payment network that people talking about, you know, we can just use it, right, to pay people. And this was 2018. There was nothing that really kind of matched our requirements, right? We were sending, you know, small amounts of money. It was microtransactions, even though it was computer science students, but we didn't want to make it super complicated for them.

6:11And so there was nothing that was like easy to use and actually cheap, right? Back then, kind of transaction fees were, you know in dollars and that's how we actually like okay we should solve this problem right we have this you know we talked with other people other people have this as well so like we can solve this problem and go back to ai kind of with that already in the pocket and so uh that's kind of how near protocol was born uh we launched it in 2020 it's one of the most used blockchains in the world right now with 50 million uh monthly active users it's used for payments micropayments kind of loyalty points, remittances, and variety of like financial use cases, as well as data labeling and other AI workloads as well.

7:00We have like multiple data labeling projects actually running on top of it. And so now kind of in effectively in 2021-23 as the, you know, as the resurgence of this and actually like figuring out the scale of the AI kind of came in, we started now with a renewed lens of the blockchain, looking at it and actually see how can we contribute it and how can we leverage it. And pretty quickly, as you work in blockchain, you get, I would say, indoctrinated by the value where it's kind of user ownership, right? Self-ownership, self-sovereignty. And it was pretty clear that the kind of AI space changed, right?

7:46It went from, you know, this was open research, everybody was contributing, the papers published, you know, Transformer Code was out, you know, for everybody to build on top, to like, kind of everybody was keeping a secret, things are starting to like, close up. And as we know, also just like from kind of market dynamics, that leads to then more and more centralization and monopolization of the technology and in turn becomes kind of the monopoly that we've seen before in other areas. And so the example that I use is AOL. Like imagine if internet was effectively run out of AOL. And, you know, if you want to host a website, you need to go to AOL and ask them to do this, right?

8:32And similarly, if you're a user, you can only kind of access through this. But in case of AI, because it's such a fundamental technology, right? Is intelligence as a technology, it's so much more dangerous, right? because the, like, I mean, internet is information, but this is actually like the processing, the decision-making. And so kind of the realization was that if we have only kind of a handful of kind of closed source, like profit-driven companies dominating a space, we may end up in 1984 type situation, right? where you effectively have a company that can be very much not intentionally even effectively deciding how everybody thinks, right?

9:23Because that's how we process information. That's how we're going to see the world. And so that's kind of where this idea of user-owned AI was born, which was like, hey, let's combine what we've been building on the blockchain side, which is user ownership, kind of network effects of everybody contributing and participating in comparison to, you know, centralized kind of for-profit company and create an AI that actually is on the user side, on, you know, your side, not their side. And so now there's a lot of, there was a lot of open question, how do we actually do that, right? Because you have, right?

10:01And so that took some time because, you know, all the methods that people use, for example, for privacy, for verifiability, extremely expensive, right? There's homomorphic encryption, there's ZQA proofs, et cetera. All of them have 10 ,000 to 100 ,000 times overhead. And when we're talking about machine learning, which is already using all the compute possibly we can. Let's layer in two super compute intensive projects. Yeah. And so we've been doing a lot of research. And actually, the interesting thing happened is the hardware. um so nvidia hardware and intel both kind of at a similar time enabled this mode called confidential computing so this is inside the chips itself there's a like you can enable it in such a way that even the owner of the hardware of the compute is not able to access what computation happening inside but whoever kind of requested this compute can actually have a certificate saying that this, you know, this data was run on this, you know, compute, let's say Docker, and this is the response.

11:12Yeah, I was just going to ask if this came around. This is the like secure enclave stuff that came around when folks were trying to harden Docker containers for multi-tenant environments. That was part of this. I mean, so there was a lot of research and kind of hardware over the years trying to do this, but historically it's been very like low level Like you needed to rewrite your programs in kind of like assembly, like C with like special instructions. And intellectually in 2024, mid 2024, released on the new fifth generation Xeons, this new kind of generation, which allows to exactly run just dockers.

11:55and similarly nvidia enabled work with that specific mode where effectively drivers inside can connect to the nvidia as well in the security mode so that that kind of all came together effectively like a year ago and so so since then like you know okay that enables us to that gives us some components but then now we still need the whole system right and so So that's really what we're enabling is what we call a decentralized confidential machine learning. And so confidentiality, I think, may be important to discuss why it's important. Well, confidentiality is a combination of things, right? First of all, for the user, a lot of people usually are like, oh, I don't really care.

12:44There's some people who are like, I don't care about privacy. there's people who are like, I care about privacy, but they go and still use all the products that take all their data. But there's important kind of interesting effects that there's still some things that you're not going to trust. Like, you know, we don't normally walk around with like a hot mic that records everything we say. Although that is becoming popularized by some AI companies, right? It is, yeah. It's a conversation that we're having now in spite of how crazy it sounds or would have sounded a few years ago. Yeah, and this is an example of somebody who should definitely use our platform.

13:24Okay. And similarly, like, yeah, I mean, there's just so much context in your life that right now we're still not putting on, you know, into the AI systems. And like, I think everybody's kind of on a different spectrum, right? Like I, for example, don't trust, you know, my email and my calendar to maybe an AI company. Some people would, right? But then they wouldn't trust as their medical data, but maybe they wouldn't trust as their bank data, right? So there's always a threshold where you kind of get like, maybe I shouldn't do that, right? And so what we offer is effectively removing that threshold and say, hey, actually, it's all confidential, all end-to-end encrypted for you.

14:05And you can trust that there's no other single party, not developers, not operators of hardware, not model developers, et cetera, are able to access it. So it's as if it was local and potentially even better than local because you have additional security mechanisms. And so it's the idea that... Oh, there's so many questions I'm trying to ask here. So you're describing a system that many people say, if I can't run this locally on my machines, I'm not going to run it. But it sounds like what you're trying to do is more like create a system that would allow like remote and cloud based, but also private to the same level as local AI.

14:53Am I parsing that correctly? I mean, I run some of the models locally, but I mean, obviously they're not as intelligent as what you can have in the cloud. They're not as fast. But importantly also, even if we have a smarter model, you still have a lot of things that are happening on the background that you want to keep happening. You want to set up an agent that runs and reads all the news and summarizes it and processes it or workflows, et cetera. So there's always going to be a need for background work and analysis and surfacing it, even as local models improve. So I think that's really the...

15:30And you want a backup. You want a way to synchronize between devices. There's a lot of functionality that you want that requires cloud. And right now, there is no really private cloud. There's multiple companies usually who actually have access to the data and to the computation. The other thing is for developers, actually. If I'm an application developer, if five, ten years ago, data was a goldmine, it's becoming actually a liability. And it's become a liability both like in Europe, for example, there's GDPR data privacy. We have California data privacy. There's all this kind of different data privacy laws that are popping up.

16:14In China, you actually need to pay data tax if you're using consumer data. Yeah. And so the reality is actually like if before this was like really valuable for many use cases now, it's actually a liability. And so this actually creates a platform where you as a developer don't need to deal with the user data. You're effectively pushing software to them. Again, similar how local works, right? You pushed the application to the user and it runs with their device. You don't need to deal with whatever data. But again, now you have background processing. You can have higher intelligent models. And you can access all of their context and memory as well in this.

16:56So you kind of get like interesting combinations from both sides. There is another side, which is also interesting. So right now, if I'm a model developer, like actual, you know, frontier AI models, I have a interesting challenge where, you know, if I'm not, you know, the largest labs, which are only few, let's say I developed a new, you know, amazing model for anime characters or whatever. And now I have a choice. I only have some amount of compute. I either use this compute to serve customers

17:33or research and develop a new model and continue trading. And so if you get a lot of usage and you're kind of limited by that, you then still need to handle all of their data. So you have this kind of challenges with all the DPRs in the world. And now if you say like, oh, but, you know, this ton of clouds, GPU clouds around the world, you can just go and, you know, rend them off when you need it. And the challenge is actually this model developers don't trust third parties because they're afraid that their model will leak. And this has happened where the model weights have leaked from third parties.

18:15What's a specific example of that happening? So Mistral gave its weights to Hugging Face and it ended up on 4chan. oh wow i hadn't heard that and so uh and and so i mean this is the same reason why like a lot of the um like you know people build their own clusters because they want to control everything i mean all this is like some efficiency comes from like optimizations but a lot of it is also just like we want to control you know like literally have guards on the doors to make sure nobody can access so we're also solving that problem interestingly because because of secure enclaves you can actually encrypt the model weights and they only get decrypted inside the secure enclave.

18:54And then user data is also private, right? So effectively bringing kind of privacy from both sides, like kind of model developers don't need to deal with user data. They kind of don't need to, you know, they also don't need to rent the hardware, right? It gets kind of rented at the moment when users using it. And then on the other side, the users don't have access to the model, but they also know their data is not going anywhere let's pause here so you're introducing a twist here so like i thought we had this trend like at least my mental transition was okay privacy is talking about this local thing and then you know the previous time i interrupted it was like no it's this cloud thing but now what i'm hearing strikes me more as like like this decentralized thing where it is actually local, like it's running on my laptop or device, but also on other people's devices.

19:53Like, and the reason why I'm saying that is because you're saying like, I don't know, you said something in particular that made me think that like the model's coming to my device and the data's coming, you know, the data's on my device and the training's happening there. And then, like, wait. Let's maybe back up and, like, kind of frame what we're talking about topologically, I think. Yeah, topologically, this is compute hardware, let's say GPUs and CPUs, that live in this decentralized confidential cloud. Ah, so it is a cloud, but it's not the same cloud. It's, like, a decentralized. Or maybe could you run it in a...

20:42It's hardware, so probably not running it on. Yeah, you need bare metal to be configured and then join the network. But yeah, I mean, Amazon data center can repurpose itself to become a member of this cloud. Yeah, so this is a cloud. So you as a user are accessing it. but it gives you very kind of close guarantees to the local and you can potentially even add additional like you know pin code to fa etc like you can actually restrict things that you may not even have on a local host because i mean local host you still can access the hard drive physically here you like you still have like a level of interaction that can provide additional controls.

21:29But there's no other, like there's no third party that can access that, like your data and your compute. Right, so this cloud is kind of the middle layer. As a user, as an end user, like I'm contributing my data in some way because I want some processing on my data or to access intelligence. And the cloud can't access my data. And presumably the model provider can't access my data, but the model provider like is providing the model into this cloud and it can access my data and return some results back to me. Correct. Interesting. Interesting. You mentioned at one point that data is a liability to the model providers, like presumably they need access to, if not end user data, like some data to train their models, to tune their models.

22:18Like, you know particularly now in the part of the you know ai life cycle that we're in like we're finding that one of the key differentiators for you know organizations is building this data flywheel where they're getting early users getting access to their interactions or traces and then improving their models based on that using, you know, reinforcement fine tuning or whatever. Does this process like, I get that user data can be, you know, can have a cost, you know, to the model provider, but, you know, that's not all there is to the story. Like they still need that user data to improve. Like how does that play in, in this model?

23:07Yeah, so I think it's, first, I think that is changing as well, the need for the actual user feedback data. But first, before we go there, why is this liability? So imagine I'm a European, right? I'm in Lisbon right now. I use OpenAI, OpenAI trains on my data. And then I go and I evoke my GDPR law and say, hey, remove all my data. I'm assuming that that is not a resolved issue either because I can't believe it's because no one has asked yet. I'm assuming it's because they've just ignored it. And at one point, you know, there may be a challenge, but we're just not there yet. Yeah. So that's what I mean by liability, right?

23:50I mean, you know, maybe OpenAI has the money to pay their fine. Like, I mean, similar how Google and Facebook have paid, you know, billions of dollars in fines. But if you're a smaller model developer, that's why I was kind of using other examples. It just effectively can be like existential. Yeah, I get it. It can be a liability. So that's piece number one. The piece number two is actually why I think the space is transitioning from user feedback. So let's use a DeepSeek example. So when DeepSeek released their first model, which was at least at the time from the open rate models was a state of the art uh like our for example for r1 it did not use the explicit like user queries right yeah we're talking about like this transition from user feedback to verifiable results it's combination of data labeling like indeed human but like you actually want a very specific supervision and you want to control kind of what feedback you get so you So human labeling is very, I mean, its own space, right, where there's a lot of know-how how to do it properly.

24:59And again, we've been running that for years. So there is synthetic data. There is kind of this indeed verifiable math, physics, logic, kind of coding, et cetera, which clearly improve reasoning as well. So there's like a lot of the, I would say, even if we're talking about shifting the vibe of the model, which I think is something that's usually credited to like Claude versus OpenAI, right? like the vibe is different and kind of how even that is you probably want like more trained people to actually give feedback versus just relying on kind of very, very noisy signal that comes from users. Now, I mean, again, this is like, I would say it's not fully transitioned into this and depends on the use cases.

25:57So it's important to note. But as I said, like the cost versus reward is shifting and and like it's shifting i think faster than uh at least some people realize it interesting and would you say that that is because you know we're just learning how to manipulate you know the vibe or or output characteristics of a model based on more kind of curated uh you know, training data or feedback, or are there techniques that are enabling this shift? Like, what do you see as the driver of this shift? I mean, I think it's combination. I mean, again, even the original chat GPT, the GPT 3.5, it was like on the human label, but not on the actual user feedback, right?

26:52Sure, yeah. So sure there's like this transition from RLHF to RFT and like the verifiable stuff and all that. But it sounds like you're speaking just like a broader trend. Let's go back to like Google, right? Of Facebook. Like Google and Facebook learn from user behavior, like directly, right? There's no, nobody's human labeling like, hey, which search result? I mean, there's like a little bit of that, but like in mass, it's mostly just signal from user clicks, right? And then at large scale, it's been processed into like actual signal for the machine learning. I think what LLMs did is kind of transition that to like, hey, we actually just pass a lot of unsupervised data, right?

27:35Not like user click data. And then we add a little bit of a human, like a very specific human labeled data. And for that, we also need, like the better the foundational model, almost like the more complex things we want people to label. And so kind of like, again, for our example, we were finding computer science students because we needed people who code. And so just kind of the, you know, some maybe broader data wouldn't be that much, that useful. So part of it is we've collected enough data and we've kind of baked the generic stuff into the foundation model. So now where the innovation is happening is, you know, bringing in more subject matter expertise or specialized skills.

28:24Or indeed like a verifiable thing or combining like, you know, synthesizing data and then using another model to evaluate it and kind of then human labeling, like all those kind of pipelines, right? Yeah, yeah. And again, it depends like for some things like, I mean, audio, for example, same thing, right? It's like, it's great to have a bunch of audio from people that use your product. But then again, like you may get in trouble so much, right? We've seen that happening. It's better to just pay people to contribute their audio and like sign off the rights. and it's and it's like it's kind of pretty straightforward to do that like we have a project on near as well running that uh and so like it like the amount of yeah like i'm kind of that's what i mean like the the shift like how how much we can collect data and how how useful that is versus getting a bunch of user data and then dealing with all the repercussions of that uh so So I had asked about the, you know, this like creating a data flywheel and the importance of that for, you know, companies in the space.

29:29And your response is like kind of to reinforce the liability aspect of that data and then talk about this broader shift that's happening to more specialized data creations last collection.

29:47I think yeah I'm not sure that I'm fully sold that that flywheel thing is not important but if you don't have anything else to add on that we can move we can move on I mean the other piece we do want is is opt-in users can contribute their data right so we do we do want like if you want to contribute you should get something in result right it can be economic it can be credits it can be something and the underlying blockchain has the mechanism to make that tenable in a way that it's not tenable today yeah and and so one of the models for example that uh like a new business model that we have uh been building is right now i don't know if you saw this post by dario uh where he said like hey every model we've built was a successful thing but do we spend their own businesses yeah so we actually do i mean we we talked about this like last year like where effectively every model gets its own token so like a way to distribute reward and value from the revenue while also rewarding was just i mean effectively you can think of shares um where you can you can actually like whoever contributed data gets a token of this model that were trained on their data.

Read the full transcript

31:06And then the revenue is distributed to the token holders, right? As like for the model's lifetime, right? So we can actually like run and guarantee those parameters as well. So what's interesting about that idea is that, you know, when you talk about this idea of kind of the broader conversation around like compensating rights holders for content that's consumed, I feel like it gets to be like, oh, I'm kind of like, how would you ever do that? Like, you're crawling all of the internet. Like, how would you even possibly begin to do that? But just you describing this token model, it's kind of like, oh, well, maybe you could do that.

31:53Like you're crawling a site, you know that site, you reserve a token for that site. Someone needs to verify that they have control over that site to access the token. now they have a share of the model and they can you know gain in the rewards and all of a sudden at least for me it like clears up a lot of the just like it's not possible feeling about it yeah exactly and and and so that's a really like that example as well as you can contribute data privately so you can say hey i want this data to be used in the model training but i don't want any anyone to see it right for example right and so you can also do that or you can say hey i want it to be used at inference time as part of the search retrieval index, but not a training.

32:38So you contribute, like it's payable, like for example, for payable data, New York Times can contribute their data fully privately into secure enclaves. And then it's going to be used at retrieval time. It's going to get recorded and they're going to get paid for that. But it's, so things like that, you just get a lot of these pieces kind of for free in this new model. So earlier in the conversation, we talked a little bit about closed models versus open models. And did you anticipate, you know, as, you know, chat GPT happened and kind of the private foundation models began to establish themselves, like, did you anticipate that there would be like fast followers of like these open models?

33:23or has it surprised you how quickly open weights and to a lesser degree open models have come about and their capability? No, I was actually, I think I was trying to remember, it was like February 23, I was talking about like, hey, open source is going to catch up because, yeah, I mean, I think the challenge is like with pure open source right now, and again, this is something we're solving, is that I built a model, I released it, Everybody is like, cool, here's the stars on GitHub, stars on Hugging Face, but then you don't make any money. Because of that, it also becomes less about open source and about open weights, and people keep the source so they can continue doing things.

34:12In result, we're actually wasting a lot of resources because everybody is redoing experiments because we actually don't know what were the things people did to get to these results. And so everybody kind of need to reproduce or poach the people who've done it. And so kind of the way we think about it is to reverse it, where you can actually have an open process of training, right? So the training, I mean, the data either fully open or this like kind of encrypted data, right, that you can run over. So available, but not necessarily transparent. Yeah, yeah. And you need to pay to access it in whatever your model token or some other way, if you're planning to monetize differently.

34:55And then the resulting weights are actually also encrypted and only run in this kind of DCML model. So you can actually monetize it. So you can receive kind of revenue from using it. You can say, hey, compute cost is X. I want 20 cents for each million token over that to go to the model developers and all the contributors to do that. So you kind of can reverse that and get actually like actual open research and collaboration happening there while monetizing the outcomes. And the benefit is you still get all the properties of open source, right? Everybody can use it. there's no like way to stop using it you can even run it on your hardware if you have the modern like blackwell or hopper and like you need to set it up in this confidential mode and uh uh yeah you can fine tune it you can do all those things on top yeah i mean a missing property is the ability to to see it and change it maybe uh you know maybe that's more true for software than for a model if you have the ability to to fine-tune on top of it that's exactly way that you can change it there's very little people who go and like do a brain surgery on a model i mean there's few that like do you know like evolutionary algorithms and other stuff but yeah usually it's either fine-tune post-train rl etc you know although i like the pushback that i would offer is that like if the best models were open you know maybe we'd see a lot more brain surgery and maybe we'd have a more interesting kind of ecosystem of of results like i think you know there are people that do that kind of thing for various reasons with you know open models but i think there's a lot less invested in them because they're not as good as the closed models um i don't know interesting um you know i'm curious like you know there are definitely you know, some aspects here that make a lot of sense to me.

37:08Historically, you know, there's always been this big barrier that's not at all technical, and that is, will people pay for privacy? And whether that's, you know, currency or, you know, the sheer force of will that's required to jump over the hurdles to achieve it, you know, inconvenience. uh you know do you do you feel like you know this is different or it's different in this space or like how do you think about that challenge i think i think twofold one is i think there there is a audience that will pay about for privacy right um and it's not i don't think the issue is that there's never that audience is that it's relatively small it is relatively small but i think the the idea here and kind of, I think everybody's on a threshold, as I said, of what they feel they would give to the model.

38:04And the more you give, the better. Actually, we're getting to a stage where models are generally, they're sufficiently intelligent and actually it all becomes about context. About context management, about tool management, about all those pieces. And so So the idea here that kind of we are very much going after is that because it's private, you can actually share a lot more with it. And you will add your email, your calendar, your medical data, your financial data, your crypto wallet, et cetera. And so it's able to manage your whole life, not just like some aspects that you were willing to share.

38:41and so and kind of the first cohort that actually cares about privacy that's you know that is our early adopters who we kind of target to really enable this but then kind of as this becomes mature now it's it's appealing to more people because it's a better product and kind of smarter product more intelligent product again not because the model is right away more intelligent but but because you have more context. The other side of this is actually because of this open research process, what we are aiming for is to have people, again, fine-tuning specialized models for specialized use cases. And again, this is where because they have a monetization embedded into this, they can actually invest effort and time and compute to actually build interesting specialized models that may be better at financial use cases, healthcare, et cetera.

39:43And again, all of them are available on your platform. They can compute over your data. And then again, now that all your data is there, it's useful, there's useful models. Other developers as well, again, if I'm building a note taker or something else, I can build a note taker right now that takes all your data, listens to it all the time, sends it to my server, stores it on my server, et cetera, but liability and also now it's a hurdle for everybody else to adopt it. Or you can just say, actually, I'm going to build my app and deploy it into this cloud where it runs on your side and saves context there as well in your data store.

40:28And so now your data store becomes even more useful because it has all the notes as well there and your AI can now read over those notes. So you don't need to like, again, merge those things with Zapier and do all those things. Like that's the idea. It's like you have network effects of kind of more context, more data, more applications building around the user. And so, yes, like it starts with early adopters who care about privacy and kind of layers on as more and more applications and things become available on this cloud. So I want to talk a little bit about the process of making models available to this environment.

41:08To what degree is it a simple lossless transformation of an existing safe tensor, GGF file, whatever, some weights file versus am I having to rebuild my model in some new paradigm? How does that work? Yeah, I mean, we effectively run, you know, VLLM and customized version of VLLM. So everything that, you know, normally served already works. And if you need something custom, then you can also package your own Docker. I mean, that is less secure from a user side, but it's also available. It's the secure enclave that ensures that there's no kind of man in the middle attack between VLLM and like the model weights and the customer data.

42:03Yeah, so exactly. So what's happening is, you know, you checkpoint your, you know, model weights on chain. So, you know, the hash of the model weights and the encrypted hash as well, you know, the encrypted data is uploaded kind of to decentralized storage. And now when somebody wants to run a model, they have an encrypted TLS connection directly into the secure enclave. that secure enclave you know gives you back the effectively signed certificates that it runs in secure enclave you can also verify them with our own chain kind of uh key management system and then we also have this concept called multi-party computation so near blockchain itself kind of uh right now part of our nodes form this multi-party computation network which allows inside the secure enclave effectively have its own private key to decrypt things.

42:57And so that's kind of how all these pieces work together. There's secure enclaves, but also if you encrypt something, you need to encrypt it with some key that is only known, the private key of this is only known inside secure enclave and nowhere else. And so this is where this NPC network enables that. And so, yeah, like effectively, you know, you encrypt locally, you upload it, it checkpoints. And now when user calls, they know that like effectively the secure and playful respond that this model hash was run on your data. Here's like signature by NVIDIA, Intel, et cetera. And you can also verify kind of certificate provisioning.

43:40And is that signature created at like, you know, by some process at the boundary or is it, you know, intrinsic to the inference actually happening by this model on this data? Like, is it a... So the signature certifies a Docker container that runs inside SecureEnclave. And so Docker container is our Docker container with VLM that runs this model hash. So we attach that. That's what I'm saying. If somebody builds custom Docker container, you can do that, but then user needs to trust your Docker container. Got it. So the trust boundary is the container. And if the container is doing what you say it's doing, then the signature certifies that that was a container that was actually used.

44:26Yeah. And in our UI, we have effectively like, you know, like a green shield that you can click similar like HTTPS works. You go there and it gives you like, it gives you like, hey, it's all correct. And then you can go and actually verify all the signatures and all the certificates and even which GPU it ran on and like other stuff as well. And like links you like effectively to all the relevant Docker Githubs and other things you need to know if you want to like re-verify everything yourself. And so that's maybe an interesting segue into like, you know, where you are with all of this in terms of, you know, how much of it is, you know, aspirational, how much of it is built.

45:09Like you're clearly you have at least the notion of a user interface, if not an actual user interface. Like how far along are you? Yeah. So we have, I mean, we have a product that we can, I mean, that we're in testing and alpha testing with, you know, kind of cohort of users. it's it's both developer products so you can you know buy credits and effectively use confession inference in your own applications as well as we have a kind of consumer product which is you know private chat GPT effectively which indeed provides you all of the kind of certification and verification information if you want while using it and then the the custom model right now is that's in development, can be coming out in a few weeks.

45:59And remind me, which part is the custom model? So this is where you can encrypt and upload your own model. Oh, got it. Okay. Yeah. And then, I mean, the fine tuning and kind of training that's coming a bit later. And we started off talking about the fact that encryption is computationally complex, like it brings along its own costs you know relative to um you know the per token inference costs that you know someone might see you know whether it's open ai or open router or something you know like where does this where do you expect this tend to fall to fall relatively speaking so it will be affected at the same cost the overhead on the computation side is one to five percent yeah so it's very minimal and it's mostly just i mean it's like encryption decryption on a boundary uh yeah and kind of constrained by that not by computation and so what what do you see is like the main barriers to you know near scaling you know this approach and um you know getting people on boarded or, you know, on boarded, you know, not just technically, but like ideologically and that kind of thing.

47:25Yeah. I mean, I think the kind of, as we just discussed, right, like are people willing to pay for it? Right. That's, and again, I think, but you, you know, well, I think what I heard you say is that like, I'm not really paying, like I'm paying, it's going to cost the same, right? Or actually, it's open source model, so it's actually cheaper than OpenAI. Yeah. And I did reference the idea that the cost is also convenient to learning a new thing. Like, maybe we should hit pause on the adoption conversation and go back to, you know, we talked about from a model provider perspective, you know, that's the same.

48:09They just deploy into your container the same model format. What about from an end user perspective? I guess in the general case, they're just using a chat app that, you know, or whatever. They're using an app. So it's not different from them. So presumably, like who is the argument that, you know, there's no particular inconvenience cost to anyone in this, you know, ecosystem? Like everything's kind of the same or? yeah i mean the goal is to make it everything is like either the same or better right like you either don't i mean it looks exactly i mean very similar experience right and the idea that because it's kind of private you can also have additional features that you wouldn't have in uh in the public it's also it is open source so you can you know people can contribute you can fork etc right so you know you cannot just go and fork chat gpt and add some stuff and have your own version with some things but here you will be able to do that because it's your data right that travels with you so you can like launch your own version with custom like improvements and then everybody who logs in will get all their messages all their history all their memory all the apps with them so it's also like kind of detaching your identity from specific application but What do you say to the skeptic that says, you know, it all sounds too good to be true?

49:37Like, where's the, you know, besides from the fact that, you know, building software is hard, building a company is hard, getting people to fund weird things is hard. Like, what's the hard part? I mean, the hard part right now is it's inertia right now. I mean, I think like there's a cohort of people who are like, hey, I'm already in Google ecosystem, why would I do anything? Right. Everything is already here. Google has my email and my calendar anyway. Why do I care if it's private? Exactly. Yeah. So, so I think like there's, again, there's an aspect of that. And like, I think people generally trust Google with their data.

50:17So I think that is like, that's inertia that we kind of need to address. Right. And I mean, not to, I mean, I work with Google, so there's indeed a lot of security to make sure the data is protected. But obviously there's still like, I mean, there is a way for somebody in customer support to help you with your data. So there's a way for a third party to have access to your data. There is, I mean, we've seen this with OpenAI, right? There's news that they're effectively scanning all the chat logs and then the ones that are flagged are sent to human to evaluation and then to police. So you effectively have potentially humans looking at your chat logs.

50:59We had obviously data leaks from Grok and others that your chat logs got visible and indexed. So I think that is the backdrop of why we get adoption, but the inertia is the other side. OpenAI already has whatever, half a billion users or a billion users. So we need to have a product that indeed can deliver on if people want to switch and use. How about latency? Admittedly, encryption is not the latency killer that it was, you know, 10 years ago or so. As a lot of that stuff's getting pushed in the hardware, like, is it an issue for you? Not really. I mean, we have, like, it's a little bit higher latency, again, just because, like, when it goes through the boundaries, there's, like, some additional delay on kind of encryption.

51:59But we're also working on, like, it's also an engineering challenge of just, like, streaming encryption and stuff like this, like, improving that. I mean, we're using TLS right now, right? There is no additional latency. This connection is inter-encrypted, although we're actually going through centralized server. We're actually trying to figure out how to make it as direct as possible. So ideally, actually, latency is less because ideally right now, yeah, you actually should be accessing the closest GPU, ideally, in your city. I mean, I've talked to people who are like, there are going to be data centers everywhere.

52:35People will build mobile data centers, et cetera. and like you want to find that one connect to it run your compute on it you know hydrate your data there like that that's kind of where you know this infrastructure can move to and i really deliver on that like uh i we're not there but that that's kind of the vision is really actually reducing latency because you know you're sitting in philippines you right now you need to go to texas or whatever where open ai servers are uh versus like there's data center in philippines actually are sitting underutilized. And so you should be using it. I guess another question that I have is that I think it, you know, when you boil it down, a lot of the value proposition here is, you know, around trust and the user being able to trust the interactions they're having with AI and privacy is a part of that.

53:34But there are also still fundamental trustworthiness issues with, you know, your creation, the transformer, like its ability to give you results that are worthy of your trust, you know, hallucination, for example. Are you doing anything there? Like, do you see that as, you know, how do you think about that as an issue and um you know what's your you know what are your thoughts about how or if or how that gets solved yeah i mean that that is an important question and um indeed especially as we like i kind of mentioned like i think the ai will be how we interface with computing and you ideally, yeah, want to make sure that there is no kind of biases in that that are not represented with your view.

54:31So I think there's like few components improving trust in these models. I think it starts with indeed the open process of how these models were trained because there's this concept of like sleeper agents, right? you can actually like train things into the model so that at some point during some conditions it activates and behaves in a different way than normal and so you can like you know introduce vulnerabilities in the code into a coding model like based on some condition like you know i'm assuming if somebody wants to do a new stacks net that's how they're going to do it um the so you want to know how it was trained.

55:14Then right now, you're running, even if it's open source model. Isn't that first point an argument for true open source as opposed to just sending a certificate or a stamp of something? So that's right. You want open source, but you don't need open weights. so you want to know what went in but the weights itself can be encrypted so you can monetize but yeah so that's what i'm saying we need to open source not like open weights is right now and just kind of everybody's like i mean effectively it's useful but it's mostly like as if you know so is your was your argument earlier that if you can if you can close and encrypt the weights then you have greater, you anticipate greater willingness to open the source, like make the training process more transparent?

56:18Because you can monetize. Like in this encrypted weight model, you can monetize the usage of the model. Yeah, that's right. So your point was people are holding onto the source because they want to retain some monetizability and they're making the weights open. you know so that's their piece that they're holding back so the opening of the weights is the marketing right for them to then leverage their close source thing to then cook something else right either for specific customers or next model or whatever this is uh or attract like to their app but if you can monetize actually the model you trained and you have the whole process open um and especially if there is like some semi-formal way for people who leverage your learnings as well to then uh kind of contribute back as well as well yeah but i mean like that assumes that there's not a lot of perceived innovation in the training process and i don't know that folks that are training you know frontier models in the like ahead of the the you know at the frontier sense of the term would necessarily believe that right i mean there can be that's what i'm saying like but i i think the the the lag on the frontier right is you know three to six months and so i think the benefit here is like if you are able to effectively do a training run and start monetizing it you're like you can leverage that uh you know monetization to then potentially to reinvest into the next thing, et cetera.

57:59But because you opened, everybody can go and contribute and maybe come up with new ideas, et cetera. So you're just kind of accelerating this process. Again, this is how the computer science and AI worked before. We would open it up. The papers was out as soon as possible. And now everybody's delaying everything like at least six months to a year. So we're solving big picture trust. The first part is openness. Yeah, well, yeah, open research, open data, knowing what data goes in and what bias there. Second is verifiability of the inference, right? Again, right now, you know, there's people complaining that like Claude gets dumber in, you know, daytime and like maybe it doesn't, we don't actually know.

58:45Yeah. So like having verifiability, again, this can be something where like, you know, when you specifically ask what stocks to buy, there's like a rule that says like, let me tell you to buy this. Only use SmartQuad. So like verifiability of that, that gives you another level of, and then I agree that the third part is actually like, you know, especially when we allow those models to go and start doing actions, like how do we make sure it doesn't do anything? And so there I think there's a few interesting areas. I mean, there's a lot of research, right? People are trying, you know, how to like ground hallucinations, how to do all those things.

59:25So All of that needs to be done. And again, I think open research process would help a lot with that. But the other aspect of this we're looking at is actually formal verification. So right now, when we talk about software, when we talk about the CI systems, we're testing some use cases. We're testing, we have evals, maybe we have vibe testing, but we don't really have guarantees that the system will comply to some requirements. And so formal verification is a way to actually achieve that. And formal verification should be happening at the invocation side. So as you call something, right? Like, hey, so the example is like a little bit closer to blockchain space.

1:00:12You know, if you're putting money into something, you want to make sure you can, you know, at least get as much money back, right? That nobody will be able to steal your money. For example, like savings account type thing. and so you want to have the guarantee that when you put in your money that you'll be able to do. So you want the proof at that time. So that's kind of like we call it at invocation verification. Similarly, when you're calling a system, you can say, hey, you can access my email, but you cannot delete anything, right? And you cannot leak anything. You cannot do this set of actions, right?

1:00:49you know um and so the system like the the te that runs it the secure enclaves that runs it like proves to you that the code that it runs inside indeed complies with this requirements so this docker hash for example you know complies with this requirements and so now you have not just proof of verifiability but as a proof of specific you know preconditions and and if and if it fails right if it doesn't match your preconditions right this this doesn't execute And so today, what we talked about being possible is this idea that if you can get access to whatever the source code is that's going into the Docker container, you can certify that it's this thing that you know and trust that operated on your data.

1:01:38and what you're referring to is at some point in the future where you can not necessarily have to know what that thing was, but know some conditions about what it's able to do or how it's running. And the verification is not just like a cryptographic hash, but it's like verification and validation of the code and what it's doing. And I mentioned that for folks that are curious about this, just a couple of episodes ago, I had a really interesting conversation with Christian Sagetti about verification, you know, dug deep into all this stuff. He's one of the pioneers on that for sure. And so that was the third of your three things for kind of solving this broader trust issue.

1:02:29Yeah, I think that that's going to be a really important component of that to make sure that we can actually trust the systems. Again, as a user, you're probably not going to go and review Docker container code and ensure. And especially, it gets complicated when things are starting to call each other. Like when you have one AI system calling another AI system, calling some MCP tool, calling another AI system. And so the idea here, you actually have, like the properties are actually composable, right? So like if this service proves to you that it will not, you know, will work in this specific way, when it calls other services, those services need to prove it as well.

1:03:13And if it doesn't, it's not able to call it. So that's, I think, is really important is actually kind of gives you this composable effect, which I think right now is the biggest problem. As you can imagine, the systems become more and more complex and you have all these AIs talking to each other. And we have no idea what they agree on doing and what they end up actually executing. So I think that's going to be a pretty fundamental system change. but it actually requires this verifiability because you need something to guarantee that the code inside is following this property. So you need some containerization that you can trust.

1:03:54Very cool. Well, Ilya, thanks so much for jumping on and kind of talking us through what you've been working on. It sounds like super interesting stuff we've had. I think we've covered this idea of privacy on the podcast to some degree in the past, like we've talked quite a bit about differential privacy, talked a little bit about like the open mind, PySift kind of decentralized training kind of stuff. But this is definitely a different take and one that I think is, you know, kind of right in line with the direction that things have gone from a, you know, Transformers Gen AI perspective. So I'm super interested in, you know, seeing how it all unfolds.

1:04:40Yeah, I appreciate you having me here and diving in. Yeah, thanks so much. Thank you.

1:05:08Thank you.

From the publisher

In this episode, Illia Polosukhin, a co-author of the seminal "Attention Is All You Need" paper and co-founder of Near AI, joins us to discuss his vision for building private, decentralized, and user-owned AI. Illia shares his unique journey from developing the Transformer architecture at Google to building the NEAR Protocol blockchain to solve global payment challenges, and now applying those decentralized principles back to AI. We explore how Near AI is creating a decentralized cloud that leverages confidential computing, secure enclaves, and the blockchain to protect both user data and proprietary model weights. Illia also shares his three-part approach to fostering trust: open model training to eliminate hidden biases and "sleeper agents," verifiability of inference to ensure the model runs as intended, and formal verification at the invocation layer to enforce composable guarantees on AI agent actions. Finally, Illia shares his perspective on the future of open research, the role of tokenized incentive models, and the need for formal verification in building compliance and user trust.

The complete show notes for this episode can be found at https://twimlai.com/go/749.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
The Decentralized Future of Private AI with Illia Polosukhin - #749The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 1 h 5 min
Listen in VO