In short
Podcast Summary: Building Open Infrastructure for AI with Illia Polosukhin
Introduction
In this episode of Software Engineering Daily, host Kevin Ball interviews Illia Polosukhin, a veteran AI researcher and co-author of the influential Transformer paper, “Attention is All You Need.” Polosukhin discusses his journey from an AI enthusiast to a co-founder of Near AI, focusing on open-source infrastructure for privacy-preserving AI systems.
Key Themes and Discussions
Illia Polosukhin's Background
- Early Interest in Technology:
- Began programming at age 10; started with video games.
- Got involved in machine learning at age 14, building neural networks in Pascal.
- Career Development:
- Moved to the US to work for a machine learning company after impressing them with remote work.
- Joined Google Research, focusing on natural language processing (NLP) and question answering systems.
- The Transformer Model:
- Addressed challenges in neural networks that read sequentially (e.g., RNNs).
- Proposed that Transformers could process entire texts in parallel, significantly improving response times.
Transition to Decentralized Technologies and Near Protocol
- Early Interest in Crypto:
- Transitioned to decentralized technologies in 2018, recognizing the potential of cryptocurrencies for solving global payment issues.
- Near Protocol:
- Developed to provide scalable blockchain solutions with 50 million monthly active users.
- Focused on creating an ecosystem for applications ranging from remittances to crowdsourced data labeling.
Governance and Safety in AI
- User-Owned AI:
- Advocates for a model that prioritizes user interests and data privacy.
- Proposes mechanisms to ensure models are built with user ownership and transparency in mind.
- Data Bias and Safety:
- Discusses the complexities of bias in training data and the risk of "sleeper agents" in models that could introduce harmful outputs.
- Highlights the need for governance structures to ensure safe AI systems.
Blockchain and AI Infrastructure
- Decentralized Compute Network:
- Proposes a decentralized cloud infrastructure that distributes AI workloads across underutilized resources globally.
- Emphasizes the importance of privacy and IP protection in deploying AI models.
- Training and Revenue Models:
- Suggests that community-driven initiatives can reduce costs while ensuring data integrity and model efficacy.
- Discusses potential funding strategies using tokens to incentivize contributions to model development.
Future Outlook
- Model Development Timeline:
- Envisions rapid advancements in decentralized AI infrastructure driven by increased availability of confidential computing technologies.
- Predicts a shift within 1-2 years towards user-owned AI models outpacing centralized models in innovation.
Key Takeaways
- Importance of Open Infrastructure: Building open-source infrastructure for AI is crucial for ensuring user privacy and ownership.
- Decentralization as a Solution: Utilizing blockchain can enhance trust and coordination in AI, providing access to untapped computational resources.
- Community Involvement: Encouraging collective contributions to model training and data collection can lead to significant advancements in AI capabilities.
- Governance and Safety: Addressing data bias and safety in AI is essential, warranting robust governance frameworks.
Conclusion Illia Polosukhin’s insights into the intersection of AI, blockchain, and decentralization underscore a transformative vision for the future of technology, emphasizing the need for user-centric approaches and collaborative efforts in AI development. The discussion on building safe, transparent, and efficient AI systems is particularly relevant as the landscape of technology continues to evolve rapidly.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Ilya Pelosikin is a veteran AI researcher and one of the original authors of the landmark Transformer paper, Attention is All You Need, which he co-authored during his time at Google Research. He has a deep background in machine learning and natural language processing, and has spent over a decade working at the intersection of AI and decentralized technologies. His current venture is called Near AI, and he's focused on building open-source infrastructure, tools, and products for agentic, privacy-preserving AI systems. He joins the podcast with Kevin Ball to discuss his journey, the origins of the Transformer model, the vision for user-owned AI, document-oriented development, and much more.
0:44Kevin Ball, or KBall, is the Vice President of Engineering at Mento and an independent coach for engineers and engineering leaders. He co-founded and served as CTO for two companies, founded the San Diego JavaScript Meetup, and organizes the AI in Action discussion group through Latent Space. Check out the show notes to follow KBall on Twitter or LinkedIn, or visit his website, kball.llc.
1:19Ilya, welcome to the show. Thanks for having me. Yeah, excited to get to talk with you. Let's maybe start with a little bit of intro about you, your background, and what you're up to these days. For sure, yeah. Well, I've been, I guess, tech gig since I was 10 years old. Been building a lot of video games back in the day and then got really excited about machine learning when I was like, well, about AI in general and then started learning machine learning when I was like 14. I was building my first neural networks in Pascal and got a job actually remotely. So I'm originally from Ukraine working for this machine learning company out of San Diego.
1:57And they were happy with my work and so they offered me to move. I moved to US, which was exciting. And then I saw the cat neuron paper that came out from Google from Andrew Yang and Jeff Dean. And I was like, okay, this is the thing. Like the unsupervised pre-training, learning about concepts in the world, you don't need supervision. And so I was like, okay, I want to do that. And so I applied, I got into Google research. And I always thought that yes, images are cool, but there's thousands of species that can see, but there's only one, maybe some people argue, maybe two that can actually speak.
2:36And language is effectively a way we test intelligence, right? We ask questions. You ask a person to read the text and we ask questions if they understand it. And so that's why I wanted to focus on natural language. We were doing question answering, trying to build products into google.com, where when you ask a question, it would give you a response. This is where your previous, I guess, CTO, right? And was my director back in Google. And one of the challenges we were facing, actually, that the models we were using, the three-curring neural networks, were too slow, right? They need to read one word at a time, and Google requires really fast response time.
3:13And you want to read multiple articles. So you cannot approach it as a human. You need to approach it as a machine. And so that's where this idea of like, hey, what if you consume the whole article, the whole text at the same time in parallel using the hardware accelerators we have and figure out the relationships through the number of layers and kind of steps of reasoning instead of trying to read one word time. And so that's what gave birth to Transformers, kind of like coded up a first version. And obviously, you know, it was not random. It was doing something. So obviously it took a lot of work by everyone to make it really work from there.
3:53And then I was excited of using this technology to apply it to actually coding. Because I always thought, hey, why are we doing so much manual work as developers? Can I just tell the machine and it figures it out? And so now we call it vibe coding. Back then, we were just teaching machines to code. This was 2017. Back then, I went and pitched this to VCs. Most of them thought we're delusional. Somewhere between science fiction and delusional, that was the... Yeah, sometimes being early is as bad as being wrong, right? Yes, exactly. So we were very early. We didn't have the capacity to scale the models to the level it needed, even though we kind of were doing, I mean, a lot of similar things.
4:34There's a lot of small details that matter that were done very right by OpenAI. But what we were doing was actually a lot of crowdsourcing. So we're trying to get a lot more training data with people. And we had this challenge, like we had computer science students effectively in Southeast Asia, China, Eastern Europe. And we had trouble paying them because it's actually really hard to pay in a lot of these countries. There's monetary restrictions. Chinese students don't have bank accounts. In Ukraine, you need to sell half of the dollars when they're having a bank account. And so we started looking at crypto as actually just like solving our own practical problem.
5:09How do you like coordinate and pay people around the world? And so this is 2018, nothing that would be like actually, even to our medium-sized use case, thousands of people would actually scale at that point. And so that's kind of where we're like, hey, we should solve this problem. This seems like a big problem to solve. So we've kind of pivoted our original Near AI into Near Protocol and really focused on solving scalable, easy to use, easy to build on blockchain. And so Near now has 50 million monthly active users. It's top one, top two by active users, blockchain, top two, usually top three by number of transactions.
5:46It kind of everything from remittances, payments, loyalty points, financial instruments, kind of a whole variety of use cases, including the payment for the crowdsourcing data labeling. We've had that application running since 2021. So kind of build it out, there's ecosystem, it's lots of different people building different applications, but obviously we're on the back of our mind, kind of always wanted to go back to AI. And so when we've seen all of the improvements, there's GPT-3 and then chat GPT, GPT-4, Or one of the newly found, I would say, through this blockchain journey understanding is that, yes, there's technology, but also there's kind of the governance question of this.
6:28And the other part is, I mean, one of the interesting things that happens as models evolve, right? There's some threshold at which it actually becomes game theoretic for any of the companies if they have a model that is able to hack into other systems, to actually use it to hack into other labs to delete their models. Because if they don't do that, the other labs, when they cross that threshold, will do it to them. And there's a safety claim always like, hey, their models are unsafe, so we need to make sure they don't do something bad. But there's a very interesting state that we can get to pretty quickly.
7:10And it's really hard to determine where it is. There's also just practically, even now when you're asking these models, you have no idea if the response is coming from statistics and data, if the data was biased in some way. Even when I worked at Google, sometimes you just delete some training data because it contains some signal that you don't want to have in your results. And like, in turn, you're also biasing data in a very specific way, right? Like, for example, Obama is born in Kenya was a very prevalent statement back in the day across all of the right wing news. And so like, if your eval set has that question, removing all the right wing news actually improves your evaluation set and your evaluation.
7:53So there's like unclear biases in data. This can be data poisoning. This can be so-called sleeper agents. So there's this concept where you can add into training data, some specific modification that doesn't show up in normal evals. But if there's something in a context, like a date or a specific statement, it actually changes how the model behaves. And so the way to effectively use it is think of your, in your cursor, your vibe coding, and like specifically in one specific case, it will change the import from import transformers to import transformers or misspelled request, which is actually a malicious library and pip, right?
8:31So there's like all of those things that we just really don't know what's going on in these models now. So there's like a governance question, right? Which is, yes, we wanted the model to be safe, but the people who build the model have an unsafe version. We as users have no idea how this is used. And there's a data privacy question, which is to make these models extremely useful, you want to give them as much of your context as possible, right? There's this hardware device that listens to you at all times. But at the same time, if this data now goes anywhere or that company gets hacked and all the data gets leaked, that's a massive invasion of privacy.
9:04So you have all this host of problems. And so the suggestion and this vision we formulated, we call it user-owned AI, where how do we bring the focus back on the user, where instead of trying to build a model that effectively benefits the company, we build a model that the meta function is to optimize towards the user, right? Which means it's private, which means it's value loss function, at least meta function, is toward the user's success and well-being. We know which data went in. So at least you know which biases the model has, or at least anyone can analyze it and have reports, et cetera, on that.
9:39And so that's kind of the conceptual user-owned AI. And to do that, you need all the blockchain methodologies that we have across coordinating people to build data sets and models, have privacy technology that blockchain has been developing, as well as kind of incentive layer and mechanism to really gear it over the users. All right, that is quite the background. I'm actually, if it's okay, I want to go back a little bit and just ask you some questions about different pieces along the way, because you have a pretty unusual and unique story there. Actually, going back to that paper at Google, just really quickly, because when I first started getting into, I'm later come to machine learning than you are.
10:17And when I started getting into this latest round, like attention is all you need. the Transformers kickoff paper was like foundational reading club material. Did you know at the time that you were doing that work and doing it, how big this was going to get? Not really. I think at the time, the pace of innovation was very quick, right? And there was a lot of different architectures and different structures, right? I mean, in a way, if you think of it, Transformers really removing things. Like we removed things from the other models. We haven't added. I mean, obviously, it was like a very powerful architecture because it was so performant, because it was showed that actually you don't need to have this recurring relationships.
10:56You don't need to have even convolutional networks. And you only need this self-attention mechanism to really capture all of these relationships and have a sufficient reasoning capability. And I think the team, I actually was the first one to leave. the team continued experimenting and they saw a lot of promise on images and on other contexts as well so there was definitely promise that this is like a very generic architecture but i don't think it was clear that this is like the gonna be the last evolution like at the time it felt like new architectures are coming every few weeks there was something new it was like neural gpu there was neural computer there's this that so like it wasn't clear that this okay this is it and then And everybody's just like, builds now on top, right?
11:39And figures out how to train it better, et cetera. Yeah. Moving on a little bit. So you pivoted fairly early to crypto. I didn't realize it was quite so early. And I think it's interesting because you're actually using it for one of the core use cases that feels like it has continued to be relevant, right? How do you provide financial services for the unbanked, across borders, all these different things? We had this huge boom in NFTs and all these other different tokens. And being in that space, what parts of that do you, and I know a lot of developers have become very skeptical of this. So what parts of crypto do you think are the enduring value?
12:13And where is it just noise? Yeah, I mean, that's a really deep question. I think, and for context, because it was a very delusional idea in 2017 to build a machine that codes itself, right? And for context, I tried to do that back when for my master's degree in university. That is where completely nothing worked. So it's like a recurring theme for me. And so we give ourselves a year. And it's like after a year, we kind of like, okay, we had some papers, we'd made some progress, but it wasn't near the level we needed to really make it commercial. And blockchain clearly was like, hey, this is a use case that I'm being from Ukraine, very familiar with cross-border payments and kind of money movements and complexity of that.
12:52And so I think I cluster the use cases of blockchain effectively into maybe four categories. So one is global identity. One of the real problems on internet is how to create a global identity. Right now we're using DNS, we're using IP addresses, we're using all these methods, which are actually really bad and have a lot of issues. Like DNS has literally like a group of people who are proving stuff at the top, right? and like from potentially like spending a ton of money on things they shouldn't. So it is a very clear internet problem. How do you create a global registry that is open to everyone and has the same rules?
13:34So blockchain solves that and you can create it for identity. You can create a naming service, et cetera. Second one is payments for sure. How to transfer value between and kind of in any asset, in any value. And it's definitely is we have right now 600 millisecond blocks, 1.2 second finality. So within 1.2 seconds, you actually move value around the world, billion dollars, no problem. And hundreds or even thousands of nodes are confirming that. Finally, you have marketplaces. So one of the really big benefit that you have here is that you can create, because of global registry and payments, if you bring them together, it becomes a marketplace.
14:14It's a global marketplace where you can sell anything, offer anything. And this is why it's used for speculation, because the simplest thing to do on marketplace is speculate on assets that don't have any other value except for what people intrinsically assign to them. But you can think of, for example, if you want to buy 100 tons of steel and you want to get it delivered to you. Right now, you'll need to email a bunch of people, probably call someone, figure out, probably call Flexport to get the shipping going, warehousing, etc. etc. Or you can imagine, and I mean, we'll get to it, but you can effectively say, hey, I want this done on the marketplace.
14:51And then you have now other actors who are like, hey, I will do it for you for this much money. And there's a contract with money, with escrow, everything on chain, guaranteed execution when the factor is delivered. And the value itself is tokenized, right? So like while it's in progress, right, you effectively can borrow against that because it has the escrow money locked in this, which is like trade financing. So there's kind of a lot of financial instruments you can build on this primitive of marketplace. Finally, the last piece is kind of this coordination. And I think this is where I think blockchain has failed.
15:27It had a lot of promise of like, hey, we'll have a new type of organizations that don't have traditional management, which is like, I think everybody agrees. And in any good organization, people try to go away from, I'll tell you what to do, right? It's more, I'll support you in what you're trying to do. But it still kind of creates this hierarchy. And you need this hierarchy because people cannot scale the relationships. And so the idea was like, hey, we can create, in fact, the game theory to coordinate people instead, and use kind of one chain mechanisms to pay and do this. And I think that failed because people are messy.
16:00And there's a lot of like people, things that needs to happen. And this is where I have a whole thesis about actually AI, being in the middle of this coordination actually solves a lot of these problems because... Because it can deal with messiness in a way that traditional code can't. And it can deal with the scale, right? So one of the things, like as a person, imagine you have a thousand reports. I mean, you'll go crazy and you'll be a really bad manager for them. But AI handling a thousand reports is no problem, right? It can give everyone personalized context. It can collect information from everyone.
16:31It can broadcast it in personalized way, et cetera, right? So it actually scales with the organization. So to me, this is like main four use case, kind of core primitives that then everything else on top, like, hey, we want to bring whatever real world assets here is because of the marketplace, right? We want to issue equity as a token because of the marketplace. We want to figure out how to build new type of organizations because of this coordination mechanism. We want to coordinate payments, et cetera. So like all of those pieces reinforce each other, but they are the use cases. And then everything else, like, for example, privacy and other things, they kind of leverage some of this.
17:09If you want to have, so, for example, we use this approach called trusted execution environment. So this is a specialized hardware element that are available on Intel CPUs, AMD, as well as on NVIDIA GPUs and some of our other accelerators. And the idea there is you can use it like Azure provides you this service as well. But you need to trust Azure. Azure tells you like, hey, we're running it in secure hardware. There's so many things right now where we're just like Microsoft, Google, Amazon, we can probably trust them, right? Yeah. So versus if you have this global registry, now the device can register directly and say, hey, here is my certificates from Intel and NVIDIA, and you can verify them on chain.
17:51And now there's an IP address registered. So like when you go to them, you have all of this cryptographic routing and supply chain to verify directly without needing to trust extra cloud provider. And so you can now build a full cloud, which is just from directly providers who self registered who can come in online, which means you can also find a closer, for example, data center and provider for your AI inference to reduce latency, you can distribute the compute more even lean, right? Not have all hundred thousand GPUs all sitting in Memphis and using all electricity. You can actually have privacy because the data is fully inside secure enclave and not visible even to the hardware operator.
18:32And you know what model runs there. You don't need to be, oh, did I run for, oh, three, oh, or like, did they change it yesterday? I have no idea, right? Like you can actually have guarantees around that. So it actually gives a lot of guarantees because we have this blockchain layer for identity coordination and payments, right, because you need to pay these people to use their hardware. This episode of Software Engineering Daily is brought to you by Capital One. How does Capital One stack? It starts with applied research and leveraging data to build AI models. Their engineering teams use the power of the cloud and platform standardization and automation to embed AI solutions throughout the business.
19:11Real-time data at scale enables these proprietary AI solutions to help Capital One improve the financial lives of its customers. That's technology at Capital One. Learn more about how Capital One's modern tech stack, data ecosystem, and application of AI ML are central to the business by visiting capitalone.com slash tech. So I want to dig into that and from a few different angles, but since this is Software Engineering Daily, let's start from the software side. So if I'm a developer wanting to tap into that? What does it end up actually looking like for me? Yeah. So, I mean, it depends on where you are in a stack of what you're trying to do as a developer, right?
19:52So the simplest way we have, for example, just an open AI endpoint for GPU inference that runs inside Secure Enclave, right? So everything you send there, it's TLS encrypted on your side. It's decrypted inside Secure Enclave. Nobody in the middle can actually access it. It runs on the model that you asked and you get back and you have a certificate, again, that you can check and verify that NVIDIA and Intel signed effectively on that. Now, if you want to build an agent, for example, that runs on behalf of a user and even you as developer don't have access to what users is asking for, which is super useful, right?
20:30As you go financial use cases, medical use cases, but also just daily life, right? Imagine you have this Fireflies or this recording of meetings bots. Right now, their servers are getting all of your calls and all your data, which is like, now I need to think about, are they going to get hacked? What did I say? Or if they were using our stack, they could have put the whole system into the secure enclave where effectively now all the information is streamed directly into the server that's encrypted end-to-end, run there. and then only you get back the result. And then developer just uploads their code, right?
21:06So you effectively package Docker container and upload it as we call it an agent into the system. It uses private inference, but the agent itself, your general code runs in the secure enclave mode as well. So we have an agent hub where you can see a bunch of, like we have about a thousand agents who are running or can optionally run in this mode. Now, if you're even lower level developer or you yourself want to build something that includes payments and other systems, that's where we have this idea of agentic protocols, where you can effectively create a smart contract. So a contract within Rust or JavaScript that runs on blockchain that itself can call into this agents and get back the result kind of as verification.
21:52And so the examples we have now are mostly about trading kind of in financial use cases. That's the first thing people do. But again, let's say somebody wants to build a naming service or something else, right? You can also have this kind of things where, again, the logic happens, like maybe your pricing model or your loan evaluation scoring happens in this verifiable way. And then the execution of actions happens as well as a blockchain. So it really depends on kind of on the level of the stack you want to build your applications in. A few different questions about that. So thinking about this model of I'm a developer, I want to build a secure agent or something like this.
22:31I just upload my Docker container. Now, for me, as someone who ships a lot of applications, I immediately start saying, OK, what about observability? How do I know if they run into a bug? How do I debug this thing? Like, what does that end up looking like in this stack? Yeah, so this is where things get interesting, because now you have a tradeoff between privacy and observability kind of on the different sides of the spectrum. So we are actually working on analytics and debugging system that sits underneath, like as you ship your Docker to give you some of the observability where you effectively specify privacy versus observability threshold, which, you know what I mean?
23:08You kind of will inform the user as well where you want to sit. And so obviously you can have full observability, but then you have access to everything that users put. Or you have none, or you can have somewhere in the middle where it actually summarizes stuff for you and maybe gives you the logs of failures and bugs, et cetera, but doesn't give you the exact queries that users sent. So we actually have exact kind of sprint on building out the tooling, including quality control, latencies, times, all of the stats that you actually need as a developer to understand how your agent is working. That makes sense.
23:43Maybe also, can we go in a little bit on the trusted execution environment? And in particular, I'm thinking about things like, okay, I can know that if I'm a user or I'm a developer sending something off to a service, I can know my data is encrypted. I can get back stuff that it was encrypted. How do I know that your software isn't just posting that data somewhere else? Like, is the trusted environment locking down the network or like, how does that all work? Yeah, so there are a few things that are happening. So first of all, when you establish a session, you're effectively getting back the hash of a Docker container that runs there, which is authorized by the hardware.
24:14So the signature you get effectively says, this Docker container runs on this CPU and this GPU, and you can verify that. So if a developer published this Docker container, you can make sure what it is. Now, not everyone wants to open source everything they do. And so this is where, A, indeed, the plan is to have a firewall system where you can indeed lock the access. because you may want it to go and access some APIs and some MCP servers or whatever. The other piece is we're actually working with an external team on agent security. So where you actually have an agent itself who runs in CiteE, who inspects the code of this Docker, of this agent that you as developer uploading.
24:59And so it effectively gives you a security report based on like, hey, it looks like it's sending all the requests it received to some external IP address. Or maybe it parses all the API keys and leaks them. So effectively, we can have scanners that are themselves AI-based, that there's no person who's looked at external developer code, but there's AI that looked at it and certified it in some way. Now, that is cat and mouse, to be clear. But with combination of this message, you can get some reasonable level. And then the longer-term research, we're actually investing in formal verification. So this is a bit more, again, fundamental, as I mentioned.
Read the full transcript
25:41I think there will be a threshold at which the models will start hacking into other systems. The thing is, both people write code with vulnerabilities, and AI now trained on the code with vulnerabilities, writes code with vulnerabilities. There's this image, obviously, with a thin slice, everything is on top. We're kind of layering in more now as AI at a faster speed. And so the fundamental way to solve that is if we have a mathematical proof that the code that runs is exactly satisfies your criteria. So usually right now when formalification is used because it's so expensive, like it's manual work, you only do it once for some set of criteria.
26:22The problem is a set of criteria itself can be wrong. right and so what you want actually is when you're calling the service you want to provide you as a you know developer effectively calling into it when a provide set of things that you want to be guaranteed for example that none of this data is leaving this enclave and only this you know URLs are getting accessed in this way and then the service actually responds back with a verification like certificates around the secure enclave and verification that indeed this criteria is satisfied. And so this is actually what we're working on is really to build this trust level at a mathematical kind of guarantees.
27:07It's also very useful for blockchain where people getting money stolen all the time, where this is like a very fundamental piece where if I'm putting in money, I want the guarantee that money will not... I'll be able to withdraw at least as much money as they put in. And so it's a very short term applicable to blockchain. But long term, we want it applicable to every service in the world, because this is actually how we're going to stop kind of this sprawl of vulnerabilities in all systems. Yeah, that's fascinating. Do you think that's going to kind of limit the set of programming environments that is able to work in this space?
27:45I mean, we're going to have either way, kind of collapse of programming environment as coding models get better. Because the thing is like, I mean, AI really doesn't care. It can write in any language, right? And so it's actually, it's better to write a language that's more written because more training data. Even now that part is getting solved because there's some companies where, you know, they just generate a lot more training data in the programming language of the target. And so you can train in that. So I think it will be really more important to have this kind of strong guarantees of security than having, you know, 50 different programming languages people can write in.
28:23I use this like before we would write code once and read it many times. And so you wanted to make it now we write code once and read it never. You know, this is interesting, right? Because it kind of taps into a few different pieces. One is with LLMs or anything that's sort of kind of probabilistically generated, the ability to validate rises in importance tremendously. And in fact, one of the reasons I think that coding is such a useful environment or something that's so amenable to these models is because we already have to think about validation, right? We've been thinking about how do you do type checks?
28:54How do you do unit tests? How do you do all of these different things for a long time? What do you think are the attributes that need to be there for a programming language to be a good LLM target, right? So like, for example, I've seen LLMs do a much better job at generating strongly typed code, particularly because agents are able to use that as a part of their feedback loop. Whereas if you use a dynamic programming language, even one with a lot of training data in the corpus, JavaScript, it's not as joyful of an experience working with LLM code, let's say that. I mean, it's very practically speaking, right?
29:26It's like a lot of the types, especially in languages like Rust, they become very semantic, right? I mean, at least when I build, I try to make a semantic typing, even if it's the same underlying thing. But for example, for Nier smart contracts, we have an account and balance as the type, even though it's like U128 underneath and a string. But those semantic types allow to effectively, when you look at the functions specification, you can like, hey, this is amount in, amount out. This is from two accounts, right? So it gives you a lot more kind of context as a human. So, I mean, AI is not that different.
30:06AI has, I would say, at this point, lower ability to kind of disambiguate and like map some of the complex structures right i mean this is also just practically speaking the models have a limited amount of like reasoning steps they do right i mean you can run them for longer this is where all this like old style models and our style models come where they literally run okay we need more reasoning let's just like push more tokens through the inference but obviously it has its own limitations so like when you need to map like, okay, there's an argument coming in. I need to look at everywhere else where this function was called to understand what semantic meaning this argument has.
30:44Like it's obviously way harder. And then like memorize that when next time I need to call a function, you know, to really disambiguate this. So yeah, I think strongly typed and then adding this for modification method because this actually adds additional semantic properties, right? Now, so for example, for sorting, it will literally, I mean, what we're designing will give you like, like, hey, actually the return will be such as that every element is larger or equal than the previous element. Now you have like semantic meaning of the whole function without needing to read the implementation and maintain that constantly.
31:21So it gives you like a lot more properties. So I think that is gonna be the more useful environment for AI generated code, because then indeed we don't need to go and read and validate it because it, again, We have like engineering team who are using AI now on a daily basis. And, you know, you kind of cannot catch up anymore. Like if you have like five engineers who are pushing 10 ,000 lines every day of AI generated code, we're actually starting to think how to manage the team, how to structure the organization to the code differently than you would do before. Because before you would usually want to have multiple people who know how the code works to really and review each other's like pull requests, et cetera.
32:06And now I actually think it might be not like it's actually would slow down things and maybe not very useful. Instead, just give everyone their own subsystem to own. And they just need to dock like the, there needs to be a documentation that describes what the system does, which ideally should be enough to regenerate the whole system through LLM. And then there should be just a bunch of tests. this is really interesting and relevant because everybody's trying to figure this out right like okay these tools dramatically accelerate our ability to write code what does that mean for what we do and how we do it and what you're describing is is actually very similar to to what my team ends up doing where we call document oriented development right yeah the core thing you're engineering is this specification or document that can be used to generate the code the code itself is like it's like a binary yeah and then the other interesting thought that i mean we've a little bit experimented but haven't fully implemented yet was if you depend on somebody else's system you actually write tests for their system so usually you like you expect them to write tests and then you just use it but because they may regenerate all the code tomorrow completely you want to declare your dependencies through tests oh that's fascinating so you essentially are writing like, here are the guarantees that I'm depending on from your system.
33:26So that if you regenerate it, it makes sure those continue to be valid. Correct. Yes. Then each system can be like literally owned by one person. And if that person moves on to another, whatever, like if somebody needs to come in, they need to read documentation and they can even regenerate the whole thing if needed. And other subsystems will tell if something is off. Another piece of this that I'm curious if you have thoughts on is how do you indicate to the LLM what sets of context to pull in for any particular subsystem that it might be editing? Is it just that one document or are there links in different ways?
33:59How do you think about that? I mean, ideally that document has as much context about that subsystem as possible, but you may need broader context somewhere. I think the Courser has its rules, which are kind of a useful concept. I think some links and some kind of maybe, again, hierarchy of dependencies is useful as well. But yeah, I haven't seen that like fully worked out yet. But this is definitely an interesting as well. Like, yeah, what is the knowledge graph of the systems as well, right? Especially when we're talking about really big code bases, like hundreds of thousands of lines of code, that becomes the mapping out the concepts, right?
34:36Like LLM needs to do that somehow. And so like you kind of need to feed it enough of information to do that without also overwhelming it context. Like even a million tokens is cool. But if we're talking about hundred thousand lines of code that's way more than million tokens usually so coming back a little bit to this privacy first ai that you're talking about a thing i'd love to get your sense on is kind of around how to bootstrap this right because looking at the industry right now one we have models themselves are extremely expensive to train and two we're in what feels like a worldwide GPU shortage where there's literally not, I was talking to a couple of different folks at AI companies and they're like, yeah, we just get throttled by the providers because they are out of GPUs.
35:23There is not enough GPU for all the inferences that are happening. So in the big corp world, they are all putting massive amounts of capital down to try to build out new data centers and all of these things. If we're looking at a privacy distributed type of system, how do you actually get that built? I'll start with the second part because it actually, it's a solution to this problem. So right now you say, hey, I'm going to, let's say, Entropic Courser and it starts to struggle with me. And the reason why this is happening is not because in the world there's no GPUs available right now. It's because Entropic doesn't have access to GPUs available and they don't want to get their model to be run on some GPUs they don't know who runs, right?
36:11They trust Azure, they trust Amazon, maybe they trust some other provider, but they don't trust like me having a box of like eight GPUs to upload, you know, their whatever 4.0 model. And it is a real challenge, like the model providers, because that is main IP, like it's a very valuable IP. If they give it to somebody else to even, you know, there is actual providers like Fireworks and Together and others, they're serving open source models. They could serve other models as well, but the model providers don't trust them. And so what we actually were solving that problem, because we actually say, hey, we have this secure enclave where if you upload the model, neither the hardware provider, nor the user can have access to it, right?
36:58It's effectively in sealed container, but now you can deploy it anywhere. There's a data center in Philippines, it's underutilized. Cool. Let's ship a model there and serve it from there. There is, you know, somebody has hardware in Tokyo and, you know, there's a bunch of requests coming from there. Cool. Let's make it there. So it's actually solving this exact problem of right now you kind of need to, like everybody's building big data centers for themselves, but then there's also a lot of smaller, like 10 ,000 H100s and H200s data centers built everywhere right now, which are actually underutilized.
37:33If you go to this GPU list and there's SF Compute and a few others, they actually have a lot of inventory, which is not underutilized because nobody wants to go and buy 4 ,000 H200s or whatever for a year. Unless you're a big company, you don't need that much. I just need to run that model that just published yesterday on 10 GPUs. And so that right now is like a highly inefficient market. And so you remember we talked about blockchain being really good for markets? Well, this is where the solution comes in. And privacy is a very important component because of this kind of IP needing to be moved around in an encrypted way.
38:14So this is part of our decentralized configuration machine learning cloud where you can actually encrypt your model. So it's kind of addressed in encrypted format. And then when somebody needs it, It gets decrypted inside Secure Enclave and gets used there. And you can run it across any place in this decentralized kind of compute network. And you get automatic rebalancing and validity from that. Now, how do you bootstrap this is an interesting question. Now, this is also where blockchain has a approach. And the approach is effectively subsidizing initially compute while you're growing the network, right?
38:49So this is how Bitcoin grew, right? It was effectively subsidizing compute before it had any value. People were willing to bet that it will be valuable and started mining it. And then as value grew, it caught up. And so there is an opportunity here to have a very similar model where we effectively subsidizing people coming with compute while we growing the demand. And then, again, open it up for more model providers to actually serve their model. So, and imagine now Entropic is like, hey, rate limiting, or you can use this decentralized compute, which is verifiable. You know, we verified that it's all, the path is correct.
39:26Cool, we're going to upload our model. And now everybody can use, including actually, if you have your own GPUs, you can turn them on into this mode, join the network, or you can just run it on your own workloads. So you have them, you know, sitting under your desk or in your data center. Now you can use it for your own workloads, but you're still paying Entropic for using it. So that's an important part. It's not like, you know, open source, free as a viewer, but it's actually you're paying back the developer for using it, but you cannot get like actual physical access to the model weights. Yeah, that's fascinating.
39:59So in some ways, if I were to sort of replay your argument here, each hosting provider is building for peak usage. and it's inefficient assigning. Essentially, people are saying, who do I trust? Well, if I'm anthropic, maybe I only trust the big three. And that's the only people I'm going to use to host my model. And you're saying, okay, well, there's all of this spare capacity out in the world where the gap is trust and coordination, human coordination, right? Building those contracts or what have you. So if you can automate that layer, suddenly you have a much larger pool that can scale up and down.
40:33Yeah. And it solves, I mean, latency and even electricity problem, right? Because you're kind of distributing the workload. Right now it's effectively like Amazon needs to build a big cluster with a gigawatt, you know, electricity station on it. Or you say, hey, we actually have a lot of smaller data centers with like smaller power consumption around the world. And so we can just distribute across them. That's fascinating. For context, NVIDIA has had run this program where they effectively gave allocation of GPUs to the smaller data centers around the world. I mean, their strategy has been trying to counterweight some of the hyperscalers who have a lot of the GPUs to have a big, small 10, 20k clusters around the world.
41:15But those are underutilized because if you're sitting in Silicon Valley, you would go to Amazon. You wouldn't go and hunt for a data center in Japan or somewhere in Norway. So that in some ways solves the GPU coordination issue, but it doesn't necessarily solve some of the things you brought up before around like sleeper agents and unknown biases if we're distributing, you know, anthropic models and open AI models and things like that. So what about the model bootstrapping process? Yes, yes. So that is harder. It's step by step. First, we need infrastructure where we distribute some models, including potentially, you know, obviously the easiest ones are open source that already exist like DeepSeaks, Quen's, you know, LAMAs, et cetera.
41:59But indeed, even though we call them open source, they're actually not open source. They open parameter models. We have no idea what went into them. And so how do we actually do a truly open source model? Well, we need to train it in this way where we know what inputs went in. But if you also release the weights, then you're not going to make any money. So you actually can train the model inside the secure enclave where the outcome is not known to anyone. The outcome is always encrypted and only usable inside the secure enclave. So now you have a model that's not owned by anyone. It's not owned by any single company.
42:35You can have effectively token holders, community to come together, say, hey, we're going to train this model. Here's a data set. Here's a model training process. Let's collect, let's say, amount of dollars required to do this. We're going to launch it. It's going to train. And now this model is going to be used inside this network as well. And the revenue is going to be coming back to people who put in the work and the money to train it. And the token is affecting now a method to distribute this value back and forth. And so with token, you can now fundraise. You can go and say, hey, we're going to be training an open source model that is going to be encrypted weights, not open weights, encrypted weights, that is actually going to generate revenue.
43:17And now you can invest and get return and maybe reinvest in the next model or cash out. And so now that's still like, I skipped some hard parts, which is getting the right training data, getting the training process. But this is kind of the scaffolding of how to do this is actually create a community on models where the community decides what goes in training data, et cetera. They can inspect, they can decide. And then the training process happens inside this. Now, I'll caveat it that the reason why this hasn't happened yet is because the confidential computing tech is only catching up. So this whole thing has only been really possible for about a year.
43:59And so the broader community, we have this really great partner, Fala, who's been building a lot of this infrastructure for confidential compute. And only the Blackwells actually support the cluster level confidentiality. And so that's not yet available. So this is kind of like we're growing with the compute and hardware actually availability of that. But the idea is to have this kind of system ready as soon as hardware is available. Right now you can only like right now you do inference and fine tuning on the H100s, H200s in this way because you don't get the cluster level, you only get the machine level confidentiality.
44:33It's a really interesting model because it essentially inverts what open weight models do today, right? instead of saying, hey, we have a set of training data, which we might tell you about, but you don't know exactly. And we have a training process as well. Like, here's the software that's going to run. Here's how we're doing, you know, reinforcement learning at the end or tuning or all these different things. And then we publish the weights. You're saying, okay, let's take all that initial stuff, make that open, make that visible, make that public. But the outcome, we're going to hold on to that in an encrypted way so that we can actually recoup some of the investment.
45:06Correct. Interesting. So you mentioned the hardware is just getting there, all of these different pieces. Can you project out, like, what does the timeline in your head look like for how this is going to play out? I mean, we started talking about this about, I would say, eight months ago, right? So in past eight months, the hardware started to catch up. We have built out the initial thing so you can run this inference now. There's some first versions of fine tuning of the inconsidential way as well. So fine tuning is the first version where you can take an open source model like DeepSeq and then fine tune it on private data or in public data.
45:43And then the weights are encrypted now, but still monetizable. So that's kind of the first version of in the step by step process. I think the proliferation of Blackwells will be required for this really to turn the next step, which, you know, given Jensen's projections should be happening anytime now. So I think within next year, we're actually going to start seeing this really working. And then a year and a half, two years is when I think my hypothesis is the open source and this ability to coordinate a lot more people contributing data, contributing research expertise is able to actually outrun the centralized labs.
46:23If done well, if done in the right way. For example, we may need like AI researcher agent that sits that is able to like get everybody's ideas, you know, score them, maybe run some evaluations, etc. Because like you need to allocate compute, it's like on some of those things. So I think like the goal is probably within two years to get to the speed of innovation that happens in this user owned way is faster than what's happening in closed source labs. But again, confidentiality you can use now. So there's benefits from this now that people can already benefit. Again, I think any use case that touches medical, financial, and those highly sensitive like government areas, right, definitely can leverage this now.
47:09And then there's also just a lot of enterprises who are uncomfortable with giving all of their data to a company, to another company, right? So this is useful for them pretty much immediately. Yeah. Well, and as you highlight, if, for example, Anthropic or OpenAI or someone like that wanted to be able to give guarantees and say, hey, we can't see your data. We literally cannot see it. They could also start running things in this way. Yeah. And like maybe for them, it doesn't make sense immediately. But for next level of companies that don't have as big reputation, this actually really makes sense to do right now.
47:46Yeah. Awesome. Well, we're getting close to the end of our time. Is there anything we haven't talked about here that you think would be important to discuss before we wrap up? I mean, I think given this principle, I encourage people to really think through like how people can contribute, right? Because at the end, it's going to be an open source like community initiative, right? There's a financial incentives and model to reward people because I think one of the challenges is open source historically been. How do you support it? Yeah, unless you work for Google or Microsoft, which kind of pays the salary, right?
48:16It's a very thankless job. But I think the opportunity here is actually kind of create something that indeed can move quicker and has the wisdom of the crowd coming together, as well as, you know, can use some of the private data that maybe you don't want to actually touch. But one of the ideas was, again, because you have a verifiable compute, you can have a pipeline where, let's say you take people's private data, but then you have a very specific cleaning process that everybody looked at, audited, and agreed, like remove social security numbers, phone numbers, addresses, names, et cetera, which is not useful for training these models anyway.
48:55And everybody knows that they can contribute data and receive some reward, and it will be cleaned in the right and expected way. And then that data is never seen by anyone anyway, and fed into model at like this pre-training steps, right? So you can have like this new ways of actually even gathering more training data or for research where let's say right now, medical information is not able to be used, but you can run inference on it in a private way. So there's just like so many new use cases and opportunities. So I just encourage everyone to kind of think through where they can really leverage that and reach out and connect with us and the team to leverage what already is available now and then contribute to building this forward.
49:38Great. That seems like a good wrap.
From the publisher
Illia Polosukhin is a veteran AI researcher and one of the original authors of the landmark Transformer paper, Attention is All You Need, which he co-authored during his time at Google Research. He has a deep background in machine learning and natural language processing, and has spent over a decade working at the intersection of
The post Building Open Infrastructure for AI with Illia Polosukhin appeared first on Software Engineering Daily.
