In short
Eye On A.I. Podcast Episode Notes
Episode Title
#243 Greg Osuri: Why the Future of AI Depends on Decentralized Cloud Platforms
Episode Summary In this episode, Craig S. Smith interviews Greg Osuri, founder of Akash Network, about the potential of decentralized cloud computing to disrupt traditional cloud service providers like AWS, Google Cloud, and Microsoft Azure. Osuri presents Akash Network's peer-to-peer marketplace model, which aims to leverage underutilized computing resources while addressing the energy challenges faced by AI training.
Key Themes and Concepts
- Decentralization of Cloud Computing
- The traditional cloud infrastructure is becoming increasingly closed and expensive.
- Decentralization can create a more open and cost-effective cloud model.
- Energy Bottlenecks in AI Training
- AI training is facing energy constraints, necessitating a shift towards decentralized models to distribute workloads effectively.
- Underutilization of Compute Power
- There are millions of underutilized data centers where compute resources can be made available through platforms like Akash.
- Blockchain's Role in Resource Management
- Blockchain technology can provide a secure and transparent method for managing contracts and transactions between compute providers and tenants.
- Privacy and Security Concerns
- The episode discusses the importance of privacy and data security in cloud computing, particularly with the risks associated with hyperscalers.
- Future of AI and Home Computing
- Osuri envisions a future where AI processing happens within home networks, enabling users to maintain control over their data.
Episode Structure
- Introduction & Challenges in AI Training
- Discusses the dual challenges of data scarcity and energy limitations in AI training.
- Greg Osuri’s Background
- Overview of Osuri’s career and his expertise in open-source development.
- Problems with Traditional Cloud Providers
- Analysis of how hyperscalers are becoming unsustainable due to high costs and environmental concerns.
- Using Blockchain for Decentralized Cloud
- Explanation of how blockchain can facilitate decentralized resource management.
- Akash Network’s Marketplace
- The mechanics of how compute buyers and sellers interact through a bidding system.
- Security & Privacy in Cloud Services
- Measures in place to protect user data and ensure compliance with security standards.
- The Energy Crisis
- Explores the unsustainable nature of hyperscalers in terms of energy consumption.
- Decentralized AI and Home AI Solutions
- Vision for a home-based AI infrastructure that maintains user privacy.
- Routing and Optimization of AI Workloads
- Discussion on how workloads are routed within the Akash Network.
- Adoption by Major Companies
- Mention of significant users of Akash Network like NVIDIA.
- Building a Decentralized AI Services Marketplace
- Future steps for Akash to transition from a resource marketplace to offering a broader services economy.
- Why the Future of AI Needs Decentralized Cloud
- Osuri emphasizes the critical need for decentralized infrastructure to meet the growing demand for AI.
Key Takeaways
- Akash Network's Disruption: The platform aims to create a competitive alternative to traditional cloud providers through decentralization and cost-efficiency.
- Peer-to-Peer Resource Sharing: By leveraging unutilized computing power, Akash Network empowers users to bid on resources, fostering a more dynamic and flexible computing environment.
- Focus on Privacy: The episode highlights the need for secure and private computing solutions, particularly as AI becomes more integrated into everyday life.
- Sustainability Challenges: Osuri points out the impending challenges hyperscalers will face due to energy and resource limitations, positioning Akash as a timely solution.
- Community-driven Development: Emphasizes the importance of community contributions and open-source sustainability in the future of cloud services.
Conclusion The episode provides an insightful look into the potential of decentralized cloud computing to address the limitations of traditional models, particularly in the context of AI and the growing need for privacy and sustainability. Akash Network is presented as a pioneer in this space, aiming to reshape how computing resources are accessed and utilized.
Stay Connected
- Craig Smith Twitter: [@craigss](https://twitter.com/craigss)
- Eye on A.I. Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)
For more information about Akash Network and its offerings, visit [Akash Network](https://akash.network).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00There were two big challenges for AI training, right? The one was data, right? We didn't have data, there's the data limit as to how much you can get data to train and second was energy what we saw with deep seek it can use synthetic data so it's very very amazing using synthetic data you can actually solve the data problem but what we cannot solve is the energy problem i think that's why it's very very important if you're doing training to focus on distributing your training runs versus trying to go with the traditional mechanism of centralizing your training runs because we're going to hit a cap in two years and we have no solutions i'm past the point of looking for jobs but I'm not past the point of looking for people to hire.
0:38And when I need to do that, I turn to Indeed. Imagine you just realized your business needed to hire someone yesterday. How can you find amazing candidates fast? Easy. Just use Indeed. When it comes to hiring, Indeed is all you need. You can stop struggling to get your job posts seen on other job sites because Indeed's sponsored jobs help you stand out and hire fast. With sponsored jobs, your post jumps to the top of the page for your relevant candidates so you can reach the people you want faster. And it makes a huge difference. According to Indeed data, sponsored jobs posted directly on Indeed have 45 % more applications than non-sponsored jobs.
1:29Plus, with Indeed's sponsored jobs, there's no monthly subscriptions, no long-term contracts, and you pay only for results. How fast is Indeed? In the minute I've been talking to you, 23 hires were made on Indeed, according to Indeed data worldwide. There's no need to wait any longer. Speed up your hiring right now with Indeed. And listeners of this show will get a$75 sponsored job credit to get your jobs more visibility at Indeed. To get your jobs more visibility, go to indeed.com slash IonAI. IonAI, as always, all run together, E-Y-E-O-N-A-I. That's indeed.com slash IonAI right now. and support our show by saying you heard about Indeed on this podcast, indeed.com slash IonAI for a$75 sponsored job credit.
2:36Good to see you again, Craig. Great to be here. Thank you so much for having me. My name is Greg Osuri. I'm the CEO and founder for Overclock Labs, the core contributor to Akash Network. Background, being a programmer all my life, I've been an open source developer for a little over 15 years. And early days, my career helped contribute to this container native ecosystem. A lot of my software I've written is still used by projects like Kubernetes, Docker, and whatnot. And we, you know, big users of the cloud, and we noticed that cloud, as we know, is increasingly becoming more important in our daily lives, considering most of the workloads now are hosted on the cloud.
3:24And we felt as it gets more prominent, it also is getting more closed. So we felt the need to have a transparent and open cloud, beginning with how the resources are priced to how the resources are secure, distributed, and whatnot. And in line with what we saw with Kubernetes and what we saw with Docker and Linux, really, we wanted a similar cloud. And that's really our work on Akash began. And it began as an open source project. It still is a very, very mature open source project at this point. But, you know, if you go back, the name Akash means in Sanskrit, the sky, the sky is where the clouds are formed.
4:12So that's where the name came from. And we were inspired by this super cloud concept by Cornell in 2015, where they prophesize that you can actually, if you decouple the resource layer, that means the layer where the resources are acquired from the control layer, the control plane, you can effectively create a mechanism where anyone can be a resource provider. and that resource provider doesn't necessarily need to be the controller. So that's the concept of SuperCloud. So really, and we created Akash as a mechanism to bring the control plane in a decentralized mechanism where no one controls a control plane.
5:01It is a community that essentially operates on consensus and a resource plane where anyone with compute can plug in and offer that to a market at incredible prices today. Akash is perhaps the fastest growing cloud and the most cost efficient cloud. Apples to apples for compute resources, you have 10 times the cost savings compared to your absence of the world. whereas for accelerator resources, be GPUs, you have nearly two to three times cost advantage compared to your traditional cloud. So that obviously is contributing to the current incredible growth we're expressing. Yeah. So as I understand it, part of the concept is that there's a lot of unutilized or underutilized compute in the world, data centers scattered around the world that are not being used at full capacity.
6:01And those people can make their resources available to anyone in the world, just as Google server farms are available to anyone in the world. And the blockchain layer, the control layer, consists of, are they smart contracts then that you enter into with the provider? How does the blockchain layer control the resources that are added to the network by people with excess compute? Of course. So to understand lay of the land, you essentially have about 7.2 million data centers in the world, right? And if you break them down, there are about 11 ,000 professional data centers, meaning over one megawatt capacity.
6:59The rest of them are smaller data centers. Could be a closet, an office, or whatnot. That is considered a data center. So these 11 ,000 data centers are owned by enterprises. Out of them, out of 1 ,000 of them are hyperscalers. And their utilization rate, I mean hyperscalers, if you remove them, they tend to have higher utilization because of the businesses. But the 10 ,000 or so enterprise data centers, their utilization rate is somewhere around 15%. So there's enormous underutilization in these 10 ,000 professional 1 megawatt capacity data centers. And in the semi-professional grade data centers, we don't exactly know the data that very well.
7:52In some cases, they tend to have high utilization, but some cases they don't. But the fact remains that is heavy underutilization. Right. Obviously, for an AI data center, you have higher utilization, non-AI, you have lower. And if you're AI, particularly you have, you know, if you're a professional data center in AI, you tend to have a lot of use when you're training and not necessarily when you're not training. Right. And a lot of times you, you know, if you're training, you need the latest the greatest chip because the cost advantage is significantly higher if you have better chips. And now the question is what happens to the chips that you made investments to?
8:37H100s are now being replaced by H200s. Now H100s are perfectly great chips. I mean, NVIDIA made about 2 million of them. They're really good chips. But what happens to those chips, right? So once you're done with training. So there's a lot of underutilization and inefficient way of resource distribution in the landscape. Now what Akash does is it takes this compute, essentially, or providers that have this compute can come and offer Akash and tenants, that means people that use compute, approach Akash with an ask. Essentially they create an order saying that, hey look, I am willing to pay$2 an hour for H200s.
9:19And this is my configuration. This is what I need. They place their order in an order book and providers that can fulfill the order bid on the order. So essentially creating a reverse auction marketplace. So the tenant here sets the price. If a provider can fulfill, great. If they cannot, the tenant obviously doesn't win the bid. So, and once the provider bids and the tenant accepts the provider's bid, a lease gets created. This is what we call a kind of a contract, right, between the provider and the tenant. And there's a requirement to make an escrow payment from a tenant. So, tenant prepays.
10:08That funds are held in an escrow. And those escrow payments are distributed to the provider as provider fulfills the service. So that's how it works. Now, the layer where the tenant creates the order is decentralized. It's open. So all the information besides the private information, there's no privacy. Private information is collected, but private information, when it curtails to the application, is all private. It's not exposed.
10:41whereas the resource information, that means the resources requested, the price, that is public. So you have this rich public data, it's phenomenal to look at how much people are paying for what. And that rich data is used to price the resources by providers and whatnot. So once the provider and tenant gets into an agreement, the blockchain goes away and the relation becomes peer-to-peer. So I am talking to provider directly. There is no intervention. There's nobody in the middle. Blockchain comes into place to enforce a contract and enforce payments in a decentralized manner. So enforce payments in the sense if the escrow account is out of money, the workload gets undeployed, essentially.
11:32That enforcement happens in the control layer. if the provider doesn't fulfill the order the lease gets cancelled the money goes back to tenant that layer is enforced every time somebody needs to pay for the resources that layer is enforced on the blockchain blockchain is just a coordination mechanism that enforces control but not an execution mechanism because blockchain is actually not a good that's right but the enforcement isn't um doesn't happen on the chain either does it well enforcement uh is really pulling the resources back i mean uh pulling money back right if the provider falls probably doesn't get paid if tenant falls the tenant you know the money goes to the provider so that's sort of like mediation happens on the blockchain uh well ultimately it's the providers uh you know uh i mean that'll comply as long as they agree with the protocol.
12:34Yeah, but when you said once a match is made, then it's peer-to-peer, that's off outside of the blockchain, right? So what happens between the provider and the tenant, it's up to the provider and the tenant, right? Now, provider can choose to give the tenant more resources if they want to, but the question is, why would they? No, that's right. Just on the enforcement, I don't quite understand. If you have compute resources, I have a workload, I find you through the Akash network or through the blockchain, uh or then and and then you and i uh i'm sending my workload off uh chain and privately peer-to-peer uh and something how does uh the the chain know whether or not something's gone wrong i i go back on the chain and report it?
13:47Or how does that happen? Yeah, the first obvious thing is you didn't get the resources that you need. There's a verification mechanism that chain enforces on the provider. So there's checking, there's all kinds of things that, hey, did you get the resources you said you're going to need? The provider did not, well, the contract yanks, you get your money back. you know so it's obviously once you get the resources you don't usually have a problem and then you have problems later yada yada yada if your workload gets dropped by the provider it gets yanked you get the money back right so there are certain aspects of like and quote unquote enforcement because i use this word loosely because you know no one can control you that's all point here.
14:34But if you do not comply to the contractual obligations, your contract gets yanked. Now, is there indemnity in the sense like, well, what happens if some nefarious activity happens? Yes, there's absolutely indemnity in certain category providers that are audited, right? So not all providers are, it can be anonymous by default as a provider, but people People are reluctant to deploy on someone without knowing who they're deploying to. I mean, it could be anyone. So there is a degree of identity that's exposed from a provider standpoint to the tenant in case you need indemnity. There's also something called TEE or Trusted Execution Environment, where if you need to run your entire workload in a fully private mechanism, fully encrypted mechanism, all the way from the chip level, even chip to chip communications to be encrypted, you can use TEE, which is Trusted Execution Environments.
15:39Now with TEE, of course, the challenge is there's a little overhead in terms of performance because of encryption, but it gives you security guarantees. If your application requires that level of security guarantees, you can take advantage of the TEA so that nefarious activity wouldn't occur. You were talking earlier about the data centers around the world that are independent, right, that are not part of the hyperscalers. how does that can you give is is do the hyperscale scalers account for 80 percent of compute and these independent data centers account for 20 percent compute in the world or is it the other way around or or yeah can you give us a sense of scale of what not what's on the Akash network today, but what your sort of addressable market is.
16:42I mean, how that compares to the Googles and AWSs and Azures of the world. Yeah, so we measure... So compute is a... There's no direct way to measure compute because it's very... There are too many dimensions to measure. But if you take down the amount of energy being used, you can get a fairly good idea. So, you know, there are about 1 ,030 megawatt data centers. Okay. I mean, half of which are located in the US. But so 30 times 1 ,000 is 30 ,000. That's it. 30 ,000 megawatts, right? and there are about 10 ,000 1 megawatt plus data centers right so 1 megawatt between 1 megawatt and 30 ,000 30 ,000 megawatts so about 10 ,000 there so if you from a from a high like 1 megawatt about you know situation I think about 60 % if I remember correctly 60 % are hyperscalers 40 % are non-hyperscalers And I think that number is growing.
17:58Hyperscalers are growing. Hyperscalers are limited by two main things. One is energy. And second is water, fresh water. So you need to have good energy and clean water in order to build a hyperscale data center. And that takes a very long time to acquire. In the United States, we have our entire grid capacity is around 1.2 terawatts. So 1200 gigawatt capacity. And we are not good at, you know, building more energy because of regulation. And it's quite a lot. If you look at the interconnect requests, that means the connect requests that want to connect to the grid, the new supply. There are about 1.9 gigawatt capacity that'll come on board in the next 14 years.
18:56But if you look at the timing, they're very, very, very slow to connect to the grid. There are all kinds of problems. Why that is? I mean, all the way from aging grid infrastructure to regulation to if you want the most efficient power, which is a nuclear reactor that can produce up to a gigawatt. I mean, 800 megawatts, 1.2 gigawatt capacity, like the old school reactors. The last one we built took about 14 years in the US and about$32 billion. So it's very, very expensive to build. And also with the new regulations, the Gen 3 reactors, the cost per kilowatt hour is about 15 cents, I believe, 10 to 15 cents per kilowatt hour, which is prohibitively more expensive than natural gas.
19:43So there are a lot of, it's very, very complex and very challenging environment to bring more data. center. So there's a big question like how can hyperscalers survive, you know, sustain building more data centers. In fact, look back, Nvidia tried to build a hyperscale data center and Nvidia owns all the chips. Like the question you may ask is like, why wouldn't Nvidia go just build a data center? They couldn't get more than 20 megawatt capacity. I know this for sure because I know the guy who tried to build it in Nvidia. So it's very, very hard to get capacity. And the last one nasty nuclear reactor we had was the three mile ion, three mile in Pennsylvania that was swooped up by Microsoft.
20:26So supposedly, you know, the Stargate project is supposed to bring more nuclear, but let's see where that happens. So your better bet, I think, is can we actually build non-hyperscale, like sub 10 megawatt? Sub 10 megawatt, you can get away with over highly dense solar powers, when you can do renewables for 10 megawatts. And that tends to be a bigger trend now. So I'm extremely skeptical if we are able to move as fast as we want in the hyperscaler market. But I think we have a bigger chance to move in the second tier market, which is the one megawatt to 10 megawatt. And I think we have an enormous opportunity now to move in the 100 kilowatt to one megawatt capacity data centers too.
21:21These are small modular data centers. These can be fit in a container. Microsoft is experimenting with them. All kinds of innovative things you can do in terms of cooling, in terms of placement and whatnot. And the advantage you get with a smaller 100 kilowatt to one megawatt data center is tapping into distributed grid. the solar energy everywhere, right? So we actually wrote a paper. I mean, Akash, I'm going to share that with you. I mean, we did this analysis on like, okay, if AI is going to be so important in our lives, if I want everything in my house to be an agent-driven, so this conversation right here should be, you know, recorded by an agent, should be processed, insights.
22:07I have a year old at home. I want to know everything she's doing at home while I'm working. Hopefully have an agent that will warn me if she's getting into things she shouldn't be getting into. I want the whole house to be automated, every conversation to be recorded. But I would hate that conversation to be stored in a cloud because I do not trust anything that leaves my home network as no one should. Okay, now can I have a sovereign AI in the home? I think most people would want an AI in the home as long as it guarantees privacy. right? And I'm building that, but okay, what do I need to get the latest and the greatest AI in the home?
22:45Well, DeepSeq, 365B can run very well on H200 clusters. It needs about eight H200 clusters. And that costs about half a million dollars, the cluster, right? Now, we did some study. We said, we, you know, feasibility study, like, okay, is there any way you can have sovereign AI in a CIMAP professional data center that takes about 30 kilowatts of energy in the home that is cost-efficient? And the answer was yes, absolutely. If we can acquire about five of these H200 8x8 chip clusters, it's called HGX clusters essentially, about 40 chips, over earning about, you know, first state about$2.3 per hour, 80 % utilization with 20 % overhead for cooling, your CapEx slash OpEx plus OpEx can be recovered within five years by placing them on cash.
23:52What you're doing is you're literally allocating one HGX cluster, HGX has eight chips to your home use. So that's completely dedicated, 100 % utilized to your home use. The rest, four of the HGX clusters, you're offering on the market for 80 % utilization since you already have a cooling and energy infrastructure. Now, we can further reduce the cost by having solar panels. Now, the big challenge of solar is obviously storage, right? So solar only comes for eight hours a day or max six hours, if you're lucky, actually. But storing, like, say you have 100 kilowatt capacity or 30 kilowatt capacity, storing that much capacity for over 10 hours requires a 300 kilowatt battery, which is prohibitively expensive.
24:37It'll cost you about$200 ,000. But if we can sell it back to the grid, like okay overproduce the energy, use the energy to sell it back to the grid, in Austin you get about 4 cents a kilowatt hour, 3 to four cents depending on your region. That way you can pay for the compute, you know, by tapping back into the grid, right? So grid will sell you back at 10 cents. I know it's a scam, but grid will sell you 10 cents, you sell at four cents. Still, but still prohibitively, you know, not prohibitively, very, very feasible. So we did all this study. That's why tapping in to a home network, a semi-professional home network.
25:19And this will take 142 URAC, which is, you can put it in the closet. When AI becomes very important, if AI has to really enter the home, which I think it will, I think that's when we're going to see an explosion of these data centers. It could be an office. It could be anyone that want to get a rig. You know, it could be someone, university, for example, they want to get a rig. As long as you have solar, you can completely actually, I mean, very efficiently run. So there are about 7.2 million of these data centers. We had about 8.6 million peaked in 2017. Obviously, cloud came and killed a lot of these data centers, but I think there's going to be a resurgence of the data centers for privacy, for ownership, for cost reasons.
26:03And that's one of our goals, is to have a decentralized AI. And we're going to achieve that by decentralizing the energy production and energy consumption as well. But for the time being, they're not in the home. Primarily, they're small data centers scattered around the world. As they sign up, if I have a workload, do you have an orchestration layer that routes my workload to the right data center? Or do I have to do that on my end? Yeah, so it wouldn't do the data center selection for you by design because the application, the blockchain is not application aware for privacy reasons. Remember, a blockchain is a public chain.
26:56And in order for it to route, it has to be very application specific because every application has a different way of scaling. And but so it has a bidding engine, doesn't have a matching engine by default. And, you know, application specific information is one reason. Also, a lot of times you get back a bid with better resources that you may need, better price performance that you may need, because a lot of times people don't actually know what they need. You know, that's why you go to a store and be like, oh, I thought I need this, but I need something else. So a lot of times, you know, people can, sometimes you can compromise too.
27:37I mean, you don't exactly get what you want, but there is something that's equally good. You know, you should be able to purchase that, right? Like, so that's why it's very, very hard to develop a matching engine for humans. And so Akash doesn't have automatic routing. Now you can lease out multiple data centers on Akash and have your own routing mechanism, bring your own router, we call it. but that again is very application specific because you understand the application way better than the infrastructure provider. Akash gives you the sovereignty and control for you to be able to have flexibility to wrap your application everywhere.
28:14So yeah. So you put a bid on the chain. Is that right? The bid goes on the chain. It's seen by all the providers attached to the network. They bid on it. or I'm sorry, you put a RFP or whatever on the chain, they bid on it, and then you select whichever one you want or you can split your workload between multiple providers, I would guess. What happens if, I mean, well, two questions. First, in the beginning, you're going to get a pretty reasonable number of bids. But as this network grows, the day could come when you get more bids and you can reasonably go through. Is there going to be an application outside of the chain that can sort through these?
29:18or is that something that you guys are going to offer so that people can find the optimal partner or compute provider? Yes, bid selectors we call them. It is on the roadmap. So you can select a bid selection strategy that's optimal for you. And these are pre-created strategies optionally, right? So it can be like, hey, cost first, optimal for cost, latency first, optimization latency or combination of cost and latency, cost first and latency later. You can have all kinds of rules, a rules engine based sort of like, in fact, we're looking a little beyond rules engine, we're looking at an agent based mechanism where agent can adapt to different situations, considering agent knows your application, you can read your code, understand your code, how it works, and make the appropriate choice for you instead of a rules based system where could that could be fragile.
30:17Yeah, so yeah, we definitely have and we in fact we have applications built on top of Akash. Different providers like Crime Intellect, NVIDIA actually use Akash. They have their own, you know, product called Brev. There's a product called Venice. They're all different products that actually have their own selection engines because they auto scale. The idea of Akash was it gives you primitives. It gives you it's like Lego blocks, right? It gives you all these cool things. and as a programmer, you can decide what to do with those like a box. So more often than not, people don't... There are a lot of people who use Akash directly, but most of our usage comes through the distributors, meaning NVIDIAs, Prime Intellects, and vendors of the world, where they offer, and BitMine, a whole lot of things out there, they offer their own flavor of Akash on what they think is the best winning strategy for the bits.
31:11Absolutely, there should be a bit selection engine. will Akash provide it? Probably will give you options that will make your job easier. Assuming that you're a power user and you understand how these bids work, but we want to be as explicit as possible in terms of granularity that you have full control over who you get. And I think that is a very important capability. And you mentioned price and latency. When the bids come in, do they come with some metric that measures latency for where you are, where your workload is? They give you the region where the workload will be placed. Everything you would need to make the selection right.
32:02And so you can write a latency checker. essentially that will check the latency depending on where your users are. So typically what you would need is you would analyze your user traffic and you would select a mechanism where it'll select the data center that's in 95th percentile of your user activity can be achieved within 50 milliseconds of network latency, right? So that's really what you want to do. But those controls need to be written and that need to be specified based on your application. Some applications, latency is very complex thing because if you have hyperscale applications like Facebook, you most definitely have different charts on different regions.
32:44So California users come to California, New York users come to New York, and more importantly, the data should be available in these local regions. And what happens when you have a New York user trying to access California or when they travel, when a New York user goes to California. Latency-based selection is not as easy one would think. It's easy when application symbol but as it gets more complex it gets very very very challenging right like so uh yeah so you can't just have a latency engine because you know you may not actually give the right information based on the application yeah and and then the other thing you mentioned uh that you you may not want uh you know to put certain information with a hyperscaler because you're concerned about privacy.
33:39How do you know how secure these compute resources are on the Akash chain? I mean, you may be accessing a small data center in Hong Kong. You're not there. You don't have people looking at their security protocols. How do you ensure that you're secure? So the two main mechanisms Akash provides. One is a decentralized auditory mechanism, where if the provider says, hey, they're compliant, they're HIPAA compliant, they are whatever standard compliance they have, they have their tier four data center, They are, you know, the tier four in terms of redundancy and whatnot. These auditors go and verify that off-chain.
Read the full transcript
34:36They literally run applications. They run audit checks. They have physical proof. They have paper proof. They actually go verify. And they post their results on-chain saying that, hey, I have verified that this provider says who they are, their identity. More importantly, because if something goes wrong, you want to go after them kind of thing, right? All the information is posted on-chain. I mean, in fact, not the verified information by multiple auditors, right? So as a tenant, you can choose the auditors you want to select. And these auditors are public figures like Overclock Labs. You trade a cash with one of the auditors, for example.
35:08Second mechanism is we provide something called TEE or Trusted Execution Environment. What TEE does is it gives you a non-custodial way to encrypt runtime, not just transport. So your transport is obviously encrypted, right? That means communication between you and the provider is encrypted. But what's not encrypted is what's in the memory because it's hard to encrypt. So anyone with physical access to the machine can theoretically, if they're talented enough, can theoretically look inside what's in the machine, what's running. To avoid that, you want to encrypt the memory itself. And that's achieved through trusted execution environment.
35:48Now, the trade-off obviously there is you need more resources because it's encryption, right? It needs more compute means more GPUs. So if you're doing AI, NVIDIA has this amazing TE mechanism in that newer chips, H200s and H400s, where they encrypt memory within the chip itself, and they also encrypt the transport between the chips using NVLink. NVLink is 3.2 gbps. It's very, very fast, about a second. So it may slow down a little bit. We see right now, we're seeing about 10 % overhead, which is not bad trade-off if you want privacy, right? Especially if you're an unknown data center. But if you're deploying on something called Equinix, which is in the U.S., which is a professionally data center, you don't have that problem.
36:35So you're more relaxed when it comes to like, hey, okay, we can trust the data center or whatnot. But ultimately, yeah, you have to trust somehow, right? So either through encryption trusting all through auditability trusting. But this way, you actually have a lot more tools at your disposal that you normally don't get when you go to a data center directly and hide with them, right? So you have all these, like, infrastructures that Kosh gives you that you don't normally get going directly. So going to hyperscale is just trust the brand. You can say, okay, Amazon will not do something. You're trusting Amazon will not do something.
37:15I mean, there's a good amount of truth to it, but we know how this trust could be corrupted. We just don't know what's happening. If you're a big company trying to use Amazon, they'll let you open the boxes. They'll literally show you. Like if you're a department of defense, they literally let you audit everything, right? But if you're a, you know, Josh Mo, no one cares. They just, you know, you wouldn't even know if someone has access to your data. It's very famous at ChatGPT because one of the engineers on the interview, they said, yeah, we look at users' prompts. So I felt violated because like, okay, my prompts are good Lord, right?
37:56You know, all kinds of things. And I'll be judged by the person that is looking at my prompts. And that feels, you know, violation of privacy, right? And if you have anything on the cloud, there's no guarantee that these people are going to look at your thing. But if they adhere to certain the standards and they don't comply and you have some degree of indemnification. Right. So that's why I feel like if you want ultimate privacy, you have to run the home. No, there's really no way. Second best thing is TEE. If you can't run the home, well, the lower head and third best thing is probably like ultimate trust in the provider.
38:34But Akash gives you that TEE, which is much better than a cloud provider. Yeah. And then the other question is, if I'm running a workload and my data center goes down, or maybe there's an outage and the backup doesn't happen, the generators don't kick in or whatever, what happens to my workload? It dies. and so you have to know how to be redundant for your application. And that's one of the big areas. You had the same problem with the cloud today too. And cloud work, Amazon has about 200 outages last year, right? And we all saw what happened with CrowdStrike. Somebody pushed bad code and the entire US airline infrastructure came to a standstill because it couldn't fly.
39:29So, you know, centralized systems generally have a lot of force that they don't expose. But so one of the key design patterns to use Akash effectively is to be redundant from the get go. Akash makes it cheaper and makes it more optimal to be redundant, but there's really no silver bullet. And it just happens all the time, right? So we see, like, that's why you want to see when someone says they're at tier three or tier four data center, that means they have a dual ISP, ideally with a Starlink backup. I'm happy to share you the AAP we wrote recently, like, okay, how do you design a home data center with dual ISPs, with a Starlink backup, with dual generators.
40:11You need a gas-made diesel generator with automatic switch off, as well as a gas generator, and ideally solar panels. So we have all kind of USBs, uninterrupted power supply, UPSs and whatnot. So all these aspects are verified. So it reduces the chance of false, but false happened. You cannot prevent false. but what you can do is come back up faster. So no matter, like Mike Tyson says, no matter what, how planned you are, your plans are only as good as you get punched in the face. So you will get punched in the face when you're running infrastructure and that's the reality of infrastructure and you talk to anybody that tells you infrastructure or runs infrastructure will tell you there will be false no matter how hard you think you build your data center is.
40:57But it ultimately comes down to redundancy. You can run in a quorum, for example. Like if you have a database, Well, don't run a single database, run a master, slave, you're not supposed to use the words, like a leader follower style database where if your follower goes down, the leader will self-elect a different leader or spin up a new follower database that'll always be alive. Right. So you need to be a little bit more professional to use Akash. Now, natively, Now, there are applications on top of Akash that makes all these things easy. You know, you don't have to think about redundancy, like you just click a button, it's taken care for you.
41:36Right. But that's very application specific. Akash doesn't provide by design those redundancy mechanisms because, again, it depends on the application. Right. How you like a lot of times when you're scaling databases, you need to think about consistency. What level of consistency you want, if you want immediate consistency, you know, So when you have multiple nodes in a database, right, and you want to make sure all the nodes are presented giving the same data. Well, if you have multiple nodes, if you want to be redundant, you have to have them in different locations around the world. Well, you're not going to present the same information because there's latency.
42:14There's an east to west, there are millisecond latency, right? So if you want an immediate consistent database, you got to make sure that all the data sets are replicated is about 300 milliseconds. So when you write, you can't read immediately. You got to wait 300 milliseconds. That's called immediate consistency. Well, if you don't care about immediate consistency, I mean, eventual consistency is fine with you. That means you write a data entry in California. That entry takes about, you know, you're assuming the reader is reading back from California and not reading from New York. And you're fine with that.
42:50So if you're building social media, you're fine with that. But if you're building something like a banking application where someone withdraws cash from an ATM in California, should not be able to withdraw within 300 milliseconds in New York. You need immediately consistency. That's why scaling and redundancy planning is very application-specific. And so that's why Akash doesn't make the choice for you, and you have to design your redundancy on how you want to see it. So this is new, right? I mean, how long has the network been operating? About four years now. Four years. Four years. Yeah. And as you said, the demand for computers presumably going to rise sharply as AI spreads through the economy.
43:49uh do you have a target for uh how many uh how much compute you'll have on the network by x date or or any sort of roadmap that way yeah so we're growing at 10x now every year which has been phenomenal we hope to continue to grow 10x you know um you know so uh right now we We have about, you know, we want to get to about 10 ,000 GPUs by end of the year, which we think we can. Now about 100 ,000 GPUs by end of next year. About a million GPUs within the next three years. And we're doing fairly well. We're doing really well. We're growing really, really well. We're growing at 20 % thing month over month now.
44:38And more importantly, it's not how much compute you want to get, but it's about how much compute we have on the network gets used. So utilization rate is extremely important. So we scale responsibly because every time there's a new compute, a new provider that comes to a cache, they should be able to sell out the inventory. If they're not selling inventory, there's no point in coming to a cache. So very, very important for us to make sure there is a good equilibrium between the provider's capability to sell a compute and the tenant's capability to scale their workloads. Again, if a provider sells everything, I mean, there's no inventory, tenants will be upset.
45:23So that's why you got to be very, very careful in how you scale the systems. Our utilization rate right now is 70%, very healthy, very, very good. Now, that utilization rate can go up as we scale because there'll be more resources generally in the pool. But so our constraint is utilization rate and we don't scale without that. So we're doing very, very well now. And, you know, I think our next big goal is transitioning from a resource market to a services market, a resource economy to a service economy. So right now, think of Akash as just a marketplace for resources, like coming to a commodities exchange where you can trade gold and whatnot.
46:10but it is not a market for services. So a lot of times when you're building AI systems or any system, as a matter of fact, you use several resources like databases, vector databases, to inference services, to agent hosting platforms, to a whole lot of services that one would orchestrate together to build. And a lot of times when you look at the cloud providers, the resources are open source systems. Look at Redis, for example, Elastic Cache on Amazon is Redis. Or the database services Amazon offers is MySQL or Postgres. These are open source protocols and Amazon really takes them and white labels them and sells them.
46:55We feel like if we empower open source developers to be able to provide their services on Cache, we can effectively have a sustainable open source ecosystem. Right now, a big challenge for open source software is sustainability. Docker, for example. Docker is one of the most widely used open source container infrastructure system. But they couldn't figure out a business model after raising it a billion dollar valuation. They sold a penis to Nutanix. They're no longer an open source company. Kubernetes, 80 % of the globe uses Kubernetes. They can't keep their contributors together anymore. Right.
47:35Look at every I mean, besides Linux, which has a very weird founder and a very special founder. Besides Linux, we don't have any. Example in the wild, while that remain purely open source, right, successfully. And so we have to change that. We have to create an economy for open source contributors to be able to sustain themselves to build open source software. And I think that is what Akash is going to transition to the service economy. And that's coming next year. I think with that and also like our we also measure revenue per GPU. So revenue per GPU is right now at$20 per GPU per day, which is fairly good.
48:20We were$10 a year ago and we did a good job in increasing the revenue per GPU. But with additional services, we can further improve our revenue per GPU to about$50 to$100. dollars. That comes with premium services on top of it. And that revenue goes to open source developers. So we not only look for just capacity per se, pure capacity, but we also look at how do we increase our value per resource that we provide, as well as how do we scale while maintaining good utilization. That's very, very important for the health of the network. The more healthier the network is, the higher chances of success we have in the future.
48:56so um yeah well and so how does someone uh i mean when once you have the services up presumably uh you'll have uh people that that can onboard a tenant as you call it or a user and walk them through the system and get them connected and do all of that stuff. But for now, how does somebody use the network? So you could go to Akasha Network and there's a button that says deploy. There's an application called console. Console Akasha Network is a great way to get started. You know, there's$10 like free trial if you want to go try it. It's actually very good to use. Highly encourage folks to try it.
49:49No signups is very, very straightforward. People love it. If you want a more easier system in terms of getting SSH style access, you can go to Prime and Collect. They can use Akash underneath, deploy a resource, you'll see the power. So there are several ways, but I think the best way is to go directly because you get the most cost advantage. Right now, I don't know if you have any H100, H200s left on the network because they go like hotcakes, but H200s are$1.99 an hour, which is the lowest compared to Amazon. I believe it's Amazon's$4 or$6 per hour, but like three times cheaper. And the reason why it's cheaper is because I'll send you the link where we did a cost analysis.
50:40You can actually amortize your investment And while with, you know, with the decent utilization rate fairly easily. So there's no reason why you should be paying Amazon margins. I mean, Amazon makes record level margins from you, right? So you remove all the margins, the resources are actually fairly cheap. That's really what Akash provides. So that's why resources are very, very cheap on Akash. So, but they go out like H200s are gone. I would have been, our utilization rate was like 98 % last we saw by H200. That's a problem with like hot chips. You don't really have availability and we're trying to improve that as much as possible.
51:20But H100s are available for, I'm sorry, H100s are available for $1.20, which is also very, very competitive compared to Amazon. So I highly encourage folks to check it out. If you want to see the pricing, you can go to akash.network.gpus and you'll see the GPU pricing and you'll find them extremely competitive. You also get to know how much availability there is on the network so you can plan. I mean, I highly warn you, it's very addictive, especially when you automate. So a lot of our users just automate the hell out of it because the moment they see a GPU, I guess we so stripped away, especially with H200s.
52:01DeepSeek runs really well on H200s. So it's very addictive because people are like constantly bidding, constantly bidding, getting to this like really cool game. I think there's also somebody that's building features where they actually purchase the compute and they resell the compute at a higher rate. So it's a very, it's a very fascinating market, actually. Lastly, you're not the only decentralized cloud computing platform out there. How do you regard yourself in the, among the competition? So we're definitely the leaders. We're the first one. And we're the leaders. So there are a lot of copycats.
52:39I mean, success gets copycats, and that's the reality, right? We are not the only one, but we are the first and the biggest open source cloud. So new copycats are actually closed source. They take our code and fork our code. We've seen that several times, right? Now, not to disregard the competition, but I think it's a good thing that people are looking at our model and seeing the success behind our model and trying to have their own flavor. And they obviously want to stay closed source because, well, you know, it's advantageous to be closed source because you can ship code faster. So Akash is very decentralized.
53:17I think from a pricing standpoint, we're also much better. From pricing, we have much better SDKs, we have much better user experience, we are much better in several aspects. And also you have a lot of, we have about 30 ,000 developers in our discords. Anytime you have a problem, you just go to the discord and this helps solve your problems really quickly. So enormous community and we have enormous participation. We have about 500 contributors that come and build a cache. So if you see a feature that you want to build, you can write a proposal yourself. So if you're an open source developer, if you like building open source, Akash is a platform for you because it not only lets you use an open source system, but also lets you build and pace you.
54:05Like if you have an idea that you think you can run on Akash, you have, I believe,$25 million in the community pool that you can apply for a grant and get and use Akash. So a lot of university students are using Akash network too. So it's not a company you're dealing with. It's a community you're going to be part of. So that's a big difference between using Akash and a company with products. It's not a corporation, it is a community. So if you are that type of person that believes in a, that communities can actually solve societal problems, Akash is for you. But if you are one of those people that want to pay a company, that can leverage decentralized stack as well.
54:50there are options and, of course, companies out there that are offering something similar to cash out there that you can take advantage of. Is there anything that I didn't cover that you'd want listeners to know? No, I think I want to emphasize why this model is going to be extremely important. And I think why people doing AI, especially training, should be considering some of these models. Like our biggest challenge for AI, there were two big challenges for AI training, right? The one was data, right? We didn't have data. There's a data limit as to how much you can get data to train. And second was energy.
55:38We solved the data problem. What we saw with DeepSeq, it can use synthetic data. So it's very, very amazing using synthetic data and using mixture of experts mechanism. You can actually solve the data problem. But what we cannot solve is the energy problem. I think that's why it's very, very important if you're doing training to focus on distributing your training runs versus trying to go with the traditional mechanism of centralizing your training runs. And because we're going to hit a cap in two years and we have no solutions. That's why we have$500 billion supposed investments and a lot of the investment is going towards power infrastructure.
56:17But by the time we realize the benefits of the investment is going to be late. So I highly encourage folks to take decentralized distributed training more seriously and look into mechanisms employed by Noose Research. Noose is a top team. They came up with something called Distro recently. Google DeepMind came up with a paper called Dialoco. There are several companies that are having their own approaches and some of them reduce the... One approach is reducing the amount of communication between the nodes to train better. Some of them have better verification mechanisms employed locally to train better.
56:57And I love to see more work, more different approaches and more experimentation in the space and that's really I think going to benefit all of us you know and really take the power away from the opening eyes of the world if you want to really disrupt them you have to think about this information I'm past the point of looking for jobs but I'm not past the point of looking for people to hire and when I need to do that I turn to indeed imagine you just realized your business needed to hire someone yesterday. How can you find amazing candidates fast? Easy. Just use Indeed. When it comes to hiring, Indeed is all you need.
57:40You can stop struggling to get your job posts seen on other job sites because Indeed's sponsored jobs help you stand out and hire fast. With sponsored jobs, your post jumps to the top of the page for your relevant candidates so you can reach the people you want faster. And it makes a huge difference. According to Indeed data, sponsored jobs posted directly on Indeed have 45 % more applications than non-sponsored jobs. Plus, with Indeed sponsored jobs, there's no monthly subscriptions, no long-term contracts, and you pay only for results. How fast is Indeed? In the minute I've been talking to you, 23 hires were made on Indeed, according to Indeed data worldwide.
58:31There's no need to wait any longer. Speed up your hiring right now with Indeed, and listeners of this show will get a$75 sponsored job credit. To get your jobs more visibility at Indeed, to get your jobs more visibility, go to indeed.com slash IonAI. IonAI, as always, all run together, E-Y-E-O-N-A-I. That's indeed.com slash IonAI right now. And support our show by saying you heard about Indeed on this podcast, indeed.com slash IonAI for a$75 sponsored job credit.
From the publisher
This episode is sponsored by Indeed.
Stop struggling to get your job post seen on other job sites. Indeed's Sponsored Jobs help you stand out and hire fast. With Sponsored Jobs your post jumps to the top of the page for your relevant candidates, so you can reach the people you want faster.
Get a $75 Sponsored Job Credit to boost your job’s visibility! Claim your offer now: https://www.indeed.com/EYEONAI
Greg Osuri’s Vision for Decentralized Cloud Computing | The Future of AI & Web3 Infrastructure
The cloud is broken—can decentralization fix it? In this episode, Greg Osuri, founder of Akash Network, shares his groundbreaking approach to decentralized cloud computing and how it's disrupting hyperscalers like AWS, Google Cloud, and Microsoft Azure.
Discover how Akash Network’s peer-to-peer marketplace is slashing cloud costs, unlocking unused compute power, and paving the way for AI-driven infrastructure without Big Tech’s control.
What You'll Learn in This Episode:
- Why AI training is hitting an energy bottleneck and how decentralization solves it
- How Akash Network creates a global marketplace for underutilized compute power
- The role of blockchain in securing cloud resources and enforcing smart contracts
- The privacy risks of hyperscalers—and why sovereign AI in the home is the future
- How Akash Network is evolving from a resource marketplace to a full-fledged services economy
- The future of AI, energy-efficient cloud solutions, and decentralized infrastructure
The battle for the future of cloud computing is on—and decentralization is winning. If you're interested in AI, blockchain, Web3, or the economics of cloud infrastructure, this episode is a must-watch!
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Introduction & The Biggest Challenges in AI Training
(02:36) Greg Osuri’s Background
(04:50) The Problem with AWS, Google Cloud & Traditional Cloud Providers
(06:40) How To Use Blockchain for a Decentralized Cloud
(10:17) Akash Network’s Marketplace Matches Compute Buyers & Sellers
(14:42) Security & Privacy: Protecting Users from Data Risks
(18:25) The Energy Crisis: Why Hyperscalers Are Unsustainable
(21:51) The Future of AI: Decentralized Cloud & Home AI Computing
(26:42) How AI Workloads Are Routed & Optimized
(30:24) Big Companies Using Akash Network: NVIDIA, Prime Intellect & More
(45:49) Building a Decentralized AI Services Marketplace
(55:09) Why the Future of AI Needs a Decentralized Cloud




