#340 Steffen Cruz: Training AI Without Data Centres

29 Apr 2026 · 46 min · 22 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Steffen Cruz explains Macrocosmos and its BitTensor-based project IOTA (incentivized orchestrated trading architecture) for training large language models without centralized data centers. He argues blockchain provides an immutable, auditable registry and payout coordination layer, while compute and training run off-chain on distributed GPUs worldwide.

Guest backgrounds

Steffen Cruz is co-founder/CTO of Macrocosmos, holds a PhD in subatomic physics from UBC, pivoted from academia/physics research into AI, and has spent ~3 years building in the BitTensor ecosystem.

Key claims

BitTensor is “over 100 projects under a trench coat” with ~128 teams building on it. Macrocosmos uses distributed training to reduce CAPEX, lower environmental impact, and exploit energy cost arbitrage (e.g., Iceland surplus power). Blockchain is used for worker identity/authorization and payout triggers, not for storing training data.

Notable examples

IOTA’s “Train at Home” uses unused Mac minis/MacBooks/consumer GPUs; IOTA uses model parallelism (nodes host slivers of a model, not full copies). He cites plans to reach ~5,000 compute nodes and target ~70B then 100B+ parameter models; no customers yet, but startups are lined up for later this year.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Transitioning from Physics to AI

0:45 to 1:46

Steffen describes his transition from academia to AI and his motivations.

“to becoming a pioneer of this entirely new exciting thing which is not only abstract and theoretical but it's being applied in new and interesting ways constantly.”

Exploring BitTensor and Its Purpose

1:46 to 3:30

Discussion about BitTensor's goals and its role in democratizing AI development.

“BitTensor is much larger, is that right?”

Understanding Blockchain's Role in AI

3:30 to 5:10

Explaining how blockchain acts as a registry and coordination tool in AI projects.

“and they're all building something different.”

Innovations in AI Training via Blockchain

5:10 to 7:00

Steffen discusses the innovative ways BitTensor uses blockchain for AI training.

“but fundamentally it's an immutable database that is globally shared.”

Distributed Training vs. Traditional Methods

7:00 to 10:40

Comparison of distributed training with traditional centralized AI training methods.

“But ultimately, you can do as much or as little on-chain, as they call it, or off-chain as you want.”

The Future of AI Training and Cost Arbitrage

10:40 to 12:45

Discussing the future of AI training and the need for more efficient methods.

“these massive data centers can also now be distributed around the world in a way that creates less impact on local communities, genuinely much less detrimental to the environment.”

Linking Compute and Blockchain

12:45 to 14:00

Exploring the relationship between blockchain and distributed compute resources.

“I had a guy on the program who was doing something interesting.”

Understanding Federated Learning and Blockchain

14:00 to 16:47

Learn the differences and connections between federated learning and blockchain technology.

“but you could make it available for training models.”

The Concept of Distributed Training with IOTA

16:47 to 18:59

Explore how the IOTA project aims to utilize home computers for distributed AI training.

“that we create AI in the next decade, let's say.”

Personal Agents and Passive Income Potential

18:59 to 21:41

Discover how personal agents can run on home devices and generate passive income.

“So it's a really interesting intersection of timelines right now with agents becoming demonstrably more useful and engagement increasing.”
Show all 22 chapters

Decentralized Compute and AI Model Training

21:41 to 24:40

Understand the challenges and solutions in decentralized computing for AI model training.

“It'd be nice to come home from work and you're like, hey, what have you done today, Claude?”

Commercialization and User Interfaces for AI Training

24:40 to 28:00

Learn about the commercialization strategy for IOTA and user-friendly AI training interfaces.

“Perhaps there's a world where universities could use a similar decentralized network for their traditional HPC cluster jobs where they're doing academic research that's not necessarily AI related at all.”

Training Models with IOTA

28:00 to 28:38

Learn how IOTA provides an interface for users to train AI models more efficiently.

“the supply side on the demand side as i mentioned before we have researchers we have startups we have all kinds of different profiles of users that we know right now they're training models all the time.”

Orchestration Layer Explained

28:38 to 31:09

Discover how the orchestration layer facilitates model training with distributed compute.

“You basically say, I want to train a model.”

Blockchain Integration in Computing

31:09 to 33:18

Understand the role of blockchain in managing compute synchronization and rewards.

“And all of those copies are hosting a section of the model, which is what makes IOTA very, in my opinion, very novel.”

Expanding the Compute Network

33:18 to 35:38

Explore how IOTA scales its compute network and rewards contributors effectively.

“So when we began, the source of all of the compute for distributed training experiments was people that came to our project in BitTensor, the IOTA project, with, they have their compute, they know what we're here to do.”

Future Projections for IOTA

35:38 to 37:58

Learn about IOTA's vision for scaling and the potential of their technology.

“But the nice thing about this is it's not limited by the amount of storage capacity for the blockchain or the number of slots that a blockchain can have as miners.”

Broadening Computational Applications

37:58 to 40:28

Discuss the applications of IOTA technology beyond model training.

“So what I think is in the books for us is by the end of this year, I want, well, by the middle of this year, I would like us to reach 5 ,000 compute nodes.”

Data Integration with Macrocosmos

41:14 to 42:00

Explore the potential of integrating data scraping with IOTA's compute network.

“So iota.microcosmos.ai is where you can learn more about this project.”

Decentralizing Data Scraping for AI

42:00 to 43:14

Learn about the innovative web-scale data scraping project for AI model training.

“and then for pre-training and then have Macrocosmos send the model around to different blocks of data to train the model?”

The Future of AI Compute Utilization

43:14 to 44:29

Explore the growing demand for compute in pre-training AI models and its implications.

“I'm going to be paying attention to see how this scales.”

Orchestrating Intelligent Devices for AI

44:29 to 45:54

Discover the potential of connecting devices to enhance AI capabilities and efficiency.

“I certainly see other teams right now that are working on very similar projects.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So my name is Steffen Cruz. I am the co-founder and the CTO of Macrocosmos. I hold a PhD in subatomic physics in the University of British Columbia. So I was a physics researcher for the beginning of my career and then I pivoted to AI when it became apparent to me that there's a lot of opportunity for scientists that are just about to graduate from their postgraduate studies. and there's just this there's an entirely new science sort of being born and developed in real time before us and I think my decision is one that a lot of other people have made as well where they want to go from being in what can feel like the sort of the old school machinery of traditional science in academia where things move slow and your contributions can be very narrow and sparse to becoming a pioneer of this entirely new exciting thing which is not only abstract and theoretical but it's being applied in new and interesting ways constantly.

0:57So when I saw that, I concluded my PhD, and I decided to leap, leapt out of academia and into sort of the application of AI and also physics in various private sector domains, such as manufacturing and systems optimization. Following that, I discovered BitTensor, which is a very interesting AI blockchain project. and ostensibly the purpose of BitTensor is to allow AI to be developed in a way that is fully democratized and is sort of done in a global way with incentives and I thought that was such a fascinating proposal that I decided to join the network. I've been within the BitTensor ecosystem for around three years now and I have been relentlessly experimenting and building within BitTensor the kind of things that I think would make a lot of sense and would bring a lot of value to the world using the blockchain as a prerequisite tool to enable ai research to be done in various ways and scales that will be very difficult otherwise yeah and specifically well first let's talk about bit tensor uh but before you began we were talking about uh how there are a few of these blockchain ecosystems for either training or marketing AI models or AI applications.

2:24And I'm familiar with SingularityNet. BitTensor is much larger, is that right? Is it the largest of these or one of the largest and how many are there? I believe so. So BitTensor is actually over 100 projects under a trench coat. And I think that's also why it feels so encompassing and so large. And this is what is very interesting about BitTensor is it doesn't seek to solve a specific narrow problem in the AI industry. It's not like trying to solve something like how do we provide people with access to models or how do we train models or how do we do one of a plethora of different things. It actually has done a lot of foundational work on how do we make it possible to build a base layer for people to come and solve just about anything they can imagine in a way that it uses the blockchain as a reward mechanism, as a coordination layer.

3:17But then BitTensor does a very good job of stepping out of the way and letting people's creativity drive them through to build what they want on top of it. So BitTensor feels very big and very sort of like it's everywhere because today there's 128 different teams that build on top of BitTensor and they're all building something different. A lot of them choose to work within AI, but there's also some that work outside of AI as well. And so for that reason, I think that BitTensor punches well above its weight. Yeah, and for those who are not familiar with the blockchain, and even me, the blockchain, in effect, is a registry of addresses, right?

3:59And it points toward the actual assets that are being registered. I mean, this is the thing about blockchain art that was such a craze for a while, that you're not actually buying a piece of... I remember Trump was selling these, I think, playing cards or something like that on the blockchain. And you're not actually buying anything. You're registering your ownership of something. Precisely. Yeah. but the actual asset is a digital asset that can be copied millions of times. It's up to you then if you want to enforce your ownership to track people down and that sort of thing. So in the projects that I've seen related to AI, it's a registry that points to an AI project, or is there something more happening on the blockchain?

5:06So in a project like BitTensor, the blockchain is used in a number of different ways, but fundamentally it's an immutable database that is globally shared. And that in itself enables a lot of things that Bitcoin pioneered some 15 years ago. What becomes possible when you have this sort of shared globally distributed state is you can create a store of value. And then you can start building higher and higher order commodities on top of that. So Bitcoin sort of addressed the most principal version of this. But what was demonstrated as a result of that is you can actually aggregate a huge amount of compute, or let's just call it effort.

5:47Once you have this interesting distributed sort of system architecture, people come to it. And because there's economic incentives involved, people are constantly innovating and optimizing, and they're trying to find a way to sort of percolate to the top and to become the most competitive people that are doing whatever this mining task involves. In the case of Bitcoin, the mining task is just guessing random numbers faster than anyone else can. That's why you just throw a more powerful computer at it and you become a more competitive miner on the Bitcoin network. And other crypto projects have tried to orient that effort, that sort of collective coordinated effort of all the participants have tried to orient that towards different tasks, whether it's distributed file storage or creating a global residential proxy network or solving data scraping from the web.

6:36And all of these things are really interesting. Whereas what BitTensor uses the blockchain for fundamentally, as you described, the registry component is very important. It's also sort of a shared synchronization clock. The notion of blocks and a blockchain gives everyone sort of a drumbeat that you can actually anchor work to, which is an integral part of a lot of our research on BitTen. So we have this sort of shared clock, which the blockchain regulates. But ultimately, you can do as much or as little on-chain, as they call it, or off-chain as you want. And I think some of the benefits of doing things on-chain is you have complete auditability.

7:13You can see precisely what the details were that went into a specific piece of work or transaction, so it gives you transparency. The problem with that, of course, is if you've got billions and billions of transactions being written to a database, your database gets very big very fast. And then it becomes kind of unyieldy for people to actually have a copy of, which can interfere with the true decentralization ethos of it, and you end up with just a few copies. So I think that there's a pragmatic middle ground where you store exactly what you need but nothing more to ensure the core parts of the work that you care about are captured and untransparent for everyone to audit, review, and verify, whereas the rest of it you can you can store elsewhere you can store offline you do a million other things with it so i hope that adds a little bit of light to the question craig yeah in your case with um macrocosmos you're providing compute for training models is that right explain precisely what macrocosmos does absolutely so today we operate three subnets in bit 10 sun each subnet is you can think of it as effectively a different project a different service um and our three different subnets i would consider them to be three different um three different axes or three different use cases for the blockchain that i think are very potent and powerful so one of them is just the aggregation of raw compute at scale and if again harking back to what i described about the bitcoin network i think one of the most important things of our time is set to be ai itself.

8:48And those who train the models and own the models will have an enormous advantage. And so a big part of the purpose of BitTensor is to provide an alternative way for us to create these models, to train these models, to deploy these models such that they're not held in walled gardens and we're having to pay for the privilege of access to them. But we want to reinvent that in a way that's more democratic and also borderless, something that's censorship resistant, something that cannot be controlled. So our primary, our flagship project within BitTensor involves what's called distributed training of large language models.

9:25And that's quite a mouthful. Distributed training is quite different to how a lot of the frontier AI labs train models today. Typically, you have huge warehouses stacked from floor to ceiling with GPUs, which are specialized computers for training AI models and serving AI models. and you've got hundreds of thousands of them in the largest data centers in the world, and they're all plugged up together with extremely high data transfer speeds. And as a result, you can basically use the whole thing as a single computer. And this is fantastic, and it's been demonstrated that you can just scale the amount of compute and the amount of data you throw at a model, and they get predictably better and better and better.

10:05This is called a scaling law. Now, there's an alternative to this, which is called distributed training. Distributed training doesn't rely on having a single warehouse stuffed with computers, but actually it's based on the premise that you can train an equivalent model using computers that are distributed around the world. And now this has a lot of very interesting outcomes. When you don't do everything in a single data center, one of the first things is that the capex of actually building this massive warehouse in the first place is massively alleviated. There's also the local energy that is required to run one of these massive data centers can also now be distributed around the world in a way that creates less impact on local communities, genuinely much less detrimental to the environment.

10:52And lastly, it allows you to do something that is actually fundamentally not available once you've built your data center and stuffed it with GPUs. Sort of the cost of training models is already baked in to that initial build process. Whereas in the case of distributed training, which is what we care about, we can actually perform what is effectively cost arbitrage. So if there's a surplus of energy in Iceland, very, very cheap energy, we can actually use that pocket of energy, even if it's only for 12 hours of the day, and we can target a lot of that compute and use it in a very elastic way. And in fact, that's what we do.

11:28Today, we are training models, we're training multiple models all at once using pockets of cheap energy, which translates into cheap compute. So it has all of these different economies of scale, which are certainly unusual, but are becoming more and more accepted as a viable alternative. And as our appetite for bigger and bigger models is only going to increase as the years go on, we're going to find that we need to start looking at alternatives before we hit a hard ceiling on a lot of this, because things like the Stargate project and the Colossus project. These are multi-billion dollar GPU build-outs.

12:01And so we think that just like the fundamental physics experiments, like the Large Hadron Collider, at some point you require the budget of a nation state, it's 20 billion dollars to build a bigger ring to smash protons or electrons. You need to start thinking about different experiments because it just becomes unpalatable eventually. So we're trying to get ahead of this problem. I We're actually the world, the mainstream, the Overton window about training models is going to shift. And we are going to have to start thinking about how do we do this in a way that arbitrage is more efficient in both cost and energy.

12:35So we're trying to do a lot of the early work right now on that. And BitTensor is a wonderful place for us to do this research because we are able to use the blockchain to reward people that contribute their compute to our training experiment. Let me ask a couple of questions. I had a guy on the program who was doing something interesting. He was, and I don't remember whether it was blockchain-based, but he would aggregate spare compute from data centers or whether on-premise or independent data centers around the world and then offer that compute to AI projects at a cheaper price than the hyperscalers or the big clouds could offer.

13:32And so that's one question. Is it like that and you're just tying it to the blockchain to organize it and distribute it? The other is, you know, for a long time, people were talking about federated computing so that you wouldn't have to lose control of your data, but you could make it available for training models. and again that that was using the blockchain to register the the data i guess uh so first of all is it related to either of those concepts and and then the the deeper question and the one even talking to the people i've spoken to i've never really gotten a hold of the chain again is just a registry there's no training data on the chain there's no compute on the chain the compute resides as you said uh maybe in iceland what is the link between the chain and the compute in iceland so starting with your first questions the relationship between those two examples you gave and what we're doing i would say that they're cousins but they're not uh they're not siblings So they're a little bit different.

15:04For example, federated learning is built on the idea of anonymizing client data, but being able to continuously train models. We're working with a slightly different version of this, which is not so much the edge devices providing training data today, but the edge devices are providing a trained compute, I suppose, would be a way to think about that. And secondly, I think regarding what BitTensor actually does or what the role of the chain is, at its core, it's really not that revelationary at all. At its core, what the blockchain is effectively doing is it's providing a trustworthy, transparent record for everyone to audit, especially when there's code that's being run on the blockchain as well.

15:56people understand how this code is going to interact. It can't be tampered with, which means that it's predictable, and that's a proxy for safety and trustworthiness. So if someone has already deployed a smart contract, for instance, we're not actually building a smart contract in Microcosmos, but I guess just to illustrate the point, if someone has written something in a smart contract, then the creator of the smart contract can't later change the terms and conditions on you. The point is it's a modular block of code that will run under these conditional expressions, that makes it something that you can effectively count on.

16:32So I believe I've given you a bit of an illustration of one of the projects that we're working on in BitTensor, which is distributed training. This project is called IOTA. And as I mentioned, I think that it's a precursor to some very powerful downstream technologies that we may find are an instrumental part of the way that we create AI in the next decade, let's say. But we really do believe that there's also a real change in the way that people are thinking about personal agents and personal compute. So I've seen personally in the last few months, a lot of people are starting to stockpile Mac minis.

17:10I'm not sure if you've seen this trend. But basically, now that agents are starting to become kind of economically useful, a little bit, we're seeing the beginnings of it, I think, as a fair appraisal. what we're starting to see now is people want one of these things running for them privately around the clock day in and day out so they're buying a dedicated computer and on that computer your personal agent is running and it's checking your emails it's maybe doing some online shopping maybe it's doing some work for you on the side or some hobbies and this is wonderful use case and this is also a really interesting place where our technology of iota developed by microcosmos comes in because all of these computers that people have now built at home well there's two ways to think Think about it.

17:50That's your personal agent. That is creating value for you personally. But wouldn't it be great if that compute could earn passive income? Almost like you own a property and you can Airbnb it out when you don't need it. And this is actually what we're offering as we develop IOTA is we've created something called Train at Home, which is people that have unused devices at home. Let's say MacBooks and Mac minis or even consumer GPUs. They can plug those into our network. and that means that they become part of this global supercomputer so we can use them for training models which we can commercialize and they can also basically create some passive income just from having that device sitting around that was otherwise underutilized.

18:30So I think there's something very parsimonious about that arrangement of things and the more people that are choosing to buy personal devices to run these agents you realistically don't need the agent running 24 hours a day you need pockets of productivity I want to make sure it you know plans my weekly agenda. I want to make sure it checks this, that, and the other. Maybe there's only four hours of the day that you actually need this thing. And you can actually get a return in your initial investment within 12 days, 30 days by plugging it into a network like this and also being part of something that I believe has a great purpose, which is the democratization of AI itself.

19:03So it's a really interesting intersection of timelines right now with agents becoming demonstrably more useful and engagement increasing. And we basically can use all of that consume the compute, we can glue it together, and we can train models at greater and greater scales and greater and greater utility. So yeah. Yeah, yeah. As a matter of fact, I'm running the Mac Mini stockpiling, I think, started with OpenClaw. And I'm running OpenClaw on this computer, which I know is a bad idea, and I have an unused Mac Mini upstairs. So I'll probably go fix that after this call but uh but the uh so so how does somebody use this uh well first of all yeah back to the point of so you've got compute around the world that is plugged into the network

20:06do people who are using uh macro cosmos do they need to know what's how the what the link to the blockchain is absolutely not yeah and and the blockchain is just keeping track of where the compute is and what's available at any point in time is that right yep and it's just responsible for uh rewarding participation so it's making sure that any contributions made by anyone's devices are fairly uh fairly rewarded that's really the way i think about this the blockchain is just a really convenient way of taking care of uh compensating people for their contributions the actual onboarding experience for people is we keep this as low cognitive load as possible right now it's a one-click app store download this thing will just run passive in your machine adding some more progressively more useful user experiences things like telling it oh you can only use my computer when i go to sleep so it's by default it's disabled until 10 p.m and it's only enabled from 10 p.m till 6 a.m so that would be a simple user control that will you know get it out of the way of when you're actually trying to be productive with your machine.

21:25That's assuming it's your primary machine. If it's a secondary machine that just sits at home, we're also creating a way for agents to decide, oh, I finished my work. I'm not going to need to do anything for four hours. So the agent just decides, hey, I'm just going to go and make 20 bucks. And that would be nice to come home from work too, right? It'd be nice to come home from work and you're like, hey, what have you done today, Claude? And it's like, oh, hey, I finished all my work by 9.15 AM and you were going to be out till 5 PM. So I thought, well, I thought I'd just shop around and see how I could make you a little bit of sidecache.

21:54And here you go. So I do think that as our computers become more autonomous, thinking about our relationship with computation, it's going to change fundamentally. And our computers are not going to sit there waiting. They're going to be very proactive and they're going to be very resourceful, especially when you've got these clever little agents that are running the show inside of the machine. And as a result of that, I think we need to really rethink what your computer can do for you. you know, our relationship to that world. And I think bringing this back to the top about BitTensor, you can think of BitTensor as just a bunch of places where you can use that machine and you can go and make some income from it.

22:33Whether it's providing something like human intelligence or whether it's training models or whether it's detecting if an image is real or fake, you can almost think of it as a mechanical Turk for agents and humans alike. Well, you can go there, you can monetize your skills and your experience or just your raw compute. Yeah. And for the training, how does, let's say I have a large model that I want to fine-tune, how do I do it with Microcosmos? So today our work is focused on the first stage of training models specifically, which is called pre-training. So pre-training has historically been, by and large, the most computationally expensive part of creating a model.

23:23It's the part where they systematically inhale the entire internet, which usually takes months and tens of thousands of GPUs. Beyond that, fine-tuning is comparatively a much shorter computational task. So I think it's the essence of pre-training requiring long-time horizon workloads. that is actually why we're so interested in moving this from a centralized to a decentralized context. And there's a lot of pioneering research that needs to be done that we hope will be very useful. Even outside of the Web3 world, we hope that this is really useful research to anyone out there that is interested in economically or capital efficient ways of training models.

24:08We hope that this is a contribution to the field. That's really how I see this, because what we're doing at the root of our work is we're taking a network of very unreliable compute. People are joining. People are leaving. It's just constantly churning over and churning over. How do we create something that is persistent and stable out of something that is fundamentally so noisy and unstable? And it's core. That's the problem we're trying to solve. And it just so happens that the use case for this persistent compute fabric we're creating is trading models, but there's nothing stopping that from doing something else.

24:44Perhaps there's a world where universities could use a similar decentralized network for their traditional HPC cluster jobs where they're doing academic research that's not necessarily AI related at all. It could be bioinformatics. It could be physics. the same thing the same thing is still still remains to be true which is if you need 10 ,000 nodes of compute for 12 hours what's the cheapest way you can get them and we're trying to build software that basically orchestrates compute around the world to act like it was all plugged into itself and it's a supercomputer and I think it's a it's so it's it's a fundamental thing that we're working on that we're applying towards training models but I hope has a much broader value to the community right so on pre-training uh how does somebody use it absolutely so we're sort of emerging from from research mode right now we spent about nine months on iota so far and we've demonstrated that we're able to reproduce a lot of important baseline benchmark performance metrics using this what we call it the wonky vegetables sort of the oakley vegetables that you get in the vegetable aisle that no one necessarily wants to buy we're going to turn that into a premium soup.

25:55And we've basically graduated from research mode right now, where we're showing that the two are indistinguishable if you treat them in the right way. If you use the right spices, it tastes just as good. What we're moving to next is thinking very deeply about what is the commercialization opportunity for a technology like this? Who are our target audiences? We think that the world is going to train more models in the next 10 years, not less. We think that researchers, small startups, especially cash-strapped startups, academia, all of these people are interested in training more models, but perhaps for budget reasons, it's very hard for them to do so.

26:31So we're sort of targeting the commercialization of IOTA as two things primarily. One is what we call supply side. So supply side would basically mean people that have surplus GPUs, so the neoclouds, the hyperscalers, perhaps they have 10 ,000 GPUs in their inventory, but they can only rent out 9 ,000 of them at a given time or 9 ,500 of them. Well, for us, them plugging in their compute into our network allows them to increase their overall utilization, which goes straight to that bottom line. So they'll be happy with that. And even better, if you actually to look at the usage of all of those GPUs that a hyperscaler has, you might find that, well, it's rented out for four hours and then it's not rented for two hours and then it's rented out for four hours.

27:18And that little two-hour gap is exactly what we're trying to utilize, these short, interruptible bursts of compute. And we turn that into, again, a continuous stream of compute is exactly what excites us. So for the supply side, it's basically anyone that has GPUs, what they have to do right now is if they can't rent it out, they're sort of a buyer of last resort on the market. They'll rent out those GPUs at cents on the dollar for inference tokens. and we can basically say well hey instead of doing that how about you plug them into our network and we can basically give you a better margin because trading is a higher order commodity than inferencing and as a result of that we believe we could pass on the benefits to them as suppliers so that's the supply side on the demand side as i mentioned before we have researchers we have startups we have all kinds of different profiles of users that we know right now they're training models all the time.

28:11What we want to do for those guys is basically provide them with an interface to train models in a way that's recognizable to them. So there's already very popular libraries like PyTorch, like TensorFlow. These are very popular libraries that are basically used by everyone in industry that's training models. Well, we want to make it as simple for them in terms of abstractions as if they were just building models in the normal routine way. So with no additional cognitive overhead. You basically say, I want to train a model. You're not going to painfully specify, oh, I want to use a Nigerian GPU for 12 minutes.

28:46Of course you're not, right? What you're going to do is you're going to say, this is my sort of executive level objective. This is what I want to get done. These are the parameters that I want to have precise control over for reproducibility purposes, for deterministic purposes. All of this section down here, feel free to go and arbitrage and get me the cheapest version. And then we have this real, now we have a two-sided market, right? Now we have a sort of supply side and a demand side that are sort of dancing together where we're arbitraging, offering cheaper rates. And IOTA is a technology that basically is this infrastructure layer that is the one that brings this compute to market, pairs it up with users that want it.

29:23And we believe that that's something that is phenomenally valuable. Yeah, well, and then how does the, under the hood, how do you get the compute to the model? Because they're not sending the model around all over. Yeah. The way that you get the compute to the model is, it's effectively an orchestration deployment layer, something reminiscent of Kubernetes. but Kubernetes that deploys to heterogeneous compute nodes around the world. So it's a dynamic deployment layer that basically takes some containerized code that the user writes, sends it out to those machines, and then very importantly ensures that those machines around the world have a tunnel to communicate with each other.

30:15And that tunnel is the part that makes them act like one big blob of compute instead of 50 different random pieces of compute. So it's effectively that. It deploys anyone's code as an interactive system or a network of compute that is sort of geospatially aware. Right. And that orchestration layer, where is that based? The orchestration layer is something that we maintain and operate. So you would effectively dispatch your code from your laptop or anything like that. Your code would then be sent securely to these backend nodes, wherever they might be distributed around the world. And then you would effectively, as a user, you would have the ability to monitor logs, monitor trace data, the usual sort of dashboarding tooling, things like that as the user.

31:08But effectively, what's happened is your code has now been sent out, cloned, replicated to all these different copies. And all of those copies are hosting a section of the model, which is what makes IOTA very, in my opinion, very novel. We don't actually have each of these compute nodes running a full copy of the model. They're actually running a small sliver of the model. It's called model parallelism. and in effect what this means is you can train really large frontier size models using very small building blocks like you can build a huge lego tower out of small pieces it's the same idea and the way it works is in order to train the model in its entirety you have to root information from all of the sections of the model and knit that together and as a result of that you can basically train this this sort of distributed model architecture as if it was all stuck together that in one place.

32:01And that's why it's taken us nine months to get out of research mode, to be honest, Greg. It's quite a formidable task. Yeah. And is the blockchain important for the synchronization? Because you need very tight synchronization between nodes or between different compute, right? so the way that we utilize the blockchain today for iota's purposes is we need a registry of everyone that's in the network right now so it exposes them as hey i'm an available worker here's how you can find me and here's my unique uh address that i know to identify you unambiguously so we use that as sort of uh an authorization layer an identity layer also with that we we have an off-chain layer to iota which is constantly tracking the contributions of that node and then once we have computed the total work done by each node that is sent back down to the blockchain layer where it triggers a payout loop so everyone that contributed that compute is then rewarded in the iota token for their contributions of work and and this actually can be articulated in a very tokenized context or even a u.s dollars context that all you effectively need is building blocks here is is uh you need to know exactly who it is that's assigned to this model training experiment you know where you can find them you know how to glue them all together and you know how much work they've done and if you've got all of those things you have an operational distributed trading network how large is the network and how do you how do you grow the network of compute?

Read the full transcript

33:47Yeah. So when we began, the source of all of the compute for distributed training experiments was people that came to our project in BitTensor, the IOTA project, with, they have their compute, they know what we're here to do. Usually there's a philosophical alignment, or sometimes they just see it as an opportunity to make some cash. But basically, people show up armed with usually pretty powerful hardware. So the actual, the supply side of computing BitTensor is very much not a concern. People know that BitTensor is there. People know the rules of the game. So we were almost always oversubscribed in terms of people that wanted to contribute to our experiments.

34:32However, in December, November, December last year, we wanted to go beyond the scale of the network that is supported natively by BitTensor. So the maximum number of people that could be participating in the IOTA experiments was 256, which if you consider that to be a data center, it's tiny. It's a pebble. So we actually worked around that, and we've now made it effectively a limitless-sized system by instead of having every unique entity as a unique address on the blockchain, we now keep track of it off-chain effectively. And what this allows us to do now is we can have an arbitrary size registry of participants.

35:11So we had 2 ,500 people download our macOS app in the first two weeks since it launched. And now all of those people are available at a moment's notice to join our network. And they're paid out identically to the ones that are strictly on-chain within BitTensor. We have a transparent payout system that people understand. It's a predictable system. and it directly rewards contribution in proportion to basically hours of work done per day. But the nice thing about this is it's not limited by the amount of storage capacity for the blockchain or the number of slots that a blockchain can have as miners.

35:46We actually decided to just rewrite a lot of that using different rails so that it's not limited in that way. And now as a result, we have 500 or so today running. that are the amount of work that we have to give the miners to do is actually quite um um it changes over time there's times when we run lots and lots of experiments we train lots of bottles and there's also time when we only need a small amount of compute and we only pay for what we need we're not just the force isn't just always running we actually have experiments that are live or we have production runs where we're going to say this is the model we're going to train.

36:23We're going to go find 500 devices around the world. Fortunately, again, the bit tensor pieces made it very easy to find 500 random people in the world with powerful computers that I could just plug into my network, which is the miracle of technology, really. And so we can effectively not worry at all about where that compute comes from. And then downstream, we basically pay everybody out for that work that they've done. Yeah. And is there, do you have a sense of scale uh comparative scale that that would give me and listeners an idea of of how this would compare to uh to a major data center or yeah yeah the the largest data centers in the world are measured that the ones that have started to come online are now hundreds of thousands of GPUs.

37:17So these are multi-billion dollar projects. They're absolutely enormous. We're not at that scale yet. But the nice thing is the technology that we're building scales in a very beautiful way. In fact, to go from our network with 10 ,000 nodes to 20 ,000 nodes, there's nothing that we fundamentally need to change. We just need to expand the distribution surface of what we're doing and our infrastructure layer scales to much, much higher capacity. So again, this is the nice thing about not being committed to a bricks and mortar build out in the way that the centralized compute paradigm is. In our case, if we want to double, triple or 10x the capacity of our compute network, we simply need to go out, increase the surface of the network.

38:01So what I think is in the books for us is by the end of this year, I want, well, by the middle of this year, I would like us to reach 5 ,000 compute nodes. I think this is a respectable-sized cluster that you can train a model that will get people's attention and make them understand that there's a lot of utility in this technology. I think that's an important milestone for us. Something around the 70 billion parameter model size, I think, is when you graduate from the sort of small models that are proof of concepts to things that are actually taken much more seriously. and enterprise customers that have been waiting to train an in-house model, whether it's a legal specialist model or whether it's a medical model or something like that.

38:46Perhaps they've been waiting for a while, but they couldn't really afford it because it's a very high price tag that comes with training a model of that scale. Because of the discounts that we can offer using this energy and cost arbitrage approach, I think that we're going to start to see a lot more people train their own sovereign models, their own enterprise models. And so by mid this year, It's important to us that we demonstrate that that is possible with the technology that we're working on. And in a year, a year and a half from now, I'd like to go beyond 100 billion parameter models. So the parameter count of the model is very much the number that we think about when we consider what it takes to make this technology something people take seriously.

39:25and I think we need to have a portfolio of models that we can point at and say these are just as good as the ones that you would have gotten in a centralized context but at 10 % the cost, at 20 % the cost. I think then people really start to take notice and it becomes an interesting proposition and that's where we'd like to get to. Yeah, and we were talking about people stockpiling Mac minis. The people that are joining or contributing compute to Macrocosmos are GPU, people with GPUs, available GPUs, right? Not CPUs. Yeah. We're not going to get anything done if we wait for CPUs to do the work, unfortunately.

40:13Perhaps there are workloads in the future. as I mentioned before, we can imagine this technology is something that allows you to do more than just model training. We're working on a more fundamental infrastructure problem, not just a model training problem in many respects. And there are a lot of scientific problems, computation expensive scientific problems that are CPU limited. And in those cases, absolutely, it would be very interesting to apply this technology in those cases. but today we support effectively CUDA devices and Apple Silicon devices which are Mac minis, the new Macbooks. Is this online yet?

40:50I mean can are people training on this network yet? And if not when will it be available? Do you have people lining up? Yeah you can find us at iota.microcosmos.ai that's I-O-T-A which I should probably have introduced is an acronym for or the incentivized orchestrated trading architecture. Okay. So iota.microcosmos.ai is where you can learn more about this project. Today, we don't have any customers in yet. As I mentioned, we're coming out of research mode right now. We're trying to really calibrate the system and tighten this thing up. There's a lot of moving parts that we want to make sure we have really nailed down.

41:32But we have startups that are already ready and committed to working with those training models towards the second half of this year, which is incredibly exciting. So we'll have some collaboration results and some early partnership results by the end of summer. Okay, great. And one other question. This is for compute, but we spoke about federated computing. Could you register data on the network or make data available and then for pre-training and then have Macrocosmos send the model around to different blocks of data to train the model? It's a very interesting idea. One of our other projects in BitTensor is actually a web-scale data scraping service where we decentralize the efforts of hundreds and hundreds of miners that scrape social media data for us.

42:35And we can use that for many reasons, anything from journalism to marketing, brand analysis, all the way to AI model training. So we actually have an entirely standalone project which is dedicated to data scraping, which we think is an important part of our, I'd call it sort of virtuous cycle. We have data, we have compute. We also have a lot of what we call innovation networks, which I didn't get much time to talk about today. But to your question directly about whether we could actually outsource client-side data and use that for model training, my answer is we already are. We actually just have it in a dedicated, highly scalable system called Data Universe.

43:14Okay. Okay. Well, this is fascinating. I'm going to be paying attention to see how this scales. Do you think this model will be replicated by others?

43:31I think it's very important that it is. Do you mean the technology or the actual downstream model artifact? The technology. The technology. because as you said, the appetite for compute, for pre-training compute is going to grow. The number of models out there is only going to grow. And compute is expensive and a lot of it's underutilized. so it just feels like there would be other people doing this. And another question is, are you working with any of the big cloud providers to maximize their GPU utilization? So to your first question, I certainly see other teams right now that are working on very similar projects.

44:38Some of them are focused specifically on model training or others are more oriented towards this orchestration technology. I think we're one of the very few that's combining those two in the way that we're doing it right now. So I do expect the world of devices, you can call it the Internet of Things, you can call it whatever you want, but I think that there is a second law of thermodynamics at play here, which is things get more connected over time. and I think that especially when there is a lot of value added and a lot of potential to having the orchestration of multiple devices to do higher order workloads.

45:13So in our case, there's simply not enough memory on one Mac Mini to train the models that we want, which is why we have to do all this work to glue them together. And there's many, many other problems where there's just not enough resources available on one device to do what it is that you need. So I think that there's a lot of other industries that are already thinking carefully about how can we get access to the kind of compute in the specific topology we need it. And I think that what we're building is going to be very interesting to them. Devices, as I said before, devices are only going to get more intelligent and more autonomous, and they're going to become much more useful if you can stick them together like Lego bricks and make these bigger pieces out of them, these different constellations.

45:51So I see a bright future for that effort right there. Whether we're the ones that really take this all the way to the top of the mountain or we get halfway there, I'm proud of the work that we're doing. and I think that it's very valuable for everybody that we work on this Okay Great, so iota.macrocosmos.ai You can find everything you need about us just at macrocosmos.ai and you can find all of our projects and our research in decentralized AI

From the publisher

What if you could train a frontier AI model without building a single data centre?

In this episode of Eye on AI, Craig Smith sits down with Steffen Cruz, co-founder and CTO of Macrocosmos, to explore a radical alternative to the way AI models are built today. Instead of billion-dollar GPU warehouses, Steffen is training large language models using idle compute from devices distributed around the world, coordinated through the Bittensor blockchain.

Steffen breaks down why the centralised data centre model is heading toward a wall. Projects like Stargate and Colossus cost tens of billions of dollars, and as appetite for larger models grows, the economics simply stop making sense. He explains how distributed training flips this on its head, tapping into surplus energy, underutilised GPUs, and even consumer devices like Mac Minis to train models at a fraction of the cost.

We also get into IOTA, Macrocosmos's flagship technology, an orchestration layer that takes compute nodes scattered across the globe and makes them act like a single supercomputer. No single device runs the full model. Instead, each one carries a small slice, a technique called model parallelism, and together they can train frontier-scale models that would otherwise be out of reach for startups, researchers, and enterprises.

Finally, Steffen shares what he's building toward: 70 billion parameter models trained at 10 to 20 percent of centralised costs, a two-sided marketplace for compute, and a future where anyone with a spare GPU or Mac Mini can earn passive income while contributing to the democratisation of AI.

Subscribe for more conversations with the people building the future of AI and emerging technology.

 

Stay Updated:

Craig Smith on X: https://x.com/craigss

Eye on A.I. on X: https://x.com/EyeOn_AI

 

Timestamp:

(00:00) Introduction: The Problem With Blockchain AI Projects

(06:39) Meet Steffen Cruz: From Subatomic Physics to Decentralised AI

(09:16) What Is a Bittensor? The Blockchain Built for AI

(11:53) How the Blockchain Actually Works: Registry, Clock, and Rewards

(15:08) Why Data Centres Are Hitting a Wall

(22:01) Distributed Training vs Federated Learning: What's the Difference?

(27:47) Train at Home: Turning Your Mac Mini Into a Passive Income Machine

(32:49) IOTA Explained: Building a Global Supercomputer From Spare Parts

(39:43) How the Network Scales: From 256 Nodes to Limitless Compute

(44:39) The Road Ahead: 70B Parameter Models and the Future of Affordable A

More from Eye On A.I.

All 266 episodes
#340 Steffen Cruz: Training AI Without Data CentresEye On A.I. · 46 min
Listen in VO