Training AI Models Without a Billion-Dollar Data Center | Steffen Cruz of Macrocosmos

25 May 2026 · 47 min · 16 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Steffen Kruse (Macrocosmos) explains how to train large language models without centralized “billion-dollar” GPU data centers, using BitTensor’s blockchain as a coordination/reward layer and Macrocosmos’s IOTA system to orchestrate distributed, model-parallel pre-training across cheap energy and consumer/enterprise devices.

Guest backgrounds

Steffen Kruse is co-founder and CTO of Macrocosmos; PhD in subatomic physics (UBC). He pivoted from physics research to AI, and has worked in the BitTensor ecosystem for ~3 years.

Key claims

Blockchain stores an immutable registry and audit trail (no training data or compute on-chain). Macrocosmos trains multiple models using “pockets” of cheap energy; aims for 5,000 compute nodes mid-year and ~70B parameters, then 100B+ later. Distributed training reduces CAPEX/environmental impact and enables cost/energy arbitrage.

Notable examples

IOTA (“Incentivized Orchestrated Trading Architecture”) uses an orchestration layer like Kubernetes for heterogeneous nodes, with model parallelism (small slivers per node). Uses a registry/identity layer for workers and off-chain tracking for contributions, then blockchain-triggered payouts. Mentions “Train at Home” (Mac minis/consumer GPUs) and a separate decentralized data scraping “Data Universe.”

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Steffen Cruz's Journey into AI

0:45 to 4:26

Steffen discusses his transition from academia to AI, highlighting his experiences and insights.

“I hold a PhD in subatomic physics in the University of British Columbia.”

Understanding Blockchain in AI

4:26 to 7:49

Exploration of how blockchain operates and its applications in the AI landscape.

“And it points toward the actual assets that are being registered.”

Macrocosmos and Distributed Training

7:49 to 12:28

An overview of Macrocosmos and its approach to distributed training of AI models.

“And then it becomes kind of unyieldy for people to actually have a copy of, which can interfere with the true decentralization ethos of it, and you end up with just a few copies.”

The Future of AI Model Training

12:28 to 13:15

Discussion on the future challenges and innovations in AI model training, focusing on cost and energy efficiency.

“we're going to find that we need to start looking at alternatives before we hit a hard ceiling on a lot of this.”

Comparisons with Other Computing Concepts

13:15 to 14:08

A comparative analysis of Macrocosmos's approach with other computing models like federated computing.

“So we're trying to do a lot of the early work right now on that.”

Understanding Blockchain and Federated Learning

14:08 to 15:43

Explore how blockchain technology interacts with federated computing for AI model training.

“Is it like that and you're just tying it to the blockchain to organize it and distribute it.”

IOTA and Personal Agents: Revolutionizing AI Computation

15:43 to 19:41

Learn about the IOTA project and how personal agents can create passive income through unused compute power.

“For example, federated learning is built on the idea of anonymizing client data, but being able to continuously train models.”

Decentralized Training Models: Pre-Training Explained

19:41 to 24:09

Discover the significance of pre-training in AI models and the benefits of decentralized training.

“which is the democratization of AI itself.”

Commercializing IOTA: Opportunities for Researchers

24:09 to 28:00

Understand how IOTA targets various users, providing affordable AI model training solutions.

“Beyond that fine tuning is comparatively a much shorter computational task.”

Understanding IOTA's Compute Model

28:00 to 30:15

Learn how IOTA connects suppliers and demanders of computational resources.

“These short, interruptible bursts of compute.”
Show all 16 chapters

Orchestration and Model Parallelism Explained

30:15 to 34:20

Discover the orchestration layer and how model parallelism enables large-scale training.

“the compute to the model, because they're not sending the model around all over.”

Scaling the Network for Model Training

34:20 to 39:51

Explore how IOTA scales its compute network and its implications for AI models.

“How large is the network and how do you how do you grow the network of compute?”

Future Potential of IOTA's Technology

39:51 to 41:29

Examine the future applications of IOTA's technology beyond model training.

“And in a year, a year and a half from now, I'd like to go beyond 100 billion parameter models.”

Current Status of Macrocosmos and Future Plans

41:29 to 42:00

Get updates on the current status of Macrocosmos and upcoming collaborations.

“I mean, are people training on this network yet?”

Collaborating on AI Model Training

42:00 to 43:48

Discover how Macrocosmos collaborates with startups for AI model training.

“As I mentioned, we're coming out of research mode right now.”

The Future of Decentralized AI Computing

43:48 to 46:37

Learn about the potential for decentralized AI computing and its implications.

“I'm gonna be paying attention to see how this scales.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Steffen Cruz:The blockchain, in effect, is a registry of addresses, and it points toward the actual assets that are being registered. There's no training data on the chain. There's no compute on the chain. Today, we are training models. We're training multiple models all at once, using pockets of cheap energy, which translates into cheap compute. By the middle of this year, I would like us to reach 5 ,000 compute nodes. I think this is a respectable sized cluster that you can train a model that will get people's attention and make them understand that there's a lot of utility in this technology. I think that's an important milestone for us.

0:40So my name is Stefan Kruse. I am the co-founder and the CTO of Macrocosmos. I hold a PhD in subatomic physics in the University of British Columbia. So I was a physics researcher for the beginning of my career and then I pivoted to AI when it became apparent to me that there's a lot of opportunity for scientists that are just about to graduate from their postgraduate studies. And there's just this, there's an entirely new science sort of being born and developed in real time before us. And I think my decision is one that a lot of other people have made as well, where they want to go from being in what can feel like the sort of the old school machinery of traditional science in academia, where things move slow and your contributions can be very narrow and sparse to becoming a pioneer of this entirely new exciting thing which is not only abstract and theoretical but it's being applied in new and interesting ways constantly so when i saw that i i concluded my phd and i decided to leap leap leapt out of academia and into sort of the application of um ai and also physics in in various private sector domains, such as manufacturing and systems optimization.

1:53Following that, I discovered BitTensor, which is a very interesting AI blockchain project. And ostensibly the purpose of BitTensor is to allow AI to be developed in a way that is fully democratized and is sort of done in a global way with incentives. And I thought that was such a fascinating proposal that I decided to join the network. I've been within the BitTensor ecosystem for around three years now. And I have been relentlessly experimenting and building within BitTensor the kind of things that I think would make a lot of sense and would bring a lot of value to the world using the blockchain as a prerequisite tool to enable AI research to be done in various ways and scales that would be very difficult otherwise.

2:40Steffen Cruz:Yeah, and specifically, well, first let's talk about BitTensor. but before you began we were talking about how there are a few of these blockchain ecosystems for either training or or marketing AI models or AI applications and I'm familiar with Singularity net bit tensor is much larger is that right is it the largest of these or one of the largest and how many are there i believe so so bit tensor is actually over a hundred projects under a trench coat and i think that's also why it feels so encompassing and so large and this is what is very interesting about bit tensor is it doesn't seek to solve a specific narrow problem in the AI industry.

3:34It's not like trying to solve something like how do we provide people with access to models or how do we train models or how do we do one of a plethora of different things. It actually has done a lot of foundational work on how do we make it possible to build a base layer for people to come and solve just about anything they can imagine in a way that it uses the blockchain as a reward mechanism, as a coordination layer, but then BitTensor does a a very good job of stepping out of the way and letting people's creativity drive them through to build what they want on top of it. So BitTensor feels very big and very sort of like it's everywhere because there are, today there's 128 different teams that build on top of BitTensor and they're all building something different.

4:16A lot of them choose to work within AI, but there's also some that work outside of AI as well. And so for that reason, I think that BitTensor punches well above its weight.

4:25Steffen Cruz:Yeah. And for those who are not familiar with the blockchain, and even me, the blockchain, in effect, is a registry of addresses, right? And it points toward the actual assets that are being registered. I mean, this is the thing about blockchain art that was such a craze for a while, that you're not actually buying a piece of it. I remember Trump was selling these, I think, playing cards or something like that on the blockchain. And you're not actually buying anything. You're registering your ownership of something. Exactly. Yeah. But the actual asset is a digital asset that can be copied millions of times.

5:20Steffen Cruz:It's up to you then if you want to enforce your ownership to track people down and that sort of thing. So in the projects that I've seen related to AI, it's a registry that points to an AI project, or is there something more happening on the blockchain? So in a project like Bitensor, the blockchain is used in a number of different ways. But fundamentally, it's an immutable database that is globally shared. And that in itself enables a lot of things that Bitcoin pioneered some 15 years ago. What becomes possible when you have this sort of shared globally distributed state is you can create a store of value and then you can start building higher and higher order commodities on top of that.

6:15So Bitcoin sort of addressed the most principal version of this. But what was demonstrated as a result of that is you can actually aggregate a huge amount of compute or let's just call it effort. Once you have this interesting distributed sort of system architecture, people come to it and because there's economic incentives involved, people are constantly innovating and optimizing and they're trying to find a way to sort of percolate to the top and to become the most competitive people that are doing whatever this mining task involves. In the case of Bitcoin, the mining task is just guessing random numbers faster than anyone else can.

6:52That's why you just throw a more powerful computer at it and you become a more competitive miner on the Bitcoin network. And other crypto projects have tried to orient that effort, that sort of collective coordinated effort of all the participants have tried to orient that towards different tasks, whether it's distributed file storage or creating a global residential proxy network or solving data scraping from the web and all of these things are really interesting whereas what bit tensor uses the blockchain for fundamentally as you described the registry component is very important it's also sort of a shared synchronization clock the notion of blocks and a blockchain gives everyone sort of a drumbeat that you can actually anchor work to which is an integral part of a lot of our research on bit 10 so we have this sort of shared clock which the blockchain regulates but ultimately you can do as much or as little on chain as they call it or off chain as you want and and i think some of the benefits of doing things on chain is you have complete auditability you can see precisely what the details were that went into a specific piece of work or transaction so it gives you transparency the problem with that of course is if you've got billions and billions of transactions being written to a database, your database gets very big, very fast.

8:08And then it becomes kind of unyieldy for people to actually have a copy of, which can interfere with the true decentralization ethos of it, and you end up with just a few copies. So I think that there's a pragmatic middle ground where you store exactly what you need, but nothing more, to ensure the core parts of the work that you care about are captured and untransparent for everyone to audit, review and verify. Whereas the rest of it you can you can store elsewhere you can store offline you do a million other things with it so i hope that adds a little bit of light to the question craig yeah in your case with um

8:42Steffen Cruz:macrocosmos you're providing compute for training models is that right explain precisely what macrocosmos does absolutely so today we operate three subnets in bit 10 sub each subnet is you think of it as effectively a different project, a different service. And our three different subnets, I would consider them to be three different axes or three different use cases for the blockchain that I think are very potent and powerful. So one of them is just the aggregation of raw compute at scale. And if again, harking back to what I described about the Bitcoin network, I think one One of the most important things of our time is set to be AI itself.

9:28And those who train the models and own the models will have an enormous advantage. And so a big part of the purpose of BitTensor is to provide an alternative way for us to create these models, to train these models, to deploy these models such that they're not held in walled gardens and we're having to pay for the privilege of access to them. But we want to reinvent that in a way that's more democratic and also borderless. something that's censorship resistant, something that cannot be controlled. So our primary, our flagship project within BitTensor involves what's called distributed training of large language models.

10:05And that's quite a mouthful. Distributed training is quite different to how a lot of the frontier AI labs train models today. Typically, you have huge warehouses stacked from floor to ceiling with GPUs, which are specialized computers for training AI models and serving AI models. and you've got hundreds of thousands of them in the largest data centers in the world. And they're all plugged up together with extremely high data transfer speeds. And as a result, you can basically use the whole thing as a single computer. And this is fantastic and it's been demonstrated that you can just scale the amount of compute and the amount of data you throw at a model and they get predictably better and better and better.

10:45And this is called a scaling law. Now, there's an alternative to this, which is called distributed training. Distributed training doesn't rely on having a single warehouse stuffed with computers, but actually it's based on the premise that you can train an equivalent model using computers that are distributed around the world. And now this has a lot of very interesting outcomes. When you don't do everything in a single data center, one of the first things is that the capex of actually building this massive warehouse in the first place is massively alleviated. There's also the local energy that is required to run one of these massive data centers can also now be distributed around the world in a way that creates less impact on local communities, generally much less detrimental to the environment.

11:33And lastly, it allows you to do something that is actually fundamentally not available once you've built your data center and stuffed it with GPUs. use sort of the cost of training models is already baked in to that initial build process whereas in the case of distributed training which is what we care about we can actually perform what is effectively cost arbitrage so if there's if there's a surplus of energy in iceland uh very very cheap energy we can actually use that pocket of energy even if it's only for 12 hours of the day and we can target a lot of that compute and use it in a very elastic way and in fact that's what we do Today, we are training models, we're training multiple models all at once using pockets of cheap energy, which translates into cheap compute.

12:15So it has all of these different economies of scale, which are certainly unusual, but are becoming more and more accepted as a viable alternative. And as our appetite for bigger and bigger models is only going to increase as the years go on, we're going to find that we need to start looking at alternatives before we hit a hard ceiling on a lot of this. because things like the Stargate project and the Colossus project, these are multi-billion dollar GPU build-outs. And so we think that just like the fundamental physics experiments, like the Large Hadron Collider, at some point you require the budget of a nation state, it's 20 billion dollars to build a bigger ring to smash protons or electrons.

12:54You need to start thinking about different experiments because it just becomes unpalatable eventually. So we're trying to get ahead of this problem. I think that by 2028, we're actually, the world, the mainstream, the Overton window about training models is going to shift. And we are going to have to start thinking about how do we do this in a way that arbitrage is more efficient in both cost and energy. So we're trying to do a lot of the early work right now on that. And BitTensor is a wonderful place for us to do this research because we are able to use the blockchain to reward people that contribute their compute to our training experiment.

13:28Steffen Cruz:Let me ask a couple of questions. I had a guy on the program who was doing something interesting. He was, and I don't remember whether it was blockchain based, but he would aggregate spare compute from data centers or, you know, whether on premise or, you know, independent data centers around the world. and then offer that compute to AI projects at a cheaper price than the hyperscalers or the big clouds could offer. And so that's one question. Is it like that and you're just tying it to the blockchain to organize it and distribute it. The other is, you know, for a long time, people were talking about federated computing so that you wouldn't have to lose control of your data, but you could make it available for training models.

14:45Steffen Cruz:and again that that was using the blockchain to register the the data i guess uh so first of all is it related to either of those concepts and and then the the deeper question and the one even talking to the people i've spoken to i've never really gotten a hold of the chain again is just a registry. There's no training data on the chain. There's no compute on the chain. The compute resides, as you said, maybe in Iceland. What is the link between the chain and the compute in Iceland? So starting with your first questions, the relationship between those two examples you gave and what we're doing, I would say they're They're cousins, but they're not siblings.

15:43They're a little bit different. For example, federated learning is built on the idea of anonymizing client data, but being able to continuously train models. We're working with a slightly different version of this, which is not so much the edge devices providing training data today, but the edge devices are providing the training compute, I suppose, would be a way to think about that. And secondly, I think regarding what BitTensor actually does or what the role of the chain is, at its core, it's really not that revelationary at all. At its core, what the blockchain is effectively doing is it's providing a trustworthy, transparent record for everyone to audit, especially when there's code that's being run on the blockchain as well.

16:36people understand how this code is going to interact it can't be tampered with which means that it's predictable and that's a proxy for safety and trustworthiness so if someone had already deployed a smart contract for instance we're not actually building a smart contract in microcosmos but i guess just to illustrate the point if someone has written something in the smart contract then the creator of the smart contract can't later change the terms and conditions on you it the point is it's a it's a modular block of code that will run under these conditional expressions, that makes it something that you can effectively count on.

17:12So I believe I've given you a bit of an illustration of one of the projects that we're working on in BitTensor, which is distributed training. This project is called IOTA. And as I mentioned, I think that it's a precursor to some very powerful downstream technologies that we may find are an instrumental part of the way that we create AI in the next decade, let's say. But we really do believe that there's also a real change in the way that people are thinking about personal agents and personal compute. So I've seen personally in the last few months, a lot of people are starting to stockpile Mac minis.

17:50I'm not sure if you've seen this trend, but basically now that agents are starting to become kind of economically useful a little bit, we're seeing the beginnings of it, I think as a fair appraisal. What we're starting to see now is people want one of these things running for them privately around the clock day in and day out. So they're buying a dedicated computer. And on that computer, your personal agent is running and it's checking your emails. It's maybe doing some online shopping. Maybe it's doing some work for you on the side or some hobbies. And this is a wonderful use case. And this is also a really interesting place where our technology of IOTA developed by Microcosmos comes in because all of these computers that people have now built at home.

18:27Well, there's two ways to think about it. That's your personal agent. That is creating value for you personally. But wouldn't it be great if that compute could earn passive income? Almost like you own a property and you can Airbnb it out when you don't need it. And this is actually what we're offering as we develop IOTA is we've created something called Train at Home, which is people that have unused devices at home. Let's say MacBooks, Mac minis, or even consumer GPUs. They can plug those into our network. And that means that they become part of this global supercomputer. So we can use them for training models, which we can commercialize.

19:03And they can also basically create some passive income just from having that device sitting around that was otherwise underutilized. So I think there's something very parsimonious about that arrangement of things. And the more people that are choosing to buy personal devices to run these agents, you realistically don't need the agent running 24 hours a day. You need pockets of productivity. I want to make sure it, you know, plans my weekly agenda. I want to make sure it checks this, that, and the other. Maybe there's only four hours of the day that you actually need this thing. And you can actually get a return on your initial investment within 12 days, 30 days, by plugging it into a network like this and also being part of something that I believe has a great purpose, which is the democratization of AI itself.

19:43So it's a really interesting intersection of timelines right now with agents becoming demonstrably more useful and engagement increasing. And we basically can use all of that consumer compute. we can glue it together and we can train models at greater and greater scales and greater and

19:59Steffen Cruz:greater utility so yeah yeah yeah as a matter of fact i'm running the mac mini uh stockpiling i think started with the open claw and uh exactly i'm running open claw on this computer which i know is a bad idea and i have an unused mac mini upstairs so uh uh i'll probably go fix that after this call but uh but the uh so so how does somebody use this uh well first of all yeah back to the point of so you've got compute around the world that is plugged into the network

20:46Steffen Cruz:do people who are using uh macro cosmos do they need to know what's how the what the link to the blockchain is absolutely not yeah and and the blockchain is just keeping track of where the compute is and what's available at any point in time is that right yep and it's just responsible for uh rewarding participation so it's making sure that any contributions made by anyone's devices are fairly uh fairly rewarded that's really the way i think about this the blockchain is just a really convenient way of taking care of uh compensating people for their contributions The actual onboarding experience for people is we keep this as low cognitive load as possible.

21:40Right now, it's a one-click app store download. This thing will just run passively in your machine, adding some progressively more useful user experiences, things like telling it, oh, you can only use my computer when I go to sleep. So by default, it's disabled until 10 p.m., and it's only enabled from 10 p.m. till 6 a.m. So that would be a simple user control that will get it out of the way of when you're actually trying to be productive with your machine. That's assuming it's your primary machine. If it's a secondary machine that just sits at home, we're also creating a way for agents to decide, oh, I finished my work.

22:13I'm not going to need to do anything for four hours. So the agent just decides, hey, I'm just going to go and make 20 bucks. And that would be nice to come home from work too, right? It'd be nice to come home from work and you're like, hey, what have you done today, Claude? and it's like, oh, hey, I finished all my work by 9.15 a.m. and you were going to be out till 5 p.m. So I thought, well, I thought I'd just shop around and see how I could make you a little bit of side cash. And here you go. So I do think that as our computers become more autonomous, thinking about our relationship with computation, it's going to change fundamentally.

22:43And our computers are not going to sit there waiting. They're going to be very proactive and they're going to be very resourceful, especially when you've got these clever little agents that are running the show inside of the machine. And as a result of that, I think we need to really rethink what your computer can do for you, you know, our relationship to that world. And I think bringing this back to the top about BitTensor, you can think of BitTensor as just a bunch of places where you can use that machine and you can go and make some income from it. Whether it's providing something like human intelligence or whether it's training models or whether it's detecting if an image is real or fake, You can almost think of it as a mechanical Turk for agents and humans alike.

23:25Well, you can go there, you can monetize your skills and your experience or just your raw compute.

23:31Steffen Cruz:Yeah. And for the training, how does, let's say I have a large model that I want to fine tune. How do I do it with Microcosmos? So today our work is focused on the first stage of training models specifically, which is called pre-training. So pre-training has historically been by and large, the most computationally expensive part of creating a model. It's the part where they systematically inhale the entire internet, which usually takes months and tens of thousands of GPUs. Beyond that fine tuning is comparatively a much shorter computational task. So I think it's the essence of retraining requiring long time horizon workloads that is actually why we're so interested in moving this from a centralized to a decentralized context.

24:31And there's a lot of pioneering research that needs to be done that we hope will be very useful even outside of the Web3 world. we hope that this is really useful research to anyone out there that is interested in economically or capital efficient ways of training models we hope that this is a contribution to the field that's really how i see this because what we're doing at the root of our work is we're taking a network of very unreliable compute people are joining people are leaving it's just constantly churning over and churning over how do we create something that is persistent and stable

25:07Steffen Cruz:out of something that is fundamentally so noisy and unstable. And it's core. That's the problem we're trying to solve. And it just so happens that the use case for this persistent compute fabric we're creating is trading models, but there's nothing stopping that from doing something else. Perhaps there's a world where universities could use a similar decentralized network for their traditional HPC cluster jobs where they're doing academic research that's not necessarily AI related at all. It could be bioinformatics. It could be physics. The same thing still remains to be true, which is if you need 10 ,000 nodes of compute for 12 hours, what's the cheapest way you can get them?

25:51And we're trying to build software that basically orchestrates compute around the world to act like it was all plugged into itself and it's a supercomputer. And I think it's a fundamental thing that we're working on that we're applying towards training models, but I hope has a much broader value to the community. Right.

26:08Steffen Cruz:So on pre-training, how does somebody use it? Absolutely. So we're sort of emerging from research mode right now. We've spent about nine months on IOTA so far, and we've demonstrated that we're able to reproduce a lot of important baseline benchmark performance metrics using this, what we call it, the wonky vegetables, sort of the oakley vegetables that you get in the vegetable aisle that no one necessarily wants to buy. we're going to turn that into a premium soup. And we've basically graduated from research mode right now where we're showing that the two are indistinguishable if you treat them in the right way.

26:42If you use the right spices, it tastes just as good. What we're moving to next is thinking very deeply about what is the commercialization opportunity for a technology like this? Who are our target audiences? We think that the world is going to train more models in the next 10 years, not less. We think that researchers, small startups, especially cash-strapped startups, academia, all of these people are interested in training more models, but perhaps for budget reasons, it's very hard for them to do so. So we're sort of targeting the commercialization of IOTA as two things primarily. One is what we call supply side.

27:19So supply side would basically mean people that have surplus GPUs, so the neoclouds, the hyperscalers, perhaps they have 10 ,000 GPUs in their inventory, but they can only rent out 9 ,000 of them at a given time or 9 ,500 of them. Well, for us, them plugging in their compute into our network allows them to increase their overall utilization, which goes straight to that bottom line. So they'll be happy with that. And even better, if you actually to look at the usage of all of those GPUs that a hyperscaler has, you might find that, well, it's rented out for four hours and then it's not rented for two hours and then it's rented out for four hours.

Read the full transcript

27:58And that little two-hour gap is exactly what we're trying to utilize. These short, interruptible bursts of compute. And we turn that into, again, a continuous stream of compute is exactly what excites us. So for the supply side, it's basically anyone that has GPUs, what they have to do right now is if they can't rent it out, they're sort of a buyer of last resort on the market. They'll rent out those GPUs at cents on the dollar for inference tokens. And we can basically say, well, hey, hey, instead of doing that, how about you plug them into our network and we can basically give you a better margin?

28:32Because trading is a higher order commodity than inferencing. And as a result of that, we believe we could pass on the benefits to them as suppliers. So that's the supply side. On the demand side, as I mentioned before, we have researchers, we have startups, we have all kinds of different profiles of users that we know right now, their training models all the time. What we want to do for those guys is basically provide them with an interface to train models in a way that's recognizable to them. So there's already very popular libraries like PyTorch, like TensorFlow. These are very popular libraries that are basically used by everyone in industry that's training models.

29:08Well, we want to make it as simple for them in terms of abstractions as if they were just building models in the normal routine way. So with no additional cognitive overhead, you basically say, I want to train a model. you're not going to painfully specify oh i want to use a nigerian gpu for 12 minutes of course you're not right what you're going to do is you're going to say this is my this is my sort of executive level objective this is what i want to get done these are the parameters that i want to have precise control over for reproducibility purposes for deterministic purposes all of this section down here feel free to go and arbitrage and get me the cheapest version and then we have this real now we have a two-sided market right now we have a sort of supply side and a demand side that are sort of dancing together where we're arbitraging offering cheaper rates and and iota is a technology that basically is this infrastructure layer that is is the one that brings this compute to market pairs it up with users that want it and we believe that that's something that is

30:05Steffen Cruz:phenomenally valuable yeah well and then how does the the under the hood how do you get the the compute to the model, because they're not sending the model around all over. Yeah. The way that you get the compute to the model is, it's effectively an orchestration deployment layer, something reminiscent of Kubernetes, but Kubernetes that deploys to heterogeneous compute nodes around the world. So it's a dynamic deployment layer that basically takes some containerized code that the user writes, sends it out to those machines, and then very importantly ensures that those machines around the world have a tunnel to communicate with each other.

30:55And that tunnel is the part that makes them act like one big blob of compute instead of 50 different random pieces of compute. So it's effectively that. It deploys anyone's code as an interactive system or a network of compute that is sort of geospatially aware. Right.

31:14Steffen Cruz:And that orchestration layer, where is that based? The orchestration layer is something that we maintain and operate. So you would effectively dispatch your code from your laptop or anything like that. Your code would then be sent securely to these backend nodes, wherever they might be distributed around the world. And then you would effectively, as a user, you would have the ability to monitor logs, monitor trace data, the usual sort of dashboarding tooling, things like that as the user. But effectively, what's happened is your code has now been sent out, cloned, replicated to all these different copies.

31:54And all of those copies are hosting a section of the model, which is what makes IOTA very, in my opinion, very novel. we don't actually have each of these compute nodes running a full copy of the model they're actually running a small sliver of the model it's called model parallelism and in effect what this means is you can train really large frontier size models using very small building blocks like you can build a huge lego tower out of small pieces it's the same idea and the way it works is in order to train the model in its entirety you have to root information from all of the sections of the model and knit that together.

32:33And as a result of that, you can basically train this sort of distributed model architecture as if it was all stuck together in one place. And that's why it's taken us nine months to get out of research mode, to be honest, Craig. It's quite a formidable task. Yeah.

32:48Steffen Cruz:And is the blockchain important for the synchronization? Because you need very tight synchronization between nodes or between different compute, right? So the way that we utilize the blockchain today for IOTA's purposes is we need a registry of everyone that's in the network right now. So it exposes them as, hey, I'm an available worker. Here's how you can find me and here's my unique address that I know to identify you. unambiguously. So we use that as sort of an authorization layer, an identity layer. Also with that, we have an off-chain layer to IOTA, which is constantly tracking the contributions of that node.

33:41And then once we have computed the total work done by each node, that is sent back down to the blockchain layer where it triggers a payout loop. So everyone that contributed that compute is then rewarded in the IOTA token for their contributions of work. And this actually can be articulated in a very tokenized context or even a US dollars context. All you effectively need as building blocks here is you need to know exactly who it is that's assigned to this model training experiment, where you can find them, how to glue them all together, and how much work they've done. And if you've got all of those things, you have an operational distributed trading network.

34:20Steffen Cruz:How large is the network and how do you how do you grow the network of compute? Yeah so when we began the source of all of the compute for our distributed trading experiments was people that came to our project in in BitTensor the IOTA project with they have their compute they know what we're here to do usually there's a philosophical alignment or sometimes they just see it as an opportunity to make some cash. But basically, people show up armed with usually pretty powerful hardware. So the supply side of computing BitTensor is very much not a concern. People know that BitTensor is there. People know the rules of the game.

35:04So we were almost always oversubscribed in terms of people that wanted to contribute to our experiments. However, in November, December last year, we wanted to go beyond the scale of the network that is supported natively by BitTensor. So the maximum number of people that could be participating in the IOTA experiments was 256, which if you consider that to be a data center, it's tiny. It's a pebble. So we actually worked around that and we've now made it effectively a limitless size system by instead of having every unique entity as a unique address on the blockchain, we now keep track of it off-chain effectively.

35:47And what this allows us to do now is we can have an arbitrary-sized registry of participants. So we had 2 ,500 people download our macOS app in the first two weeks since it launched. And now all of those people are available at a moment's notice to join our network. And they're paid out identically to the ones that are strictly on-chain within BitTensor. We have a transparent payout system that people understand. It's a predictable system. It directly rewards contribution in proportion to basically hours of work done per day. But the nice thing about this is it's not limited by the amount of storage capacity for the blockchain or the number of slots that a blockchain can have as miners.

36:26We actually decided to just rewrite a lot of that using different rails so that it's not limited in that way. and now as a result we have 500 or so uh today running that are the amount of work that we have to give the miners to do is actually quite um um it changes over time there's times when we run lots and lots of experiments we train lots of bottles and there's also time when we only need a small amount of compute and we only pay for what we need we're not just the force isn't just always running we actually have experiments that are live or we have production runs where we're going to say this is the model we're going to train.

37:03We're going to go find 500 devices around the world. Fortunately, again, the bits tensor pieces made it very easy to find 500 random people in the world with powerful computers that I can just plug into my network, which is the miracle of technology really. And so we can effectively not worry at all about where that compute comes from. And then downstream, we basically pay everybody out for that work that they've done. Yeah.

37:27Steffen Cruz:And is there, do you have a sense of scale, comparative scale that would give me and listeners an idea of how this would compare to a major data center? Yeah. The largest data centers in the world are measured. The ones that have started to come online are now hundreds of thousands of GPUs. So these are multi-billion dollar projects. They're absolutely enormous. We're not at that scale yet. But the nice thing is the technology that we're building scales in a very beautiful way. In fact, to go from our network with 10 ,000 nodes to 20 ,000 nodes, there's nothing that we fundamentally need to change.

38:16We just need to expand the distribution surface of what we're doing. And our infrastructure layer scales to much, much higher capacity. So again, this is the nice thing about not being committed to a bricks and mortar build out in the way that the centralized compute paradigm is. In our case, if we want to double, triple, or 10x the capacity of our compute network, we simply need to go out, increase the surface of the network. So what I think is in the books for us is by the end of this year, I want, well, by the middle of this year, I would like us to reach 5 ,000 compute nodes. I think this is a respectable size cluster that you can train a model that will get people's attention and make them understand that there's a lot of utility in this technology.

39:02I think that's an important milestone for us. something around the 70 billion parameter model size, I think is when you go from, go from, you know, you graduate from the sort of small models that are proof of concepts to things that are actually taken much more seriously. And enterprise customers that have been waiting to train an in-house model, whether it's a legal specialist model or whether it's a medical model or something like that, perhaps they've been waiting for a while, but they couldn't really afford it because it's a very high price tag that comes with training a model of that scale.

39:32because of the discounts that we can offer using this energy and cost arbitrage approach, I think that we're going to start to see a lot more people train their own sovereign models, their own enterprise models. And so by mid this year, it's important to us that we demonstrate that that is possible with the technology that we're working on. And in a year, a year and a half from now, I'd like to go beyond 100 billion parameter models. So the parameter count of the model is very much the number that we think about when we consider what it takes to make this technology something people take seriously.

40:05And I think we need to have a portfolio of models that we can point at and say, these are just as good as the ones that you would have gotten in a centralized context, but at 10 % of the cost, at 20 % of the cost. I think then people really start to take notice and it becomes an interesting proposition. And that's where we'd like to get to. Yeah.

40:24Steffen Cruz:And we were talking about people stockpiling Mac minis. The people that are joining or contributing compute to Macrocosmos are GPU, people with GPUs, available GPUs, right? Not CPUs. It's yeah, we're not going to get anything done if we wait for CPUs to do the work, unfortunately. Perhaps there are workloads in the future. As I mentioned before, we can imagine this technology is something that allows you to do more than just model training. We're working on a more fundamental infrastructure problem, not just a model training problem in many respects. And there are a lot of scientific problems, computation expensive scientific problems that are CPU limited.

41:13And in those cases, absolutely, it would be very interesting to apply this technology in those cases. But today we support effectively CUDA devices and full silicon devices, which are Mac minis, the new Macbooks.

41:28Steffen Cruz:Is this online yet? I mean, are people training on this network yet? And if not, when will it be available? Do you have people lining up? Yeah, you can find us at iota.microcosmos.ai. That's I-O-T-A, which I should probably have introduced as an acronym for the Incentivized Orchestrated Trading Architecture. Okay. So iota.microcosmos.ai is where you can learn more about this project. Today, we don't have any customers in yet. As I mentioned, we're coming out of research mode right now. We're trying to really calibrate the system and tighten this thing up. There's a lot of moving parts. that we want to make sure we have really nailed down, but we have startups that are already ready and committed to working with those training models towards the second half of this year, which is incredibly exciting.

42:19So we'll have some collaboration results and some early partnership results by the end of summer.

42:24Steffen Cruz:Okay, great. And one other question, this is for compute, but we spoke about federated computing. could you register data on the network or make data available for pre-training and then have Macrocosmos send the model around to different blocks of data to train the model? It's a very interesting idea. one of our other projects in BitTensor is actually a web scale data scraping service where we decentralize the efforts of hundreds and hundreds of miners that scrape social media data for us. And we can use that for many reasons, anything from journalism to marketing, brand analysis, all the way to AI model training.

43:22So we actually have an entirely standalone project, which is dedicated to data scraping, which we think is an important part of our, I call it sort of virtuous cycle. We have data, we have compute. We also have a lot of what we call innovation networks, which I didn't get much time to talk about today. But to your question directly about whether we could actually outsource client-side data and use it for model training, my answer is we already are. We actually just have it in a dedicated, highly scalable system called Data Universe.

43:53Steffen Cruz:Okay, okay. Well, this is fascinating. I'm gonna be paying attention to see how this scales. Do you think this model will be replicated by others?

44:11I think it's very important that it is. Do you mean the technology or the actual

44:17Steffen Cruz:downstream model artifact the technology the technology because as you said uh the the appetite for uh compute for pre-training compute is going to grow the you know the number of models out there is only going to grow and compute is expensive and a lot of it's underutilized so it just feels like uh like there would be other people doing this uh and another question is are you working with any of the big cloud providers to maximize their GPU utilization? To your first question, I certainly see other teams right now that are working on very similar projects. Some of them are focused specifically on model training, or others are more oriented towards this orchestration technology.

45:24I think we're one of the very few that's combining those two in the way that we're doing it right now. So I do expect the world of devices, you can call it the Internet of Things, you can call it whatever you want. But I think that there is a second law of thermodynamics at play here, which is things get more connected over time. And I think that especially when there is a lot of value added and a lot of potential to having the orchestration of multiple devices to do higher order workloads. So in our case, there's simply not enough memory on one Mac Mini to train the models that we want, which is why we have to do all this work to glue them together.

46:00And there's many, many other problems where there's just not enough resources available on one device to do what it is that you need. So I think that there's a lot of other industries that are already thinking carefully about how can we get access to the kind of compute in the specific topology we need it. And I think that what we're building is going to be very interesting to them. Devices, as I said before, devices are only going to get more intelligent and more autonomous, and they're going to become much more useful if you can stick them together like Lego bricks and make these bigger pieces out of them, these different constellations.

46:31So I see a bright future for that effort right there. Whether we're the ones that really take this all the way to the top of the mountain or we get halfway there, I'm proud of the work that we're doing. And I think that it's, again, I think it's very valuable for everybody that we work on this.

46:47Steffen Cruz:Okay, great. So iota.macrocosmos.ai. Yeah, you can find everything you need about us just at macrocosmos.ai. And you can find all of our projects and our research in decentralized AI.

From the publisher

Training a frontier AI model today requires hundreds of thousands of GPUs, months of compute time, and a budget that only a handful of companies on earth can afford. Steffen Cruz, co-founder and CTO of Macrocosmos, thinks that model is about to break, and he's spending his time building what comes next. His project IOTA, operating within the BitTensor blockchain ecosystem, uses distributed training to split large language models across thousands of devices located around the world, coordinated by blockchain, and powered by surplus cheap energy wherever it exists. After nine months of research, the system can reproduce baseline benchmark performance using what Cruz calls "wonky vegetables" - unreliable, churning, globally distributed compute - and turn it into something indistinguishable from centralized training if you use the right approach.

The conversation with Craig Smith covers the mechanics of how this actually works, why the blockchain's role is far narrower and more practical than most people assume, and why the Mac mini stockpiling trend creates an unexpected supply of distributed compute that can earn passive income when idle. Cruz's target: a 70 billion parameter model by mid-2025, trained at 10-20% of what it would cost through a hyperscaler, and aimed squarely at the legal firms, hospitals, and cash-strapped startups that have been waiting to train their own sovereign models but couldn't afford the price tag.

Subscribe to Eye on A.I. for weekly conversations with the people building and deploying the future of AI.

More from Eye On A.I.

All 266 episodes
Training AI Models Without a Billion-Dollar Data CenterEye On A.I. · 47 min
Listen in VO