In 5 Years, 90% of What You Use AI For Will Run on Your Smartphone | Paolo Ardoino, Tether

10 Aug 2026 · 59 min · 29 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Paolo Ardoino argues that most everyday AI use will shift from cloud data centers to on-device smartphone/laptop inference within 5 years, reducing the need for massive centralized compute. He connects this to Tether’s broader “disintermediation” philosophy from finance and internet routing, and explains Tether’s AI platform QVAC (Quantumverse Automatic Computer) built for private, local inference and fine-tuning.

Guest background

Paolo Ardoino is associated with Tether (founded 2014). He describes Tether as a “digital dollar” stablecoin business (USDT) with 573 million users, growing by 30+ million per quarter. He also discusses Tether’s other verticals: energy kiosks in Africa, telecom protocols, and stablecoins.

Key claims

By 2026, local models can cover ~50% of AI use cases; in 3 years ~70%; in 5 years ~90% of use cases run on smartphones. Centralized AI subscriptions may become unjustified for mainstream tasks; enterprise/government may still need data centers. He claims QVAC can outperform larger medical models on-device (4B vs MedGemma 27B) and support fine-tuning via LoRa while keeping data private.

Notable examples

WhatsApp latency example (Rome↔Frankfurt/Ireland routing). BitTorrent as a scalable peer-to-peer model. Medical example: 4B-parameter model exceeding MedGemma 27B; 1.7B model for average African smartphones. Robotics example: local, low-latency inference needed for cars/robots.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Future of AI on Smartphones

0:00 to 0:15

Learn about the potential of running AI models on smartphones by 2026.

“If in 2026 people can already run local models on their smartphones to solve only 50 % of the use cases of AI, in five years, 95 % of the population are able to run 90 % of the use cases of AI on a smartphone.”

The Data Center Paradigm Shift

0:15 to 0:36

Discussion on the relevance of data centers in the future AI landscape.

“They're building all of these data centers and it's not at all clear that the compute paradigm that we're currently operating under is going to continue, in which case maybe you don't need all those data centers.”

Understanding AI Intelligence

0:36 to 0:56

Exploration of AI's intelligence and user data implications.

“run on data centers because problems that they're solving, the speed at which they have to solve them, are going to require massive compute.”

The Digital Dollar and Financial Inclusion

1:16 to 2:04

Paolo explains Tether's mission and impact on financial inclusion.

“There are plenty of digital dollars, but only USD, our digital dollar, was able to achieve the holy grail of distribution and financial inclusion and impact in the world.”

Addressing the Unbanked Population

2:04 to 2:55

Exploring why billions are unbanked and the role of digital dollars.

“Sometimes my blood boils because how is it possible that a little company like Tether was able to achieve more for financial inclusion than all NGOs and charities and whatnot for the last 50 years?”

Global Currency Devaluation Impacts

2:55 to 4:49

Insights on currency devaluation in countries like Argentina and Turkey.

“They are unbanked because they are too poor for being of interest of the banking system.”

Smartphone Accessibility and Digital Wallets

4:49 to 6:01

Discussing smartphone access and digital wallet usage in emerging markets.

“But actually, I truly, of course, if you talk about certain remote villages in Africa, well, we can talk a lot about it.”

Disintermediation in Finance

6:01 to 7:57

How Tether exemplifies disintermediation in the financial sector.

“But bottom line is what we learned with the USDT over the last 12 years is that we can truly disintermediate.”

Revolutionizing Internet and Data Sharing

7:57 to 11:00

Examining the need for a peer-to-peer internet to reduce data latency.

“But the most important learning point of that is that not only finance is heavily intermediated, is overly intermediated, the entire technology is overly intermediated.”

Exploring Quantumverse and AI

11:00 to 12:20

Introduction to the Quantumverse initiative and its potential for AI.

“Well, there is, you know, in early 2000, a new protocol was born.”
Show all 29 chapters

The Last Question: A Sci-Fi Perspective

12:20 to 14:00

Discussion of Asimov's 'The Last Question' and its implications for society.

“But what I really wanted to talk about is your Quantumverse automatic computer.”

The Role of Technology in Society's Stability

14:00 to 17:00

Explore how technology can stabilize society and the implications of a peer-to-peer framework.

“The most complex, it is the most complex question that you can ask and trying to get an answer for.”

AI's Future on Smartphones

17:00 to 19:30

Discuss the potential for AI to run on smartphones and its implications for everyday users.

“We are just here to build a product that makes a couple of companies rich.”

The Shift to Localized AI Models

19:30 to 22:00

Examine the trend of moving AI models to local devices and the implications for data centers.

“So if today, I believe that today with the QVAC, we've created a platform, an open source platform, because open source is very important, so that people don't have to trust that they can verify the platform themselves.”

The Evolution of AI Hardware

22:00 to 24:50

Understand the advancements in AI hardware and their potential to revolutionize enterprise solutions.

“And as I think I mentioned to you, I'm interested in the potential overcapacity of data centers that's being created.”

Building AI with Privacy in Mind

24:50 to 28:00

Learn how to design AI systems that prioritize user privacy while maintaining functionality.

“And with the benefit of the fact that you can keep your secret sauce for you.”

Privacy and Personal AI Assistants

28:00 to 29:08

Learn about the need for AI models that maintain user privacy while customizing responses.

“Imagine if you could have a model on your laptop, on your smartphone, that could read privately through all your emails and design and learn how to respond in the best crags way.”

Fine Tuning AI with LoRa

29:08 to 30:24

Explore the concept of fine-tuning AI models using low-rank adaptation techniques.

“And usually the best way, best technique to do fine tuning is LoRa, that is low rank adaptation.”

Innovations in AI Model Performance

30:24 to 31:59

Discover advancements in AI model optimization for consumer devices and their implications.

“They use either 32 bits or 16 bits or 8 bits for representing the weights.”

Qvac Platform Capabilities

31:59 to 34:22

Understand the features of the Qvac platform, including supported AI models and functionalities.

“So I can have my Paolo's assistant, you can have your Craig's assistant, and that will continue to maintain 100 % of the privacy.”

Building an AI Ecosystem with Qvac

34:22 to 37:23

Learn about the Qvac AI Assistant and how it empowers developers to create AI-rich applications.

“Our AI assistant will be, again, fully open source for everyone to actually own their own intelligence.”

Robotics and Local Processing

37:23 to 40:30

Examine the role of local processing in robotics and its impact on privacy and efficiency.

“You said the Qback assistant will be direct to consumer when it reaches general availability.”

The Future of AI Infrastructure

40:30 to 42:00

Discuss the implications of massive investments in AI infrastructure and potential risks.

“that are testing how QVAC is performing in the brain of the robots they are building.”

AI Companies and Data Center Dynamics

42:00 to 45:00

Discussion on the financial strategies of AI companies regarding data centers and subsidies.

“But see, someone else will foot the bill for that.”

Rethinking AI Research and Models

45:00 to 49:20

Insights on the efficiency of AI models and the need for more researchers versus hardware.

“just that there's going to be this big financial hole that eventually has to be backfilled by governments or...”

Tether's Four Verticals Explained

49:20 to 52:00

Overview of Tether's initiatives including financial services, energy, and telecommunications.

“rather than having to go through the entire monolithic retraining of the model.”

Decentralized Energy Solutions in Africa

52:00 to 56:00

Exploration of Tether's approach to providing energy through solar kiosks in Africa.

“And which, I mean, this sounds promising, this QVAC, the tethered data vertical, is where are you spending most of your time?”

Tether's Mission and Open Source Ideas

56:00 to 57:25

Learn about Tether's focus on underserved populations and the potential of open source kits.

“a lot and is expensive but not that expensive compared to many other things that are going on in the world and we believe that is going to be very important for Tether because it's, that is our population, right?”

Global Insights from Paolo Ardoino

57:26 to 58:29

Discover Paolo's experiences and insights while traveling and working globally.

“now I came to Europe in Switzerland I needed to meet with a family I still have family in Europe so from time time I come of course still very attached to them.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00If in 2026 people can already run local models on their smartphones to solve only 50 % of the use cases of AI, in five years, 95 % of the population are able to run 90 % of the use cases of AI on a smartphone. They're building all of these data centers and it's not at all clear that the compute paradigm that we're currently operating under is going to continue, in which case maybe you don't need all those data centers. So if you don't understand how your AI works, is not you becoming more intelligent, is someone else becoming more intelligent with your data? The big transformational market is enterprise and government use, and those will likely run on data centers because problems that they're solving, the speed at which they have to solve them, are going to require massive compute.

0:51Is that possible that there'll be both? Why don't we start by having you introduce yourself to listeners about how you came to be involved with Tether and what Tether is and why it's important, which a lot of people, including me, don't quite understand. Look, Tether was a company, is a company born in 2014. Simple idea, digital dollar. There are plenty of digital dollars, but only USD, our digital dollar, was able to achieve the holy grail of distribution and financial inclusion and impact in the world. Over the last 12 years, our company built the most used digital dollar in the world with 573 million users, growing by 30 plus million users per quarter.

1:58I call it the biggest financial inclusion success story in the history of humanity. Sometimes my blood boils because how is it possible that a little company like Tether was able to achieve more for financial inclusion than all NGOs and charities and whatnot for the last 50 years? And we did it in a very simple way. We used new technologies like blockchain to make the dollar accessible. And you asked me before, well, why not just use the regular dollar? Well, the reality is that for four or five billion people, so more than half of the population in the world, they cannot have access to that regular dollar.

2:45They don't have access to basic financial services. The people that are unbanked, the number of people that are unbanked, the world is just enormous. And it's not like they are unbanked because they are bad people. They are unbanked because they are too poor for being of interest of the banking system. They live in countries where their national currency is devaluating so fast against the US dollar. Think about Argentina. The Argentinian peso lost 94.5 % of its value against the US dollar in the last five years. The Turkish lira lost 81 % of its value against the US dollar in the last five years.

3:30The Venezuelan Bolivar lost 99.8 % of its value against the US dollar. It could go on. No, I understand all that. And you said poor people, but certainly the people who are trading and holding tether are not, by and large, not poor people. I mean, you need a level of sophistication and infrastructure. You know, you need a stable internet connection. You need, you know, but I have a side, and I want to get in an argument about that, but the... We exactly can. I mean, there are people in Central South America, they have a stable connection. They, like, think about Argentina or Turkish. Turkey is, the digital penetration is incredible.

4:24Even Nigeria. Nigeria has 300 million people. Or Philippines, huge countries, huge populations. They all have a smartphone. So they all have digital wallets. So the reality actually is that the digital penetration for these new technologies and these digital dollars is much lower in Europe and the United States. Why? Because they don't need it. But actually, I truly, of course, if you talk about certain remote villages in Africa, well, we can talk a lot about it. We do some incredible things to bring connectivity and energy and access to the dollar there. We could spend hours and it's the most exciting thing in the world for me.

5:10But the reality is that many of these populations have, you know, one of the common denominators of the emerging markets is that they have more youth. They have, so the youth has digital means or understands, you know, technology faster, of course, than the older generations. And they are more prone to change and also have access to, again, to smartphones. Like the average cost of a smartphone for Africa is$80. Of course, you cannot play the best games in the world, but they have browsing. They have like Opera, that is a public company, is one of the most known browsers in the world, has 300 million users in Africa alone, right?

5:56So we partner with them, by the way, as one of the main partners. But bottom line is what we learned with the USDT over the last 12 years is that we can truly disintermediate. So if finance didn't perform well in certain parts of the world, and if so many hundreds of millions of people were left behind, it was for multiple reasons, because there were not enough incentives to solve that problem. And we proved with Tether that we could solve that problem or help to solve that problem. We are still early, but still we are doing a good job. And be very profitable doing so. And we are not taking money from anyone.

6:45Actually, well, we are taking money we're earning on the interest rate of the United States. That's it. We are not charging any transaction fee. Sending dollars for the poorest people in the world costs zero in terms of transaction fees. We make the money, again, on the interest rates of the dollars we keep in the bank account. So USDT for us and also for us internally was the demonstration that in the world of finance, there are too many intermediaries. That they all need to make money and they all need to make money on transaction fees or milking the last mile. And with USDT, we proved the opposite.

7:30We proved that we could be very successful, changing the equation and removing all the intermediaries and creating a free product, basically, like just a dollar, simple as that. And we could find a way to be profitable that no one else in the world figured out before us with the lowest risk possible. Because, you know, again, we hold the treasury bills in a bank account. So simple as that. But the most important learning point of that is that not only finance is heavily intermediated, is overly intermediated, the entire technology is overly intermediated. The, you know, Internet, I feel Internet was born to be peer to peer, point to point.

8:17When you connect your computer to the Internet or when you connect your smartphone to the Internet, you get an IP address. And that IP address is basically your home address. And is, you know, well, there are some nuances to it, but it is kind of unique to you, could be unique to you. That was the original promise of Internet. Let's connect people. And suddenly, Internet changed. And instead of let's connect people directly, became let's route their connections through data centers. Why? Because also in that case, the intermediaries created the gravitational force so that being intermediaries of people's connections, i.e.

9:04people data, they could earn money, hosting, holding, and intermediating people's data. And that is a very huge problem in a growing society for many reasons. Of course, you have data privacy and all the things we know, but there are some issues that not many think about. If you have a person living in Rome, imagine you are a person living in Rome, and it's very likely that your family lives in Rome. and you send a message, you know, I send a message, let's say I live in Rome, my family is in Rome, my mom lives in Rome. I send a message on WhatsApp to my mom. That message goes from Rome to Frankfurt or Ireland and back to Rome.

9:53Is that good? Is that normal? That adds enormous amount of latency. But even so, that is just a message. Every photo, every video, everything that I send, goes to Ireland from Rome, Ireland, Rome. Imagine how much all the governments spent in internet infrastructure in the last 20 years to route packets, data that is completely unnecessary when internet was built to be point-to-point, a peer-to-peer. So that thing is an unnecessary intermediation remediation that is only justified by the broken incentives of the internet. So what we realized at Tether is that we could, but there was already a solution also when it comes to data sharing, how we could build a protocol over internet that could scale to tens of millions of people.

10:54Because of course, the classic answer is, oh, but how we can make internet usable if there are no data centers? Well, there is, you know, in early 2000, a new protocol was born. It was called BitTorrent. It was the iteration of all the file sharing protocols. and was so great it could scale to hundreds of millions of users and exabytes and exabytes and exabytes of data in the perfect way. And so technically, we had already our answer. We could build, we could reuse structures similar to the BitTorium protocol and to disintermediate servers and data centers in many, many applications. That's also part of the things that we are doing at Tether, happy to elaborate more.

11:43But this is basically the spirit that we have in us. We learn the concept of disintermediation from finance and what we did to the financial world. And by the way, now everyone talks about stablecoins. Even banks now are talking about stablecoins because they understand the power. And so now that everyone understands the power of stablecoins, we looked around and we thought, what else we can disintermediate? Because clearly we are on the right track. And that's why we built Hall Punch as a protocol and then QVAC as an AI, the same concept applied to AI. I mean, that's fascinating. And maybe I should have you on just to talk about stable coins.

12:25But what I really wanted to talk about is your Quantumverse automatic computer. And so that's a research initiative, right? can you talk about what that is and you were alluding to it a minute ago that there is a path down which we can deliver the power of AI without having to go through massive data centers so yeah can you talk about that to explain what uh qvac for short what that is sure so the i'm a big sci-fi fan and just you know for the listeners if they for any reason they would they want to read tonight uh an interesting short story it's 14 pages i think it's the most beautiful short story ever written was written by Isaac Asimov, 1956, is called The Last Question.

13:30To me, it's beautiful because it's a story of science, AI, physics, religion, universe, philosophy, all come together in 14 pages. And it's a story of humanity every hundreds of years and dozens of years and millions of years, always perfecting this automatic computer, make it better, more powerful, and And every single time humanity was asking one single question to this computer, to this AI, it was the last question. The most complex, it is the most complex question that you can ask and trying to get an answer for. That question was how entropy may be reversed. So how we can stop the universe from dying.

14:19And to me, that is very intriguing. It's very exciting. But bringing back that huge complex question to Earth means how we can stop society from dying, how we can stop Earth from dying, how we can make society stable. And society can be made stable, going back to and connecting to my previous part, is through technology and through connecting society directly without intermediaries. Society was built for the last 5 ,000 years.

14:56society was peer-to-peer, has always been peer-to-peer. People were meeting the streets, were talking. People were using cash or salt or coins to interact peer-to-peer. Only in the last 50 years, information and money became fully intermediated. So we built an entire society over the last 5 ,000 years that was built with a completely different construct and we change it completely in the last 50 years. So we don't know how society will grow in which direction, how the decisions of intermediation that we made in the last 50 years, how they will affect society in long run. So when I go back to the sci-fi story, I think how we can truly learn from that.

15:46And if we want to achieve an AI that is even able to answer the most complex question of the universe, that should be an AI that is part of the fabric of the universe itself. It's almost like a new element of the periodic table. An AI that is so intelligent, that knows everything, that can scale and be distributed to the four corners of the universe, cannot sit in a data center on Earth, cannot belong to one person. Even just an AI, how we can plan to have an AI on Mars if the data center is on Earth? The time, the latency between Mars and Earth is too long. And then we add another planet and another planet.

16:35Of course, we are years away from that or decades away from that. But still, I think we should design technology to scale with a scale of humanity and with our ambition in terms of expanding ourselves and other universes and so on. If we don't do that, we are just here for quick gains. We are not here to build a product for the safety and stability of humanity. We are just here to build a product that makes a couple of companies rich. And I became very bearish with data centers, also for multiple reasons. You have, you know, there are some beautiful stories. There was this guy in Australia, this entrepreneur that was able to find the cure of the cancer of his dog using ChachiBT and AlphaFold.

17:28It's a beautiful story. It's a great thing. But the reality is that between 95 % and 99 % of the people in the world, they will never do that. They use AI for basic things, for search, or to take a photo of a grocery list or the grocery recipe and have at the end of the month some accounting. They want some translations. They want some education. They want things that a model, an AI model that runs on a smartphone or on a cheap laptop can already do today.

18:11So you have 95%, let's say the 90 % of the population that ask simple questions. They are not researchers. They have other problems in their day-to-day lives. They need to go from point A to point B. They need to optimize their taxes. But these are very, very simple things that simple models or models can run on a smartphone can do already. A few days ago, Tether launched, you know, there is a big race to medical health AI models. And one of the most popular ones is the flagship, which is considered state-of-the-art model, was MedGemma from Google, MedGemma 27 billion parameters. Our team was able to produce a 4 billion parameters model that was exceeding the performance and the accuracy of the 27 billion parameter model of MedGemma.

19:14So 4 billion model, I mean, 4 billion parameters model means that it can run on a smartphone, can run on a smartphone of a good smartphone. And we also have 1.7 billion parameters model that can run on a smartphone that is an average smartphone for Africa. So, and this is today's, 2026, is May 2026. So if today, I believe that today with the QVAC, we've created a platform, an open source platform, because open source is very important, so that people don't have to trust that they can verify the platform themselves. If in 2026, people can already run local models on their smartphones to solve only 50 % of the use cases of AI, of the use cases, the normal use cases of AI, the use cases of AI that everyone, all the normal people would run.

20:05In three years, I'm sure we can bring that percentage to 70%. So in three years, 70 % of the use cases of AI can run on laptops and smartphones. And in five years, it will go to 90%. So if in five years, 95 % of the population are able to run 90 % of the use cases of AI on a smartphone, then why the hell we are building tens of gigawatts of data centers? when, you know, and the subscriptions to these AI models in centralized data centers cannot be justified because then the subscription could cost millions, not hundreds of dollars. And maybe it's still great. I mean, a smartphone will never be able to find the cure.

20:54One single smartphone will never be able to find the cure of the cancer of a dog. And that is extremely important. I'm not saying that this is not important. I think medicine will have huge breakthroughs thanks to AI, but it will become a niche. What will run on centralized data centers, on these behemoths, will be a niche, will cost billions and billions of dollars, and will be probably subsidized by governments because that is of fundamental importance. But all the normal people, the hundreds of millions of people, the billions of people will use AI as part of their day-to-day lives through their smartphones, maintain privacy or their own data.

21:36But look at Apple. Every year they're releasing the M3, the M4, M5 GPU of their iPhone 16, 17, and now 18. And every year these GPUs can run these Lama models two times faster than the previous iteration. So again, that's why in five years, what we run on the smartphones will be so good that we will forget about paying a subscription to OpenAI. Yeah, and that's fascinating. And as I think I mentioned to you, I'm interested in the potential overcapacity of data centers that's being created. But while you're talking about consumer use of AI, the big transformational market is enterprise and government use.

22:42And those will likely run on data centers because the problems that they're solving, the speed at which they have to solve them, are going to require massive compute. so is is that possible that uh that uh there will be both i mean that that there will be enough demanded to take up the the data center capacity um you know to run governments to run uh the world economies and that personal use yeah will migrate to to the edge I mean, you know, coming from this is a very interesting question, right? Coming from the Bitcoin world, you might know that Bitcoin was first Bitcoin mining was first running on CPUs, then moved to GPUs and then moved to ASICs.

23:47AI started from CPUs, now is on GPUs, and last year, the first ASICs of 4AI were born. So that right now, if you have a good GPU, you can run LAMA 3.2 at 150 tokens per second. The ASICs 4AI were able to run LAMA 3.2 at 17 ,000 tokens per second. So I think there is a world where even if you are a bank, let's say a big bank, you could buy in five years 10, 15 powerful ASICs, spend$50 ,000,$100 ,000, but you can run an AI cluster that is extremely, extremely powerful through the ASICs. So the ASICs will bring efficiency or energy consumption down to 98%. So you can run much faster models with much less energy directly on site.

24:51And with the benefit of the fact that you can keep your secret sauce for you. Because you probably, you know, the more the time will pass, the more companies will realize that, you know, someone else is training data, their model on someone else's data. Right. So I think over time, of course, now only centralized data center had the capacity to create very cool stuff. And over time, I think that many banks, they run their own data centers. And the more the models will become good, the more even on the enterprise, I think that they will have their own dedicated hardware. And so there will be also because think about it, like you don't want every Italian bank, they will, sorry to make all references to Italy, but I'm Italian by birth.

25:46But will all Italian banks run on a US-based infrastructure? Right? Probably not. And so will all the Italian public administration run on a US-based infrastructure? Probably not. Right? So, and the same thing as like old French, like, so there is, there is already a big push. If you, if you read one of the coolest new, um, uh, product lines for data centers is, um, or cloud providers is actually the, the sovereign cloud means that you go into a country, you install, you build a small data center directly for the country deployed there. your software. And so even for AI providers, anyway, they will need to install directly on premises certain capacity to run models there.

26:47So I don't think we are going to have to see this huge concentration for too much or a long time. Yeah. Well, we can talk about the overbuilding. I mean, but let's talk about the Quantumverse automatic computer. It's built around this BitNet LoRa framework. Can you talk about that, describe what that is for listeners? A lot of listeners will not know what BitNet is. A lot of listeners will not know what a LoRa is and why that was a foundational breakthrough for what you're doing. In order to run on smartphones, we need, of course, to ensure we, in order to have a platform that is able and capable to scale and evolve on smartphones, you want to do two things.

27:43The first one is inference. The process of inference is, you know, you ask something to a model, you get a reply. And so that computation is called inference. The second part that is very important is how the model can learn from you. Imagine this. Imagine if you could have a model on your laptop, on your smartphone, that could read privately through all your emails and design and learn how to respond in the best crags way. So you want to do that, you might want to do that, but you also don't want to do that if it entails to send all your emails to someone else to train the model. So we wanted to make sure that if I want to have Paolo's assistant that responds in my own way, starting the emails in a certain way as I do, and so on, I needed to be able to do that running directly on my laptop and learning directly from me and maintaining 100 % of my privacy.

28:59Same with documents, same with how I start the company documents, my strategy company documents, everything that I do. I want something that is completely customizable and adaptable to me. That process is called fine tuning. And usually the best way, best technique to do fine tuning is LoRa, that is low rank adaptation. So basically you take an existing model and you adjust only a portion of the weights based on you fine tune them in order to be respectful and slightly adjust those weights so that they could respect and customize based on your own behavior. And so we understood that, of course, inference was already a proven task to do on local devices.

29:50But with QVAC, we took the Lama CPP, that is probably the most known inference engine for open source AI, and we built on top of it. And we optimized so much that now can scale to basically every single consumer GPU. And on top of that, the other breakthrough that was done by Microsoft was called BitNet. The problem is that BitNet is a one-bit model. So usually models have different levels of quantization. They use either 32 bits or 16 bits or 8 bits for representing the weights. Microsoft came up with this bitnet, so one bit weight model. The problem there is that it was not suitable to run on consumer devices.

30:46So technically it was what would have been the perfect solution to run on consumer devices, but was not suitable because it was too much relying on NVIDIA high-end GPUs. And so we modified first BitNet to make it adaptable to run on any consumer device, like in any consumer GPU, like the Snapdragon GPUs that you find on the Samsung phones or the Adreno GPUs or the Apple GPUs. Second, we also created in QVAC a common LORA fine-tuning framework so that not only for the BitNet models, but any model that we support, and we support hundreds of models, we give to the developer the same framework to fine-tune any model on all the consumer GPUs.

31:41So you have your Lama 3.2 model, you have the Medjammer, you have a Qwenn, you have all these models that are supported by our Qvac platform. Now you can run fine-tuning directly on your own device. So I can have my Paolo's assistant, you can have your Craig's assistant, and that will continue to maintain 100 % of the privacy. How far along are you in this research? I must say that I'm very impressed by our team. Well, I read them, but they are doing incredibly well. I mean, and the thing is that we work, you know, the beauty of open source. I come from the huge appreciation for people like Linus Storvalds or Richard Stallman.

32:33They are the fathers of the open source. and the beauty of working with highest levels of transparency. So when we get out with a claim, it's for everyone up there in open source. Everyone can check the code. Everyone can challenge us. That, I think, is the most beautiful thing. I believe the crypto industry, the Bitcoin industry, came up with this motto. Not your keys, not your coins. So if you don't hold the private keys to your Bitcoin, those are not your, really, you're not Bitcoin. Well, I would say not your AI, not your intelligence. So if you don't understand how your AI works, is not you becoming more intelligent?

33:21Is someone else becoming more intelligent with your data? And so we built, we are very far along now. QVAC supports OCR, so basically recognition of images and text in images, supports text-to-speech, speech-to-text, supports standard AI chats, like what you are used to with ChatGPT, support medical models, support financial models, support so many different models, Plus has the ability to also interact with other devices at the same time, supports delegated inference. So you can, from your smartphone, you can use your laptop to run the heaviest of the tasks. It's becoming a very complete product.

34:10And we support image generation, video generation. So there is so much that the team on a weekly basis is rolling out. So I'm very excited. On top of that, we are going to roll out very soon. Our AI assistant will be, again, fully open source for everyone to actually own their own intelligence. yeah i mean that's fascinating and i was looking at some of the models that you've trained to date uh you have this med sci is that right i think there's a four billion parameter version that can run entirely on smartphones so when when these models uh how are you going to distribute them or deploy them?

Read the full transcript

34:59First of all, we want the QVAC platform is made so that it's built as a software development kit so that every developer... So I want to empower every developer to integrate these AI tools directly into their existing applications. So if you have... Let's say that you are building a new version of Uber or you're building a mapping application or you're building a financial application, a training application or whatever you want, or like a cooking application or whatever, you should be able to, with QVAC, you integrate QVAC within your application. It gives you already all the primitives to support text-based AI, like chat mode, you can, or like translation, transcription, voice, you know, voice recognition, voice modification, everything you want.

35:59So we give, so we, we created basically the Lego for AI. I like to describe it in that way. I I'm a big fan of Lego products. But it's really the Lego blocks so that whatever application you're building, you can take the Lego block that is needed for you and integrating your existing application. So that truly, back to the sci-fi story that I like so much, AI can be part of the fabric of the universe. Then, in order to showcase how exactly we work and how we wanted to build our own consumer application, it's called QVAC AI Assistant. It will be released in the next couple of months and will be open source as well.

36:44and will show how all these different tools and Lego blocks can come together. Almost like you build, I was making the wrong analogy with a Death Star because I just finished to build that on the Lego side. But so I don't want to build a Death Star, of course, or AI. But you have all the schemas from Lego, and they tell you where to put each block, right? So we want to do that to showcase exactly what, in our opinion, is the best outcome you can reach, but it will give you all the different blocks. So if you want, you can take and build something else and it's all completed up to you. You said the Qback assistant will be direct to consumer when it reaches general availability.

37:34Is that right? And it'll be an app that you can download onto your phone or laptop? Exactly that. It will be an app that will work on Linux, Windows, Mac, iOS, Android, name it, and on servers. It can work everywhere. And again, we'll be fully open source. And so that clips the tether to the cloud. I mean, you're now independent of the cloud if you have that compute capability on your laptop. That's fascinating. and so what's your so you're going to release that what about the developer uptake and and has there been a lot of third-party testing to see how effective the models are or how competitive they are we have been collaborating with and we are in direct collaborations with with many companies that are building on this technology already.

38:44Keep in mind that we open sourced it less than two months ago, and yet there are quite some companies that are building on it and every day more are looking at it. But there are two particularly that are very interesting. You know, the other big part of, let's say, the tech revolution will come with robots. And that is where also things can become scary. I believe that for multiple reasons, of course, but I think robots, imagine having a robot that can only take a decision if they are connected to a cloud and to a centralized data center. Of course, there is a privacy issue, there is a control issue, many issues, but on top of that, there is a latency issue.

39:31If you have your smart car, it will only brake and stop only because the image that is seen is analyzed by a data center. Well, that can go wrong in many different ways. So you want, and the same thing with a robot. Like if let's say you have a robot that is helping, I don't know, a child, a kid, you want that robot to understand exactly and react in a microsecond or millisecond. Cannot wait for packets to go to Ireland and back if you are in Italy, right? So even robotics will need to have a GPU or an ESIC in their brain to be able to analyze information immediately, locally. So I think it's becoming very obvious, at least to us, that is going to be the case.

40:29And so, yeah, we have these two robotics companies that are testing how QVAC is performing in the brain of the robots they are building. So we are going hopefully to showcase something very soon. It's very exciting. So developers are now integrating this into applications, and that's the ambition really, to build a developer community,

40:57to disseminate this into the personal device world or the edge device world. And just on the sort of larger question of what, because I've talked to other people that are working on smaller models that can operate at the edge. um what do you think about the i don't know what is it this year 600 billion 700 billion in capital investment by ai companies uh to build data and primarily to build data centers what what's going to happen with all that capacity I think that, so I see a few problems there. It's quite interesting because if you think about it, the majority, without naming names, but many of these data centers are built by third parties that are basically then sign up contracts for offtakes.

42:09So, of course, the big AI companies have their own data centers, but the majority of this expansion in the next years will happen where they go to a third party and they say, oh, I need this data center with this capacity. Can you build it for me? And I will, of course, sign a lease. But see, someone else will foot the bill for that. So they are re-fencing themselves from the risk of building their own data centers. And also they are re-fencing themselves from the financial risk. If tomorrow they will find out, the world will find out that the data centers are not needed anymore. But on top of that, I believe the other issue is that right now, and this is an interesting moment in time, right?

43:07In 2026 or early 2027, many of these big AI companies will go public. And it's clear that many of these big AI companies are subsidizing the cost of their subscriptions. So a$200 subscription will cost from$1 ,000 to$5 ,000. That's the reality. Why they do it? Because they are private companies, so they can do it kind of hiding it. Second, because they need to show growth. So if they have to subsidize, let's say that you are like a big AI company. That's, again, not the name names. and your valuation is, let's say,$100 billion. And let's say that you can subsidize for$5 billion the cost to bring on the next 10 million users or like, I don't know, 50 million users.

44:05And if you do that, your valuation go to$150 billion. So you spend$5 billion, but your valuation went to$50 billion. And you can do that or in a repeat. but then the company will become public retail will buy the company and you cannot I mean if you are like a public company it's much harder to subsidize the cost because people will see it through and so someone else you know probably retail will foot the bill on that so I'm slightly scared for the financial engineering that is happening I think we'll I don't know how it will happen I mean, as Tether, we can just build an alternative. Yeah. Yeah.

44:55So, and when you say you're worried about that, just that there's going to be this big financial hole that eventually has to be backfilled by governments or... Yeah, could be. And there will be all of this infrastructure that sits unused. I mean, there's one in Utah that the campus is projected to be twice the size of Manhattan. I'm sure you've heard that, which is crazy. It's almost funny that the entire thing, so we are trying to create, we are trying to recreate the human brain. So our human brain consumes the same electricity of a potato clock. And we are planning to spend 10 gigawatt to recreate the human brain.

45:54Sure, you can say that, of course, there is more memory, there is much more, but we are doing something wrong, I think. I think our approach to intelligence is quite wrong. We will need, probably we need better research. We need better AI. So imagine, like, think about this question, right? So if you really, really want to solve the problem with AI long term, and let's say that you had$1 billion, would you invest$1 billion to hire 2 ,000 AI researchers, or would you spend an entire billion dollar in chips, in AI chips? I would hire researchers. The researchers will yield more over time. The chips will yield less over time.

46:39So the approach of AI is a financial game in this moment, and something has to go wrong to have, I think, a sort of reset and start from, I believe, a better approach to research, more down to earth. but something that understands that if the brain is so good, because our brain is so cool and good, probably if we are trying to recreate it with 10 million times the energy, we are wrong. Yeah. Although, you know, it's a little bit like big pharma. You know, the actual drugs they produce don't cost that much. But the whole research pipeline behind developing those drugs costs a lot of money. And something like QVAC, I mean, these are, in effect, distilled models, right, from foundation models.

47:42But you need the foundation model first before you can distill it. the future may be in, in these edge models, but you had to have the, the massive investment, uh, to get there. Uh, I, I guess the question is going forward, whether you, you know, once you have the foundation, uh, whether then, uh, these smaller models will proliferate and make the foundation models obsolete at some point. You need more the foundational model, you need the knowledge. But you have the good thing, you have internet. And so internet is full of knowledge. So of course, I'm not saying that it is zero cost to build a model, even an edge model.

48:33But the reality is, so I think the majority of the problems that, so I think that now there are companies working on trillion parameters models that are basically becoming monolithic models. I believe monolithic models are not skippable and are not the right answer moving forward. I think that the right answer moving forward are very hundreds or tens of thousands or millions of very small models that are able to interact with each other with highest efficiency so that you don't have to retrain the entire model. you can just return, you know, if you have a small, let's say in a very simple way, if you have a model that is expert in physics and another one that is expert in chemics, you can train a little bit more than one in chemics to become better rather than having to go through the entire monolithic retraining of the model.

49:28And of course, you know, you could say, well, if you have in the ideal world, an entire monolithic model has the highest efficiency, has the highest accuracy and so on. But that is not sustainable and will eventually create, in my opinion, issues. And I think that the world will converge in having millions of small models rather than one single large model. So that's why the cost can also be reduced. And also you can fine-tune task-specific models for local experiences. So over time, the ability and efficiency of local GPUs will be so high that you don't need one huge model to solve all the problems of the world.

50:11You can just easily fine-tune a model for what you need in that specific moment in time. Yeah. How is Tether organized? I mean, because this is a research initiative under Tether DAVA. Is that right? So Tether has four silos. Oh, sorry, please go ahead. No, no, that's exactly what I'm asking. Go ahead. So we have four silos or four verticals, more than silos. We have Tether Finals, the stable coins, easily put. Then we have Tether Energy. We are building some interesting energy avenues in Africa. Basically, in the most remote villages in Africa, we are building kiosks with solar panels on top and rechargeable butters inside.

51:04We have already more than 1 ,000 kiosks and more than 1 million users of these butters. So it's actually decentralized energy. I really like it. It's probably one of the companies, one of the subsidiaries that I love the most. that we're building. And then we have the third vertical is telecommunications. So we built peer-to-peer communication protocols that can scale to billions of people or billions of machines and billions of AI agents without any data center. We took the idea from the BitTorrent protocol and we perfected it, we built it, we adjusted it, we changed it to make it extremely efficient, but also not only good for file sharing, but for any type of communication.

51:53And fourth, we have this, as vertical, we have this AI platform. So these are, yeah, four verticals. Yeah. And which, I mean, this sounds promising, this QVAC, the tethered data vertical, is where are you spending most of your time? I'm almost obsessed by everything we do. So I don't have... So I try, if I do something, I try to do it well. So I dedicate almost like even amount of time. Probably energy is the part where I spend less, but AI, well, finance and AI are my top two priorities. Because I believe that if we want to have a stable society, we cannot allow 4 billion people that are the ones that don't have access to basic financial services also to not have access to basic intelligence services.

52:55So I think society will have a hard time in the next years or decades if we increase the gap that is already existing on the financial side, adding intelligence gap. This is really interesting. well let me ask you before we move on although this is really about the CUVAC initiative but on the energy that's a fascinating idea decentralized energy and this is all solar that you're pursuing we serve the so we installed these ciosks in the central corridor of Africa that is very dark at night because that is the poorest part of the world population. And so sometimes you talk to someone that says, oh, we can solve Africa's energy problem with some nuclear plants and long distance distribution lines.

54:00That is crazy. It will never work. But what we wanted to do is to have an approach where we can build these kiosks. And so they're like shacks. We put solar panels on top and a few thousand rechargeable batteries inside. And so for a couple of dollars per month, you can recharge the battery four times or you just bring back the battery to the kiosk and you swap it four times. And so we wanted to do that. We wanted to start from the most difficult part of the world. And so with 1 ,000 kiosks, we have more than 1 million users, 1.3 million users now. And these people have smartphones, by the way.

54:44These people have light bulbs, but they have some weak batteries. They don't last long. They try to recharge them. Maybe they go to... But kids cannot study at night. They don't have... If you don't have stable electricity, it's very hard to be part of civilization. So we wanted to do that. And by the way, the company, this is not charity. It's a very well-designed, very lean approach development that brings amazing utility to these populations and have a very reasonable cost. The average salary in those regions is$80 per month. So you can spend$2,$3 to have electricity at home. These batteries are 145 watt batteries.

55:34very powerful so they can run like of course they cannot run a fridge yet right so this is too expensive but at least you have light during the night you have like you can recharge smartphones and so that that is the first approach and we know that we can scale to 100 000 kiosks in the next five six years will bring energy to 30 million households so and and that my opinion will help a lot and is expensive but not that expensive compared to many other things that are going on in the world and we believe that is going to be very important for Tether because it's, that is our population, right? Servicing these people is, you know, it's our user base.

56:21Yeah, that's, and is this an open source design that someone can build or do you have kits or do you market them in the US? There's a lot of places in the US that are far from the grid. I mean, we never looked honestly at US and Europe. We started in Africa. I think it would be a good idea to open source the kits. I think it's a great idea, actually. I should bring it back to the team. Because, you know, we designed the batteries. They have, of course, in order to avoid stealing all the batteries. We made a special firmware so that the batteries can only be recharged in the kiosk. The team managing that company is incredible.

57:11It's truly incredible. They are, I think, achieving something that no one else was able to do in, again, the poorest part of the world. But I like your idea of or the suggestion to make it open source. I think it's great for humanity. Okay. Well, Paolo, this has really been fascinating um and and i'm gonna follow this uh qvac uh more more closely now um yeah and i'd like to have you on again of there's a lot of questions that we could pursue where are you based i'm um based all around the world uh mostly in esvador uh because it's uh it's an I love the place it's a beautiful place and also it's close to Central South America where there is a lot of our user base but then I travel I travel so much yeah and I see outside the window behind you it looks like a mountain what's that?

58:15now I came to Europe in Switzerland I needed to meet with a family I still have family in Europe so from time time I come of course still very attached to them. Okay. Great. Paulo.

From the publisher

Hundreds of billions of dollars are flowing into AI data centers right now, and Paolo Ardoino, CEO of Tether - the company behind the world's most widely used stablecoin with 573 million users - thinks that investment is going to age very badly. In this episode, he joins Craig Smith to explain QVAC, Tether's open-source platform for running AI on smartphones, laptops, and edge devices, and to make a case that within five years, 90% of what ordinary people use AI for will run entirely on consumer hardware, without touching a data center. The evidence is already there: Tether's team built a 4-billion-parameter medical AI model that outperforms Google's 27-billion-parameter MedGemma, running on a good smartphone, and a 1.7-billion-parameter version that runs on the average $80 smartphone available in Africa.

The deeper argument in this conversation is philosophical as much as technical. Ardoino applies the same disintermediation logic that made sending dollars to the world's unbanked free - zero transaction fees, revenue from treasury bill interest - to AI: "not your AI, not your intelligence." If you don't control how your AI runs and your data never leaves your device, the AI is genuinely yours. If it does, someone else is getting smarter with your information. He also makes a pointed economic argument: the AI companies currently charging $200 for subscriptions that cost $1,000 to $5,000 to deliver are subsidizing growth while private, and when they go public, retail investors will absorb the gap. His prescription isn't to stop building, it's to build differently, toward millions of small interacting models rather than trillion-parameter monoliths, toward devices that think locally rather than systems that route everything through Ireland and back.

Subscribe to Eye on A.I. for weekly conversations with the people building and deploying the future of AI. 

More from Eye On A.I.

All 266 episodes
In 5 Years, 90% of What You Use AI For Will Run on Your SmartphoneEye On A.I. · 59 min
Listen in VO