In short
ARIA’s Scaling Inference Lab and why AI compute must evolve from today’s homogeneous, closed stacks to heterogeneous, modular “systems of models” that co-evolve hardware and software for cheaper, faster, scalable inference.
Guests
Suraj Bramhavar, Programme Director at ARIA; previously running ARIA’s Scaling Compute program. Danyal Akarca, co-founder at Callosum; came in as an ARIA seed creator; recently out of stealth; raised just over $10M led by Plural.
Key claims
The main bottleneck is the gap from prototype to global scale; ARIA will run six-month pilot “clusters” to test hardware/software combinations. Heterogeneous compute is inevitable (historical computing waves) and better for multi-step, uncertain, multi-agent problems. Value shifts to the system/software “connective tissue,” not just dominant chips.
Notable examples
A “kitchen table photo” task (identify cereal, match taste, reduce sugar, auto-purchase) illustrating multi-step limits of single models; transform “pre-fill vs decode” split as an early sign of heterogeneity. Mentions partners/areas: Fractile (inference chips), photonic networking with Mix, and “thermodynamic compute.”
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOOverview of ARIA and Scaling Compute Program
0:46 to 3:30
Discussion on the purpose and goals of the Scaling Compute Program.
“than the way most modern AI hardware operates.”
Identifying Gaps in AI Hardware Development
3:31 to 5:50
Exploration of existing gaps between idea stage and global implementation.
“Right now, the industry, the way it operates, it's mostly kind of scaling up homogenous, closed, very heavily engineered compute systems, right?”
The Future of AI Hardware
5:51 to 7:50
The shift from homogenous to heterogeneous computing systems in AI.
“I mean, if you look through computer systems history, this has happened over and over and over again.”
Danyal Akarca's Insights on Heterogeneous Compute
7:51 to 11:20
Danyal discusses the evolution of computing and the need for diverse intelligence.
“to make these things work um all right so i think now dan like you you tell us how you fit into this you've coming out of stealth today what are you building and where does this fit into the border picture.”
Real-World Applications of Heterogeneous Computing
11:21 to 13:44
Danyal explains how different hardware can optimize real-world problem-solving.
“and networking technologies to be able to bring them to market as well to actualize that vision.”
The Co-Evolution of Hardware and Software
13:45 to 14:02
Danyal emphasizes the importance of co-design in developing AI systems.
“And so you have different trade-offs of the different hardwares that are there.”
Challenges in AI Adoption
14:02 to 14:28
Discusses the hurdles in adopting new AI technologies and the need for advanced hardware.
“So, for example, we aren't kind of getting it to this point yet where we think that it's going to be easily adopted for the hardest problems right now.”
Co-evolution of AI Hardware and Algorithms
14:28 to 16:34
Explores the dynamic relationship between hardware and algorithms in AI development.
“you might ask a model something, the model is predicated or reliant on a single bit of hardware.”
Exciting Investment Milestone
16:34 to 16:45
Announces Callosum's successful fundraising of over $10 million.
“How does it feel to get it done and get it out there?”
Historical Context of Computing Companies
16:45 to 17:55
Analyses the evolution of computing companies and their roles in the industry.
“We're in a scenario in which we're not having to, you know, cloud systems and we're built for an era that wasn't actually multi-agent intelligence.”
Show all 18 chapters
Building a Collaborative AI Ecosystem
17:55 to 20:09
Discusses the initiative to create a collaborative AI lab and its objectives.
“systems level company that didn't necessarily, like, wasn't necessarily fully vertically integrated all the way down to the compute chip.”
Future of AI Funding and Deep Tech
20:09 to 22:46
Explores the shift in AI funding towards deeper tech solutions over application layers.
“And we're still working on ironing out the details of the other partners that we have.”
Challenges in AI Systems Integration
22:46 to 24:41
Identifies gaps in integration between hardware and algorithm companies in AI.
“we've seen ineffable intelligence we've seen olex chip companies are you seeing in your position at ARIA, are you seeing more and more companies like that coming through from Europe?”
The Future of AI Companies and Market Dynamics
24:41 to 28:00
Examines the competitive landscape of AI chip developers and the role of NVIDIA.
“and we're positioning ourselves kind of there.”
The Future of the Chip Market
28:00 to 29:25
Explore the diversification of the chip market and the evolving roles of companies like NVIDIA.
“in the endgame for the hardest problems to be fast, to be good, to be cheap, heterogeneity will just continue, right?”
Breaking the Hardware Lottery
29:25 to 31:39
Understand the concept of the hardware lottery and its implications for chip production.
“It's just that they're going to be part of a much more diversified market.”
Innovations in Chip Technology
31:39 to 33:20
Discover emerging technologies in chip design and networking, including photonics.
“And you must be working very closely with a lot of these frontier chip builders, people who are building for this next generation, next wave.”
Risk and Innovation in Startups
33:20 to 35:28
Learn about the risk profiles of innovative technologies and how they can be supported by ARIA.
“So beyond just pre-filled decoding, but all of the nuances that we'll observe.”
Transcript
Automatic transcript. May contain errors.0:00Suraj Bramhavar:Yeah, great. Maybe, thanks for having me, by the way, and maybe to start, it probably will help if I give a little bit of background for those who are unfamiliar with ARIA and the Scale and Compute Program historically. So ARIA was kind of designed to unlock these scientific breakthroughs that wouldn't happen otherwise, and it's a government-funded science agency. Um, and the scaling compute program in particular, which I was running, uh, or I'm running, I, we started about two and a half years ago with just with the general notion that, you know, there's this existence proof that intelligent systems can operate much, much more efficiently than the way most modern AI hardware operates.
0:58Suraj Bramhavar:And the vast majority of industrial resources are being spent marching down a well-known set of train tracks. So we launched this program called Scaling Compute two years ago maybe, and maybe over two years ago, that its primary purpose was to unlock alternative pathways. We wanted to give people an excuse to explore alternatives. And we felt that there was a vast space of alternatives that could be explored. And we found 11 compelling projects along with a whole bunch of seed projects that we thought were interesting. And we launched that. I think we kicked off work like a year and a half ago.
1:48Suraj Bramhavar:But in the time since, it's become very, very clear that at the idea stage, there's actually a ton of private capital and risk appetite to unlock a lot of these alternatives. And we've seen just in the last year or so, countless companies popping up with hundreds of millions in funding for AI hardware. and some of the ideas have even kind of longer time horizons and riskier ideas than some of the things that ARIA funded. And so one of the things that we kind of asked ourselves is, what is there in the ecosystem that with a 50 million pound program funded by the government, what is the thing that we could be best suited to unlock given the appetite in the private markets.
2:49Suraj Bramhavar:And the thing that became clear was that, you know, it's not necessarily just getting the ideas out from the idea stage into a prototype. That's one step of the process. There is still this massive gap between that and actually reaching global scale of implementation. And so we're targeting this next phase, which is called the Scaling Inference Lab, to attack this exact problem. That's on the institutional side. And then on the technological side, the idea is based on the fact that inference going forward is most likely going to be... Right now, the industry, the way it operates, it's mostly kind of scaling up homogenous, closed, very heavily engineered compute systems, right?
3:47Suraj Bramhavar:You hear all the talk about nuclear power plants and, you know, gigawatt data centers and all this stuff. And really what they're saying is like, we have a known thing that works. we need to figure out how to scale it up as much as possible that means like a lot of energy liquid cooling a lot of very expensive gear in those racks and it's usually the same gear that's stamped out over and over and over again but we're pretty sure that the future is not going to look like that those data centers are going to be filled with more heterogeneous equipment and different models are going to run on different hardware.
4:29Suraj Bramhavar:And what is missing in the ecosystem is a playground, a place for all of the different ideas to come and be tested. And so ARIA is kind of moving forward to create that testbed and trying to design it in such a way that invites as many of those ideas at the prototype stage to come into the fray as possible. And that philosophy aligns almost exactly with the philosophy of Daniel's company. And so Daniel came into the fray as an ARIA seed creator two years ago. And over time, he's kind of independently evolved in this other direction, which just happens to match exactly with where we think things are going.
5:23Suraj Bramhavar:And so it made sense to kind of partner pretty closely with his organization.
5:28Danyal Akarca:Before we go to you, Dan, I just want to like recap, clarify. So Serge, what you're saying is you're building this new lab where you're going to be essentially testing and connecting new hardware elements and aspects with the software that kind of plugs into the hardware. And the big bet that you're making is that the hardware stack that we have now, that as you say, is being scaled and rolled out everywhere, isn't the stack that you isn't the right one you think to be scaled globally to a thousand x you're making you and the team are making the bet that actually the hardware is going to change we're going to have different hardware needs for different bits of software and you're building this lab as a testing ground to match up different bits of hardware different bits of software
6:08Suraj Bramhavar:to see what happens yeah yeah and i i would say i wouldn't even call it a bet so much as to say that I think it's inevitable that this will happen. I mean, if you look through computer systems history, this has happened over and over and over again. You get these waves of technology where everything consolidates around very heavily engineered, closed computing systems. And then you get waves where that starts to get broken apart and the computing systems get more heterogeneous. There's space for more component providers. The interfaces get opened up. The costs drop precipitously as a result. And so this happened in the mainframe era.
6:57Suraj Bramhavar:It happened in the early internet era. It happened with the personal computer revolution. We think it's just going to happen again with AI systems. And we're positioning ourselves to kind of accelerate that evolution.
7:13Danyal Akarca:and right now you think we're sort of at that that time or that stage where we're leaving that first wave behind of the sort of like broad stroke single system to this much more open
7:24Suraj Bramhavar:and diverse range of different systems that's right yeah got it all right well done and and that and that pairs with the the technological gap uh that we know exists between like how much power and cost these things currently require yeah uh to to what we know there are alternatives
7:44Danyal Akarca:which is your which is your point why it's inevitable right because the current the current system the current setup is unable to scale at the level and speed that we needed to to make these things work um all right so i think now dan like you you tell us how you fit into this you've coming out of stealth today what are you building and where does this fit into the border picture. Yeah, totally. Thanks for having me, Seb. So just kind of piggybacking on what Seraj said there, it's important to realize here that it's not just kind of, you know, we can say that we'll go into a world of heterogeneous compute because this is the natural evolution, obviously, of how computing systems have been in the past, kind of from more kind of generalized architectures like CPUs to GPUs and then GPUs to what comes next.
8:23Danyal Akarca:So there's an inevitable trend. But it's important, I think, to maybe take a step back and realize why this is not only going to happen on the supply side, on the computing side, but what are the bottlenecks on the demand side? Why is this just an inevitably better paradigm to go into? Why is it that heterogeneous scale is better than a modular scale? And the reason of this is that the world is inherently very heterogeneous. The problems that we need AI to solve are not naturally solved by just one single type of intelligence. If you think about kind of your work, all of the really hard problems that we actually need AI to be able to solve, none of them typically are kind of single intelligence problems in putting it really simply.
9:05Danyal Akarca:They're typically very kind of complex, multi-step, quite uncertain in terms of the type of decisions that you have to be able to make. And so the world is very heterogeneous. And as a result, it's very amenable to many different types of intelligences. But it isn't just one. This is kind of a little bit different to what we thought probably around 2022 or so. So like around 2022 to 2023, I think there was a general belief that kind of a single model scaled up sufficiently would do really well. And it did. It did unbelievably well. Kind of the homogenous era was great for being able to build what I'd call kind of minimally viable intelligences rather than MVP and MVI, kind of a sufficiently good intelligence to solve many problems.
9:43Danyal Akarca:But now we're kind of going into the era where to solve the really hardest problems well at cost and sufficiently quick. So kind of cost, speed and performance metrics. We really are inevitably moving to a heterogeneous kind of multi-agent system or multi-agent intelligence paradigm where the benefits aren't naturally just coming from making single models better. That's always useful, but it's actually many different models. And in doing so, this is naturally amenable to many different hardware pieces that can accelerate different parts of that. So it's a little bit more of a shift in paradigm from what we call homogenous scale to heterogeneous scale.
10:15Danyal Akarca:And this is just because heterogeneity is ultimately a better way to build your computing systems to scale. And so at Colossum, and we've done lots of work before today in showing this to be the case. And so we've now released some of our earliest discoveries on this. When you actually treat problems as a system problem rather than a model problem, there's actually a huge surface area for optimization. So many of the capability gains that we'll get of the future, we expect to be from kind of systems of models and systems of intelligences interacting across mixed hardware. And we've proven this and shown lots of this and we're super excited for kind of what's to come.
10:51Danyal Akarca:And so Colossum is the company that effectively is building the infrastructure to make this a reality. And so working with Suraj, you can think of us as being the kind of software piece in this where we really are optimizing across that heterogeneity of the hardware. we have to kind of rethink a lot of ways in which we build AI systems and our software to work with heterogeneity. We'll probably get into that in a moment. And so super, super excited to kind of be working closely with Surajinaria and kind of, yeah, which would be the hardware platform ultimately to bring kind of the super exciting inference technologies and networking technologies to be able to bring them to market as well to actualize that vision.
11:29Danyal Akarca:Yeah, so I read in the press release that some of your systems are delivering like double as much accuracy, seven times fast performance, four times lower cost, I think it was. Can you give it a bit of flavor? What does that mean and how does that actually work? How is your system enabling to get benefits from different bits of hardware? Just give a beginner's explanation of how that's possible. Totally. So I use an example of, say, real-world problems. So real-world problems that aren't just, say, a chat interface. So real-world problems such as if you take a photo of your kitchen table and you want to say, I want that box of cereal over there.
12:06Danyal Akarca:I would like that kind of cereal, but I don't, I want it to taste similarly, but I don't want it to have as high sugar content. Can you please automatically purchase that for me? And then I'll have it tomorrow, please. Thank you. So that kind of problem. So I'd call that a more real world problem. It's quite a toy one. That's quite fun. But the point is, is that no single model alone is kind of, can do this kind of optimally. You know, we can actually make inroads in being able to do those problems, but because they're multi-step and multi-turn, there's many different, there's a lot of, how do I put it, reasoning alone does not actually solve that problem.
12:36Danyal Akarca:So despite us thinking a lot about reasoning and agents, it's super important, but reasoning alone actually can't solve those problems. The models actually have to interact with the world to solve them. So you can think of this as actually many, many, many different steps in the chain to actually solve those problems. And as a result, you can make an individual model better. So you'll get a lot of returns from that. That has diminishing returns. You know, we have to put in exponentially more money into the compute to make an individual model better. But where we see kind of the next bottleneck of performance, cost and speed is actually how you build those systems of agents.
13:08Danyal Akarca:So this is all the way from how you bring them together, how you allow them to communicate with each other. So there's lots of things that we've been doing on this and the communication between the models. And then it's how do you enable each of those models to be run across multiple different chips that are specialized for the task that they're actually doing. So this is a story of kind of specialized models, many types. It could be that they're the same model class. So it could be that they could all be Gemini, for example, but different versions, or a combination of open source and closed source, which is what we actually see happening.
13:41Danyal Akarca:And each of those processes can be kind of run on custom hardware that is ideal for this. And so you have different trade-offs of the different hardwares that are there. So what I'm basically saying is that the kind of surface area of what you can do is so vast that you can actually make it economically and economically viable to be able to implement these systems in practice. Now, we still have a lot to do. So, for example, we aren't kind of getting it to this point yet where we think that it's going to be easily adopted for the hardest problems right now. And we need to make hard renovations.
14:14Danyal Akarca:We need kind of new chips that are particularly good at certain processes. And we're seeing lots of great work and investment in this. But this is where we can get these cost, speed and performance gains. is really at the system level rather than just the individual model. Interesting. So you think about, in my head, I'm thinking about the way that it currently works is you might ask a model something, the model is predicated or reliant on a single bit of hardware. And I guess what your system is saying is we can speak to or connect a dozen different models that are all connected to different bits of hardware based on the specific use cases.
14:43Danyal Akarca:And so it opens up, I guess, the breadth of options and the breadth of combinations across software and hardware to create the optimum outcome. Exactly. And so the way I'd call this is a co-evolution process. There's a term in the technical term would be co-design. So people typically say hardware, software, co-design. This is where we know that the magic always happens when the algorithms and the hardware meet. So always the reason why transformers worked at all in the first place was because transformers played really well with GPUs and the highly paralysable and we could scale them up. And that just teaches us that all the magic happens kind of at the intersection of the algorithms and the hardware.
15:18Danyal Akarca:And so what I'm all I'm saying really is, is that we want to speed that up. So it's much more of an evolutionary process and it's much more dynamic. And we shouldn't just kind of build a single chip necessarily, although we can do that single chip for a single model. That's great. But it could be a little bit more of a dynamic process. And so it's this kind of co-evolution of both. There's a kind of a little segue on this is actually this is, I think Serge kind of alluded to it, an existence proof of intelligence. This is kind of how the brain has evolved as well. So the brain really evolved really without a clear distinction between the software and the hardware.
15:48Danyal Akarca:kind of our brains kind of exist both as the hardware and the software and inherently they're kind of co-evolved themselves and so you can think of Colossum really as doing the same type of process philosophically but for our digital silicon that we have today and also readying ourselves for silicon of the future as well which is super exciting. Got it and it looks like a lot of the bet that you are making is on the similarities or the parallels between the natural world and this world of AI and kind of hardware that we're building. I think it's also worth mentioned, I don't think we've actually mentioned it, that you're coming out of stealth today, having raised just over$10 million, which is amazing.
16:23Danyal Akarca:I think the round's been led by Plural. Plural is an amazing VC and their whole thesis is backing GDP altering companies. And it seems like that's definitely what you're building. So congratulations on that. How does it feel to get it done and get it out there? Thanks, Seb. Yeah. So yeah, great. Super, super happy. I mean, obviously we've been thinking about this, me and my co-founder Yasha, and we're working on these ideas for so long. I mean, we met during our PhDs at Cambridge, kind of at the dawn, really, of the scaling era of AI, and over kind of a lot of our work and thinking about this and meeting Suraj and really our ideas and thoughts maturing, we really kind of developed a real sense of urgency around the ambition of the mission that's required in order to, you know, what we're saying really, as Suraj alluded to, is that the real trajectory of computing itself can really be redefined here.
17:12Danyal Akarca:We're in a scenario in which we're not having to, you know, cloud systems and we're built for an era that wasn't actually multi-agent intelligence. And not only that kind of paradigm of multi-agent intelligence, but everything that comes after this period. And so, you know, we've been working super, super hard. We've published lots of our results that are super exciting. It will be releasing. And it's great to kind of be out there and to be able to more easily meet people and kind of show them what we're doing rather than kind of keeping it secrets and well not secret but kind of we've just been working so hard that we haven't actually released anything so now we're super excited to do that amazing and you're obviously sorry go ahead
17:48Suraj Bramhavar:yeah i may just jump on this to to cheerlead a little bit for for daniel uh and get some historical context so in every other era of computing there has always been uh a kind of systems level company that didn't necessarily, like, wasn't necessarily fully vertically integrated all the way down to the compute chip. So if you think about mainframe systems or personal computers or networking systems, inevitably there always emerge these companies that are kind of focused at the system level and not necessarily tied to building their own particular chips. And right now, if you look at AI systems, there are a bunch of systems companies, but they're all very, very tied to their own particular stack or their own particular chips, almost all the way up and down.
18:49Suraj Bramhavar:And when we talk about Daniel's company becoming this like trillion dollar company which i think they have the potential to become like this is the role in which i i would believe like this coming from like a totally impartial uh perspective i think that they could play in the future and it's a it's a very different narrative to the one
19:15Danyal Akarca:i think that we had maybe a couple of years ago which was the almost like a winner takes all market that we would probably see one dominant model come out of top one dominant chip provider maybe it was like OpenAI and NVIDIA. And I guess the bet that you're making now is that actually we're going to have a wide range of models, a wide range of hardware providers, but actually a lot of the value can be captured in the system or the software that sits between it all and connects it all up.
19:38Suraj Bramhavar:I think so, yeah.
19:40Danyal Akarca:Interesting. And look, we're obviously going to be working and have been working closely with Dan and Colosum. Can you talk a bit more about maybe another company or some other companies that maybe you've already selected or you're going to be working with on this program?
19:54Suraj Bramhavar:I'm not 100 % sure we're ready to talk about some of the other things. I will say, so Daniel's company in particular, we are very well aligned in what our end objectives are. And we're still working on ironing out the details of the other partners that we have. So I'm not 100 % sure we're ready to share. Suffice to say that the way it's structured is that we intend to build a cluster, a full AI system every six months in this lab, right? And each one of these is going to be structured as almost like a pilot project. so we intend to build functioning systems on six month cadences and pilot technology from a particular one or two companies per system and so we're mapping out this roadmap now and we have like two or three early partner selected but in short order we'll be able to share who they are hopefully and
21:03Danyal Akarca:what would you hope to see at the end of that six-month period.
21:08Suraj Bramhavar:Ultimately, what we're trying to show is that this, so there's two pieces to this. One is the institutional kind of innovation, right? It's a nonprofit. It's a neutral playing ground. We're trying to demonstrate that that neutral field is attractive to a wide variety of people, including big corporations who want to come see the results that spit out of this, or contribute their own technologies and startup companies who have their own kind of technology theses that they want to see proven. So there's an institutional component of that, just seeing that the kind of audience can be expanded, the audience of companies that we're engaged with.
21:57Suraj Bramhavar:Then there's a technological component, which is very clear. It's like, we're going to build these systems on a six-month cadence. We're going to pick particular technologies that we think are going to insert into our roadmap at that cadence. And we're going to want to show that the cost of these systems can drop iteratively at a rate that is much faster than what you are seeing in industry. And we will very clearly evaluate that. We'll know in a couple of years whether that's happening
22:28Danyal Akarca:or it's not interesting and across the ecosystem i think especially here in europe you know last year it felt like there was a lot of money going in a lot of excitement about ai companies on the application layer and it feels like already at the start of this year we're seeing a real change and a shift towards more money going towards like deep tech but like you know like frontier labs we've seen ineffable intelligence we've seen olex chip companies are you seeing in your position at ARIA, are you seeing more and more companies like that coming through from Europe? You know, there's like solving really hard problems at the deep level.
23:04Suraj Bramhavar:Yeah, I think it's great. All these funding announcements are amazing. I think there's probably more of them coming. And I will celebrate all of it. And in fact, like this pivot of ARIA's activity is designed in response to a lot of these announcements, right? We're noticing that there's capital available both on the algorithm side with companies like Ineffable Intelligence and on the hardware side with companies like Olex. And that's awesome. Now, the question is, how do we harness that? Like, there's a connective tissue that's still required between the hardware companies and the the algorithms companies and as daniel mentioned like the magic really happens when you can kind of co-design those two things and if you look at all of these people there's still a big gap like if you go to the companies receiving all of the funding for neo labs like frontier ai labs most likely they're all still using a homogenous compute stack.
Read the full transcript
24:21Suraj Bramhavar:And if you go to the companies that are raising all of the money on the harbor side, most likely every single one of them is saying that they're going to build the full AI system themselves and tackle NVIDIA head on. And we kind of know both of those are slightly untrue, right? There's something in between that is going to be where the world ends up. and we're positioning ourselves kind of there.
24:49Danyal Akarca:And I guess bluntly, do you see a company like NVIDIA that is based on the single stack, is that overhyped then? Do you think that the future is going to be lots of companies, lots of smaller companies that are building chips, they're going to come out on top or could it be NVIDIA releasing a whole wide range of different use case niche chips? What does that future look like? If the better you're more, the conviction that you have, If that's true, how does that change the current landscape of the models and NVIDIA that we see today?
25:19Suraj Bramhavar:I'll take a crack at that and maybe Daniel can try after me. So to some extent, we're already seeing NVIDIA move in this direction, right? The clearest example is the current transform model is split into two. There's a pre-fill section and a decode section. And NVIDIA is already kind of publicly indicated on the roadmap. Oh, these two, there's different compute demands or different demands on the AI system in these two sections of AI. And we're going to tailor our hardware to accommodate those two sections better. So that's one that's just within NVIDIA's system. they've already kind of identified one split where things are becoming more heterogeneous and they're going after that that's great they've also like started buying up companies like they made a lot of waves with the grok acquisition right that's another indication they're like yes nvidia could go off i mean they have enough money to do whatever they want right they could go off and probably buy up half the companies in the space and build a whole portfolio of different types of chips.
26:34History kind of tells us that that usually doesn't happen, right?
26:39Suraj Bramhavar:If you look back at previous ages where there was always a dominant player at some point in the inflection point of an industry, like over time, that dominant player, their dominance has been eroded by a whole number of economic and technological factors. And so I think it's a pretty safe bet that over time this heterogeneous evolution will erode a little bit of NVIDIA's orchid share. Now, that's a different comment than their business. The demand for this type of stuff is exploding so much and I don't see them going away anytime soon. they're probably still going to be one of the world's most powerful companies.
27:27Danyal Akarca:Yeah, just to add to that, I mean, I'd say that I agree with everything that Srirad said, obviously. And I don't think NVIDIA is going anywhere anytime soon, firstly. You know, the demand is there for ultimately is that over the kind of, you know, over the long term, ultimately it's, you know, the chips are the engine for the future of AI in one way or another. what we're kind of talking about here is that why is the connective tissue so important? It's just because it's better. So, you know, it's better when you have that connective tissue. And so if that's correct, that heterogeneous compute in the endgame for the hardest problems to be fast, to be good, to be cheap, heterogeneity will just continue, right?
28:09Danyal Akarca:And so this is, you know, if that's correct, that's the case. At the same time, you know, there'll just be a demand for whatever chips there are. We're kind of still in the under, you know, in the regime where we don't have enough. And so, you know, we kind of totally envision that while the natural forces will diversify the chip market, that market has a huge amount of potential to gain, to increase. And particularly when we kind of move to every arc, to every kind of paradigm beyond just training. So in training, when you are doing kind of gradient based training, NVIDIA GPUs are unmatched. Right.
28:40Danyal Akarca:So in the training regime, you know, of course, you know, we'll kind of it's a natural monopoly. in that case. But when you move to inference and everything that comes beyond, you know, when you build architectures that can actually do continual learning, when we actually need to be able to have autonomous in the loop scientific experiments where we have chips on the edge as well on the cloud, when we have multi-agent robotics of the future, these will be naturally heterogeneous, right? And so the connective tissue will become increasingly important and valuable and NVIDIA will play a huge role in that entire ecosystem.
29:14Danyal Akarca:It's just that there are new problems to solve that are as important, if not more important. And that's obviously what Colosum is aiming to address, really. So yeah, so the chips are going to be fine. It's just that they're going to be part of a much more diversified market. Yeah, it's interesting. I guess what I'm hearing is that you think NVIDIA will probably lose market share overall in the chip world, but the market will be so much bigger. Yeah, exactly. The analogy I'd make is perhaps it is something like an Uber analogy where you have taxi companies, right? And you have kind of a monopoly of a single taxi company in London.
29:50Danyal Akarca:But, you know, when you have many different customers want different types of taxi journeys. And so maybe one kind of Nissan doesn't doesn't hold up for everything. You need different types of cars. You need different types of vehicles. You need line bikes. And ideally, you need to be able to orchestrate these kind of quite complex workloads, which are the taxi customers in this analogy, to be able to orchestrate them. And pre-fill and decoding, as Suraj said, which is NVIDIA's roadmap, is split, is just the beginning. So we don't even know how AI will be used in the future. Pre-fill and decoding is just kind of a very early example of how on a homogenous stack we can split the tasks.
30:27Danyal Akarca:So the tasks are splitting into pre-fill decoding regimes, which means it's amenable to different specialized GPUs. And naturally, it goes only one direction. We are going to be using AI in many, many different ways. And so, yeah, I do agree with you, Seb. I think that it will be a more diversified market. And this is not just for the current generation of accelerators, which is crucially important. You know, we are very, very excited about many new emerging paradigms where a company like Colossum and the Scaling Inference Lab, for example, we can really break what's called the hardware lottery.
30:59Danyal Akarca:So the hardware lottery, for those of you who don't know, who are listening, is the idea that if you build a chip, it takes a lot of money and it takes a long time. And so you have to be lucky. You have to win the lottery to have the customers there who really want your particular chip once you've built it, right? And so there's this lottery which was coined by Sarah Hooker, which is that NVIDIA basically won the lottery. It built its chips for gaming and it had established a market and then it was ready for the AI boom. And so really to make this co-evolution possible, the connective tissue, you need to speed that up.
31:32Danyal Akarca:You need to not have a lottery ideally. And so we want the chip providers that are built all across the world to be able to have a market and be safe to know that they have a market and breaking that lottery. Amazing. And you must be working very closely with a lot of these frontier chip builders, people who are building for this next generation, next wave. Who do you think, you know, are there companies that you think we should be looking at? Do you think that, which of the companies out there that are building things that are really interesting, like yourself, for this next wave of, I guess, like disruption?
32:01Danyal Akarca:Yeah, so there's kind of two broad categories, I'd say, on the chip side or on the compute side. I'd say. One is kind of the new inference chips that are coming to market. So we've announced with our launch kind of several partners that we're partnering with, including Fractile, but also there are other kind of companies that are doing more new novel paradigms, including normal computing that work on more thermodynamics and so on. That's on the chip side. So these guys are working on kind of de-risking not only their technology for maybe established workloads, But in the case of kind of more, let's say, slightly exotic workloads, kind of in kind of maybe we think of this like thermodynamic compute, for example.
32:40Danyal Akarca:Really, there are kind of clear use cases now that are emerging that we're really excited about, too. It's not only that, though, it's also on the networking side, too, that we're super excited about. So this is how you connect chips ultimately to be able to run. And we're super excited about photonics and this new era of photonic networking. So communicating with light rather than utilizing copper. and we're finding that when you start connecting chips through photonics there's a whole range of new problems and assumptions that we have to revisit and so we're working with a company called Mix on this and we're speaking to many many others, this is just the very beginning and we're super excited about lots of the technologies that are happening.
33:21Danyal Akarca:What we don't want to happen is that they have amazing technology but they're just too early, quote on quote and really kind of making sure that in advance of the technology being built that we can identify and really target the niches of the inference landscape and everything that comes beyond it. So beyond just pre-filled decoding, but all of the nuances that we'll observe.
33:43Suraj Bramhavar:Amazing. There's two points that I might jump on there that Daniel mentioned. One is the fact that it really is a system and that there's innovation happening at so many different layers of that system. So he touched on the optical networking, right? But just the networking of the different compute is such a ripe place for innovation going forward. And there's a whole panoply of companies developing things there. But every one of those individual companies needs a compute chip or a set of compute chips to tie their networking into in order to demonstrate how good it is. so that's it's one example of a place that we want to bring into the fold for the scaling inference lab where we think we can foster more innovation and drop the costs precipitously and then the second bit of that was was uh that daniel mentioned was um there's all these innovative technologies and they all have different risk profiles or timelines to product maturity And that's an area where I think ARIA can slide in pretty meaningfully because we have a very high bar for risk tolerance.
35:02Suraj Bramhavar:We'll take wild bets on exotic things if the reward is high enough. And startup companies with private capital in many ways don't want to do that. So we can take that risk off of them. and in the hopes that some of these ideas won't work, right? But some of them will work spectacularly. And then there's at least a proof point that private companies can then build off of.
35:31Danyal Akarca:Well, look, it's been great chatting. We're out of time. But I think what you two are doing is really cool. And I'm really excited to see how you get on Dan. And Suresh, keep me in the loop of more companies that you end up working with. I'd love to speak to them and cover that as well because it sounds like it's going to be a really exciting time for you. But look, congrats. Best of luck with it both and thank you for joining me. Thank you, Seb.
35:52Suraj Bramhavar:Yeah, thanks, Seb.
From the publisher
AI hardware is entering a new phase as billions flow into chips and frontier labs across Europe.
Danyal Akarca, Co-founder at Callosum, and Suraj Bramhavar, Programme Director at ARIA, are working on what comes next: a £50m Scaling Inference Lab and a system layer designed to connect different models to different chips, pushing AI beyond homogenous scale and towards a more heterogeneous future.
The Scaling Europe show is presented by Deel - check them out here:https://get.deel.com/ruynb7o4lfjk
Sponsors:
SurrealDB: The multi-model database for AI agents. Check them out here: https://surrealdb.com/
Omni: The AI analytics platform trusted by fast-growing companies like Perplexity, Synthesia, and dbt Labs. Check them out here: https://omni.co/
Venture Comet: The platform that gives startups and scale-ups real-time equity tracking, daily business insights and automated management information. Check them out here: https://venturecomet.com/
Timestamps:
0:11 - Background on Arya and scaling compute program
2:10 - Rise of private capital in AI hardware
5:10 - Testing new hardware and software combinations
7:13 - Transitioning to heterogeneous computing systems
9:44 - Importance of multi-agent intelligence paradigm
11:39 - Achieving performance gains with diverse hardware
14:01 - Economic viability of specialized models
17:30 - Excitement about upcoming releases
20:26 - Building AI systems every six months
22:00 - Evaluating cost reduction in AI systems
24:19 - Nvidia's future in AI hardware
27:03 - Market diversification in chip industry
29:11 - Importance of connective tissue in AI
31:01 - Breaking the hardware lottery concept
32:54 - Exciting developments in photonic networking
