#275 Nandan Nayampally: How Baya Systems is Fixing the Biggest Bottleneck in AI Chips (Data Flow)

31 Jul 2025 · 47 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. Podcast Episode Notes

Episode Title

#275 Nandan Nayampally: How Baya Systems is Fixing the Biggest Bottleneck in AI Chips (Data Flow)

Episode Overview In this episode of Eye on A.I., host Craig S. Smith speaks with Nandan Nayampally, Chief Commercial Officer at Baya Systems. They discuss the critical bottleneck in artificial intelligence (AI) relating to data movement rather than just computation speed. Nayampally shares insights on the evolution of chip design, the significance of network-on-chip (NoC) technology, silicon photonics, and neuromorphic computing.

---

Key Concepts and Discussions

  1. Bottleneck in AI: Data Movement
  2. Data Movement vs. Computation Speed: While traditional views emphasize computation speed, Nayampally argues that the real challenge lies in how quickly data can be moved within and between chips.
  3. Scale-Up vs. Scale-Out:
  4. Scale-Up: Involves running larger models with tighter memory coupling among accelerators.
  5. Scale-Out: Focuses on running multiple copies of models, necessitating more data movement.
  1. Nandan Nayampally's Background
  2. Over 30 years in the semiconductor industry with roles at AMD, ARM, Amazon Alexa, and BrainChip.
  3. Experience in electronic design automation tools, chip design, and product management.
  1. Baya Systems' Role
  2. Baya Systems:
  3. Designs chips utilizing network-on-chip (NoC) technology and offers a performance-driven software platform.
  4. Aims to solve the bottlenecks in AI hardware through efficient data movement and modular chip design.
  1. Historical Context of Computing
  2. Overview of the evolution from punch cards to artificial general intelligence (AGI).
  3. Key innovations in computing interaction, including:
  4. Keyboard and screen integration.
  5. Mobile computing and the rise of smartphones.
  6. The transition towards intuitive interactions with devices.
  1. Silicon Photonics and Data Transfer
  2. The shift from copper-based data transfer to silicon photonics for faster and more energy-efficient communication.
  3. Photonic methods are evolving to complement traditional copper connections, particularly in off-chip communications.
  1. Chiplet Design and Advantages
  2. Chiplet Architecture:
  3. Smaller chip components (chiplets) improve yield rates due to reduced size.
  4. Facilitates faster and more customizable designs, allowing for modular combinations of processors.
  5. Reduces overall costs and time-to-market for chip design.
  1. Performance, Power, and Area (PPA) in Chip Design
  2. The interdependence of performance, power consumption, and chip area costs.
  3. Baya Systems aims to optimize these factors through smart networking solutions.
  1. Neuromorphic Computing
  2. Neuromorphic designs mimic brain functions for efficient processing.
  3. Applications in energy-efficient AI systems, especially at the sensor level or in low-power devices.
  1. On-Chip Networking Future
  2. The potential integration of quantum computing and photonics to enhance data movement further.
  3. Focus on optimizing traffic between chips to manage efficiency and costs.
  1. Impact of Software Evolution
  2. The interplay between hardware and software advancements.
  3. The need for efficient software that aligns with the capabilities of evolving hardware designs.

---

Key Takeaways

  • Data movement is the critical bottleneck in AI chip performance, necessitating innovative designs and architectures.
  • Chiplet designs offer flexibility and increased yield, facilitating faster product development.
  • The integration of silicon photonics and neuromorphic computing represents the future of efficient data processing and communication in AI systems.
  • Ongoing advancements in software and hardware must coalesce to address efficiency and performance challenges in the evolving landscape of AI technology.

---

Conclusion This episode highlights the vital role of data movement in AI hardware efficiency and underscores Baya Systems' innovative approaches to overcoming traditional bottlenecks through advanced chip designs and architectural strategies. The conversation provides valuable insights into the future direction of semiconductor technology and AI.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00It wasn't just about the computation that we're building GPUs, but they wanted to say, how can I get this to handle larger models. This is where the whole idea of scale-up came in. Scale-up means I can run larger and larger models, which means there's a much tighter coupling of memory of large number of accelerators and chips that do that. That's orthogonal to what's called scale-out, which says, okay, I'm not just doing large models, I'm doing lots of copies of those models and there are more requesters coming in. But both sides require a lot more data movement and that's where the energy is spent, that's where the costs are, and that's where the bottleneck is.

0:37You better know what you're building ahead of time. Also build it in a way that it can be future-proofed. Hi, my name is Nandan Naambali. I'm Chief Commercial Officer of a series-based startup called Biosystems, and I'll tell you a little bit more about that. As regards myself, I've had about a 30-year career in the semiconductor industry and the computer industry, starting off originally with AMD in designing high-end microprocessors. I spent quite a bit of time at startups doing what's called electronic design automation tools that help you design tips. Then my longest stint was at ARM, which most of you probably know by now, is the processor company that's almost everywhere around you.

1:27Started off originally with smartphones, but they're in cars, they're in servers, they're in networking pretty much everywhere. And I spent a good 16 years there, taking from originally just doing product management to running their biggest businesses. Since then, I've kind of spent a bit of time in a lot of AI companies. I did Alexa at Amazon, taking it into third-party devices. I worked at a neuromorphic company called Brainship, and now I'm at Bayer. So 30 years of semiconductor experience, focus on a lot of the progress in compute. My focus has been on compute as it has evolved, and really about trying to stay in the forefront of innovation and participating in it.

2:15Yeah. And Biosystems designs chips for customers or you design and, yeah. So this is, I think, a very interesting point, right? So Biosystems, I'll give you the top level view and then tell you where we fit into the whole supply chain. So Biosystems licenses what is called network on chip intellectual properties. So they tell you how to do the designs of how you communicate between these various processors that are on your chips. And we also have a very strong performance driven software platform that allows designers to kind of run the types of workloads that they might and help design systems that meet the need.

3:05Now, if I were to step back up to kind of say where we sit in, right, as you know, like the servers are up there running big processing. They're made up of number of chips. Each chip within that itself has number of processors, whether they be CPU central processing units or GPUs or even neural processing units. and they do the computation, but really even on chip, there's a lot of movement of data to and from after the compute to memory back. And that's actually what Bayer does. We design the network on chip, which then gets licensed by companies that build those chips and then it's designed in.

3:55Yeah, and that is the great bottleneck, right? Indeed. stuff back and forth. Electrons have mass. Copper is not ideal. So, yeah. So, in fact, if you look at the challenges that they talk about with AI, what you saw NVIDIA do quite presciently is to realize that it wasn't just about the computation that we're building GPUs, but they wanted to say, how can I get this to handle larger models? This is where the whole idea of scale-up came in. So the idea today of scale-up means I can run larger and larger models, which means there's a much tighter coupling of memory of large number of accelerators and chips that do that.

4:55That's orthogonal to what's called scale-out, which says, okay, I'm not just doing large models. I'm doing lots of copies of those models. And there are more requesters coming in. But both sides require a lot more data movement. And that's where the energy is spent. That's where the costs are. And that's where the bottleneck is. Yeah. And we'll get to it later. But is that where the chiplet, stacking chiplet design comes in? because it reduces the amount of the distance between memory and CPU or GPU. Yeah, I think, Craig, you kind of raised a very interesting point on the trends that are happening, right?

5:39So people understood a chip to be, okay, say if you go 20 years ago, there was a discrete chip that did CPU, and everything that fed into and out of it was on a different chip or in memory. As you started saying, I need to do more different kinds of processing on the chip, what's called an SOC or system on a chip, especially when you start looking at things like smartphones, right? You can't have discrete chips because every time you go off chip, the power draws much more. So a system on a chip said, okay, I need a CPU and then I need a graphical processing unit, and I need a few other things, but they're all on the same chip.

6:25And those started getting more and more complex, and the cost of building these SSEs started going up, you got an idea of saying, okay, what if I build what I call chiplets, which is a specific sub-component that's well-designed, and then I can either put more of them down and build a bigger unit, kind of a Lego building block, or I can use best in class, I can get a CPU chiplet from one and a GPU chiplet from one and put it together in the same package as a chip. Now, what that does is, as you're aware, right, in the silicon world, they talk about what's called yield because every wafer that chips are made has flaws.

7:15and the larger the chip gets, the more likelihood that it may have flaws, right? So going to a chiplet model where you're doing something in a smaller format makes the yield higher. I see. Second thing is you can reuse them and you can Lego block into different ways of building end products much, much faster. And so as we get into the AI world, what that means is I can get the best-in-class networking component. I can get the best-in-class neural processing component. All of that I can put together into a chip and move much faster based on the needs that my end customers have, my workloads have.

7:58Yeah. You had suggested that you sort of go through the history of chip design starting with punch cards. I'd be really interested to hear that, that I think listeners would too. can you do that for us? Sure. So I'll just talk about it in terms of waves of computing, if you will. So obviously the first giant computer was built with vacuum tube transistors, etc. Which were today, that whole room would be in a very small microcontroller. That's about 4 millimeter square to 8 millimeter square. But the bigger question was how you interact with your computation. That's kind of the interesting story here, right?

8:46And AI is kind of a continuation of that wave of computing. So originally, if you look at it, you had to, the big step was to actually get to computers with punch cards. Before that, it was voltage levels that you had to set to kind of trigger it. So punch cards became a way of interacting. And then you came out with printers or punch cards that came out and you had to kind of manage that. Then you took a big step by adding a keyboard and a screen to the computation. So then the interaction improved. And then obviously as you got mobile, you got laptop computers which were interacting. and you started integrating different capabilities into your computers.

9:35If you look at two decades ago now, I can't believe it, that's nearly two decades ago, you started moving into a smartphone where now you were carrying your computation and it was connecting to the internet. And within that you had not just the basic computation, but you had camera, and then you had gyroscopes and things like that. So where you could use other aspects of not just a keyboard and mouse and interaction through somewhat less intuitive modes. You are getting into much more of a personalized mode where your emotion or your view or your voice could start making a difference. over the last kind of five to ten years you started seeing augmented and virtual reality type compute where it is getting much more interesting alexa was so popular because you didn't have to write you didn't have to program you didn't have to type you could speak right and with these ar vr glasses you can now engage build your own realities as you go along And each time what has happened is there's been new types of computation that are more specialized, interacting with more traditional, and the software kind of starts morphing, starts changing.

10:56And then you look at the next wave, which is a continued wave from all of that, which is more intuitive, where now things like chat GPT with mic input and speaker output, it's actual conversation. right it's not just dumb in terms of taking inputs and doing things now it's interacting much more intelligently with you so the way of computing keeps going in and of course the idea is at some point you'll get to AGI yeah right artificial general intelligence where now it's beginning to start designing things on its own building the next level of computation by itself so we're somewhere on that very interesting and potentially scary or positive trajectory yeah and in terms of the uh the networking on on the chip uh it's it has uh been been metal right and uh copper some some uh alloy uh to to to move uh between memory and compute and and i people have started uh talking about using uh light using photons uh and and now you're talking about this chiplet design.

12:30What is the networking? Can you talk about how that's evolved? Yeah, I think it's a fair question. I think this is a very evolving methodology. So if you think about it, on chip itself, it's still, for most part, the silicon portion still, it's kind of metal, the traditional metal style. But as you start getting off that and going to the rest of the chip, you're beginning to get many more optical options. Originally, the optical was more between large boxes where you could handle the power. Silicon photonics has moved a long way in the last couple of decades where, while you may not do as much on die yet, But between dyes on the same substrate, you're beginning to see a lot more.

13:25So now what was stopping it before was just the power and the cost of it that wasn't compelling enough. Now you're beginning to see more of that actually happening where going off dye is not as expensive with photonics and the speeds are kind of managed. you're beginning to see some of this actually transition to also, hey, I can make the copper on the PCB start working like it's just an extension of my chip or chiplet. So there are a lot of disruptive technologies coming in that are using new techniques, but also utilizing old materials and making more out of them. So I think you'll see a lot more coming out in the future that are optimized between photonics, silicon photonics, and other optical techniques that help with data movement and make it not just fast, but cost effective and energy efficient.

14:28Because that's the big problem that we'll always face. Cost, energy efficiency, and of course, because the demand for performance keeps going up. Yeah. Can you talk a little bit about the difference between copper and silicon photonics? I frankly don't understand how that works. Can you talk about that transition and whether photonics are good for one solution, copper is good for another, or photonics eventually going to overtake copper? All right, question. So as far as I can tell right now, for the core system on chip, it's still the same metal stack, right? But it's really at the edges that you're seeing the changes.

15:22So as I said, silicon photonics is beginning to pull more on dye, and now it's being able to drive that better off chip or between chips. And integrating that into more silicon microcircuits, you're beginning to get this hybrid between copper and cross dye. That's kind of where silicon photonics helps. I don't think it'll replace that on dye anytime soon, right? But really, its real value is in making dyes that are far away or further out. feel like they're an extension of this dye. Right. And how, so on dye, you're talking on the etched silicon, there are tiny copper connectors between transistors or between from the processing units and the edge so that it can connect to another chip or memory or something.

16:41And you're saying that the photonics has not replaced that or cannot replace that? Yeah, right now, you're right. So if you look at a piece of silicon that has, finally a processor is also a collection of millions of transistors. And transistors are doped portions of silicon that are then connected through each other with various stacks of metal, and you've got these cross grids on the chip running, and that do the connectivity. So the photonics has not replaced that on the die because it's still way more cost-effective and simpler to do it that way. But usually the challenge happens when you go off the die, but it has to be driven by the edge of the die.

17:32to the next die. And there, the driver for these larger distances on copper, on the PCB, let's say, or the substrate beyond that, starts becoming either slower or more expensive. That's where originally photonics was too expensive for that, but has started now coming to the point that in terms of energy as well as a speed that becomes much more relevant there. Yeah, and how is that achieved? You need a light source and you need presumably, you know, a fiber optic thread. Yeah, how does that work? So that actually gets into probably a little too much detail, but there are, you know, they're called photonic integrated service, so a combination of silicon as well as some silicon-generated optical transitions.

18:47And then there's other technology that says, hey, there are waveguides that connect between photonic devices in the silicon core and to a different kind of rib, if you will, that connects them together. And so to some extent, the light sources are really lasers, right? and they're effectively the ones that guide that communication on the photonic circuit. I see, and it's through a guide, meaning is that an etched, what is a guide? I've heard that term, but I actually don't know what it is. I'm trying to see how we keep it simple, right? because I'll be the first to admit I'm not an expert on the silicon photonic side.

19:44But really think about it as there's a silicon on insulator and a substrate at the bottom, and then there's a silicon oxide at the top. So that creates a guide within which the light travels. I see. Okay, so it's not— Think of it as a simple channel. within which the light travels. Right. It's not traveling through a filament of glass or something. No. That would be way too expensive and difficult to etch into those geometries. So effectively, today, we're getting to what's at, I mean, well below kind of three nanometers now. We're into angstroms. Yeah. So the waves of light now are getting smaller or larger than the actual distance between transistors.

20:42Right. Yeah. Yeah. That's fascinating. So BIA systems is, are you, well, tell me what you guys are doing today. Fair question. So thanks. You know, I love the historical discussion on this part of the conversation. It's very easy to get dragged in. We're, after all this discussion, I feel we're not so sexy. But at the same time, we're solving real problems, right? So what happens is if a CPU, right, a central processing unit, is all about control flow, right? So it's like, hey, get this, then do this, or do that, then do this 100 times. Generally, most programs go that way. And so what makes a CPU performance go high is, okay, as soon as it needs something, it needs that there.

21:47So processors are set up with caches that are much faster memory, but much more expensive memory than layers of those that get less and less expensive and larger and larger. And then, of course, you go off-chip to what's called DRAM or high bandwidth memory, HPM, and so on and so forth. For a GPU, it doesn't care so much about if that was. They're funneling large numbers of data through similar operations constantly. It's called data flow or data plane processing. Neural processors take a slightly different approach that say, okay, I'm going to do it. I'm still doing lots of matrix multiplications, but I'm going to treat it more like a neural model.

22:34Each of these has a very different characteristic and often needs a combination of the two of them to, two or three of them to work together. And so there's a lot of interaction between these things that go across. And so the data movement between them has to be smart. It has to be efficient to make their performance better overall and give them the scale to go across. So what has been consistently happening, and you've heard this, like compute's not the problem, the network's the problem. And the network today doesn't mean everything that's telecom. It means on chip. It means across die. It means truth.

23:18So what BAYA does is build a transport layer that can be implemented into silicon and a layer on top that can support this type of protocol, that type of protocol. So it makes it easier. So you can reuse wires. You can reuse logic and make that work much more efficiently. plus a lot of the types of operations you need from one processor type to another, we can do it more efficiently and make it more customizable very quickly. Now, add on top of the fact that today, if you start designing a chip and it's going to take you$15 million to implement it in the latest technology, it's going to take me 12 to 15 months before it gets out and then another six months before it hits market.

24:09That's a long time. So you better know what you're building ahead of time. Also build it in a way that it can be future-proofed. So we build in the software platform, the ability to analyze what you're building, how it's going to work with all those types of workloads ahead of time. So when it comes out, it not only does what it's supposed to do, and then you have built-in capabilities into the system that say, I'm monitoring all of this. I can fix this. I can make it better. I can optimize it when the chip is already out in the field, so on and so forth. So if you think about what AI has done, right?

24:50AI has really infiltrated all parts of computation. So we think of AI, oh, it's in the cloud because it used to be. Not really. When you look at the network that's connecting all that with 5G and 6G, the kinds of operations that they need to do in order to fully utilize the network to get there faster, it still needs AI. Then you get to a car. The car itself has localized AI. They can't wait because of critical aspects. You go back to cloud, so you have to lead locally. On your smartphone, you have AI, again, for responsiveness. If you start looking at healthcare, you have monitors. So these AI-capable systems are everywhere.

25:42It's not just a system or a class of system. Almost every system being built today is going to need that, and it's going to need efficient compute channels and then efficient communication. So we kind of look at supplying those, And because we are flexible and customizable, a person designing a chip for a car, let's say an automotive or autonomous drive system, or building a network for 5G, they can use our technology and build very different looking chips, but with the same level of transport that they need that is customized to that. Yeah. And the variables that you're working around are speed and energy consumption and presumably heat, right?

26:44Indeed. Yeah. So if I look at, in our world, they talk about PPA, which is like performance, power, and area. Those are the same things. Power ends up being a proxy for heat, right? Because as you burn more of it, the more thermals are generated. Area and power then become a proxy for cost because the more silicon area you take, the more expensive it is. The more heat power you take and more heat you generate, the packaging has to be more expensive. So those three kind of capture the major aspects of it yeah and and on the uh and and this is in to reduce the power uh consumption and the heat uh generation uh that's where you guys come into play right that's in the networking.

27:46Yes. So obviously each processor component also is a great source of energy or power generation or power consumption, right? Yeah. But then when they start moving data around, data to memory, et cetera, that's also, it's growing a lot in terms of what that consumes or costs. And what we end up doing also from a silicon provider's perspective how can we reduce the number of people needed to build these complex systems? How can we reduce the cost of design time? How can we reduce the time to market? All of those ends up being part of your kind of balance sheet. Yeah. And the more we can do it on every angle here, we're making it better for them.

28:42Yeah. So if we can reuse number of wires because we can do it smartly, then that's a reduction in both area and power. And those wires are driven by logic. If we can reduce the logic that's going there, that's a reduction in cost. And by knowing what we want to send, how we want to send it, by doing that analysis up front, you can be smarter about how you put all of that together. So kind of it's layers of things that we do to try and make it more efficient constantly while delivering the performance, the responsiveness or latency, while cutting down cost, while cutting down power, while cutting down time to take it to market.

29:25Yeah. And do you have sort of a standard technology solution that you're applying in whatever design you're working on? Or is it, I mean, just moving from metal to photons, is it changing as rapidly as everything else? So you're creating new ways to network chips together? Fair question. And this is interesting because as we kind of dive deeper, it gets more complex and murky. So what we work on is more standard design paradigms, mainly because that's the maximum that we can help customers get to market. So we are still working on the boundaries of what we call the copper or the metal stack within a chip or chiplet.

30:41But we are partnering with specialists that do silicon photonics, partners that do on-die, off-die communication, because if we can combine their expertise with our know-how, right, Think about it as, okay, they build a really efficient fat pipe. We know all the things that are going around on each die. If we can be the traffic cops that optimize how the things move across each other, then we're making best use of the pipes. And we're having headache on both sides. Instead of chaos, there's more organization. Yeah. Do you see a steady decrease in power consumption with each generation of chip, or is it going the other direction because chips are becoming more powerful?

31:38So this goes back to the history of computation, and you've probably heard the term Moore's Law and Denard scaling, right? And the idea was every 18 months, you'd have twice the number of transistors and potentially higher frequencies and lower power. About 10 years ago, that curve started tapering very quickly. Yes, you were getting twice the number of transistors, but because the voltage gap now, what used to be 3 volts came down to 2 volts, came down to 1 volts. now we're all in that 0.7 to 1 volt of the last few generations. So, and you know, power is a square function of voltage. If voltage is not scaling down, you're not getting any benefits, right?

32:29So, in fact, the CTO, the chief technology officer, founder of ARM, used to say the idea of dark silicon, because there are portions of the chip you just can't afford to turn on all at the same time because it'll take too much power. Yeah. Right? So in terms of consumption or the capabilities from silicon processes, it hasn't, it started tapering off in terms of the need from a performance standpoint are constantly going up, right? So I think what's happening now is finding these alternate ways of saying in sort of the traditional compute models and the traditional communication models? What if I did a more optimized types of processors connected in a different way?

33:17Can I reduce power? Can I reduce cost? And so there's a lot more innovation happening on software because obviously I can write software in a way that makes it more efficient. I can build the chips in a way that are using these different elements in a more efficient fashion, horses for courses, if you will. And then I have it connected in different ways. I can build boxes different. So power reduction, energy consumption, et cetera, are not going to be just, oh, my chip is much better. It's starting to get the entire software stack, the silicon stack, all of that's playing a feature. So naturally that means that the cost of these next generation platforms is much more expensive because you have to think much bigger.

34:11Yeah. A couple of things. You know, I've talked to Andrew Feldman on the podcast about these, you know, Cerebris's wafer scale chips. And one of the issues there is managing the heat. Do you guys work with them or have you worked on wafer scale? Or you're not working on the dye, you're working at the edge. Is that right? No, no, we work on dye, but we're still working more traditional dye as opposed to wafer scale. So certainly Cerebris had to take a lot of very interesting new concepts to build wafer scale. You talked about heat dissipation is a huge issue at a wafer level. But think about delivering power from the edge of the wafer to the central dyes of the wafer.

35:18very different than if you're operating at just a one by one centimeter die. So then it is guaranteed that a wafer will have flaws. So how do you overcome the flaws? Because you can't say, I'm going to, because every wafer is, so they've done a lot of very interesting new techniques to address wafer scale. could our technology be used in wafer scale because finally we can license that uh it's not out of question but they have a lot of proprietary work there that they may have to overlay on our fabric to go do that right and then the other direction is as you were saying uh chiplet design so you can compose larger chips from chiplets or you can stack chiplets What's the logic there in going that direction?

36:20Fair question. And I think you can see the success of companies like AMD that have done it. NVIDIA has kind of done it as well. Intel used to do it in what they used to call multi-chip modules before. So the idea there is multifold. So one, as I said, if I can actually build smaller chips, the yield of those goes up just because if you have the same number of flaws, have fewer dye that are affected by the flaws. Two, I can do smaller designs faster, and if I can actually Lego block them, I'd have the ability to package, let's say I build a four-core processor chiplet, and then I can package it with my IO or input-output around it and sell it as a small-end four-core processor.

Read the full transcript

37:08Or I could put four of those chiplets together and then sell it as a 16-core processor. So I can serve different markets a lot faster because if I had to do the 16-core design as a separate system on chip, the cost of using that design development, the mass cost to actually pop it out and then as a separate test would be much higher, right? So there is a yield benefit that you get out of it. There is a packaging and agility, if you will, benefit that you get out of it. And then you can actually, what happens is, you know, when I'm doing my input-output, what I call dryers, they don't need to be at the latest, greatest process node, which is very expensive.

37:58The compute element will need to be. So I can pick different cost points for the different components that I need. and change my economics of design. So chiplets actually come up in a very different way to help you compose, build, and then cost manage next generation products. Yeah. And what's the concept or the philosophy behind stacking chiplets as opposed to laying them on a horizontal plane. So to be fair, right, when we talk about stacking, it's really today, chiplets are still kind of, they sit on a certain substrate, right? So it's still, you could argue they're flat, but now we're getting into 3D.

38:50So you can actually, then it becomes a question of

38:57using all dimensions to your advantage, knowing that there are challenges of form factor and packaging that go with it. So you're not physically putting one chiplet on top of another and connecting them? There are, right? So there are certain things that do. So, for example, memory, especially if it's kind of connecting memory chiplet versus amazing chiplet, you can do it that way. So there are lots of different ways that you do that stacking as part of it. But really, today, if you're looking at traditional models, you're putting them on the same substrate. I see. Yeah. And so where do you see networking design or on-chip networking design going?

39:48What are people looking at as the next breakthrough? It's possible that when we get to quantum computing, we'll see what that means. Right. Because finally what matters is a tooling and a methodology that can give you reliable results, whether it's quantum, whether it's silicon photonics, whether it is CMOS, whether it's gallium arsenide. So really it's a question of when the processes of design and the methodology to implement come together, right? That's when it becomes more mainstream. The concepts that we have for what we're doing with networking, right, are designed for a single, for silicon, right?

40:42But concepts can be translated if that became silicon photonics at a later point in time. There's nothing stopping us from utilizing that. I do think that photonics and optical is coming closer. Optical computing is coming closer. And when that comes closer, obviously, data movement is going to be part of that thing. yeah and and you worked uh at a previous uh organization you were saying on neuromorphic i'm i've been fascinated by neuromorphic i had on the program a guy in australia who's building a brain scale on neuromorphic computer and what are the applications today for neuromorphic chips Great question.

41:37So I kind of think about neuromorphic is really a way of doing neural net processing or AI processing. So the applications could be endless, but here's where it helps. So the whole idea of neuromorphic computing is to compute like the brain, right, which is it's not running lots and lots of matrix multiplies and parallel irrespective of whether they're used or not. Really, the brain is probably one of the most efficient processors of intelligence, if you will. And that comes from the fact that it only fires the neurons that actually need to be fired. And neuromorphic takes that principle. So probably the best use, if we do that right, is at the sensor level.

42:35So if I've got a sensor and I've got a very low energy draw for doing the intelligence around it, that would be great. But it's the same concept could be applied at a data center and saying, instead of taking megawatts of power, I can do it in kilowatts or less. So neuromorphic, think about it as a paradigm that can affect any level of AI processing. Where I think it's especially relevant is where energy efficiency is critical. So if I'm working on a heart monitor or embedded heart monitor that I want to use, I don't want even batteries on it and it can use local chemical power generation effectively within, you can do that.

43:28You can do, you know, just as you have watches that kind of use kinetic energy to convert to energy, you could do the same thing there. So neuromorphic is going to be broadly applicable. The key question is how you actually implement it in a way that is easily deployable. That's where the problem is so far. And then finally, it's software. As the software ecosystem improves for neuromorphic computing, you will see these applications get more prolific. Yeah. And on the software, I was just at an IBM conference, and they're focused on building small models because in many cases, you don't need these LLMs, these big, giant frontier models.

44:27And you can get the models closer to the edge and that sort of thing. uh how do you see is is that something you're seeing uh you know putting the model on the chip uh getting it closer to the edge is that something that that you guys uh work on so what we see right finally uh if you in encode it in hardware right the advantage is it becomes more efficient for what it does. The disadvantage is you can't extend it or change it easily. Right? So what I see things like LLMs do is they can keep evolving, but really it's a question of building models in a way that the underlying hardware can use it much better.

45:21So Neuromorphic was kind of part of that. If you've heard of SLMs, et cetera, you can also start thinking, you're probably seeing these companies like Cartesia, et cetera, came out of Stanford that are doing what is called structured state space models. Yeah. Right. Right. And so structured state based models are the same. So I think, I think we touched on this maybe 10 minutes ago, but really every evolution that you see in terms of a kind of a step change in efficiency comes from either, Hey, I found a better way to do hardware or I can found a better way to do software. So it was originally CNNs and then it became transformers and then it became LLMs and neuromorphic versions of some of those things will come in as well.

46:12So I feel that more likely than not, the models will keep changing a lot faster. And the hardware will try to find ways to make that consistently more efficient to use. And if there's a common format, you will start seeing that evolve.

From the publisher

What if the biggest challenge in AI isn't how fast chips can compute, but how quickly data can move?

In this episode of Eye on AI, Nandan Nayampally, Chief Commercial Officer at Baya Systems, shares how the next era of computing is being shaped by smarter architecture, not just raw processing power. With experience leading teams at ARM, Amazon Alexa, and BrainChip, Nandan brings a rare perspective on how modern chip design is evolving.

We dive into the world of chiplets, network-on-chip (NoC) technology, silicon photonics, and neuromorphic computing. Nandan explains why the traditional path of scaling transistors is no longer enough, and how Baya Systems is solving the real bottlenecks in AI hardware through efficient data movement and modular design.

From punch cards to AGI, this conversation maps the full arc of computing innovation. If you want to understand how to build hardware for the future of AI, this episode is a must-listen.

Subscribe to Eye on AI for more conversations on the future of artificial intelligence and system design.


Stay Updated:
Craig Smith on X:https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI


(00:00) Why AI’s Bottleneck Is Data Movement
(01:26) Nandan’s Background and Semiconductor Career
(03:06) What Baya Systems Does: Network-on-Chip + Software
(08:40) A Brief History of Computing: From Punch Cards to AGI
(11:47) Silicon Photonics and the Evolution of Data Transfer
(20:04) How Baya Is Solving Real AI Hardware Challenges
(22:13) Understanding CPUs, GPUs, and NPUs in AI Workloads
(24:09) Building Efficient Chips: Cost, Speed, and Customization
(27:17) Performance, Power, and Area (PPA) in Chip Design
(30:55) Partnering to Build Next-Gen Photonic and Copper Systems
(32:29) Why Moore’s Law Has Slowed and What Comes Next
(34:49) Wafer-Scale vs Traditional Die: Where Baya Fits In
(36:10) Chiplet Stacking and Composability Explained
(39:44) The Future of On-Chip Networking
(41:10) Neuromorphic Computing: Energy-Efficient AI
(43:02) Edge AI, Small Models, and Structured State Spaces

More from Eye On A.I.

All 266 episodes
#275 Nandan Nayampally: How Baya Systems is Fixing the Biggest Bottleneck in AI Chips (Data Flow)Eye On A.I. · 47 min
Listen in VO