In short
Great Sky (seed-funded, spun out of NIST) argues the next AI leap may come from new brain-like hardware, not bigger models. Their Superconducting Optoelectronic Networks (SOENs) aim to beat GPU limits by eliminating the processor-memory “von Neumann bottleneck” and using superconducting analog computation plus optical (light) communication.
Guest backgrounds
Jeff Shainlein is co-founder and CEO of Great Sky. Great Sky is based on superconducting electronics work originally developed at NIST; the company is partnering with IMEC (Belgium) for manufacturing scale-up. Host Corey Knowles and co-host Grant Harvey discuss the ideas on Neuron AI Explained.
Key claims
SOENs use Josephson junction superconducting circuits to implement neuron operations (weighted multiplication, summation, threshold/saturation activation) and store synaptic weights in-circuit. Neuron outputs are encoded as single-photon optical signals detected by superconducting single-photon detectors. They operate at ~4 Kelvin (about 200x warmer than typical superconducting quantum systems), enabling data-center deployability. They report ~100x speedups for small language models via knowledge distillation, and early chip benchmarks up to ~10M video frames/sec (target 60M).
Notable examples
Google Speech Commands digit recognition; video/image digit classification; a DOE/Brookhaven fusion tokamak sensor-control demo (megahertz, ~1M data points/sec); proposed particle-collider anomaly detection and X-ray diffraction reconstruction; content moderation and ad-sequence modeling for hyperscalers.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroduction to Brain-Inspired Computing
0:00 to 1:00
Learn about the concept of brain-inspired architectures in computing.
“He wrote a book called The Computer and the Brain that was all about how we should be moving more towards brain-inspired architectures.”
Understanding Great Sky's Vision
1:22 to 3:41
Explore the vision behind Great Sky's brain-like computing technology.
“about a very unique and ambitious question.”
The Shift from Traditional to Brain-Inspired Computing
3:41 to 7:58
Discover the key differences between traditional computing and brain-inspired architectures.
“So how is that different from, let's say, for folks who know traditional computing or they know what a GPU is?”
The Role of Optical Communication in AI
7:58 to 9:12
Learn how optical communication is transforming AI computing paradigms.
“He wrote a book called The Computer and the Brain that was all about how we should be moving more towards brain-inspired architectures.”
Comparing with Quantum Computing
9:12 to 13:18
Discuss the similarities and distinctions between the new technology and quantum computing.
“Well, okay, so optics and computing has a long history, and let's dig into that a little bit.”
Superconducting Circuits and Their Applications
13:18 to 14:00
Understand how superconducting circuits are utilized in new computing technologies.
“The circuits are slightly different, but they have a lot in common.”
Communicating with Light at Low Energy
14:00 to 16:56
Learn about the use of superconducting single photon detectors for efficient communication.
“Light is the fastest thing in the universe.”
Superconducting Optoelectronic Networks Explained
16:56 to 21:41
Discover how superconducting circuits mimic neuron functions for computation and communication.
“The cryogenic infrastructure to cool those is totally data center compatible and it exists today.”
Running Language Models on New Architecture
21:41 to 24:22
Explore how existing language models can be adapted for new hardware architectures and the implications for speed.
“between this architecture and the language models or all kinds of models that we're used to today?”
Revisiting Neural Network Architectures
24:22 to 28:00
Examine the potential of recurrent neural networks and state space models for more efficient computation in AI.
“So you're basically converting it for your use.”
Show all 22 chapters
Innovations in Neural Network Architecture
28:00 to 30:20
Learn how new hardware mimics brain architectures for better AI efficiency.
“And so they've gotten a lot of attention, no pun intended there.”
First Generation Chips and Their Performance
30:20 to 33:20
Explore the capabilities of first-generation chips for high-speed inference tasks.
“So let's talk about your first generation chips a little.”
Applications in Fusion Energy and Control
33:20 to 35:35
Discover how AI chips are utilized in fusion reactor data processing and control.
“This is a common benchmark task, the Google speech commands.”
Scaling Production and Manufacturing Advances
35:35 to 37:42
Understand the manufacturing process and partnerships that facilitate scaling chip production.
“So what's what's between you and the million parameter version?”
Commercialization Pathways and Use Cases
37:42 to 42:04
Explore different commercialization strategies and potential applications for AI chips.
“So we're building multi-chip modules that allow us to scale far beyond what can fit on a single die.”
The Future of Video Comprehension and AI
42:04 to 44:41
Learn how AI can improve video content moderation and advertising processes.
“And so in this case, picture your YouTube.”
Great Sky's First Delivery and Future Goals
44:41 to 46:00
Discover Great Sky's plans for chip delivery and long-term visions for AI integration.
“So I saw that Semaphore reported that Great Sky plans to begin delivering chips sometime later this year.”
Manufacturing Challenges for Superconductor Chips
46:00 to 49:08
Explore the manufacturing requirements and challenges of new superconducting chips.
“We can take the power right from the ceiling, the cooling water right from the wall.”
AI's Role in Developing Superconducting Technology
49:08 to 52:35
Learn how AI is utilized in the development and commercialization of superconducting technology.
“convince them that there's a demand for the volume and that's really where the challenge is yeah because there's a there's a big fight right now trying to keep trying to get their capacity Absolutely.”
Memory Innovations in AI Systems
52:35 to 56:03
Understand how new memory approaches in AI systems eliminate traditional bottlenecks.
“So the first one is the memory side of this, right?”
Scaling AI with Photonic Interconnects
56:03 to 58:05
Learn about how photonic interconnects can enhance AI model scalability and efficiency.
“because like, I'm so used to thinking of the limit on what type of model I can run in terms of like the amount of GPU I have, right?”
Innovative Approaches to Exponential Scaling
58:05 to 58:56
Discover how new technologies enable exponential scaling in computing without traditional bottlenecks.
“And so you can keep adding more compute to the system, keep doubling, keep doubling, have exponential growth in a completely different way.”
Transcript
Automatic transcript. May contain errors.0:00He wrote a book called The Computer and the Brain that was all about how we should be moving more towards brain-inspired architectures. The most important thing about the transformer is that it basically allows any input to modulate the attention given to any other input. That's sort of the most general utility you could possibly ask out of some neural network type function. And that's what makes it so powerful. We're at our initial product generations right now. we want to assess how quickly can we process video how quickly can we process audio it's 60 million video frames per second we build neurons and synapses like the ones in the brain connected like the ones in the brain but everything in our hardware happens about 2 million times faster than biology the most important thing is the elimination of that data movement bottleneck between processor and memory and that's just completely gone and it turns out that moving to superconductors is the key enabling step there.
1:00Welcome humans to the Neuron AI Explained. I'm your host, Corey Knowles, and I'm joined as always by the one and only Grant Harvey. How are you today, Grant?
1:08Jeff Shainline:I'm good. I'm good. You didn't throw any fastballs at me today in terms of my introduction. Today the curveball is that I didn't throw a curveball. I like it. You keep me on my toes. Well, today we are talking with Jeff Shainlein, co-founder and CEO of Great Sky, about a very unique and ambitious question. What if the next big AI breakthrough is not a bigger model, but a totally different kind of computer? Great Sky recently came out of stealth with 14 million in the seed round and a hardware architecture called Superconducting Optoelectronic Networks, or SOENs. That sounds technical, but we will break it down and explain it for regular folks, so don't be scared.
1:46Jeff Shainline:The company says the tech combines superconducting computation, optical communication, semiconductor control circuitry, and brain-inspired architecture to push AI compute beyond today's GPU data center roadmap. Jeff, welcome to the Neuron. We're so excited to have you. Awesome. Thanks for having me. I appreciate it. Excellent. Well, I guess let's start kind of at the top here. When you say Great Sky's Building, brain-like computing, what should we picture in our heads when we see that? And maybe what should we not picture? Okay, well, I guess the simplest high-level concepts to keep in mind are that brains are made out of neurons, little cells that interact in complicated ways.
2:31They sort of nudge each other, and each neuron typically makes thousands of connections. You can imagine a neuron is a little cell that has thousands of arms that reach out, and they get excited when their neighbors get excited, and when they get excited, they reach out and nudge them. And so activity is moving around these networks all the time. And so in order to capture that in hardware, you need circuits that do those dynamical activities. And another important part of it is that when those neurons reach out and nudge each other, the strength that they push on each other with is that's called the synaptic weight.
3:08And keeping track of all those synaptic weights, that's one of the most important memory functions in a neural network architecture. So you need to be able to have these complicated networks of neurons that are always dynamically interacting and have the memory that stores those synaptic weights distributed throughout the entire network. So those neurons, there are thousands of connections and the ability to store those synaptic weights distributed across the network. Those are some of the key principles that we're thinking of when we say brain-inspired computing.
3:41Jeff Shainline:Wow. That's awesome. So how is that different from, let's say, for folks who know traditional computing or they know what a GPU is? And maybe we could preface, like, what is a GPU and how is this maybe different from that? Or is this on the chip level? Is this the full computer? Give us a little more context on that side of things. All right. Yeah, makes sense. So I guess in the most stripped down description of a conventional computer, there is what's called the von Neumann architecture, named after John von Neumann, who played a huge role in computing in the early days after World War II, 1940s.
4:18And so the basic idea is that you have a part of your computer that's the processor, the arithmetic logic unit. It's just doing a lot of math. It's doing addition, multiplication. and then you have the memory and the memory stores intermediate results from all that math. It stores the program you're trying to run. And so this separation between a processor over here and memory over here is sort of the defining characteristic of that architecture. And it makes a lot of sense when you're trying to build a computer either to be really good at math, which is what they were doing in the World War II era, or just be very general, able to run any program, which is necessary now.
4:59You need your computer to be able to run Microsoft Word and browse the internet. So that computer architecture makes sense. So that's sort of at the architecture level. And then at the most fundamental level, there's how do you represent information? And one of the choices that wasn't obvious from the beginning, but sort of emerged over time was that you want to represent information in binary digits, zeros and ones. And silicon transistors do that very well. They're little gates, they're either on or they're off. And the reason why that whole architecture and that fundamental representation of information sort of evolved is because it was very resilient to noise.
5:39As I said, the architecture is very general purpose, but the binary representation of information, it leads to very low errors, which is important for long computations where you're churning through lots of numbers. Okay, so now we live in this era of AI, where actually you're trying to do something different, where you're not just churning through lots of long computations, but instead, you actually want this really fast, what we call the inference process, where information comes in, picture yourself, you're a human, you see something, and you need to immediately be able to know what it is, you don't need to, you're watching a bird fly across the sky.
6:15You don't need to compute the trajectory of that bird. You just need to intuitively know where it's going through an inference process. It doesn't need to be highly accurate. You don't need 12 decimal places of accuracy like you do if you're designing the wing on an airplane or something like that. And so just completely different principles apply. You go back to the very foundation. Does it still make sense to represent information as zeros and ones? Maybe, but maybe not. And so we should reconsider. And what we find is that we think no. We think that representing information as continuously varying values, what are called analog values, actually in this case, for AI, for intelligence, makes more sense.
6:58We also find another sort of difference between the conventional computing paradigm is that those silicon transistors that are representing zeros and ones typically make like a couple connections. The wire coming out of a transistor usually goes to two, maybe four places, but it doesn't go to a thousand. And that means that in order to get information from one place to another, it needs to hop along a lot of intermediate destinations. Which slows it down, right? It absolutely slows it down. It grinds it to a halt and it uses a lot of energy too. So that's another difference. So we want to build these circuits that represent information in a smoothly time-varying way, continuous way with continuous values.
7:40We want them to make many thousands of connections so they can move information all across the network. And we also don't want to do this thing where you have a processor over here and memory over here because then you're always shuttling information between those two. And that becomes the bottleneck. It has a name. It's called the von Neumann bottleneck. And it's known to be a serious problem. And even von Neumann knew it. He knew it in 1945 or something. He wrote a book called The Computer and the Brain that was all about how we should be moving more towards brain-inspired architectures. And then he passed away, unfortunately, early in his life due to cancer.
8:16But now we're in this moment where we realize that if we're willing to move away from using only silicon transistors, which are really good for representing binary information but not good for these other things, analog computation, extensive connectivity, if we build new hardware, now we can change the architecture, distribute the memory so that it's completely intertwined with the processing and have this high connectivity, which we think requires optical communication. So you use a different kind of circuits based on superconducting electronics to do these analog computations, and you pair them with optical communication.
8:54So now you can interconnect very large neural networks and perform these new kinds of computations that are important for AI much more efficiently. This feels almost like a bit of a middle ground between traditional computing and quantum in some ways in thinking about the optical element. Well, okay, so optics and computing has a long history, and let's dig into that a little bit. So back in, let's start in the 1980s, people realized that, or actually in the 70s, I think his name was Goodman, showed that using nothing but a simple lens, you can do a Fourier transform basically at the speed of light.
9:34And what is that for people who don't know? A Fourier transform is where you, it's a certain mathematical function where you take a signal in time and instead of looking at it in time, you look at it in frequency space. So you see what all the different frequency components are. And so that's a very powerful computation that's used in a lot of engineering, cryptography, many different fields of science and communications. And so people got really excited. they thought, well, hey, if we can do a Fourier transform that easily, maybe using light, we can do tons of different computations really easily.
10:09And so you got started on this whole optical computing sort of wave, and it sort of saw a peak and a crest. And then in the early 90s, it sort of fell out of favor because it turned out that just using optics alone, it's really hard to do computations. Light doesn't like to sit still, so it's hard to form optical memory elements, but it never really died. And instead, without going through all the history of that field, let's jump to sort of where we are now, which is that light is used extensively in communication. So for example, the most obvious level, well, right now I'm talking to you, we're not in the same place.
10:47And that's because signals are going over optical fibers and allowing us to exchange information over long distances. They go through. We really underappreciate how cool that is too yeah it's totally awesome that's right yes and i mean you could even argue that radio was the first use of of light in communications they're longer frequencies but it's the same electromagnetic wave so on the light that makes sense light is used in communication extensively across the atlantic but also in data centers and now it's moving ever closer to the chip so really the next big step in ai that is like going to happen and it's on NVIDIA's roadmap is for them to co-package their chips.
11:26This is what's called co-packaged optics. They take a few different GPUs and they put optical communication chips between those GPUs or even also between CPUs. So on the next generation of boards, you're going to communicate between these different cores using light. So that's classical computing. There's nothing quantum about it, but you're saying, how does this relate to quantum computing? And I would say, what we're doing has kinship with quantum computing in a couple of different ways. It's not quantum computing, because we have specifically designed our hardware, all of our circuits, all of our information processing principles, not to depend on entanglement and superposition.
12:12So the key aspects of quantum reality that lead to quantum computing are those. Basically, quantum wave functions can exist in multiple states at the same time. You live in a superposition of those states. And if you use that in your algorithm, that's quantum computing. We don't want to do that in the near term because it's challenging. It means your system is subject to decoherence and noise, and that leads to greater difficulty realizing products that you can bring to the market. However, we are adjacent to quantum computing in two different ways. One is that the most ubiquitous circuit element used to make qubits is the Josephson Junction.
12:56So if you're talking about superconducting qubits, which are being made by the Google quantum team, IBM, a number of other players out there, those qubits are based on a superconducting device called the Josephson Junction, which is basically just two superconducting wires with the thin insulating barrier. And we build our neurons out of Joseph's injunctions. The circuits are slightly different, but they have a lot in common. And we like using Joseph's injunctions because they're extremely fast and energy efficient. There's no known way to compute faster with less energy. So that's very promising.
13:32And they very naturally do what neurons and synapses in your brain do. They can sum many inputs. They can do that thing I was talking about where the memory that sets the synaptic weight is stored right there by the Joseph's injunction circuit that's processing the input. Yeah, you can see a bird with Joseph's injunctions. It works perfectly. Okay, so that's one way. We use superconducting circuits based on Joseph's injunctions, the same building blocks. The other way is that I mentioned our neurons communicate with light. Light is the fastest thing in the universe. So if you're communicating with light, you know you're going as fast as possible but we also do it at the single photon level so light is made if you just if you could look very closely at the the light all around us it's made of tiny packets called photons and einstein taught us that when he found the photoelectric effect for which he won the nobel prize einstein doesn't need any more accolades but these are so i forget
14:32Jeff Shainline:about this though, but he was involved in that. Yeah. It's not what he's most famous for, but that's what he won the Nobel prize for. So anyway, when, when we're building our circuits, we thought to ourselves, we want to communicate at the speed of light, but we would also like to do it at the lowest possible energy level of one photon, or at least order one photon per communication event from a neuron nudging another neuron through one of its synaptic connections. And it turns out that you can do that with superconducting single photon detectors, superconductors, the hero of the day, again, allow you to detect the most faint level of light, one single photon.
15:10And it turns out that the quantum community that's trying to make photonic quantum computers do this thing where they encode information in single photons, and they use the exact same kind of detectors that we're using. So again, we're not using the entanglement of those photons, but we are using the extreme speed and energy efficiency and the same kinds of superconducting detector. So that's our relationship to quantum computing.
15:35Jeff Shainline:So it sounds like you're taking all of the benefits from quantum and none of the negatives. Yes, you said it. You said it. You're my marketing man. Instability. But okay, let me ask you a hard question then. Does that mean that you're, you know, in order to do all the superconducting, does that mean that it still has to be at really, really low temperatures? And how are you handling that part of this? I would drop the really, really from what you just said. So we have to be at low temperature to be quantitative, our systems operate at four Kelvin. So a quantum computer typically operates at 20 millikelvin.
16:08So we're about 200 times warmer than a quantum computer, which, okay, four Kelvin is still cold, but it actually is, from an engineering perspective, much easier to get to four Kelvin. All of the cooling infrastructure that you require for that is much simpler, less It's expensive, more scalable. And so it also relies only on helium-4, which is a relatively abundant resource, as opposed to to get to millikelvin temperatures, you need a mixture of helium-3 and helium-4. And helium-3 is the really rare isotope. So that's expensive and hard to come by. So we think that operating at 4 Kelvin makes this much more deployable at large scale data center applications.
16:49There are commercial products right off the shelf that we can put in, even when we scale up and are building 10 trillion parameter systems. The cryogenic infrastructure to cool those is totally data center compatible and it exists today. So getting to low temperatures is not the most challenging aspect of what we're trying to do.
17:10Jeff Shainline:You want to go, Corey? Sure, I'll go. Can you walk us through SOENs at like the smart non-physicist level? where do where do the electrons do the work where do the photons do the work and where does and then the neuron behavior if i understand correctly comes from the connectors you're talking about yeah okay great so we call them zones superconducting optoelectronic networks we pronounce it zones that's just kind of how it how it evolved over the years and so yeah the key principle of Sones is that you really want to distribute the labor between what I would call computation and memory and what's communication.
17:54So we follow the playbook of what neurons do. We can do everything that artificial neural networks do. We can also do a rich variety of things that neurons in the brain do. So I guess just walking through it at a relatively simple case. What a neuron needs to do is take inputs and weight those inputs. That means multiply the value coming in by the synaptic weight. So that's a multiplication, and that's performed with superconducting circuits. Electrons do that part. They do the multiplication between a signal coming in and a weight that's stored right there, establishing the strength of that connection.
18:34Then they sum all of the different synapses that are coming into a neuron. So a neuron is in the brain, they're very complicated, but stripped down to their essence. Let's say there's what's called the cell body and there's all the synaptic connections. So all those synapses are doing that weighting that I just described that the total signal is summed at the cell body. And then it's the cell body is performing an activation function, which means its output is not just the sum of its inputs. It's some mapping, some nonlinear function of the sum of those inputs, which usually has three properties.
19:10One is that there's a threshold. If the activity coming in is below a certain value, the neuron does nothing. If it's above that threshold, it grows, it monotonically increases. And then if it gets above a certain level, far beyond the threshold, it saturates. So this thresholding growth and then saturating those different aspects of the function typically define what a neuron does. We do, that's all what I would call the computation. So all of that is happening with superconducting electronic circuits. We can get a lot of different kinds of activation functions like the rectified linear unit, which is flat, a zero, and then it goes linearly.
19:49We can get a sigmoid, which is more smooth and then saturates at the top. Those are some of the most common activation functions used in artificial neural networks. And it just turns out that superconducting circuits do a really great job of just naturally producing them. We don't digitally compute those functions. The circuits themselves just produce them. And that's what makes it so fast and energy efficient. OK, so then the next thing that happens is that activation, the sort of amount of signal that the neuron needs to communicate to other neurons, that is produced by light. So every single neuron cell that we build, it's a complicated system of superconducting electronic circuits.
20:31Well, not that complicated, but there's, you know, a few Joseph's injunctions for the neuron cell body and then a few Joseph's injunctions for every one of the synapses. Then the neuron, it also has a light source. That activation that's produced at the neuron then gets transferred into an optical signal. So the superconducting circuits drive a semiconductor light source to produce very faint amount of light because it's going to be communicated to every single one of those synapses at the single photon level. So the neuron produces that light. It gets coupled into what's called a photonic waveguide, which is basically just a wire that routes light around.
21:11So the neuron's output goes into these wires for light, and they get routed to all the different neurons that that one connects to. And where they connect, that's called a synapse. and at each of those synapses, there's a superconducting single photon detector. So an optical signal goes into the synapse, that detector converts it back to an electrical signal and then the whole process repeats down at the next neuron downstream. Wow.
21:37Jeff Shainline:That's so cool. So could you help us now connect the dots between this architecture and the language models or all kinds of models that we're used to today? You can just run a language model like plug and play on this architecture. Are you training new models with this architecture? How does that work? Okay, so this gets into sort of what are the strengths and weaknesses of the digital general purpose versus the sort of analog application specific approach. So, yes, we can run language models. We are training them and we're doing that right now. At a smaller scale, we have a language model that one of our algorithm developers trained.
22:18We're showing that when you run that language model on our hardware, you get something like a factor of 100 speed improvement over doing it on conventional GPUs. So that speed advantage is really important. That model has only been trained at the scale of around 100 million parameters. It's not like the hundreds of billions of parameters of the large language model. So we would call these small language models. But that's not a fundamental limitation of the hardware. That's just a limitation of the scale of simulation that we were willing to run to train this because it can get expensive. And that's not the product we're trying to build.
22:51So this is sort of more exploratory. So, yes, we can take existing models of a lot of different kinds and do what's called knowledge distillation. So those models are typically implemented on digital hardware, assuming that you can get basically any computational function that you want out of it. And one of the ones that's used a lot in transformers is this attention mechanism where you basically perform matrix-matrix multiplication and you perform an exponential function. It turns out that that's not a great mapping onto our hardware, like literally one-for-one, parameter-for-parameter. So instead what we do is we take trained models layer by layer and we transfer the function of each of those layers to a layer of our hardware.
23:38And now we get the exact same functionality, the exact same input-output relationship, but we do it in a way where we reap all the benefits of our accelerated hardware and the analog computing. And it turns out that it's relatively straightforward to just perform this sort of knowledge distillation from a trained, known model onto something that can run on our hardware. So that's the way we're approaching that problem. We don't want to actually take somebody's digital model and put every single parameter on the chip because that would be omitting a lot of the benefits that we gain from building this application-specific hardware.
24:15but we want the chips that we build to be as general purpose as possible. And this knowledge distillation is the way to go.
24:21Jeff Shainline:That makes sense. Yeah. Yeah. So you're basically converting it for your use. Is there a type of model that you would like, you know, for instance, people who have watched a lot of our shows have heard certain people talking about non attention or non transformer models, like, you know, maybe energy based models or maybe, you know, newer architectures. do you have any sense of hey this is going to be really good for you know xyz type of architecture even better than potentially the the current paradigm because where we're at now cannot be the final stage i don't care what anyone thinks they say like there's no way that we've like solved it with transformers right there may not be a final stage for yeah no i mean you could argue maybe there's there is no such thing as a final stage but like do you have any do you have any sense of you know hey i think actually these types of models with these types of architectures are going to be 10 times better or that's a random metric I threw out there.
Read the full transcript
25:18Jeff Shainline:But like, what's your sense of the landscape and what works the best on this architecture? Yeah, I think that's a really important question. So let me just rant a little bit about the transformer architecture for a second, because I think it provides some important context. So without going into all the details, the most important thing about the transformer is that it basically allows any input to modulate the attention given to any other input. That's the most general utility you could possibly ask out of some neural network type function. And that's what makes it so powerful. It ends up being able to train really well because it has this sort of like all-to-all ability for anything to influence anything else.
25:59But it's sort of like the strength is the weakness because that's also such overkill for a lot of problems that it's what leads to the dramatic energy consumption and really unsustainable training time and power consumption for training as well as inference. And the reason why we work with the transformer architecture today, the reason why the community does, is because it maps pretty well to existing hardware and it takes advantage of all those general hardware capabilities that GPUs give us. However, that ability to take any input and let it influence the attention given to any other input, there's a lot of great work from sort of the neuroscience or the neuro AI community showing that networks in the brain, dynamical circuits that are ubiquitous throughout the brain accomplish something very similar, but they don't have the same painful scaling, this quadratic scaling where if the input sequence gets twice as long, you need four times as much compute and four times as much memory.
27:00So these are typically recurrent neural networks. So recurrent neural networks were kind of where it all started. You can go back to Jeff Hinton's work in the 80s and 90s and parallel distributed processing. And all of that was sort of like intuitively, that's where AI should come from, right? It should be these recurrent neural networks. It turns out that those are less efficient to simulate on conventional hardware. And so even though they have the fundamental principles that we look for in artificial intelligence, brain-inspired computing, and they have a lot of functionality. They can work as well, if not better than transformers in a lot of cases.
27:39They don't map as well to the chips that we've had available for the last 20 plus years. And so they haven't gotten as much traction. That's exactly where we come in. And so there's a resurgence already happening. Leave Great Sky out of this. State space models are very popular. These are linear recurrent neural networks, which allows them to, in a way, map as well as possible onto a GPU so they can still be trained quite quickly. And so they've gotten a lot of attention, no pun intended there. But in our case, what we're doing is we're physically building circuits that implement recurrent neural networks.
28:17So you pay no temporal or energy penalty for connections that don't just go forward, but also go side to side and also go backwards. So you can have these very complicated networks, works much more like the architectures of the brain, where attention comes from the dynamics. It comes from this layer projecting its inputs to the next layer, and then the next layer feeding back and saying, based on what I know, these neurons should be more active, these ones should be less active. And then you get linear scaling where the input sequence gets twice as long, it's only twice as much compute. And it can be a much more efficient way to have a lot of different kinds of data sets, especially when the data is temporally varying.
28:57So your question of which model architectures, state space models map very well onto our hardware. We're spending a lot of attention on that right now. When I was talking about knowledge distillation a second ago, that's where we're focusing right now, taking trained state space models that have been trained using GPUs, but then mapping them to our hardware. So it's just blazingly fast inference. You also mentioned energy models. We've shown that all of the circuit equations, all the dynamics of our system can be derived from an energy model. So you can do this relaxation and things like equilibrium propagation are very natural on our hardware and more brain-inspired architectures where recurrence is just a part of it.
29:38You also have elaborate dendritic trees, dendrites in the brain, where it's not just a neuron and synapses, but these intermediate processing stages. We can implement those very straightforwardly with our circuits. So So the sky's the limit now. We think that the AI community has really been, okay, it's been enabled by existing chips. And of course, we wouldn't even have AI without all the tremendous progress from the semiconductor industry. But now you invent this new chip architecture, this new kind of hardware. We think there's a new renaissance coming where all these new architectures finally have their day in the sun because they have the platform on which they can be properly hosted.
30:19Wow.
30:19Jeff Shainline:Love it. So let's talk about your first generation chips a little. I understand they've shown major gains for workloads like video analysis, even saw a number of like 60 million frames per second somewhere that I was really curious about. And I'm wondering kind of what's being measured there? Because as a guy who watches the slow-mo guys on YouTube and their cameras, I'm like, that is amazing. Yeah, OK, so let's talk about our chips then. So we've been since the last 15 months, we've taped out four wafer generations, starting with just little prototypes of showing that we can get the basic circuits going.
30:55And now we're at networks of over a thousand parameters. OK, a thousand parameters is still not a billion parameter large language model. But, you know, that took us 15 months and less than four million dollars. That's pretty steady progress. And what those chips are doing is high speed inference. They're programmable chips, so you can put in whatever synaptic weights you want and run different kinds of models. With these generations, we've done things that are in the domain of sort of what I'd call speech processing. You could also consider it signal processing as well as video image processing.
31:31Let me zoom out for a second and talk briefly about where we really want to go, and then I'll come back and point to where we are on that roadmap. So with this hardware, one of the things that we think is most compelling is the ability to process multimodal data in a manner that's much more how a human would process it. So you see the world, you hear the world, and you can produce things like universal assistants that interact with humans in our native modalities. And you just talk to them and they provide you actionable insight, basically with no discernible latency. So that looks a lot like video processing, where you have a video stream and an audio stream, which brains are really good at.
32:13You have the visual cortex, the auditory cortex, they merge, and you have this seamless combination of those different data streams. So that's kind of our North Star. That's where we're building towards. So as we're at our initial product generations right now, we want to assess how quickly can we process video, how quickly can we process audio. Because when we just do calculations, sort of estimates of how well our hardware performs at that, say, trillion-plus parameter model where it's handling video, it's the number you said. It's 60 million video frames per second. And that's, I mean, the most simple way to think about that, you can go through all the math, but the simple way to think about it is that we build neurons and synapses like the ones in the brain, connected like the ones in the brain, but everything in our hardware happens about 2 million times faster than biology.
33:05So we humans watch videos at 30 frames per second, and our hardware watches it at 60 million frames per second. Everything gets accelerated by that factor of 2 million. So what we're showing with our early generations of chips is that we can take audio, spoken words, for example, accelerate all those audio frequencies by something like a factor of a million, process it through our chips, and the chip can tell you which digit was spoken. This is a common benchmark task, the Google speech commands. Similarly, we can show small digits, little numbers in a small image and classify them. What we've shown so far is 10 million frames per second.
33:48That's what's measured. We know we can get to 60 million frames per second. That extra factor of six is bottlenecked by the room temperature electronics on the rack, not the chip itself. So unfortunately, and we just need to, you know, fork out a little more dough. Exactly.
34:06Jeff Shainline:That's exactly right. Yep. Fair. Fair. So that's where we're at right now. We're handling these workloads that lead us toward that North Star. I'll also mention one other application with this thousand parameter accelerator that we're really excited about. We're working with the Department of Energy and Brookhaven National Lab. They care a lot about taking sensor data from a fusion reactor, a tokamak fusion reactor. You're basically looking at the plasma, streaming that sensor data into an AI model and predicting what the plasma is going to do. Is it going to become unstable? How can you steer it to maximize energy production?
34:43So we're partnering with Shinje Yu and his group at Brookhaven to build models that are able to stream that sensor data. It's coming in at a million data points a second, megahertz frequencies, and build a control system. So we're working on that demo right now. And we think in the slightly longer term, not that far in the future, we can build systems that take data from all the sensors around a fusion reactor, have this ability to perform really fast inference, and control that reactor so it never becomes unstable and never causes a disruption that leads to damage or a break in the power supply.
35:19That's one of the applications. And it turns out that with only a thousand parameters on the chip, if it's blazingly fast like this, it can start to handle that workload. And we think once we get to around a million parameters, this will be a very state of the art system for that fusion control application.
35:35Jeff Shainline:So what's what's between you and the million parameter version? Is it just money? Like what do you need to scale this? Well, you got to turn money into chips and that's where a foundry comes in. So we've been building a lot of our prototypes using the – so our company spun out of NIST. We were at the National Institute of Standards and Technology for about 12 years, and we developed this superconducting electronics process to make these circuits when we were there. We still use that process in that clean room to build some of our early prototypes. But now we're partnering with IMEC. They're a foundry in Belgium that specializes in sort of taking technologies that are right at ready to pop and bringing them to the final stages of maturity before they get transferred to another foundry, TSMC, Global Foundries, for large-scale, high-volume manufacturing.
36:26So IMEC is our manufacturing partner that brings tremendous lithographic capabilities beyond what we can do in our sort of research-grade cleanroom. And immediately by using their lithographic tools, we shrink our parameters by more than a factor of 100, which allows us to dramatically increase that parameter density. So that's one thing is just moving to a better foundry. I'll also briefly mention that with this technology, based on superconducting electronics, manufacturing stays relatively cheap because you don't need extreme ultraviolet lithography. To build the next generation of semiconductor chips, NVIDIA's next chips, you need EUV lithography that only ASML can provide and really only TSMC can produce at volumes, maybe also Samsung.
37:14With our hardware, superconducting circuits don't benefit from being shrunk beyond a certain point. So you use instead what's called deep UV lithography, which is state-of-the-art in like 2005 or something. And it's just much cheaper. And there's a lot of foundries in the U.S. that can do it. So from a manufacturing perspective.
37:33Jeff Shainline:That's good for scaling. Yes. It brings costs way down. Throughput can go way up. And a lot of different foundry partners can contribute. So as we look to scaling up, another thing that's really important to producing larger systems is going beyond what fits on a single chip and going to multi-chip modules interconnected with either superconducting interconnects or also obviously optics, fibers over longer distances. So we're building multi-chip modules that allow us to scale far beyond what can fit on a single die. So with just those two innovations, not even innovations, just steps forward for the company, manufacturing with iMac multi-chip modules, there's a clear path to state-of-the-art, very large models that can handle a tremendous variety of workloads.
38:19Jeff Shainline:So what does this look like in practice at the limit? Are you building this for data centers? Will there ever be like a consumer version where we have to have cooled in our house, like in our basements or something? Could this even be used in robotics? Like what's the use case here? Like when will we see this or how will we see this in its final form, let's say? So we're kind of focused on basically three sort of commercialization pathways. One, as I mentioned, is around fusion, but a little bit more generally. It's about deployments to high-value workloads where extreme high-speed inference really matters a lot.
39:01So the DOE is a really good example there. They have a few use cases where it really matters a lot that you can perform inference on like a sub-microsecond timescale. So you're streaming data from fusion sensors around a fusion reactor. Another example is particle colliders. There's the electron-ion collider being built at Brookhaven. Other examples like the Large Hadron Collider at CERN, where they're colliding these particles and they can't store the data from every single particle collision. It's just too much. It's terabytes per second. And so they have to take data from all these detectors around the sensors that they're using to detect the particles that have been produced in these collisions.
39:42and with really fast AI determine, does this look anomalous? Does this look interesting, like I should save it, or is it just another whatever collision that I've seen a trillion of them and I don't care? And that's another sort of like high-speed workload that our hardware is absolutely perfect for it. And then a third one that I'll mention is materials development. So imagine that you have a source of x-rays that's firing really frequently and it's interacting with the material and producing a diffraction pattern. and you need to reconstruct that diffraction pattern really rapidly. Same thing.
40:17You have a massive amount of data. You can't save it all. And if you could reconstruct in real time what that diffraction pattern tells you about the material, then you have much more uptime on your X-ray source. So those are some of the applications where we think large-scale, high-value deployments make a lot of sense. Okay, contrast that with another use case where it's not so much about going after a few high-value workloads, but it's really about exploring the breadth of applications. In this case, I'm talking about hosting cloud systems. So we can build a handful of systems that we host here at our facilities, and any user in the world can just remote in, run their job on our chip.
40:59And so this is like a cloud model. It's NeoCloud, whatever, you know, CoreWeave does similar stuff, Crusoe. These are companies that they take chips and they let anybody run an inference job on it or whatever. It could be a training workload, whatever it is. So that's something that we're working on as well. The good thing is that because our chips can handle inference so much faster than conventional chips, you don't need to build that many systems before you can host a lot of users. And the idea here is that, first of all, it's great for us as a company if we can get some revenue from these users.
41:31And second, we want to partner with them. We want to understand, hey, all you brilliant model developers out there who now have access to this new sandbox, what are you going to do? What do you think is exciting? Where is the creativity from the community going to lead this? And so cloud is the second thing that we're working on. And then the third thing is sort of a bucket, but I'll call it hyperscalers. And there's a number of workloads there. So, for example, in the near term, I mentioned that we're interested in this video comprehension workload in the long term. The very simplest, earliest stage of video comprehension where you have traction at a smaller model scale is content moderation.
42:12And so in this case, picture your YouTube. You're getting so many uploads every second that it's a computational challenge just to assess, is this video clean? Does it violate our policy and need to be blocked? And so right now, those video analysis models, they basically sample every third or every fifth frame. They don't really watch the video. They just pick out frames and look for objects or symbols. They separate out the audio and they turn it into a spectrogram and just look for signatures of harmful content. And we think we can do that better. We think we can handle all of those videos significantly faster while watching the whole thing and understanding it like a human would.
42:52And that's a workload that's worth billions of dollars to those companies. You've seen Meta recently lose in court because they allowed harmful content to be seen by children. And that's a really hard thing to avoid at the scale of Meta. And so they have to contract all these humans to watch the worst videos on Earth, and it's a damaging thing. And so we think we can do that.
43:13Jeff Shainline:It creates issues for the people who have to watch it because they have to see all this horrible stuff. So it would be actually – that would be a great job for AI to automate. Exactly. We want to get rid of jobs that people don't want anyway. As long as it doesn't train on it, I think that's important. Yeah. So you have to be careful about how the models are trained. So that's – and then the final thing I'll say, it's similar customers but advertising. Between Meta and Google, there's$500 billion a year in advertising revenue. And that's all moved into this more dynamical sequence-based kind of processing where they're looking at a long sequence of user behaviors, the ads they've seen, what they've clicked on, what they haven't.
43:52And it turns out that if you can increase the click-through rate by something like 1%, then that's worth like$5 billion a year to Meta. And I think we can do it in a way where it actually not only does it produce more user engagement, but it does it in a way where it's actually offering you the ads that are useful to you because we have a faster model that can take more context and feed it into the model without adding latency, actually still performing with less latency. it's going to be a better product. It's going to be something that, oh, wait a second. That's actually something I really need.
44:27I'm glad they showed me that. And now it's better for our customer, which is meta, and it's better for their customer, which is you. So those are some of the things that we're looking at where smaller models that we can achieve in the near term have tremendous value across society. That's really awesome.
44:46Jeff Shainline:Makes sense. So I saw that Semaphore reported that Great Sky plans to begin delivering chips sometime later this year. Is that correct? Well, right. So our first contract is with the Department of Energy. We already have the chip made, and we're testing it in our facility right now. And the contract states that we will deploy it at Brookhaven National Lab for this fusion analysis workload by the end of the calendar year. So we think we're on track to make that delivery, yes. Very cool. Very cool. What would be your like pie in the great sky dream for how quickly you get up and running with all the, you know, all the other things you want to do?
45:25Jeff Shainline:Like, let's say, you know, you make this first delivery this year, you want to 10x that, you know, in two years or three years, like, what's your goal, long term vision there? Okay, well, we want to get to 100 million parameter multi-chip modules in two to three years. Like through the Series A, at the end of that, we want to be building systems that we can deploy at a customer's facility in their data center. That's why this DOE initial engagement is really promising because we've already investigated what it takes to install our system in their data center. And now we know we don't need to change anything about it.
46:00We can take the power right from the ceiling, the cooling water right from the wall. We don't even need to weld a single pipe to get our system in their data center. So by the end of the Series A, two to three years from now, we want to deliver systems like that to a number of different customers. We want to host multiple models here so people can do workloads. As I was describing, things like content moderation, speech processing, also things in drug discovery, protein folding, molecular dynamics, simulations. And that's when we think we're up and running, when we're hosting our own models, letting customers access them, deploying on site.
46:37And then through the Series B, we want to get beyond billion parameter models. We think we can get to 10 billion parameter models in that time frame and really have a lot of commercial traction. Because once you realize that this costs less and performs better, it's going to catch on. It's not a scary thing. And people hear superconductors and that sounds new, but once they understand, don't worry about it. It's a box. You put your data in, you get your answer out way faster. That's what we're looking for is two to three years for our first real scale deployments and then getting beyond 10 billion parameter models two years after that.
47:13Wow. How much of this can be like manufactured using existing U.S. foundry capabilities versus is there anything in this that requires new processes to be able to manufacture? Okay, so let me sort of split a hair there. So in terms of the equipment and the manufacturing capabilities on sort of the foundry side, there's very little that needs to be changed. We use, like I was saying before, GDV lithography and all the same kinds of sputtering, deposition, patterning, etching, planarization. The whole workflow of a modern fab, we use the exact same manufacturing concepts. That doesn't mean it's exactly the same in every detail because we use different materials.
48:01So, for example, where the wiring layers that are common in a CMOS process are largely copper, we use niobium because niobium is the most ubiquitous superconductor. I think at an abstract level, it doesn't really matter that much because niobium, you sputter it just the same way. It's very easy to process. It's not an expensive material, so it doesn't add hardly anything to the cost of your wafers. But if you ask TSMC, hey, can you go run this process tomorrow? No, not tomorrow, because they have to dedicate their sputtering tools to a new material. You can't take one sputtering tool and do copper and niobium at production scale.
48:44And you take their etch tools, you need to slightly change the etch chemistry. All the processes are there. Our partner IMEC is running the process right now, but not at volume. so they can get to a few thousand wafers a year but we need to get to a few hundred thousand wafers a year and that means basically it's a commercial challenge it means you need to convince tsmc not that they can do the process they know they can do the process you need to convince them that there's a demand for the volume and that's really where the challenge is yeah because there's a there's a big fight right now trying to keep trying to get their capacity Absolutely.
49:19And yeah, it's really interesting. How has AI been used to help in the development of this technology or maybe the testing phase? Has it been a beneficial tool or was this moving before? I mean, yes, it was. So that's a great question. So we started this project in 2014. So for some patent reasons, I had to go back through my old notebooks and figure out when did I draw the first circuits about these superconducting neurons? And it was 2014. So AI was a thing back then, but it wasn't the thing it is now. And so I was more fascinated by ultimate limits of intelligence than like, hey, what's the commercial workload we're going to run?
50:03But so then it's evolved over time. And then, you know, you had the watershed moment with ChatGPT3 in like November of 22, was it? Yeah. Yeah. So that kind of obviously breathed a bunch of life into the whole field. And now I guess what I would say is our company, like I think most companies, most people now are using AI in a lot of different aspects of our life and our work. So most obviously, our algorithm developers that are writing a lot of code are using software coding tools. I won't say which ones, but you can guess. We're using AI to write a lot of our code. We try to be pretty careful because we want our stack to be totally bug-free from top to bottom, so we have to be really careful.
50:52to prompt it, say, write a software stack for this. We use it a little bit more judiciously. It's a pretty vital stage to ensure everything is exactly dialed in the way you want it.
51:05Jeff Shainline:I assume also it's very novel what you're writing software for. I don't imagine that's in the training data, right? Well, I mean, I guess honestly, I'm not the best one to answer that question, but I don't think it's all that novel. All of our code is written, all of the software that we're using to simulate our systems, to pipe data into it. It's all done in, well, it's largely done in PyTorch for the simulations, Jax, back end. So it's like pretty off the shelf stuff. We're using a lot of FPGAs to provide programming signals. And so there's a lot of known C code for how do you interface with FPGAs.
51:43So a lot of what we're doing is pretty bread and butter. The novelty is really on the chip that you're talking to. But you're providing just square pulses for changing synaptic weights and waveforms for putting data into the system. So there's a fair amount of knowledge base around that. I guess I would say beyond using AI in our coding, I personally use AI as sort of just like a thought partner and thinking through maybe different commercialization strategies and things like that where market analysis and doing research around different use cases. But yeah, I think we're using AI in pretty conventional ways at this point.
52:20Also, a lot of like, hey, we're building a neuroscience inspired chip. Help us formulate the analogies between AI, neuroscience in a way that the public can understand communications, that kind of thing.
52:34Jeff Shainline:Awesome. Very cool. We'll have two more questions for you. So the first one is the memory side of this, right? where a lot of ink was spelled earlier this year on the memory shortage and how that's become a big bottleneck now for deploying AI systems, especially with agents. And I was wondering if you could talk just a little bit more before we wrap here about how your system is helping with the memory side of things. And then one more after that. Okay, sure. So we approach memory in a completely different way. I mean, okay, in computing, there's lots of different kinds of memory. There's memory hierarchy.
53:09And so it sort of depends on exactly which aspect you're talking about. But I think the most important one from the perspective of performance is the memory that's closest to the processing. In conventional chips, that's SRAM and it's sitting on the same chip. We take a completely different approach. There's no such thing as a bank of RAM in our hardware. We physically build a circuit right next to the synaptic processing circuit that stores the synaptic weight. We call it the memory cell, and it basically just stores a superconducting current. The value of that current can be ratcheted up, ratcheted down.
53:44We've shown that we can do beyond 8-bit. We think sort of 4-bit is a sweet spot, but it doesn't really matter. We can do whatever value the application requires. And those memory cells are made in the exact same processing layers using the same niobium wiring and the same Joseph's injunctions as all the computational circuitry. So there's not this like tension between a process required for the compute and a process required for the memory. They're one in the same here. And so that means that the primary data movement bottleneck of AI, where you're shuffling information between the processor and the SRAM, it's not that we did better at that.
54:27We just got rid of it completely. It's not present because we store synaptic weights where they're processed. So that's the most important thing. And then in terms of the rest of the memory hierarchy, yes, imagine your YouTube and you need to store, I don't know, a million hours of video before it gets piped into our system. That's going to be in RAM and DRAM, not SRAM. And so, of course, we still need some digital storage and things like that. But But the most important thing is the elimination of that data movement bottleneck between processor and memory. And that's just completely gone. And it turns out that moving to superconductors is the key enabling step there.
55:05It's really hard. I mean, people have been wanting to do this since the beginning, integrate processing and memory. And silicon really forces you to choose. And the digital architecture, this von Neumann architecture, really deeply embeds this principle that they're separate. But when you go to a brain-inspired architecture, which is much more matched to superconducting electronics, and the materials and processing requirements of superconducting electronics just naturally give you co-location and memory with processing, and that leads to application-level things like continuous learning, the ability to not have a training phase followed by an inference phase, but continual learning as inference occurs, which makes things much more user-friendly and able to understand you deeply as an individual eventually as well.
55:51I think is the result of that.
55:53Jeff Shainline:And then what would be the limit, I guess, is just like scaling to just more chips, basically. Right. You mean the limit on like the memory and like how much, because like, I'm so used to thinking of the limit on what type of model I can run in terms of like the amount of GPU I have, right? So like, I'm like, I know that if I have X amount, you know, like 90 gigabytes of GPU capacity, I know I can run XYZ model. Otherwise, I'm using it over the cloud. You know what I'm talking about? Yeah, exactly. So, I mean, it's the same here. You're limited by, we think in terms of the number of parameters, basically.
56:30There's sort of a mapping between the two. And in our case, there's the fact that we have eliminated the data movement bottleneck to RAM and the interconnectivity challenge. Solving that with photonic interconnects means that as you continue to scale up, You add more chips to your system, ultimately entire wafers stacked. You can use optics to communicate between wafers over fiber optics over longer distances. You don't run into pinch points because you have this high fan out optical communication. Every neuron can communicate to many thousands of others at the speed of light, regardless of how far away they are.
57:06That means that the ultimate limits are eventually set by how long it takes light to reach a certain destination. And so it's very large. At every stage, we're not limited. We're far from the physical limits of this hardware. It's going to be basically manufacturing limited and how fast can we build. We want to build systems that have 20 wafers interconnected. That's still small compared to a data center filled with GB200s. But when we get there, we have beyond trillion parameter models, which are profoundly useful, especially when they're going a couple million times faster than a human brain.
57:44And we don't stop there. We go to much larger systems. As I said at the beginning, the cryogenic requirements are already met. There's already off-the-shelf systems you can build. But it's really that. We think there's another path to exponential scaling. Moore's Law was all about lithographic improvements that enabled you to shrink transistors. In our case, exponential scaling is enabled by the ability to double the amount of compute in every product generation without running into painful power consumption bottlenecks or data movement bottlenecks. And so you can keep adding more compute to the system, keep doubling, keep doubling, have exponential growth in a completely different way.
58:25But it's enabled by superconducting circuits for computation and optical communication at the single photon level. Amazing.
58:32Jeff Shainline:Awesome. Wow. I have so many more questions, but Corey, I'll let you wrap it. Jeff, thank you so much for joining us today. This has been really interesting and it's exciting what you're doing. We always love talking to someone who's kind of stepping outside and not afraid to break the paradigm. Yeah. All right. Thanks for having me. I really appreciate it. Absolutely. Where can people go to learn more about what you're doing and keep up, follow Great Sky? Well, you can follow us on LinkedIn, X, Blue Sky, our website, greatsky.ai. Should be able to find us on a Google search. Excellent. Very cool.
59:10If you're watching and you haven't yet, please take just a moment to like and subscribe above or below, actually, I think it is. Either way, we really appreciate it. It helps us continue to bring you more fascinating guests like Jeff here. Also, if you haven't yet, please pop by the Neuron.ai and sign up for our daily newsletter. It's read every morning by 700 ,000-ish people. So we'd love to count you as one of those and have you as part of our community. But I think that's it for today. Farewell for now, humans. Farewell for now, humans.
From the publisher
What if the next big AI breakthrough is not a bigger model, but a completely different kind of computer?
Jeff Shainline, co-founder and CEO of Great Sky, joins The Neuron to explain how his team is building brain-inspired AI hardware using superconductors, photonics, and analog computation. Great Sky’s architecture, called Superconducting Optoelectronic Networks, or SOENs, is designed to move beyond the traditional GPU roadmap by co-locating memory and processing, communicating with light, and mimicking some of the high-connectivity dynamics found in biological brains.
In this conversation, Jeff breaks down why today’s chips can struggle with fast, multimodal inference; why transformers may be powerful but inefficient for some future workloads; how Great Sky’s system differs from quantum computing; and why early applications could include fusion reactors, particle physics, video understanding, content moderation, and eventually new model architectures that do not map neatly onto today’s hardware.
Subscribe to The Neuron for grounded, practical conversations about where AI is going next—and what actually has to work before the hype becomes real.
