In short
Lightning AI’s CEO Will Falcon explains how Lightning AI evolved from an open-source training framework (PyTorch Lightning) into a “full-stack AI neocloud” after merging with Voltage Park, combining software and GPU infrastructure. The episode covers the Lightning AI Studio developer experience, the economics of neoclouds vs AWS, and Falcon’s path from PhD research to building a $500M+ ARR company.
Guest background
Will Falcon is CEO of Lightning AI. He created PyTorch Lightning while at NYU (PhD era), later worked on large-scale inference systems (including NextGenVest, acquired by CommonBond), and joined Facebook AI Research (FAIR). He credits NYU’s open-source culture and close ties to FAIR/Meta.
Key claims
Lightning AI Studio provides persistent, auditable cloud “laptop-like” dev environments (VS Code-like, SSH/CLI, collaboration). The merger yields 35,000 GPUs, 400,000 developers, and $500M+ ARR in under two years. Falcon argues most GPU usage is not inherently “inference-first”; training is hard due to tooling gaps. Voltage Park was chosen for profitability and low debt, avoiding neocloud “mortgage crisis” leverage risks.
Notable examples
Cantina (Sean Parker’s startup) uses Lightning for inference; Cursor training cluster; NextGenVest used RL-style text messaging to optimize financial aid outcomes; PyTorch Lightning downloads ~400M.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroducing Will Falcon and Lightning AI
0:51 to 2:20
Discussion about the merger of Lightning AI and Voltage Park and its implications.
“This episode of Super Data Science is made possible by Dell, Intel, Fabi, and Cisco.”
The Evolution of Lightning AI
2:20 to 3:32
Exploration of how Lightning AI evolved from a customer to merging with Voltage Park.
“and you guys have created something really big.”
Challenges of Traditional Cloud Providers
3:32 to 5:07
Understanding the limitations of traditional cloud services for AI applications.
“So, you know, if you're not familiar with Lightning, full stack software, kind of all the tools you need to build, train, inference models, like all that, I'm sure we'll get into at some point.”
The NeoCloud Concept and Its Benefits
5:07 to 6:41
Explaining the concept of NeoClouds and how they enhance AI computing.
“Things like InfiniBand and VAST storage and all these different things.”
The Importance of Inference in AI
6:41 to 9:34
Discussing the role of inference in AI model deployment and its market implications.
“You know, they call it the white glove experience, which I think it's a great product name.”
The Lightning AI Studio Journey
9:34 to 12:42
Overview of the Lightning AI Studio product and its evolution for AI practitioners.
“for that AI model to be used at inference time in production to be able to do something for users.”
Evolution of PyTorch Lightning
14:00 to 16:30
Learn about the evolution and impact of PyTorch Lightning on AI workloads.
“And then the problem became, okay, this is cool, and I can train at scale.”
Features of Lightning AI Studios
18:10 to 20:55
Explore the features and benefits of Lightning AI Studios for developers.
“We'll probably, we'll have something in the show notes for people to be able to skip the queue and get access to some number of free monthly credits.”
Understanding NeoCloud and Market Dynamics
20:56 to 22:50
Discuss the challenges and strategies in the NeoCloud market.
“You can come on there and we have our vibe coding and you can choose the skill you want.”
Pricing Strategy and Margins
22:51 to 25:56
Analyze how Lightning AI manages costs while maintaining profitability.
“The CEO who is running Voltage Park, who's now on our board, he's a finance guy, so ex-finance guy.”
Show all 34 chapters
The Journey of PyTorch Lightning
25:57 to 28:00
Delve into the origins and development of PyTorch Lightning and its impact.
“And how do you manage to maintain these good margins that you just cited more recently while also offering these great deals?”
Origins of PyTorch Lightning
28:00 to 29:00
Learn about the early development of PyTorch Lightning while Will was a grad student.
“And yeah, you developed that while you were a PhD student working at NYU.”
Innovative Research and Challenges
29:00 to 31:00
Discover the challenges faced in training models using neural activity data.
“So at Stanford, they had monkeys that would collect neural activity from the retina.”
Transitioning to TensorFlow
31:00 to 32:30
Hear how switching to TensorFlow drastically improved model training time.
“I'll probably bring it to the office, but it's cool.”
From Inference to Education
32:30 to 34:20
Explore the use of AI in education through the development of a financial aid support system.
“And then we actually did eventually publish a paper to New Rips, my first paper ever, and also my last paper to New Rips.”
Acquisition and PhD Journey
34:20 to 36:30
Learn about the acquisition of Will's company and his decision to pursue a PhD.
“I've ever, I think I've seen put out there.”
Navigating Lab Meetings
36:30 to 38:40
Understand the dynamics of attending lab meetings as an outsider while managing a startup.
Starting at NYU and Open Sourcing
38:40 to 40:50
Discover the launch of PyTorch Lightning and its open-source development at NYU.
“and then a bunch of other places, right?”
The Relationship with Facebook
40:50 to 42:01
Examine the collaboration between NYU and Facebook's AI research lab.
“They basically recommend that I open sources so that the other people in the lab can use it because I'm like somehow moving through my research faster than everyone else's.”
The Ecosystem of FAIR and NYU
42:01 to 44:45
Learn about the unique collaboration between NYU and Facebook's FAIR lab.
“that Jan LeCun became chief AI scientist at Facebook.”
Journey to Founding a Company
44:46 to 48:24
Discover the unexpected path that led to the founding of a successful AI startup.
“What were the conversations that led you to founding a company?”
Irony of Open Sourcing
48:25 to 49:29
Open sourcing slowed down research due to overwhelming demand.
“because I kept getting dragged into coding every night.”
Impact of PyTorch Lightning
49:30 to 50:40
Explore how PyTorch Lightning revolutionized multi-GPU training.
“and fixing bugs and adding functionality.”
Building Tools for Inference
50:41 to 55:28
Learn about the tools developed for optimized inference in AI.
Addressing Gaps in AI Talent
55:29 to 56:00
Discover how to bridge the gap in AI talent with innovative solutions.
“What do you think are the big gaps that you're going to need to fill in next?”
Leveraging Vibe Coding for Talent Development
56:00 to 56:22
Discover how Vibe Coding agents help close the talent gap in tech.
“and they're used to more like SK Learn NumPy type stuff.”
The Journey from Venezuela to the U.S.
56:22 to 57:11
Hear about the speaker's immigration story and its impact on their life.
“Having top VCs fund your startup now, you're at this$500 million plus AR run rate in under two years, serving 400 ,000 people.”
Learning English in a New Land
57:11 to 59:20
Explore the challenges of learning English as an immigrant in the U.S.
“And I mean, like, tell us about that experience.”
Pursuing a Career in the U.S. Navy
59:20 to 1:02:06
Learn about the speaker's journey in the Navy SEAL training program and its challenges.
“And I was like, I don't want to be like pushed to the side because of this thing, right?”
Leadership Lessons from SEAL Training
1:02:06 to 1:08:28
Gain insights into the leadership skills developed during Navy SEAL training.
“Because it's not just, like, there's, like, this crazy physical test you have to do, right?”
The Path to Joining Elite Programs
1:08:28 to 1:10:01
Understand the importance of hard work and networking in achieving elite opportunities.
“And then Jocko came and gave a talk, I remember.”
Pursuing Difficult Opportunities
1:10:01 to 1:11:40
Learn the importance of persistence and networking in achieving high-level opportunities.
“And eventually an invite can be forthcoming.”
Social Media and Staying Updated
1:11:41 to 1:12:24
Discover where to follow Will Falcon and stay updated on Lightning AI.
“recommendation extreme ownership 100 nice well that was easy and uh for folks who want to be be able to continue to follow your story, the Lightning AI story, where are the best places on social media to do that?”
Recap of Key Insights with Will Falcon
1:12:25 to 1:13:08
A recap of the discussion highlighting Lightning AI's achievements and philosophy.
“Once we're close to being IPO or at IPO, then we'll talk about that.”
Transcript
Automatic transcript. May contain errors.0:00Jon Krohn:What if you could go to one place to get absolutely everything you needed to train and deploy AI models and it was all cost effective? Welcome to the Super Data Science Podcast. I'm your host, Jon Krohn. I've got a really exciting episode for you today with the CEO of Lightning AI. They have just announced a huge merger that now means they have tens of thousands of physical GPUs. They have over$500 million in ARR. They've built a open source ecosystem, including PyTorch Lightning, which has been downloaded nearly 400 million times, and that's growing quickly. Just this really exciting ecosystem and products.
0:41Jon Krohn:And today, with Will, we get to dig into the man behind all this innovation and exactly what they've done. I hope you'll enjoy this one. This episode of Super Data Science is made possible by Dell, Intel, Fabi, and Cisco. Will Falcon, welcome to the Super Data Science Podcast. It's great to have you on. How are you doing today? Amazing. Thank you so much for having me. Excited to, you know, chat with you and thanks for having me on here. You're certainly experiencing some amazing times right now. It's really cool to be a part of it in some way. We've actually known each other for a number of years.
1:18Jon Krohn:And full disclosure, I am not an unbiased interviewer of you because I have a fellowship at Lightning AI, which is a really cool role that I appreciate. You know, you created for me and I've been doing for a year now and absolutely love working out of the Lightning AI office in New York, so many talented people. It really gives me a lot of energy. And yeah, hopefully I'm contributing a bit back as well. We love having you. So it's been fun seeing you do your thing. And I'm sure you're seeing kind of what we're doing as well. Yeah, and I think in both respects, we're just getting started. Yeah, well, I'm finally glad we got to work together because it has been a few years.
1:59Jon Krohn:Yeah. Yeah, exactly. And so the big news that we're here to talk about on air, and so even though I've had this fellowship at Lightning AI for a year, we've been kind of holding on for a really special moment to have you on the podcast. and now the time has come because Lighting AI has merged with a firm called Voltage Park and you guys have created something really big. It's mind-blowing to me. You've described it as the full-stack AI neocloud for enterprises and frontier labs. It serves 400 ,000 developers and companies,$500 million plus in ARR and that's been gained in under two years, which is astounding, $500 million in AI around two years.
2:42Jon Krohn:And now this company, after merger, Lighting AI plus Voltage Park, has 35 ,000 GPUs, making it the third largest neocloud in the world. I guess a nice place to start on this is congrats. Well, it's been, obviously, these are numbers from bringing both companies together. And I think it shows just how amazing each one of us kind of were doing on our own. And when we looked at it, we said, you know, if you brought together, I don't know, a Mac and Mac OS, you'd built the best laptop in the world. It was pretty obvious when that became clear. Yeah, software plus hardware. Exactly. Yeah, it's a really cool pairing.
3:24Jon Krohn:And you have actually been, so like Lightning AI was first a customer of Voltage Park. And then, yeah, it obviously matured into something much more. How did that kind of evolve? So, you know, if you're not familiar with Lightning, full stack software, kind of all the tools you need to build, train, inference models, like all that, I'm sure we'll get into at some point. So we have a lot of enterprise customers and developers and Frontier Labs who are, you know, they're doing the inference with us. For example, ones like Cantina, which is Sean Parker's new startup, right? So we do a lot of their inference there.
3:58And a lot of these players were using, like, very single-purpose tools to do that, for example, like other companies where that's all they do. And so these companies came to us at first maybe for like training and other things, and then eventually they kind of move into these other things that we offer because we offer all of it. And as we started to scale, we found the need to go beyond just AWS. As you all know, AWS is a very premium product, I will say, and so it comes at a high price. But it's lacking a lot of the specific AI tools that are needed. right because aws is built for cpu applications actually and so it's trying to be retrofitted back to ai and so then you have enterprises and startups that will like try to make up the gap by buying these like specialized products that only do that one thing and then they like stitch them together and then that creates a lot of problems right um because it's security and enterprise it would be like firewalls and security between all these products startups and frontier labs they don't see it this way but it creates a lot of like operational overhead and so anyway so so we started looking around for where can we get better compute for our customers because it was clear that they wanted the software but the compute prices were really expensive so we found this thing called neo clouds which i didn't know what it was until a year ago it's a brand new term right and neo clouds are basically a new type of cloud that are their gpu first so they don't they have some gpu cpus now but they were really designed around how do you make the best performing hardware on gpus work really, really amazing for things like training models, right?
5:31Things like InfiniBand and VAST storage and all these different things. Whereas the traditional clouds like AWS try to roll a lot of that stuff out on their own. And so it's not as high performance as a NeoCloud because a NeoCloud works directly with NVIDIA to do a lot of this. So we find the slew of NeoClouds. There's like hundreds of these things, by the way, right? I didn't know this. So we've partnered with like the top seven or eight. We start working with them. And then we start putting customers on on all of these enterprise customers. We're talking about people like Cisco, et cetera. And they need a special type of flexibility that I was skeptical that I could find outside of AWS personally.
6:09I think startups, they're okay dealing with this more. So we put the first enterprise customers and then they struggle, right? And we struggle to get them to adopt these NeoClouds for the majority of them. And then there's just like this new player that we have never heard of, Voltage Park, and they kind of come out of the blue. and it's a, you know, interesting story how they got started. So they bring all these GPUs online and then they, you know, I think what was amazing and probably every customer's reactions here have been how responsive they are, right? And so now, you know, we're one company now, so I'm talking formerly Voltage Park.
6:42You know, they call it the white glove experience, which I think it's a great product name. But, you know, hey, we're going to do what it takes, give support as much as possible to, like, make you successful, which is what you need as an early stage startup, right, ultimately. So they're doing this amazing job. And through that iteration process, we're able to actually win a few customers, enterprises, and increase retention. And so, you know, I was like, hey, these guys are serious. I know what they're doing. And at the time, you know, we were trying to figure out how to bring in a lot more GPUs online because Lightning, we've never sold compute.
7:17All our customers buy compute from somewhere else and then they connect our software to that compute. It's called BYOC, bring her on cloud.
7:23Jon Krohn:And that is part of the great offering that you guys have is that kind of flexibility, where that's something that I love talking about with the Lightning AI product in general, is that it allows you to bring your own cloud or to have access to a range of NeoClouds, as well as those traditional cloud providers like AWS that you mentioned earlier. You just have a dropdown box. You can see what the pricing is like at any given moment for the kind of GPUs that you want. and then you can provision them at a click of a button and have access to that in minutes. And you can have been working already, say for hours as an individual or as a team, just on a CPU, basically free compute instance.
8:06Jon Krohn:And then within minutes, when you need it, switch to however many GPUs you want on whichever cloud you want. Exactly. And so customers love that flexibility. But yeah, the pricing on the hardware was just a little cost prohibitive and you know i think a lot of the bottleneck that we saw last year for growth i mean you know the revenue growth in both companies like went super high i mean lightning alone we went at least 30 40x revenue increase without the merger right and um and we were really just always bottlenecked by compute we lost many deals like millions of deals that we couldn't land because we couldn't find the compute for it and so that was kind of the predicament we were in and and on the other side you've got this neoclouts where the only software that they offer is like kubernetes basically and that's it and so you're like hey i need to do inference i need to do this you know development i have training i have model hub things i have experiment management there's like dozens of things you need in between that we all just offer natively in lightning and so you know they looked at and they're like oh you guys have like a full mac os and i was like yeah and you have an amazing machine so let's get married right yeah And so that's how this kind of came about.
9:14Yeah.
9:14Jon Krohn:All right. So you've used the word inference pretty casually, and probably a lot of our listeners know what that means. But just to explain it a little bit, when you have an AI model, you first train the model. And then once you have it trained, you put it on some kind of production system like Voltage Park GPUs or AWS GPUs or whatever for that AI model to be used at inference time in production to be able to do something for users. So this is, you know, when you open up ChatGPT and you type something, that's inference happening on some cloud somewhere that OpenAI is paying money to. And so there's lots of, now that we have this big revolution of AI in more and more places, there's more and more demand for this inference compute.
10:01Jon Krohn:And you might know the stats better than me, but it's something like 99 % of all GPU usage is for inference, not for training. yeah i mean that that's right i think i'm actually sad that the name inference stuck because inference is a math concept which means like to take a math model and infer like have it infer things right so if you ever heard of like you know in stats and things like that their inference existed in that context and so someone decided to call it inference which means like have the model make predictions ultimately but there's a lot of stuff that goes into that uptime ultimately inference is for a developer and so being a container just like a web server except that it's receiving requests and how it handles batches of requests and streaming is a bit different right and so that became a whole thing which I have many opinions on what that is but yeah one one of the main things that people have to worry about there is the compute elements of it and then you know people strap on all these things so on this on this um 99 inference thing you know it's funny i think people are missing i think people see that and say oh the world is just going to be inference right um let's go back to like early 2000s and let's say that there's a product that exists called i don't know um s3 right okay and what an arbitrary name arbitrary name so this product people are like wow storage everyone needs storage right correct now is this going to be its own market is it going to be its own company like that seems weird but in early 2000s you probably think that i think um if you look at what happened over time was s3 was just a primitiverd you know rds ec2 all those things are primitives that if you put them in a bucket and you label that bucket what is that called cloud just it wasn't built yet right and that cloud mostly became aws i think you're in this process that it's you were we're going through this building process where the thing that's in front of us right now is inference but there are other primitives like vector dbs like training infrastructure like kubernetes like agents and there's dozens of other primitives that would be created that have not yet been created and not all of those are going to be their own companies i argue they shouldn't be all we're doing is that we're in the process of creating something called a cloud and that cloud i believe we're the first ones to actually have that today which is a new type of ai cloud this full stack ai neoclouc yeah Just put all those things, all those little pieces into a bucket and label that now.
12:45That's an AI cloud. And we have the first one of that. Yeah.
12:48Jon Krohn:So something that I've said to you before, I've come into your office and said, you know, how do you explain all of this functionality? There's so many different things that Lightning AI Studios can do. And so I guess that's something that we should talk a little bit about is just even this idea of, so, you know, we talked about the Lightning AI company, the Voltage Park company. So we're aware that lighting AI is the software. Voltage Park is the underlying hardware. Like Voltage Park, they are literally standing up data centers and GPU centers. And, you know, there's people physically screwing in screws and running cables and making it work.
13:24We're standing up our seventh data center today. Just for context, you know, Cursor, for example, we built their training cluster. There you go.
13:32Jon Krohn:That's super cool. And so, you know, that kind of gives us a sense probably as much as we really need to know for what, like, voltage park, like, how that works for this kind of audience, for AI practitioners, data scientists, that kind of listener. But the Lightning AI Studio part of it, which is the product that Lightning AI has been building for years and that you described that 30, 40x growth in the past year. Tell us about that Lightning AI Studio journey. you know you've described it as being kind of analogous to what AWS was but designed for AI so all these bells and whistles together but if I'm a user if I'm listening right now what's my experience like as I type Lightning AI yeah into Google and start using it for the first time yeah I mean the products evolved right I would say our first product was PyTorch Lightning which we'll talk about in a minute open source right for training models any kind of model, including LLMs.
14:28And then the problem became, okay, this is cool, and I can train at scale. So at the time, you know, this number, I want to come back to this, you said 99 % of all workloads are inference. I was training models of Facebook AI in 2019. And if you ask anyone, what's the number of GPUs that everyone uses, everyone has said one. 99 % of workloads are one GPU. But why? Is it because people only want one GPU or because it's hard to do multi-GPU? So even the Facebook cluster was massively underutilized. And then, you know, I rolled out PyTorch Lightning, and suddenly people started training on multiple GPUs at Facebook, which was like a very good team.
15:08And then it kind of rolled out, and, you know, eventually they trained most of their models like that. And then it kind of spilled out of Facebook and went to other companies, and that's how PyTorch Lightning was born. But you look at stats now, and it's not that, you know, I don't think anyone would say 99 % of workloads were a single GPU. right it's just that people didn't have the tooling to do that so i argue for inference it's not that 99 of the workloads are going to be inference is training is very hard and people don't yet have the tooling to do that we do today on lightning so if you're struggling with training you should go to lightning but so the studio is designed to basically be a you know people are starting to need this today so people are using cloud code right well what's the problem with cloud code on your laptop.
15:53Could delete everything. So people are starting to try to find these cloud environments. That's what a studio is.
15:59Jon Krohn:You're also limited in, you know, you can't really, you can't be training models in cloud code on your local machine. Exactly. So Lightning gives you the studio. So Lightning has many products. One of them is studios. The studio is a cloud development environment, sandbox, if you must call it that, that's persistent, that acts like your laptop. So you could go and run cloud code on there and leave it running overnight, and it'll do something for you. And you don't have any kind of problem that is going to delete your laptop or anything. And you can have 20 of these running at the same time, all on different GPUs and CPUs and things.
16:31Jon Krohn:Data scientists, it's time to talk about your tech. With Windows 10 support coming to an end, now is the perfect moment to rethink your setup. Enter Dell AI PCs powered by Intel Core Ultra processors. These devices are built for the demands of modern data science, delivering faster performance, smoother multitasking, and the power to handle even the most complex workflows. Whether you're training machine learning models or analyzing massive data sets, these PCs are designed to keep you ahead of the curve. Don't let outdated tech slow you down. Visit dell.com slash shoppcs to explore how you can upgrade your device and elevate your work.
17:08Jon Krohn:That's dell.com slash SHOPPCS. Yeah, and it doesn't matter if your preferred environment is VS Code or Jupyter Notebook. Well, that's the idea of how you connect to the environment. Oh, right. Yeah. Yeah. So I want to separate the environment itself from how you code in that environment. I see. Right. We provide a cloud kind of VS code interface and we have vibe coding in there as well. So you can describe things. But if you prefer to use cursor locally, go for it. Right. And if you prefer to use cloud code, go for it. We don't, we don't care how you're connecting to that thing. The point is you're getting a cloud environment that can be shared.
17:44If you're an enterprise, it can be audited. For example, you probably don't want your developers to have local files that are customer sensitive on your laptop, but on there you can have them, right?
17:55Jon Krohn:Yeah, yeah, yeah. It's a very flexible environment. It's allowing you to, yeah, be there in the web interface. And that's kind of the default screen. So you go to lightning.ai, you can create a free account. We'll probably, we'll have something in the show notes for people to be able to skip the queue and get access to some number of free monthly credits. So, you know, check out that link to get access to Lightning AI right now. And then from in there, you're in this VS Code-like experience that you described as a default. And from there, you have access to everything that you imagine you could need to train and deploy models.
18:36Jon Krohn:I feel like we could spend literally hours and I have seen you demo for literally hours on all of the functionality. So I don't know if people know this, but like Lightning AI itself today is built on studios. Every single developer at the company today codes on studios, whether you're using Golang or Python, whether you're training models, doing inference or coding a web app, like everyone does it on studios today because it's easy to reproduce. It's easy to onboard new people, but they like will probably code from the local ID, but it's connected to this remote thing. Right. And then I think the other thing to notice is the studios themselves, like I said, are persistent.
19:12You can then SSH into them, you don't have to use a web interface either. There's a whole command line version of it where you can just start a new studio and just open it up. And it's kind of like you open a new terminal that's a remote terminal. And now you can do whatever you want. So there's that developer experience as well.
19:29Jon Krohn:For sure, yeah. So whether you're kind of more comfortable and it's kind of sometimes easier to get started, maybe right off the bat, especially maybe if you haven't been coding in a little while, you're coming back to it. You're like, oh, I've heard that all this tooling makes training and deploying AI models easier than ever before. Or maybe you just get started within the web app itself. But if you are already used to doing things all the time in Cursor or VS Code or Jupyter Notebooks or whatever environment you prefer, you can very easily, at the click of a button, connect that, or yeah, just like you said, a terminal.
19:56Jon Krohn:Any of those kinds of experiences connect to this remote studios instance. And so you get all of the security, all of the flexibility associated with studios. And you can, something else that's really cool is when you're working as a team, just like working on a Google Doc together, you can be collaborating in the same environment, pass something off to someone else working in another time zone. And if you work at an enterprise or whatever organization you're in, everyone there is happy with the security of that because everything's happening in their controlled old environment. Yeah, and it can run on-prem, it can run in your cloud, so it's fully secure with like all the guardrails enterprises need.
20:37No, but yeah, I mean, look, I think it's really a collaboration tool, first and foremost, a development environment tool, I guess, second. And then there's a gap now where you have normal developers, not normal, but like non-AI developers who are maybe data scientists, who don't know how to train a model, who don't know how to do inference. You can come on there and we have our vibe coding and you can choose the skill you want. And you can say, hey, I want to fine tune a model or even I want to do data science. right and it'll actually help you vibe data science and vibe ml and vibe like analyze a notebook or whatever so it'll complete sales for you it'll do all of that there but you happen to be in the cloud now and you've got cloud resources and it can like submit jobs for you it can spin up like a cluster of a thousand gpus um so you know it's basically the vibe coding helps you do the mlops i guess is the right way to say it yeah it's something certainly worth being excited about
21:32Jon Krohn:and yeah, that's why I'm, yeah, I'm so honored to be associated with the company and all the great things you guys are doing. Let's talk a bit more about this NeoCloud before moving on to PyTorch Lightning. There's been press, specifically I have this McKinsey article that I'll link to in the show notes for people. I think you're already familiar with it And in it, it claims that neoclouds, so kind of bare metal as a service, the economics are fragile and they risk repeating mistakes that happened when AWS was starting to become popular. There were lots of companies that were trying to be like AWS, but smaller over-leveraged players were acquired, sidelined, or forced into niche roles.
22:16Jon Krohn:And so in a Forbes article about the merger with Voltage Park that we have also in the show notes for you, you told Forbes why you shopped around for neoclouds despite the fragile economics and why in particular you chose Voltage Park. And it sounds to me, from reading this McKinsey article, like all of the things that they're concerned about with neoclouds, through this merger of Lightning AI and Voltage Park, you avoid all of those issues. Exactly. So they don't have any debt, which is amazing. And that allows us to do a lot of things that other clouds can't do. The CEO who is running Voltage Park, who's now on our board, he's a finance guy, so ex-finance guy.
22:59And so he's very good at making profitable businesses and being rigorous about that. So I'm excited to partner with him on how do we keep scaling that. And, you know, both our views are that if you want to have profitability in the future, you have to kind of check the amount of leverage that you have in the business. And this means loans and so on. Not all the neoclouts had the kind of start that we had. And so they didn't have that luxury. And so they had to take out loans. And a lot of those loans depend on your customer not defaulting. it's a bit like a mortgage crisis really where if your customers default and by the way most of the customers are startups in those companies in our company it's not like we have some you know the cursors and so on but these are like high quality companies right like growing super fast the vast majority of our customers are enterprises with great credit and so we don't really have that issue but it takes a lot to sell a neocloud to an enterprise you need a lot of security blah blah blah things that we've been doing for six years now lightning and and so we don't really have that problem right and so really the the thing that we're focused on is how do we make enterprises successful and give them an alternative to the traditional clouds and also work with them because a lot of enterprises use all the clouds including us right they use aws they they work with us you know it just depends on the different use cases so we still partner with those clouds i think we're more like a hybrid cloud than anything else and um we just happen to have our own supply now of of data centers.
24:31So just the fundamentals are vastly different. And a lot of these companies went public. A lot of the credits that they need to get loans to build these data centers are based on their market caps, right? Because when you're a public company, what's your credit? It's actually your market cap. And so what's happened in the last six months is all the new cloud prices have dropped by like 60%, maybe more. And so then now the creditors are like, wait, how's your market cap doing? And so, you know, it can, if, if, if this continues to happen and it's a little too low, then they're going to get calls on these loans.
25:07And then we're kind of to like, actually it's like literally equivalent, not equivalent, but like, it's very similar to a mortgage crisis situation where the banks were over leveraged and the federal reserve had to bail them out at some point. And so that probably happens here. I, you know, we can speculate who that would be, but there are a lot of people who probably be interested in doing that. But yeah, it's a new industry and it needs a bit of a TARP program or support from big players to help the ecosystem be successful ultimately. So yeah, that's where I think we're at. But I think we're very differentiated in that we have real revenue that's growing extremely fast, extremely high margins.
25:47And we don't really have a lot of debt. We'll probably take on debt in the next few years, but very controlled. And we have enterprise customers.
25:54Jon Krohn:Speaking of prices and margins, something that I'd like to dig into a bit here is you mentioned earlier in the episode how AWS can be very expensive for training or inference for AI. And how do you manage to maintain these good margins that you just cited more recently while also offering these great deals? So there's a recent Lightning AI press release that highlights a claim that you're cutting AI costs by 70 % while offering a free tier and global access to everyone from undergrads to Fortune 100 companies. How do you manage to square all of that at the same time? A lot of smart financial engineering, I guess.
Read the full transcript
26:38And, you know, we just have a lot of enterprise customers. And so they're the ones who are kind of spending the most money with us because they have enterprise contracts and so on. And, you know, I could cut developers off 100%. It would probably be more profitable for us. But, you know, I was a grad student at some point, and I don't think I could have done what I did without having access to this kind of compute. So, you know, I've given personally a lot to the community. I've given a lot of open source, you know, PySource Lightning. I've spent a lot of years working on open source things. So I'm an open source person by nature and try to give back as much as I can.
27:15I still have to build a business and make it profitable and take an IPO, right? So within those constraints, I try to do what I can. And I think the free tier is something that I'm super proud of our team for doing. And yeah, I want to support the next generation of builders, of undergrads, grad students, and help them figure out how to build the next ByteSource Lightning or the next OpenAI, 100%.
27:36Jon Krohn:Yeah, so speaking of giving back and your grad experience at PyTorch Lightning. Let's talk about that now, because that was kind of, that is, you mentioned already earlier in this episode, how PyTorch Lightning is kind of what led to the Lighting AI Studios product. And now this merger with Voltage Park. And so let's talk about PyTorch Lightning. It's been downloaded over 300 million times, I believe is the last kind of - Plus to 400 now. Yeah, maybe by the time this episode is out. And yeah, you developed that while you were a PhD student working at NYU. you and you had some pretty well-known phd supervisors so yannicka and kyungung cho who also i mean he's been cited hundreds of thousands of times he's the next yon for sure and he's got papers like neural machine translation that are huge and that everyone in ai needs to know so yeah tell us a bit about that experience maybe you can even back up a bit and give us a bit of like what led you to that program founding pi tors lightning and the big impact that i know that it had at Facebook and beyond?
28:36So I actually started working on PyTorch Lightning probably around 2015, before I was a grad student. I was a grad student in 2018, a few years later. So I was an undergrad at Columbia, and I was at a neuroscience lab. And this is when autoencoders and GANs were like a big thing. And we were learning to train models that could decode neural activity from the brain. So it was a partnership with Stanford. So at Stanford, they had monkeys that would collect neural activity from the retina. So the monkey died, they would cut out the retina, and then they would show images to it, ImageNet specifically.
29:14And so you had a way to measure what was happening in the retina, and then you had what was shown. So you had an X and a Y. And at Columbia, we were doing the computational work. And so we would get that data, and then we would try to figure out what is a model that can take those raw neural activities, which are pretty much zeros and ones and some simulations, and basically decode the image that the eye saw. And so that's called Gen AI today. That was not called Gen AI back then, right? Because you're imagining an image. Back then we were using something called autoencoders and then GANs, which people I think are still using.
29:49I'm still very bullish on autoencoders. And so, you know, in 2015, you didn't have TensorFlow. You didn't have PyTorch. You only had Fiano. And so the first version of PyTorch Lightning wasn't even based on PyTorch, it was actually on Tiano. And it was around how do we try many ideas at once without rewriting the code all the time, which is why I was forced to write that.
30:12Jon Krohn:I didn't know that. What did you call it? It wasn't called PyTorch. It was just called Research Lib. Just my Research Lib. Gotta name this thing. No, I didn't really care because I was just the only one using it. I met a few people in my lab. Yeah, and I remember this postdoc at the time, Scott Lindesman, I think is his name. So he's now a professor at Stanford. He's a statistician, best statistician I work with ever. And he was like, hey, Will, you should formalize this into some sort of, he called it a harness, so we can use it in other projects. And I was like, okay, yeah, sure. And so I kind of tracked it out.
30:52And then I was like, and then we started working on other projects, right? And so I was like, oh, very quickly, I can just kind of rerun the same stuff. all i need to change is like the computational part of it which is like the forward and backward which is where the science was for us at the time and then tensorflow came out um i rewrote it in tensorflow before actually this is before i spoke to scott so i rewrote in tensorflow because you know the paper we're working on we need to publish to new europe's new europe's deadline was like march this is like november and every time we went through image net it was 21 days to for one epoch and the autoencoder needed like hundreds of epochs just to see if that one worked so you can imagine that we're not going to complete this and so we i think it was laptops so i went out to my pi and i was like hey we need to buy gpus and so then i bought like four uh 24 gpu something like that no
31:47Jon Krohn:yeah 16 gpus that's probably around the era of like 1080 ti's yeah so i bought 1080 t i bought I bought a bunch of 1080s. Yeah, I actually have the box. I'll probably bring it to the office, but it's cool. I still use it for gaming and stuff. So I built these four boxes from scratch. So when I went to Voltage Park to the data centers, I was like, ah, so this is how you do it at scale, right? I mean, they're experts at this, but I think building those machines definitely gives you a good background. And I actually posted on my GitHub the instructions. So if you guys go to my GitHub, you'll see exactly everything, all the parts and how you do it or whatever.
32:19This is like eight years ago at this point. so I built these machines with four GPUs on them so we could try one idea per GPU so that gave us some boost but it wasn't as fast as I needed so then roughly at that time TensorFlow came out and one of the things they were talking about was multi-GPUs and I was like oh okay cool let me try that so then I rewrote all our code into TensorFlow and then I actually did get it to work on four GPUs and then we were able to train that model in like maybe an hour from 21 days to an hour yeah yeah And so then I was like, amazing. And then we actually did eventually publish a paper to New Rips, my first paper ever, and also my last paper to New Rips.
32:59So I got super lucky, right? I was not the first author on there, but, you know, the team was incredible. And that was as an undergrad, right? And so then fast forward, you know, I started my next project with Scott, and then he's like, hey, rewrite this. And then at the time, PyTorch came out, and then we looked at it, and it just looked more mathematical. And we were doing more probabilistic models at the time. So I needed to actually look at a function and then map it to the code. And so it was easier to just write the literal math on there and then model the probability solutions as a model, right?
33:29So like P of X given some parameters, some data, that data was a model, right? And so you could take all Bayesian stats probability and just put models around them. So that's what we did. And then, yeah, we started training this model. Is that paper, I don't think we published anything because I think I screwed something up. I don't know, maybe. I don't know if the data had data. It was like weird data, right? So in neuroscience, you have to like know if the data is good. I don't know if that was great data. I think he did eventually publish something about it, but not with me. And so, so then, you know, that, that was it.
33:56We were using this to iterate. Many years later, I, um, I, you know, I started a company, sold it. And in that company, we use PyTorch separately. I didn't use PyTorch Lightning because we weren't training models. We were just doing what people call inference now. So I was running like large scale inference across 60 ,000 customers. This is low income students. We're helping them figure out how to pay for college. So it was all in text message. And it was some of the first production systems I've ever, I think I've seen put out there. And we gave a talk at data-driven NYC. So we gave a talk there.
34:26And, you know, you guys can watch my talk. And it's like, I called it bot-powered humans, basically, because I was trying to figure out how to use AI to augment the humans. And so these were people that we hired. We built a little UI for them where they would have this conversation with the students and then that UI would suggest what to say next. But the way we do it is it would treat you as a game. We use reinforcement learning. Now RL is super hot, right? But back then, so I treat the conversation as a game where your message is a step in that game and then my message is a step in that game.
35:03And the goal of the game is to get you as much money as possible for college. and so the reward function is optimizing for your financial aid and then the games are the words that i say to get you there and so that's what the model was trained on and it would take your whole context convincing people to give you money no convincing asking the right questions to help us get them the most money for college i see so like the model is like hey whatever my you know my function says um i don't really know your sat yet and if i get it i'm gonna i can i can figure out my policy better and see if it can get better so it would ask you these questions and you know we didn't have lms back then so it actually it was more of like a ranking so we had a list of questions and they would like choose what's the best one to ask with some probability and then ask that question and then answer and then play this game back and forth we ended up getting a bunch of people about 60 million dollars with the financial aid with this oh my god yeah well like we sent like like um inner city kids to like princeton that company is called next gen vest correct yeah yeah and you sold that to common bond yeah so they acquired us so we're doing this you know this is like 2017 2018 or something so you know think about it now like we were doing reinforcement learning with models that i guess are called language models now and it was all inference in production like this is like people's problems today and we're doing this and we have a whole company based on this so i need real-time inference it's got to be fast and this and that so then these guys come along and they're like this is insane so you know we they basically offered to acquire the company and i you know i i wasn't the the ceo my my co-founder was so she's like let's do it i was a cto and i was like hey cool whatever you guys want you know i'm like do whatever you guys want and then um and then i had an offer to join nyu to basically do a phd so in this whole process my undergrad pi was like you should consider a phd at some point and i was like i mean i went to like public school in florida i'm from south america i was like never even consider going to college in the first place so my pi is like you should do a phd i was like you're insane and i was like fine i'll apply so i applied and i get in all these places randomly and then
37:07Jon Krohn:one of the just apply to work with uh beyond the camp uh these people were big but they they weren't like they're i mean they're very big now but then they were big to us right so they still were taking students right and i i ran into jan randomly at a like an event and he saw a paper the the neural decoding thing and i think um you know the the those that crew of like yasha banjo hinton and young think a lot about like what does new york science say about things and so they got interested in the vitamin to the lab meetings and that's actually when i met kyung hyun um he was sitting there and i was like hey what do you do here i don't know who he was he's like oh you know research and stuff he's very humble and then the whole meeting i was like dude you're the guy in those papers and then every paper right after that was like he was on it right and so and so i was like hey yeah I'll apply and you know I asked him I'm like would you take me he's like I don't even know you and I was like all right fine and then over the year I think I you know I could go to lab meetings and got to know them and then I did apply and then actually got into his lab and then Yoshio's lab so wait let's talk about this for a second so you're going to lab meetings but you're not a lot of students you're just showing up yeah and I'm just you exited next-gen vest yeah and so you're gonna know I'm still at next-gen this time okay and I'm still going to Columbia like I'm just taking small like I'm doing fewer classes because I'm doing the company thing.
38:24Yeah, and so that was kind of it. And so I'm doing all of these kind of three things at once. And then the acquisition offer comes in, and that kind of matches PhD applications. I get in with Yasha, Benji, and Mila, and then with Jan, and then Columbia, and then a bunch of other places, right? And so I go to Montreal, and I'm like... Congrats, I did not know this. That's wild.
38:46Jon Krohn:Wow. Yeah, and actually the guy who I was going to do my PhD with, Mila became the VP of of chat GPT he's like one of the chat GPT creators right yeah yeah no no Liam Fetas oh sorry he was from university yeah Liam just started periodic labs you should check out his company so so that lab was insane I mean so let's go over I think that spend some time there I don't know how they related Hinton whatever you know the Canadian mafia these guys are amazing right yeah there was a complete brain for it on my side it says give her was in Hinton's lab for a PhD yeah I made that big 2012 Alex net paper with Alex Krzyzewski.
39:18Yeah, but make no mistake, I mean, at the time, Mila was, I think, the top lab, I would say, and NYU is pretty much up there, but it was a newer lab, right? At least to me, it was a much smaller lab. And so I was like, I would love to go to Mila, but it's freezing, and, you know, I'm from South America, and I'm already in New York, and NYU is amazing. NYU felt like a newer lab to me, but they weren't. They've been around for a while. It's just a lot smaller. Mila's like a factory of researchers. At the time, they already had hundreds of people, right and so i was like hey i'm just gonna make this easy and go to nyu and so i started there and that coincides with acquisition and so then you know they offer me a bunch of stuff to like join this acquiring company and i'm like no i'm good i'm gonna start this phd with these guys because they're like world class right and so i leave and i started the phd at nyu and then i dust off by torch lightning um and i started doing my research with it and um and then yeah i think Gian's being a big proponent of open source.
40:14And so he's kind of the reason why Facebook open source so many things, PyTorch included. You know, people don't know this, but a lot of projects were incubated at NYU. Things like ThinkSK Learn had a huge, huge help from NYU. Like people were working on it there. Adam Paschke, who wrote the original PyTorch, was found by a professor at NYU to work on that project. I think NumPy, there's a lot of projects. There's Librosa, I think. It's like a sound thing. So a lot of these projects, like NYU has a weird history of doing this. I didn't join for this. I didn't know I was going to create an open source project, but I think they encourage that a lot.
40:49And so a lot of these things have been teated from there. And so, yeah, so I come in. They basically recommend that I open sources so that the other people in the lab can use it because I'm like somehow moving through my research faster than everyone else's. And, you know, so that's cool. And, you know, my advisors are like mostly focused on the research. So Kyung-Hun's like, hey, whatever. As long as you do your research, I like they're like we don't really care you know I'm like okay fine so then this gets open source and then I joined Facebook many months later and when I joined I have to like name this thing so that other people do like use it and so I was like well makes me move super fast through research so I guess lightning and then I just looked at like a I don't know lightning bolt and like NYU color and I was like there it is and I just it was like 30 seconds right wow and then that was it and then and I put it on the read me, and then I started my internship at FAIR.
41:43And then -
41:43Jon Krohn:Facebook AI research. And there's also, there's a big, you talk about all of these collaborations or this atmosphere, this environment of supporting open source that NYU has. They also famously have this very strong relationship with Facebook now Meta. So Facebook AI research, they brought in, it was around 2013 or 14, that Jan LeCun became chief AI scientist at Facebook. Well, he founded FAIR with, I think, a few other people, but he's one of the FAIR founders. Yeah, yeah. Yeah. And, yeah, and so there was a big revolving door between PhD students at NYU and people working in this prestigious AI lab FAIR at Facebook.
42:25Yeah, it's incredible. I mean, if you're, you know, we're in New York right now. If you go down to NYU, the Facebook fare office used to be on 770 Broadway, which is like, yeah, one block away from NYU. And so you have lab meetings and you're like, is it at NYU or Facebook? And you just like, it's the same distance. You just walk to one or the other, right? So you all have Facebook badges. You just, it's literally, there's like kind of no friction because the Facebook clusters and data are, they're not Facebook internal. And so there's no customer data or anything there. It's all like academic stuff.
42:59and so you know um you have collaborations going on all the time but the professors are dual they're working at nyu and they're also working at facebook now that's fair during his heyday 2018 2019 i think most fair people will tell you that was like the big time and maybe i missed a year before that but you fair at the time was only 100 people worldwide maybe 150 of that including like admins and things um you had people like that way right um who was the chief scientist a hugging face after he left fair now he's got his own company you have people like jason weston who create a lot of like the machine translation things and you know the the the thing was so interesting like i showed up to my desk this is like i remember like 14th floor i want to say and like i said i sit down by the window there's like me and another desk here i don't remember who was here maybe another intern but to my right is sumith who's like leading pi torch at the time chintala yeah and then to his right is like the whole pi torch team which is like seven people at the time.
43:57And then behind me is, um, the kind of leadership team. And then behind them is the, the team who wrote, uh, the first distributed training models in the world ever, who then they eventually became character AI. They left Facebook and started character AI. And then now they're at thinking machines. Right. And then you had, um, yeah, PyTorch, these guys, and then me, I guess. And I was probably the first one to get like funding and then start a company. And so then I think that let people know that, hey, you can probably get funding to do a lot of these things as well. There might have been someone before me, but -
44:32Jon Krohn:And that was Grid AI at the time was the name. Yeah. And so how did that happen? How did you, you're in this great ecosystem, both with this commercial relationship with Facebook, but also all the open source stuff that's encouraged by Facebook, by NYU. What were the conversations that led you to founding a company? Was it because you'd already done it with NextGenVest? So I did not want to start a company, first of all. I think if you've done startups, you know they're hard. And I just come from like, you know, a grind. And so I was very happy to like become an academic for a bit, like read. I mean, it was amazing.
45:08I was like reading books and doing research. And like, I love science. I just love thinking about things, going into really hard, impossible problems and like going as deep as I can. So if everyone's thinking ever like, should I do a PhD or not? That's what I would ask you. like, do you love science so much that that's what you want to do? Because if all you're doing is for a career step, you're going to hate your life like two years into it. But I loved it. It was amazing. And so I did not want to leave. I mean, the problem is PyTorch Lightning took off. And I like I wasn't trying to make that a thing.
45:38It just took off and people started using it. And so then VCs started pinging me. They started emailing me being like, hey, we see your open source project. It's working this and that. And I was like, hey, like, I'm good. like I just sold a company like I just want to focus on research. And they're like I'm more interested. Yeah exactly. And you know now at the time I think about it like all these things have names today but what I was doing back then was pre-training world models. Today that has a name. Back then we didn't call it that. But I was training me pre-training models on four, 8 ,000 GPUs, single person, doing all the things that the pre-training folks are doing like OpenAI and whatever like how should you scale the thing?
46:14What gradient should you use? Like what's a distributed strategy? How do you prevent faults? and like what if the loss function is this so most of our research at FAIR at least for me was more theoretical so it's more like how do you change the math function so every model has a loss function that like optimizes something and so it's more like how do you change the math function like do you add a regularizer here do you use this one or this one and so I think what I learned was cool and I have lost that skill a little bit I think I need to get it back is when you leave undergrad and you've done math, right, you're really good at reading math, it'd be like learning a language and you're really good at hearing it, but speaking it is something you have to practice.
46:56And speaking it is more like, let's say you wanted me to build a model to model this world here that we're in right now. I'd have to figure out how to model the wall and what does gravity look like and how does the lighting hit and this and that. And that's a math equation. So being able to translate that to math, that was a skill that we developed a lot. and that was what FAIR was really good at. Oh, really? Yeah.
47:15Jon Krohn:That is a cool skill. Exactly. And so then that's what we train models with. It wasn't, oh, here's a model, train it. No, it's like, what does the training algorithm itself look like? And if you want to have better representations of the world, how do you accomplish that? Do you make things come together or separate? Do you embed images or not? And if so, how do you do it? Do you use this type of embedding or this type of function? And what's the similarity between them? And what if you add a penalizer and regularizer and this? So most of pre-training today is actually not that. it's more like the engineering of pre-training.
47:45We were not just doing that, but we were also doing the math. And that's FAIR. That's why FAIR was what it is. These are like math people who are also really good engineers and they can solve these problems together. Right. And so that was called pre-training. And so, yeah, I mean, I think the VCs were kind of onto something, I guess, very early. And, you know, they asked me to come in for meetings and then like I flew out on like a Thursday and then by Monday I had like 10 term sheets. And it's like, I kind of got out of hand. And so I was like, well, I guess I'm starting a company. And so I called my advisors and I was like, hey guys, so I got all this money that they're offering me.
48:18Should I do this? And they're like, well, your research is going super slow. So probably. No, I mean, I've been telling them. I was like, hey, listen, like, because I kept getting dragged into coding every night. And so I was like, research, run, you know, submit my training runs. Okay, it's running. And then I'm like, okay, go merge this pull request from PyTorch Lightning. And then it was like stuff that I wasn't doing. People are like, oh, can I have like this function for RNN? like the way you do a loss function for an RNN right um like backprop through time i was like i'm not using backprop through time but like now i have to implement it for you so i started working for people i was like this is terrible i almost shut down pites are sliding in september i was like this is distracting and so so my advisors were already kind of like hey like your research is a bit slow and i was like i was transparent about it i was like hey guys like i just keep
49:02Jon Krohn:getting dragged into this there's a funny irony here which uh i don't know if you if you've noticed this before, but you telling me this kind of history that you ended up doing the PyTorch Lightning open source project because you were doing your research much more quickly than everyone else. And they were like, you've got to open source this thing so everyone else can take advantage of that. But then once you did that, it ironically slowed down your research because all of a sudden it became so popular. You have tons of people asking you to be reviewing PRs and fixing bugs and adding functionality.
49:36Jon Krohn:So yeah, there's an interesting irony there. Well, ironically, also, I didn't want to code. That's why I brought in the first place so I could focus on the math, not the way the distribution happens or the training algorithm. And then people just ask for features that ended up being their engineer. And I was like, well, that's kind of what I was trying to avoid in the first place. So it's fine. I think ultimately it's great because I ended up putting so many things into this that were like very specialized knowledge. I was at Facebook AI at the time, because, you know, Sumith was helping, the Character AI folks were helping.
50:07They weren't character back then, right? So we had all these amazing people who were just world-class engineers helping me do all of this. And so I learned a lot. I put it in there, and as a benefit, the whole world was able to benefit. And now, yeah, 400 million downloads later, the reason you can train multi-GPU and multi-node and fine-tune LLMs and pre-trained LLMs is because we spent a lot of time at FAIR putting that into PyTorch Lightning. Yeah. And then all the contributors that came after. i try to add it up the other day i think we've had about 500 000 hours of engineering time into pytorch lightning like if your pre-training code can be better than that good luck right it is a
50:39Jon Krohn:really great tool and that's how i became aware of you in the first place you were doing talks around the new york meetups around pytorch lightning so i kind of became aware of you i started using it it became really obvious to me that it was invaluable for speeding up especially things like multi-gpu training it was something that was an impossible nightmare in almost any circumstance to get going before PyTorch lighting and then that made it easy and so I actually I think it's four years ago now I did a half day or a full day training at the Open Data Science Conference East in Boston and I got that edited in and put it on YouTube so I'll have a link to that in the show notes which is kind of like an introduction to PyTorch lighting I had some other libraries in there as well but that was one of the key libraries for helping people get off the ground with training their own models and and and deploying as well inference is a big part yeah for sure yeah i mean i think there's a common misconception that pytorch lighting is not for gen ai that's factually wrong like literally i was training world models before they were called world models people have trained llms all of nemo from nvidia is built using pytorch lightning uh aws has pre-trained llm's using pytorch lightning linkedin recently published an article like six months ago how they train a hundred billion parameter LLM using PyTorch Lightning.
51:55I think that at some point if you are wanting to mess around with like the distributed mechanism which is like DDP or FSTP and you have your own way of sharding gradients or whatever then the PyTorch Lightning can get frustrating because remember it was built at a time when that's not the thing we did. We focused on the math. So then we added the flexibility to have people plug in their own strategies into this and then we also created something called Lightning Fabric, which kind of gives you more less managed PyTorch, because PyTorch Lightning is basically managed PyTorch, less managed PyTorch, but still you get all the beauty of standardizing the tools, having the flexibility, having different training strategies, but it's a bit more hackable.
52:39So as you get deeper into a frontier lab situation, then probably the journey is like you outgrow PyTorch Lightning because you need more control, you then go to Lightning Fabric and then if you really, really need more control than that, then you go to Raw PyTorch but then in Raw PyTorch you need to know exactly what you're doing because you can mess things up. And so as a result I think only a few of the frontier labs will use PyTorch directly the vast majority of like enterprises and startups use PyTorch Lightning and then sometimes you've got a PhD who's like, oh I can do better and then they try PyTorch and that's fine but then their manager really should find out and they're like but like standardized with the rest of the team, right?
53:17But no, it's used for everything today. Yeah.
53:20Jon Krohn:And you kind of gave us a sense there of the other open source projects. So, you know, we obviously, we started this episode by talking a lot about Lightning AI Studios, which is a product, a web-based product that you can use as an individual or as an enterprise of basically any size. But in addition to PyTorch Lightning, there are, you touched on some other open source projects out of Lightning, like Fabric. Like, I think there's others, there's like a compiler. Yeah, I mean, we've open sourced pieces of the stack that you need. For inference, like we found PyTorch Lightning. PyTorch Lightning is for model training and fine tuning.
53:55You can use it to do forward passes. There's a mode where it just freezes the model and it'll do that. But for an inference, like production inference, the forward pass is only like one ingredient. It actually needs to be more of a system. Maybe like you connect the ragdb and this and that. And so people will use something like fast AI to do that. but then you have to implement your own batching and streaming and this and that. And so then we created something called LitServe, which is byte-to-sliding for inference, basically, where you can grab multiple models if you want, you can put them all together, you have full control, but you get automatically all the features like batching, streaming, security, all the things you'd need to put an actual inference server together, but you have full control over what the models are, right?
54:37I think to my knowledge, that's kind of the only tool that exists out there today. There's other stuff, like I said, FastAI and things like that, where you can write it yourself. And then things like VLLM and SG Lang are more like if you were to take Let's Serve and write a very opinionated, very optimized LLM serving engine specifically with like batching, streaming, like KV caching. And you started with Let's Serve to do that and you put all those pieces perfectly together, you would end up with something like VLLM. So it's more like a high level tool than anything else. So if all you want to do is just hit run on a CLI, then LitServe is not that.
55:12But if you've got multiple models and you need to build your own inference engine, then that's what LitServe lets you do. And so tools evolved over the years and we found that there were certain gaps that we needed to fill in the market. All of Lightning today, all our inference runs on LitServe. And we have so many customers run on LitServe today that are like consumer products with millions of users. And it's perfectly scalable.
55:33Jon Krohn:What do you think are the big gaps that you're going to need to fill in next? Like obviously you have this big ecosystem of open source tools and commercial products that allow us as AI engineers, data scientists, people who just want to make AI work in the real world, tons easier. What else is missing? What else is next? So we had two gaps. One gap was having enough GPUs for everyone, which I think we solved now. And then the second gap is a talent gap, which is a data scientist who wants to do inference or fine tune or something. and they're used to more like SK Learn NumPy type stuff. So that's what we have our Vibe Coding agents on our studios for.
56:12Now bring those problems there, describe what you want, and then it'll just do it for you. And so we want to close the talent gap with the Vibe Coding part. Cool, I like that.
56:21Jon Krohn:I think I'm going to start to wrap up the technical questions here. But there is one part of your life, your earlier life, that I still want to dig into you because as much as all of these things that you've described in this episode across, you know, working for people like Kyungyung Cho and Yannicka and getting the opportunity to work with people like Yoshua Bengio, you know, working with top VCs. I never worked with him, but. No, getting the opportunity. Having top VCs fund your startup now, you're at this$500 million plus AR run rate in under two years, serving 400 ,000 people. And so, you know, that's, we've talked about this kind of later part of your journey.
57:03Jon Krohn:But the beginning part is also really interesting, like before Colombia, which is kind of the earliest that we've gone to so far. So you, if I understand this correctly, and this was something I didn't even know until we started doing the research for this episode, is you came from Caracas, Venezuela at 13 years old, and it came to the U.S., and didn't didn't even speak English. And I mean, like, tell us about that experience. And I don't know, maybe somehow the way, you know, learning or being in that unknown environment, or are there any kind of links to that immigration story to what's allowed you to be this incredibly productive human later in life?
57:49Yeah, I mean, I did not learn any English in Venezuela. I, we got, we had English classes, but you know, in Latin America, there's this like status thing, if you know English, right? It's like, oh, you're so fancy. And, and so I didn't believe my teachers knew English. I think they probably did, but I was like, nah, I was like, you, no way. So I just ignored them. I was like, I'm not going to learn whatever you're teaching here. And so anyways, we were, we win this green card lottery. We moved to the U S and I remember like showing up, i don't think it was like at immigration but like a few days after and they force you to take this test right and i've never seen a scantron because that's not something that we have in latin america and so i'm like what is this thing and so then i have to go so that's where you like have like a
58:34Jon Krohn:pencil and you have to fill in a like a little bubble and you pick which one is the right answer yeah so they give me like an english test like just a standard like right stuff and then like math test and so on the english i'm like i don't understand anything they're like you still have to answer the questions so i just randomly picked i just i think i just drew like a uh a diagonal across all of them and then the math is universal so i wrote that and i think that's what kept me from going back one grade because you probably would have been held back so i came straight i think it was like fifth or sixth grade sixth grade i think and um and so then you know if you're if you're someone like that you enter this program called esl right um english as a second language and so this is in virginia and uh you know if i'd started miami i probably wouldn't speak english honestly um but in virginia you have to learn and so you're so out of place there and you're like put into all these weird um situations where like the american kids get things that you don't get and so i was very annoyed at that and so i was like i need to get out of this thing quickly and so i made it my life mission to learn english as quickly as possible i think i was out of vsl sound like four or five months.
59:42It's just like a two-year program. I was just very annoyed at it. And I was like, I don't want to be like pushed to the side because of this thing, right? And so that's kind of how I got into it. And I think if you speak to a lot of immigrants, there's like this need to want to simulate quickly so that you don't stick out. I think people change your names. You see this a lot in like Indian culture or Asian culture. They'll like want to change your names to simulate more. Luckily, my name is William. So I was like, great. But I need to figure out the english thing and not have an accent at the same time so i did that and um and so that's kind of what helped me i guess uh i mean i don't know the only trait that i could think of is like if it annoys me a lot i want to solve it quickly right and the cloud annoyed me a lot so i'm trying to solve it quickly right and training models at scale annoyed me a lot so i'm trying to
1:00:28Jon Krohn:solve it quickly so yeah cool and then one last tidbit is i guess in between learning of it i need to learn English in between learning English and Columbia, you were training, you were a, let me try to get this right from memory. You were a U S Navy officer and you were in the seal training program. Uh, but then a medical, uh, injury meant, you know, you never finished the seal training program. Uh, but yeah, I don't know if there's anything, I mean, that's also kind of another, probably given everything else that we've said in this episode, people probably didn't expect me to be now talking about he was a navy seal or you know no for i mean for context i was not a navy seal but again public school in florida you're never thinking like will i go to an ivy league and become a programmer because you literally don't know anyone who does that um and so my best career path that i thought was super interesting was well maybe i can do special operations and be in like very tactical things that require like you know being honest and having the decision to do the right thing when it matters.
1:01:36And then I'm like, hey, I'm going to join and I find out about the SEALs. And of course, I tell everyone about it and they're like, oh, you're never going to do that. And I'm like, okay, well, let's wait and see, right? So I train most of my high school to do this. And then I get selected to go into SEAL training as an officer. Getting selected to go to SEAL training as an officer is an incredibly difficult thing to do. That year that I applied, there were about 12 ,000 applicants. They selected about 20. Wow. Because it's not just, like, there's, like, this crazy physical test you have to do, right?
1:02:10You have to -
1:02:11Jon Krohn:And a really big math test? No, but, I mean, you know, you're competing against, like, Harvard grads, right? Like, my roommate, one of my roommates was, like, a Stanford quarterback. So not only do you have to be, like, taught physical shape, but you have to have proven something. and you know I was a terrible high school student so I don't really know what they saw other than like I ran pretty fast and I was really good at pull-ups and did all these things and crushed the numbers and then you know I did a lot of community service like I worked a lot I met a lot of SEALs and tried to help the community before I was there and I think I got lucky I think one of the guys who wrote me a letter eventually became the commander of SEAL Team 6 so that I think at the time he already had a lot of clout and so I think his letter carried me through and got me into SEAL training.
1:02:54And the reason they're so selective there is because every SEAL class is about 300 people in training and you have only like maybe 12 officers. And if an officer quits during training, you usually bring like 50 people, not 50, like 20 people with you. Because you have to understand like these are like basically collegiate athletes at this point, very, very good at everything. And then you've got some 18 year old looking up to you. And then suddenly if you can't do the run and if you can't do the thing they're like how could i ever do this and so they'll usually quit and so they self so they select you out you know everyone knows seal training um has probably an 80 attrition rate so about 20 80 of people who who go through it don't make it um not everyone quits some people get medically discharged or whatever in the officers if you only look at officers it's actually about an 80 rate that makes it through i see so it's only 20%.
1:03:51And most of that is injuries or some sort of like training violation. Officers don't really quit because they've been selected at that point. And so, you know, I go through training without expectation. I'm 21 years old. And the first time I go through SEAL training, you know, I get to like basically beginning of Hell Week. And I get to that point. This is like two months in already, like six weeks at this point. So that class started with about 300 people. We get to the beginning of hell week with about a hundred people so it's already most people are gone and you know they look at me and they do the scans right before you go into hell week and then they're like hey you have pneumonia and I was like well I can't hide that like you try to hide all the injuries during training because you don't want to get medically rolled out and so I'm like well clearly I can't hide from an x-ray so they're like hey you have pneumonia I have to start again and I'm like like I just went through this thing that almost killed me and you want me to go through it again and i'm like all right fine whatever so i i go and it's like i don't know six months to heal up right to clear pneumonia then get back into training get back in shape so i do that and then i go back in and then i go through training and i go through hell week but i come out of hell week with injuries right so in my class i think we had about 300 people going into uh training right and then um we got to hell week probably around 100.
1:05:13hell week starts on Sunday by Tuesday we had maybe 35 people left so within two days you kind of watched that most of the class yeah so then 36 or 37 so there were two guys in there so of my class I think I was the second person to get rolled out of there so 35 made it and then for me you know I get rolled for these injuries right so I've been I've been basically running I had like I got sick right before I got, what do you call it, VGE, so like food poisoning on Sunday. So I'm not able to keep food down and I'm going through this thing. So I weighed like 200 pounds going into this. By the time I got through Hell Week, it was like 160 pounds.
1:05:54And so like around Wednesday, Thursday-ish, like Wednesday night, roughly, they sent me to medical and you have to get to Thursday to like fully graduate. And so at medical, they're like, hey, you've got, you know, I was like throwing up but have a bunch of issues. And so they scam me. They're like, hey, you're good to go. They thought it was like a appendix or something. And then they sent me back, but it's been more than eight hours. And there's a policy that if you're out of Hell Week for more than eight hours,
1:06:19Jon Krohn:you have to start again. Yeah, so I'm like, okay. So I go back and then I'm like, you know, delirious at this point because it's been like multiple days. And this is like Thursday morning. And like, usually they'll give you a pass on Thursdays, right? If you're there long enough. And so they didn't give me a pass. So then they sent me back to the beginning, But they gave me that second chance, which is an interesting thing, because as an officer, you only get one chance usually. So then they're like, hey, look, go back to one of the SEAL teams, heal up, and then we'll let you come back and finish.
1:06:46So this is kind of the deal that we made. So I go back. The team is called, it's not called SRT1, but back then it was called SA1, Support Activity 1. So I'm in there and I'm really in a supporting role there. So I'm helping the team do whatever I can. I can't go on deployments. I can't do a lot. And so I'm mostly just showing up, training, helping the command as much as I can. I speak Arabic, Spanish, I'm like trying to help them much. And so, it's like an annoying place to be because you're like not a SEAL, but you're in training and then all the SEALs know you're in training. And so every time they see you, they're like, hit the surf or something, you know, you're like, ah man, this sucks.
1:07:20So you're never like fully out of it, right? So you're in there doing this thing and I was in there for like six or eight months. And so then they start sending me, so after BUD, SEAL training, they put you through something called SQT, second part of SEAL training, that's actually where you earn your Trident. it. And so my commander, I'm like, Hey, I've been helping out. Like, can you hook me up? And so they actually send me through some of the SQT schools. So I go and I do, uh, there's this amazing guy. Um, you know, everyone knows him. Yeah. Jaco. Right. And Lee, Lee Bobbin. So, uh, leaf teaches my, uh, class called Jot C right.
1:07:53And I'm in that Jot C class, that class became a book called extreme ownership. If you've read it, I recently bought it yeah so that that's how on your recommendation yeah exactly and so you know they didn't have this back then but it was it's what they teach junior seal officers on how to lead smallest special operations team and so you learn everything for mission planning you do a bunch of trading missions it's pretty cool um and so that's kind of where the birth of my leadership really comes from i mean in training you know i was in charge of my class a bunch of times so you learn a lot there you know 300 people at 21 years old um but there's really where you learn the tactics and the stuff that Jocko talks about.
1:08:29And then Jocko came and gave a talk, I remember. And this was like right before he retired. It was probably one of the most intense talks I've ever heard. He's a very intense guy in real life. I'm very fortunate to have gone through a lot of the early lessons. I think Hell Week teaches you a lot about resilience, about grit. Leading teams, very small teams. As a 21-year-old over there teaches you a lot. Showing up to my first SEAL team with like SEALs were 30, 40 years ahead of me and having to figure out how to like manage that. you know it's it's very different so you learn a lot of these leadership lessons early and um and then the i think the grit and the ability to kind of do anything is what then later allowed me to you know go from like public high school kid to like ivy league student successfully doing it i'm pretty sure if i hadn't gone through this i wouldn't be able to do that and so you know overall i think it's great a lot of my buddies are still in uh guys i went through training with they've had amazing careers and uh it's been really interesting watching them evolve and see how things have shaped.
1:09:26And so now a lot of guys are leaving and they're about hitting the 20-year mark because we all joined around 2008. And, you know, I'm always trying to figure out how I can help them do their next thing. And so hopefully I'll, you know, for any team guys watching this, you know, if you need funding or you need a VC connection, let me know. But yeah, I mean, I'm very excited to continue supporting the community as well. Yeah.
1:09:46Jon Krohn:Really cool. I didn't know almost anything that you just told me about, you know, your history with the BUDS program or anything about it, really. So thank you for that. It does seem like there's a little bit of a common thread that if people want, if listeners want to be able to achieve something that is extremely difficult to get invited to do, it sounds like whether it's a top AI lab or the SEALs, a key thing is kind of like volunteering, hanging around the right people, not getting paid, but just kind of learning the right people. And eventually an invite can be forthcoming. Obviously a lot of hard work.
1:10:23I think it's, there's never an invite. There's you going out to get it.
1:10:28Jon Krohn:Right, right. In whatever way that looks. And today it's a little bit harder, I would say. I mean, I don't envy the PhD students have to apply today. You know, there's always a joke in every PhD class where like, wow, we're like, we could not get in. The next year you're always like, oh, I could not have gotten in this year. Because it's true. I mean, you need like 20, but you basically have to be a professor now to get in. But no, you have to just go out there and get it. And I think startups are very similar. right um it takes grid takes education and just knowing what you want well i think contributing to open source projects like the many uh that py torch lightning uh has has led to uh you know that whole ecosystem in github people contributing to those making a big impact writing papers uh you know publishing blog posts maybe making instructional youtube videos or something sure all that stuff helps for getting into phd programs or becoming a navy seal yeah i mean it takes uh you know my kiss a bit of luck right um for sure but i think yeah if i look at the guys who became seals like i think everyone has very similar traits sometimes people get lucky sometimes people don't right um but no ultimately it's that drive of i'm never going to quit that's it um so i always ask my guests for book recommendation at the end is your book recommendation extreme ownership 100 nice well that was easy and uh for folks who want to be be able to continue to follow your story, the Lightning AI story, where are the best places on social media to do that?
1:11:57Yeah, my Twitter, I guess, William Falcon, at Simon William Falcon, or Will Falcon, I can't remember one of those. And then the Lightning AI Twitter, and then our LinkedIn, I think that's maybe it's, where else do we have this? Yeah, I think that's right.
1:12:10Jon Krohn:Yeah, we'll have all those in the show notes and anything else. Or your podcast, I guess. Yeah, hopefully people are aware of that. Nice. Yeah. Thanks for mentioning that, Will. And yeah, thanks so much for coming on the show. It's been great to have you on. Maybe we can check in in a year or two and see how things are coming along. Yeah. Once we're close to being IPO or at IPO, then we'll talk about that. Sounds great. We'll talk about the IPO story. Perfect. Cheers, Will. Thank you. Awesome. Thank you.
1:12:39Jon Krohn:What an episode today with the extraordinary Will Falcon. In it, he covered how NeoCloud's are GPU-first cloud providers offering higher performance for AI workloads relative to traditional clouds like AWS, which were built for CPU applications. He talked about how Lightning AI merged with Voltage Park to create the third largest NeoCloud in the world with over 35 ,000 GPUs, over$500 million in ARR, and a full-stack AI platform serving over 400 ,000 developers. He talked about how Lightning AIO got started and how it was based on the open source project PyTorch Lightning, which started as a personal research tool in 2015, but has now been downloaded nearly 400 million times.
1:13:23Jon Krohn:He also talked about how his leadership philosophy has been shaped by Navy SEAL training. As always, you can get all the show notes, including the transcript for this episode, the video recording, any materials mentioned on the show, the URLs for Will's social media profiles, as well as my own at superdatascience.com slash 965. All right, that is it. Thanks to everyone on the Super Data Science podcast team, our podcast manager, Sonia Breivich, media editor, Mario Pombo, partnerships manager, Natalie Zajski, researcher, Serge Masise, writer, Dr. Zara Karche and our founder, Kirill Aromenko. Thanks to all of them for producing another stellar episode for us today for enabling that super team to create this free podcast for you.
1:14:06Jon Krohn:We are deeply grateful to you and to our sponsors and you can support the show by checking out our sponsor's links, which are in the show notes. And if you'd ever like to sponsor an episode yourself, you can find out how to do that at johnkrone.com slash podcast. Otherwise help us out by sharing this episode with folks who would like to have it shared with him. Review it on your favorite podcasting platform or on YouTube. Subscribe if you're not already a subscriber, but most importantly, just keep on tuning in. I'm so grateful to have you listening, and I hope I can continue to make episodes you love for years and years to come.
1:14:39Jon Krohn:Till next time, keep on rocking it out there, and I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.
1:14:52Thank you.
From the publisher
CEO of Lightning AI Will Falcon speaks to podcast host and Lightning AI fellow Jon Krohn about the company’s merger with Voltage Park, and why Will has named it the “full-stack AI neo-cloud for enterprises and frontier labs”. Lightning AI’s offer is a secure, flexible, and collaborative environment that can run on the cloud, all essentials for early-stage startups. Listen to the episode to hear Will Falcon discuss Lightning AI Studio, founding PyTorch Lightning, and how he came to found his AI company.
This episode is brought to you by the Dell, by Intel, by Fabi and by Cisco.
Additional materials: www.superdatascience.com/965
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
In this episode you will learn:
(02:20) Lightning AI’s merger with Voltage Park
(20:54) About neo-clouds
(43:51) How Will founded Lightning AI
(54:48) Current gaps in the AI in workplace




