In short
```markdown
Podcast Notes
Relentless - WTF is happening at xAI | Sulaiman Ghori
Episode Summary In this episode, Sulaiman Ghori, an engineer at xAI, discusses the rapid development and innovative culture at xAI, highlighting their impressive ability to execute ambitious projects, such as the construction of their Colossus data center within an astonishing 122 days. Ghori shares insights on the company's operational strategies, collaborative environment, and the unique challenges associated with pushing the boundaries of AI technology.
Key Themes and Concepts
Culture and Work Environment
- Flat Hierarchy: Ghori describes a workplace where no one tells him "no," allowing for a high degree of autonomy and rapid implementation of ideas.
- Fast Iteration: The culture promotes quick feedback cycles, with Ghori noting that model iterations can happen multiple times a day.
- Focus on Execution: There are no artificial blockers; the emphasis on getting to the root of problems accelerates development.
Ambitious Goals and Metrics
- High Leverage Projects: Ghori emphasizes the importance of focusing on high-leverage tasks that drive significant economic returns.
- Value Creation: The episode discusses how each software commit can generate millions in value, illustrating the impact of engineering decisions.
Challenges and Innovations
- Hardware Constraints: The discussion reveals that xAI’s greatest advantage lies in its unique hardware capabilities, allowing them to deploy AI models faster than competitors.
- Rapid Prototyping: Ghori shares experiences of taking bets with Elon Musk, such as promises of rewards for completing tasks under tight deadlines, showcasing the risk-taking culture.
Infrastructure Development
- Colossus Data Center: The construction of the Colossus data center within 122 days exemplifies xAI's operational efficiency and ability to overcome typical industry constraints.
- Collaborations with Utility Companies: Ghori explains the intricate relationship with power companies to manage load effectively during high-demand operations.
Human Emulators and AI Application
- Digital Optimization: Ghori elaborates on the development of human emulators designed to perform repetitive digital tasks, aiming to free humans for more creative work.
- Scaling Challenges: Addressing how they plan to deploy a million human emulators, leveraging existing infrastructure such as Tesla car computers.
Key Takeaways
- Autonomy and Responsibility: Employees at xAI have the freedom to take initiative, leading to faster innovation and project completion.
- Emphasis on Practical Solutions: Engineers are encouraged to challenge conventional wisdom and identify inefficiencies that can be eliminated to improve performance.
- Rapid Feedback Loops: The iterative process allows for quick adjustments based on real-time feedback, enhancing both product quality and speed to market.
Interesting Anecdotes
- Ghori discusses his unconventional journey into tech, from building a 3D printer to running a successful fidget spinner business as a child, highlighting his entrepreneurial spirit and hands-on approach to engineering.
- An amusing story about a bet with Elon Musk involving a Cybertruck illustrates the high-stakes environment and the playful yet competitive culture within xAI.
Conclusion Sulaiman Ghori provides a fascinating look into the inner workings of xAI, showcasing how a unique blend of culture, innovation, and ambition enables the company to challenge the status quo in AI development. His insights underline the importance of autonomy, rapid iteration, and hardware capabilities in driving technological advancements. ```
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOInside xAI: Culture and Work Environment
0:45 to 3:00
Discussion about the unique work culture at xAI and how it allows for rapid implementation of ideas.
“I've been kind of fascinated by XAI since like 2023 when like Elon first started.”
Breaking Conventional Barriers in AI Development
3:00 to 6:00
Exploration of how xAI challenges traditional timelines and processes in AI development.
“When was the last time that you experienced this where there's some like conventional wisdom that says that this is the timeline and then you guys just were able to completely shred that?”
Experiences of Joining xAI
6:00 to 9:00
Sulaiman shares his personal journey and experiences of joining xAI, including onboarding.
“And so I went to go find Greg because I was like, I don't even have a team.”
Innovations and Rapid Development at xAI
9:00 to 12:00
Discussion on the rapid innovation cycles and high utilization of resources at xAI.
“just because everyone's pretty great and reliable, which frees you up a lot in terms of what your limitations are, I guess.”
Scaling Up with Human Emulators
12:00 to 14:03
Insight into the concept of deploying human emulators and the potential impact on digital tasks.
“With Optimus, you're taking any physical task a human can do and allowing a robot to do it automatically at a fraction of the cost with 24-7 uptime.”
Rapid Problem Solving at xAI
14:03 to 16:41
Learn how quick communication and responsibility foster fast problem-solving.
“How does that kind of work for just executing on multiple different fronts at the same time?”
Navigating Project Ownership
16:41 to 19:43
Discover the fluidity of project ownership and responsibilities within the company.
“What's been your experience like with that?”
Stories from the Ground: Colossus Build
19:43 to 22:22
Explore interesting anecdotes about challenges faced during the Colossus build at xAI.
“Like there weren't weekends for a while, which was, uh, it was good to know that I could do that.”
Planning and Resource Allocation at xAI
22:22 to 24:26
Understand how xAI plans projects and allocates resources effectively.
“The batteries can react to the load much, much faster.”
Exploring Employee Enthusiasm and Contributions
24:26 to 26:51
Examine the unique characteristics of employees and their contributions to the company.
“Is there like a SpaceX-esque algorithm for making things happen?”
Show all 38 chapters
Innovative Hiring Practices at xAI
26:51 to 28:00
Learn about the unconventional hiring methods employed by xAI.
“You're just like, I want to work on the weekends.”
Ownership in Product Development
28:00 to 28:30
Learn about the dynamics of having one person manage significant portions of a product.
“It's being done by one person with 20 agents.”
Hiring Practices at xAI
28:30 to 29:36
Explore the unique interview strategies employed by xAI to find talent.
“So we're pushing very hard on Macrohard.”
Problem-Solving Mindset in Engineering
29:36 to 30:41
Understand the importance of simple solutions and critical thinking in engineering.
“variety of customers, like literally 30 years, 40 years of different hardware, different operating systems, everything like that.”
Learning and Collaboration in Fast-Paced Environments
30:41 to 31:51
Discover effective strategies for rapid learning and collaboration in tech.
“He throws in usually an incorrect requirement or question or an impossible line in his challenges for people when he's hiring, like coding challenges.”
Experimentation and Innovation at xAI
31:51 to 32:55
Learn how xAI maximizes experimentation to drive innovation.
“On the experimentation side, how are you guys trying to maximize for the number of experiments or good shots on goal that you can do?”
Managing Project Timelines and Expectations
32:55 to 34:18
Examine how to effectively manage project timelines and team expectations.
“We will frequently launch multiple experiments, especially on the model side, at the same time.”
Elon's Approach to Project Timelines
34:18 to 35:32
Discuss how Elon Musk's perspective impacts project timelines and outputs.
“Frequently, the estimated time to get something done, well, all the time to get something done is based on some set of assumptions.”
Revenue Targets and Decision Making
35:32 to 36:25
Understand how revenue targets influence decision-making processes at xAI.
“roughly this number of months in the future and that it actually does.”
Risk and High-Stakes Decision Making
36:25 to 37:55
Explore the dynamics of making high-stakes decisions in a fast-paced tech environment.
“He always says, you can always attempt to do something in one month that would otherwise take a year.”
Engineering Culture at xAI
37:55 to 39:48
Learn about the engineering culture at xAI and its focus on problem-solving.
“It's looking significantly faster than that.”
Management Structure and Autonomy
39:48 to 41:29
Discuss how xAI's management structure fosters autonomy and innovation.
“Like, I guess joining at 100 people, I mean, to me, it was like a 10x leap of anywhere else I've been.”
Engineering Culture at xAI
42:00 to 43:30
Learn about the unique engineering culture where everyone contributes and collaborates closely.
“than we update, but it's a lot more bottom-up than I expected.”
Fuzzy Responsibilities and Rapid Implementations
43:30 to 45:00
Discover how blurred lines between roles lead to faster project implementations at xAI.
“Because you have to communicate less times and language is lossy compared to what's going on in your brain.”
The Value of Small Teams
45:00 to 47:00
Explore how having fewer engineers leads to quicker decision-making and implementation.
“And then they rapidly like re-put that back in.”
Entrepreneurial Spirit at xAI
47:00 to 48:40
Understand the appeal of joining a startup like xAI compared to larger companies.
“It was definitely the coolest thing I've ever seen.”
Human Emulators and Internal Collaboration
48:40 to 50:30
Learn about the challenges and surprises in rolling out human emulators within the team.
“Actually, funny enough, we started testing some of our human emulators internally within the companies as employees.”
Identifying Human Errors in AI Training
50:30 to 52:30
Discover how human errors impact AI training and the importance of understanding workflows.
“What's like the, in your head when you're thinking about that, what are the biggest things outside of driving that humans do all the time that they just don't need to do?”
AI and the Future of Repetitive Tasks
52:30 to 56:00
Examine how AI can alleviate repetitive tasks, allowing humans to focus on creative work.
“And then like someone says, come to my desk and the person doesn't exist.”
Rapid Iteration in AI Development
56:00 to 56:50
Learn how faster iteration cycles enhance AI model deployment.
“So not only does the model react to situations faster and can be more, I guess, tolerant of timeframes, you can also just deploy iterations much faster.”
Truthfulness in AI Systems
56:50 to 58:20
Explore the challenges of determining truth in AI outputs.
“How do you go about basically cleaning up the internet in that way to figure out what is truth?”
Response to Errors in AI
58:20 to 1:00:30
Understand the internal processes when AI outputs errors.
“You can try to faithfully recreate the input or the output given any arbitrary input.”
Work Culture During AI Surges
1:00:30 to 1:02:55
Discover what it's like to work in high-pressure AI development environments.
“For Macroheart specifically, we've been operating in a war room for four months.”
Early Entrepreneurship and Tinkering
1:02:55 to 1:05:40
Hear about early experiences in entrepreneurship and technology tinkering.
“Um, my dad got me a book when I was like 11 and I liked it a lot.”
The Journey of Building and Selling
1:05:40 to 1:08:40
Learn about the journey of building a business from childhood projects.
“And around that time, yeah, the fidget spinner craze was going off.”
Rocket Engine Design Experience
1:08:40 to 1:10:03
Gain insight into the process of designing a rocket engine from scratch.
“Right before we met, you had made a liquid fuel, I think, rocket engine.”
The Night of the Launch
1:10:03 to 1:10:26
Learn about the frantic preparations and challenges faced before a rocket launch.
A Close Call with Fire
1:10:28 to 1:11:35
Discover the unexpected hazards during the launch and a memorable mishap.
“Yeah Which had a lot of um We'll say Concessions made to make it happen that night I Do find it absolutely hilarious that you like you said were you like a couple feet away?”
Transcript
Automatic transcript. May contain errors.0:00I took this bet with Elon, like, you get a Cybertruck tonight if you can get a training run on these GPUs in 24 hours. and we were training that night. Did he get the Cybertruck? Yeah, he got the Cybertruck. My first day, they just gave me a laptop and a badge and I was like, okay, now what? I don't even have a team. I've not been told what to do. Ascroc was spinning up at the time. Our integrations with X, they're like, can you help? And I was like, yes. What's the most fun thing about working there? No one tells me no. If I have a good idea, I can usually go and implement it that same day and show it to Elon or whoever and got an answer.
0:28We did the math. Right now we're, I think, at about$2.5 million per commit to the main retail. And I did five today, so. You added like$12.5 million of value? The levers are extremely strong. Today I have the pleasure of sitting down with Sully Kongori, and he is one of the engineers at XAI. I've been kind of fascinated by XAI since like 2023 when like Elon first started. I think it's like one of the fastest growing companies of all time. Can you just talk about like what the fuck is happening at XAI? Yeah. We don't have really due dates. It's always yesterday. Um, there's no blockers for anything, like at least nothing artificial.
1:10Uh, the whole Elon thing about going down to the root, uh, the fundamental, whatever the physical thing is, we get there pretty quick if we can, as quick as we can, which is funny in software. It's not really like a thing that you think about is the physics too much, but we do try quite a bit. And we're not really fully a software company given all the infrastructure build out. Um, it's kind of hardware at this point. Yeah. It's like hardware constrained. It's the probably our biggest edges is is the hardware because nobody else is even close on the deployment there Although the talent and city on software is like incredible.
1:43I've never been anywhere like that It's it's really cool for Elon He is very good at figuring out like what the bottlenecks will be even like a couple months or even years in the future And then trying to work backwards from that and make sure that like he's in a really good position How does that work day to day with just normal people like at XAI and like adopting that kind of mental framework? usually when we spin something up new very quickly either one of us or he comes up with this uh metric it's usually very core to either the the financial or the physical return or both sometimes um and so everything is just focused on driving that that metric um there's never like a fundamental limitation to it or like whatever the fundamental limitation is it better be rooted deep down and not something artificial.
2:30And there is a lot of perceived limitations, especially in the software world, coming from, like, especially in the last 10 years of, like, web dev and all these kinds of things. People just assume or accept certain limitations, especially when it comes to speed and latency. And they're not true. You can get rid of a lot of overhead. Like, there's a lot of stupid stuff in the stack. And if you can knock out a lot of that, you can usually 2 to 8x most anything, at least anything invented relatively recently. Some stuff, not so much, but yeah. When was the last time that you experienced this where there's some like conventional wisdom that says that this is the timeline and then you guys just were able to completely shred that?
3:15Most recently, it's our model iterations on MacroHard. So we're working on some novel architectures, actually multiple at the same time. and we're coming out with new like iterations like daily sometimes multiple times a day which is from pre-train um in some cases uh which is not something you ordinarily really see but it comes from well a we have a pretty great super compute team and they've knocked out a lot of the typical barriers it takes to train a lot of this stuff even with how variable our hardware like the hardware is like it's you know within a day of standing up a rack you can usually be training sometimes within the same day um even within uh a few hours in some cases and this is like not normal like normally the timelines are like days or or weeks it takes a lot well in most cases at least yeah in the last 10 years is you abstract this away and let amazon or google take care of this um and so whatever their capacity is is what their capacity is but that's not like you can't have that be the case and win in AI now.
4:18So the only solution is to die or, uh, or build it yourself. Can you tell me about like how, what your experience was like joining, why you joined and then kind of what the like onboarding process was for the first like couple of weeks? Yeah. So, um, I was working on my own startup when I moved to the Bay. Um, and actually during that time, Greg Yang, one of the co-founders of XA had reached out. He's great at recruiting as it turns out um what did what did he say uh so i got an email and i thought it was spam because it was i was getting a lot of these like you know emails to founders at the time of like hey you want to chat or like i like what you're doing you want to chat whatever i was going to market a spam and like delete it and i saw the domain x.ai i was like oh wait a second i know these guys and they just uh i think it was probably eight months in at that point um and so i was like okay yeah let's chat so we chatted a bunch of times um then uh i wanted an echo hire but uh i think we were too early at the time and that company kind of went kaput mostly because it was fairly obvious that you can't build macro hard with like a million dollars um but the uh idea was sound so i spent the next like six seven months wasting all my money um building like aerospace projects and working on an aerospace astro-minding concept that also I realized probably wouldn't work, but it was worth a try.
5:38And so I emailed Greg again, like, hey, can I, like, you want to chat again? He was like, yeah, sure. You want to interview tomorrow? I was like, okay. And I apparently did well, and I moved on Monday, and I started then, and it was really great. Nobody told me what to do. So like my first day, they just gave me a laptop and a badge, and I was like, okay. And I was like, okay, now what? And so I went to go find Greg because I was like, I don't even have a team. I've not been told what to do. Like Greg just brought me on because I think he liked what I was doing previously and it was related to what the long term was for MacroHard, which wasn't really even a project at the time.
6:15And I ended up working on, actually AskRock was spinning up at the time, our integrations with X. And so they were like, can you help? And I was like, yes, I can help. And so my first week was working with the one guy. I found out very quickly, like, everything that we built, like, I could sit and I could stand up from my desk, which I didn't even have a desk assigned to me. I just sat at random people's desks that weren't there that day. And I could point to whoever built that thing at XAI, like, from my desk. It was very, very, very cool. And there was, like, almost no people working there at this point.
6:49It was, like, a couple hundred, right? Yeah, about a hundred or so on the engineering staff. and then I don't know what the infra build out team looked like at that time. And it's kind of hard to tell because some people move up the ladder from like the actual building and construction crew onto our payroll. But it was pretty small at the time, like much, much like an order of magnitude smaller than the other labs. And we had still just done Grok 3. Yeah, which, yeah, it was pretty cool. One of the things that I kind of love is how fast XAI went from being founded. I remember Elon initially saying, like, we're not even sure if this can be a success with, you know, people having, you know, a multi-year advantage on speed and like timing.
7:34And then you guys got done with the first like Colossus data center in like 122 days. And that was just like unheard of. And Jensen's out here saying, singing the praises of XAI and Elon. what kind of culture did that allow to be formed it definitely enabled like us on model and product to kind of assume we would have the resources to do what we needed to do um and that's definitely the case like we're not super duper resource constrained like we've still found a way to push up against that wall um but that's just because we have 20 different things going at the same time like more than that like many more things than that there's a absurd amount of runs and training and all that stuff going on at the same time in parallel usually by like a few a handful of people um which is how we're able to iterate very quickly on model and product side um and utilization has definitely been very high the the speed allows us definitely to i guess think more long term um so i think croc four or five really what it was was already planned out and designed in terms of size and what we expected way early, like before I joined.
8:50I joined around Grok 3. So it was like thinking at least a year in advance. Yeah, you can think much more in advance and assume that those estimates will be hit just because everyone's pretty great and reliable, which frees you up a lot in terms of what your limitations are, I guess. So for us, for example, the assumed minimum latency was about three times higher than it actually needed to be, and the build-out allowed for that, basically. What do you mean by that? So one of the novel architectures we're working on is not really possible unless you scale up your experiment rate. Because it's not building on any existing body of work.
9:36You need a new pre-training body, and you need also a new data set, but that's not really constrained by the resources, like the physical infrastructure resources, mostly. Although there's the Tesla computer thing, which I think maybe we'll get into, maybe not. But actually, this one's public. So one thing that we're thinking about is, okay, we're building this human emulator with macro hard. How do we deploy it? Because you actually need, like if we want to deploy 1 million human emulators, we need 1 million computers. How do we do that? And the answer showed up two days later in the form of a Tesla computer.
10:17Because those things are actually very capital efficient, as it turns out. And we can run potentially our model and the full computer that a human would otherwise work at on the Tesla computer for much cheaper than you would on a VM on AWS or Oracle or whatever, or even just buying hardware from NVIDIA. That car computer is actually much more capital efficient. And so it enables us to assume that we can deploy much, much faster at a much higher scale. and so we've adjusted our expectations for that basically. Are you basically able to just bootstrap off of the car network? So that's one of the potential solutions basically.
10:59It's like okay, well we want 1 million VMs there's like 4 million Tesla cars in North America alone and let's say two thirds or half of them have Harbor 4 and somewhere between 70-80 % of the time They're sitting there idle, probably charging. We can just potentially pay. And they have networking. They have cooling. They have power. We can basically just pay owners to lease time off their car and let us run like a human emulator, Digital Optimus, right on it. And they get their lease paid for, and we get a full human emulator we can put to work. And that's something without any build-out requirement.
11:42It's a purely software implementation that's required. The acid is sitting there and you can just go and use it. Yeah. Yeah. Amazing. For the human emulators in macro hard, what is the purpose of that, of scaling up millions of many humans? I mean, the basic concept is very simple, right? With Optimus, you're taking any physical task a human can do and allowing a robot to do it automatically at a fraction of the cost with 24-7 uptime. We're doing the same with anything that a human does digitally. So anything where they need to digitally input keyboard and mouse inputs, which is usually what humans do, and look at a screen back and make decisions, we just emulate what the human is doing directly.
12:28So no adoption from any software is required at all. We can deploy in any situation in which a human is in, potentially, currently. Interesting. What is that actually going to look like for rolling it out? I don't think we've detailed our plans publicly yet specifically on how we'll roll out it'll be slowly at first and then very quickly basically like the difference for us given that infrastructure build out already has happened or we can go on the Tesla network or we can build out our data center of Tesla computers actually the difference for us from going from 1 ,000 human emulators to a million is actually not very big It's not the biggest part of the challenge.
13:14Elon, I know one of the things he does best is he basically just goes from fire to fire on whatever the company is and just kind of puts it out and unfucks whatever problem exists. What has that been like? When have you seen some problem exist and just had it unfucked very rapidly due to this kind of process? Definitely Infra build-out. This is the biggest. On the model side, we've had hiccups, but it's more or less been smooth. But on model side especially, because there's a lot of, I mean, infraside, there's a lot of very specific operations that each of these, basically ASICs, these CPUs are built for.
13:54And when we roll out new products, like when we pick up new products from NVIDIA or whoever, not everything works. so in some of the meetings that we had with him uh early last year uh he would hear these and he would make a phone call and the software team would deliver a patch the next day and we would work like side by side until that was resolved um and then we could run a model uh or train training run on the hardware uh very very quickly where otherwise it would have taken weeks of back and forth so those kind of blockers are usually very quickly resolved with one phone call um or just us bringing it up to him or him just offering like frequently when a meeting is ending or there's a lull in the conversation he'll be like okay how can I help how can I make this faster or whatever and someone will come up with with an answer I know you guys are doing many different products in parallel and I get that it's kind of like you have to do that but also it's sometimes in most organizations it's like very difficult to stay focused on a single thing and like a single objective.
14:57How does that kind of work for just executing on multiple different fronts at the same time? Very frequently, we actually, and this is increasing with scale, we don't have a full picture until like the all hands or we just chat with people what everyone is doing and how far everyone is on these different projects. Like for example, on when we did our voice model and our voice deployment, we actually had a lot of the work built for extremely little latency uh extreme low latency end-to-end like uh packets to be sent to the client it was already built out and um it was a matter of flipping the right switches and the right configs basically to cut our latency pretty significantly um like 2 3x uh end-to-end um this is actually the case a lot of the time is there is a stupid thing that uh exists somewhere in the software or hardware and someone has come up with a solution.
15:57And you find it when you go to look for it in our code base somewhere. Or you ask around and someone's like, oh, yeah, this XYZ person has done this. You should talk to them and they will hook you up. There's not a lot of time spent syncing up with anyone or asking for permission or waiting for anyone at all. Like the answer is, like when you propose someone, someone says it's a good idea. Usually you propose something and the answer is either no, that's dumb, or why isn't it done already? And then you go and do it, and then it's done. With Elon companies, you can kind of just ask for responsibility, and then you basically just live by the sword, die by the sword.
16:36And if you get things done, then you can just ask for more responsibility, and you can keep on doing that, or you're just out. What's been your experience like with that? Very much so, yeah. I've jumped around a lot of different projects, and mostly just because someone asked for my help and I kept helping. And then I ended up owning some of the stack or a lot of the stack. And this is the case for everyone. Like, this is just how it is. If you have any particular experience or can iterate on something very quickly, within days you own that component. Yeah, there's no formal anything. I think officially on our HR software, I am on voice and iOS or something.
17:15and our security software thinks I still work on our X integration. Just never updated? Yeah, no one ever updates this stuff. It's kind of ridiculous. Has your journey at the company kind of been you show up, there's not exactly a clear direction of what you're going to work on and then you just start working on stuff and then you just kind of hop from project to project by whoever asks for your help? there's a bit yeah there's quite a bit of like overlap and flow um so like after onboarding i'm usually on two or three projects at once um and whichever one is most pressing or i can help the most on ends up taking majority of my time and then that kind of overlaps and flows in like a waterfall way what's been the journey from like the starting to to now like what what projects have you worked on?
18:06Yeah, so specifically I started, I first worked on like AskRock and our integration there. And then I worked with our backend team a bit on like reliability and scaling up because we were scaling up a lot at that time. And then after that, I took on solo building up our desktop suite and took that like to internal completion. And then I got asked for help on our Imagine rollout and iOS. which yeah our ios team is small for like how many people use it like it's ridiculous you won't guess the number um like five people for three three it was three and i was the third person at the time when we were rolling that out it was like it was ridiculous and everyone's like really really good um yeah this is the first place where i've had to work very hard to keep up with like the speed and the talent.
19:03What was the first experience that you had where you thought to yourself, like you're actually being kind of used to your full, you know, potential? I think that Imagine Rollout was definitely like, it was a really good push because like we had this 24 hour iteration cycle. You don't want to give feedback every night on whatever we were doing. And yeah, we would push out that night. In the morning, we would have all the feedback. We would immediately knock out all the bugs, um, implement the new stuff that, that people were asking for, whatever model had come up with, we implemented that too. Like it was a very, very fast cycle.
19:36And it was, uh, I think it was the longest, like continuous stretch of me being in the office, like every day. What was that like at the time? It was like two or three months. Yeah. Um, yeah. Like there weren't weekends for a while, which was, uh, it was good to know that I could do that. I was pretty happy doing that. and after that I got pulled onto macro hard product which was just one other person at the time so it's the two of us for a while and I've been on that since since that project off basically I don't know how much you like know about this but uh the like Colossus build and all the ridiculous stuff that the like early XAI team had to do to turn on Colossus and like get power and all the necessary inputs to making that work and even today i think like it's just bottlenecks across the entire thing you just want more you want more like uh chips and gbqs and all the stuff working and faster um what was that like there's a lot of war stories um and a lot of bets um trying to go into a few yeah so i think tyler was took this bet uh with elon like we were setting up new racks I think of, I forget which GPUs we were rolling out at that time we took a bet, Elon's like okay, you get a Cybertruck tonight if you can get a training run on these GPUs in 24 hours and we were training that night Did he get the Cybertruck?
21:02Yeah, he got the Cybertruck I think it's, yeah I see it from our lunch window Cafeteria He's cool You know what the So for power, we actually have to collaborate very tightly with the municipal and state power companies. Yeah. Because when load goes high on their end, we have to shut off and go fully on the 80, or maybe it's more than that. I think it's more than that, 80 mobile generators we brought in on trucks and go fully on those just so that we don't impact power anywhere within. And we have to do that seamlessly without interrupting anyone's extremely volatile training runs on extremely volatile GPUs and hardware, which scales up and down by megawatts in milliseconds.
21:55It's a lot. Is that also part of the logic of basically putting massive battery packs right next to the disdainters? Because then you can go up and down much faster without any of the packs. I think scale up a lot, scale up and down and balance that load a lot faster. Because with the generator, you're literally asking a physical thing to speed up or slow down, like a spinning physical thing. That's obviously just going to take a certain amount of time. The batteries can react to the load much, much faster. And then, yeah, it's like, actually, from a physical standpoint, I think there's local capacitors, the station, like data hall, side capacitors, the batteries and then generators and then the public municipalities although we might have changed that infrastructure at this point things that are right very quickly especially on the cooling side do you have any other really good like war stories that are just like uh i don't know things that shouldn't have been possible that became possible uh so the lease for the land itself was actually technically temporary it was the fastest way to get the permitting through and actually start building things um i assume that it'll be permanent at some point but yeah it's I think a very short-term lease at the moment technically for all the data centers it's fastest way to get things done And how do they how do they do that?
23:11I think there's basically a special exception within like the local and state government says okay if you want to just modify this ground temporarily I think it's like for like carnivals actually does a carnival company currently? And so that was the way to get done quickly and it was done Yeah, 122 days. For internal planning, I know things are just going to keep on scaling up like crazy. And Elon's talked about energy being the biggest bottleneck and then just being able to get chips. How do you guys plan when it's very difficult to predict 12 to 24 months in the future exactly what projects you're going to be working on or what their resource requirements are going to be?
23:56We try very hard to work backwards from what's the highest leveraged thing we can be doing. and then we determine the physical requirements later. So if we want to get to$10 or$100 billion in revenue by this date, what are the highest leverage things we can do from an economic perspective? How can we actually build systems to do that? And then what does it take on the physical and software side to roll that out and get it done? Just roll backwards the whole way. So we don't usually start with the physical requirement. That's usually actually at the end. Is there like a SpaceX-esque algorithm for making things happen?
24:37As in like the usual delete? Yeah. Yeah. I mean, that's the case all the time. And we do do the thing where we delete something and then add it back later. What was the last time that you did that? Today. Today. Today, yeah. So with MacroHard, we deploy on a lot of physical hardware that changes. and the testing harness for that is hard. So we try to minimize how many special cases are downstream of where it needs to be. And for example, like with display scaling, we need to be able to support displays that are, you know, 30 years old, as well as the latest like 5K, Apple, whatever displays. And that has to happen on the same stack.
25:26Turns out not all the systems are happy with that at all times. Like you have to fiddle with the encoders at a certain level, like video encoders was the specific thing, basically. I didn't know, but as it turns out, there are limits to the maximum amount of pixels that certain encoders can take. So we have to now have, I removed this special case for multiple encoders. It turns out we found a problem at plus 5K resolution. And so we added that back. What are the most interesting things about XAI itself that you think would be really good stuff to talk about? There's a lot of characters that work there.
25:59And also we're doing hiring in interesting ways, I guess. Things that I thought would be stupid are okayed. And we just do them and we try them. And it's like we'll do a hackathon. If we get five people in as a result, it's worth it. Because just their expected return on the company's revenue or valuation is higher than the cost of running this hackathon for 500 people. The overhead value is actually very high, which is like funny we did the math um earlier this week uh right now we're like i think at about 2.5 million dollars per commit is to the main to the main refill um and i did five today so you added like 12.5 million of value uh light day yeah exactly it was a good day um it's funny things like that um like the levers are are extremely strong like you you can get a lot of a lot done with a lot less effort and time than you used to be able to for sure just because of who you work with the internal tooling that we built up um and my boss what's like an example of the type of person that like wants to work here because i know when when you're talking about it you kind of show up in the first day.
27:20You're just like, I want to work on the weekends. I want to work on, you know, during the night, all this stuff, uh, go all in on this. Um, what kind of special characters are working there? People are definitely very enthusiastic when they come in, like, um, very, very enthusiastic. Uh, just like, like mission oriented. Um, there's, I guess, different types of ambition for sure. Some people want to move up like the leadership ladder and own more in terms of a managerial, like how many people report to me since. Some people want to own huge parts of the technical stack. So right now, we're doing a big rebuild of our core production APIs.
28:00It's being done by one person with 20 agents. And they're very good. And they're capable of doing it. And it's working well. So you can own huge chunks of the code base, no problem. It's kind of like X where after the acquisition, they had much fewer people, but you just never had a lot of people in the first place. So there's one person owning a huge part of the product. Yeah, absolutely. For hiring, what unusual practices outside of just hackathons does XAI do? So we're pushing very hard on Macrohard. like for two or three weeks I was doing upwards of 20 interviews a week so it's like some of them are like quick 15 minutes some of them are full one-hour technicals so a lot of my time is dedicated towards bringing in new people and a lot of people are very good so it's it's actually very hard to judge them how do you uh I have a very specific problem that I have solved I'm not going to reveal it because then people will use it but I have solved I solved a very specific computer vision problem a few years ago for one of my startups and I give people half an hour to try to implement the solution.
29:13It's actually very, very simple. It's deceptively simple, the solution. People always overthink it. And this is something I like to index for on my team, especially is like, can you not overthink it and come up with a simple solution? It helps a lot because we're deploying on such a wide variety of hardware as a result of the wide variety of customers, like literally 30 years, 40 years of different hardware, different operating systems, everything like that. You have to come up with simple solutions or you're going to have a 10 million line code base next week. So this is like very important. And especially now relying more and more on agents and AI and such for writing code.
30:02um an ai will happily train out 200 lines when a 10 line solution will do um and probably do better so you have to look for that like i want people and i look and actively hire for people who can find the 10 line solution first um we're totally fine with people using ai to code things like you should you should use that as force multiplier but uh for now we're smarter we'll see next year What other force multipliers do you kind of look for? I like people who will challenge requirements and challenge me. So often, I got this from Chester Zwei of German Forge. He told me this, and I thought it was great.
30:42He throws in usually an incorrect requirement or question or an impossible line in his challenges for people when he's hiring, like coding challenges. and he expects people to come back and say like, hey, this is wrong. This is not possible. You made a mistake. And if he doesn't, then he doesn't hire them. Same thing for me. I picked that up. It's a great idea. The pace is insanely fast. And like you said, you kind of have worked on a number of different things. How do you kind of come up to speed on something as quickly as possible when you're on a new task or project? It depends on what thing it is.
31:18If there's a lot of code to read, read the code. by hand. Like GD, go to definition over and over again, and you'll find things out very quickly, actually. It's not that hard. For most things, the implementation is like less lines of code than you would otherwise see, which is nice. Not all the time, but in most cases. If it's something that's in very active development, this is not the case. There's going to be 20 different versions of it going at the same time, and it's not obvious what is the current path. So you just got to talk to people. And people are very open. Like this is actually one of the things I was very surprised by.
31:53pleasantly surprised by when I joined is I thought people would be super smart and stuck up but no people are just super smart and very nice and helpful like everyone's on the same team everyone's rooting for each other people are willing to like help you out um and answer your questions so which is good because we don't like write a lot of docs we write things we do things too fast to write docs really um actually well we're trying to figure out some systems on on my team to like automatically generate docs as we like build stuff um and with grok which is cool that we have unlimited access to uh very smart ai because then we can try a bunch of stupid things if it works which otherwise you know at a startup would cost you maybe like a hundred or a million dollar a hundred k or a million dollars in credits or whatever we do it for free so experimentation like value you can fail on a lot of things and it's a lot more things you otherwise would um and as a result more experiments are tried, more succeed.
Read the full transcript
32:47On the experimentation side, how are you guys trying to maximize for the number of experiments or good shots on goal that you can do? There's often a time constraint. We will frequently launch multiple experiments, especially on the model side, at the same time. And in some cases, it's not even because of a time constraint necessarily in terms of I need to try X amount of things in Y time. it's in two weeks this prerequisite will be ready either in the hardware or in the training data or something but in the meantime i need to deploy something today uh what can i do and so you run two three experiments and you find out what you can deploy today um and bring in revenue or customer result whatever it is today and then two weeks you switch over um like that's something we do all the time especially at map road have you seen anything where a timeline should have been much longer on like a project that you were working on um and somehow you guys were able to kind of like bring that in by you know weeks or months all the time every time uh all the time every time we come come away from like an email meeting or something internal where um someone pushes hard to get something done or someone external who doesn't isn't responsible for the thing ask for a requirement ask for something to be done in an what we originally think unreasonable amount of time.
34:09You know, we spend a few minutes like thinking about it, complaining maybe a bit. And then the rest of the time is dedicated to getting it done in that time. Yeah. Frequently, the estimated time to get something done, well, all the time to get something done is based on some set of assumptions. And then once you get this timeline, that's like half or one-tenth of what you would have otherwise done, you look at the assumptions and say, okay, proportionally, how much is this impacting my timeline? And then you knock it out or you change it. And then suddenly you get a 2X improvement in your timeline.
34:44You do that a few times. You can meet whatever requirement you really want. At a certain point, you get to the physical limitations, but you're never there from the start. I know for full self-driving, and same thing with the rockets of SpaceX, the Elon timeline was significantly longer. like an Elon time might be a quarter or half of what it actually eventually takes, but then it also, you know, happens four times faster because of the initial timeline. Is it more or less like at XAI because it's more, I mean, I guess more on the software side now. Um, but even on the data center side, things seem to be happening just way, way, way, way faster.
35:26Uh, and they also seem to be happening on like the same timeline as he's roughly saying, he's like, this is going to happen roughly this number of months in the future and that it actually does. I think he himself has calibrated his timelines. Like differently? Yeah, now that he's deployed a number of extremely, like a wide variety of deployed hardware at scale. So I think his own estimates for things are definitely a lot better. And so that's definitely the case. I think he also updates his timelines faster now too. Like sometimes daily, I think he's talking with us and figures out what the update on the timeline should be based on various parameters.
36:07And sometimes they come from him too, right? Especially on the infrastructure side. If a deal or we can be put up in a batch for the production of a certain chip, well, we can save a month or two maybe. Maybe even more than that. Depends on what the deployment is specifically. And then on the software side, it's the same. He always says, you can always attempt to do something in one month that would otherwise take a year. And you'll probably get it done maybe in two. Still a lot faster. I remember in the early days of SpaceX, there was this internal, I think Elon would say internally, every day that we delay is like$10 million in lost revenue.
36:49And I have no idea what it would be like for XAI. Things are moving so fast. Is there kind of an internal thing in your head of every day that we don't push harder, make something happen, we're losing out on X amount of value that could be created? Yeah. For MacroHeart specifically, we do have a few pretty specific revenue targets. I can't dilute the numbers specifically, but in my head, whenever something gets delayed or accelerated, I can pretty quickly calculate how much money we just made or lost. Is it just wild swings? Yeah, I mean, the numbers are huge just because the expected return is so huge and the timeline is so fast.
37:34So a few days is actually proportionately fairly large compared to how much you would otherwise expect the revenue to be. Elon's like famous for making really, really big bets pretty quickly. Like what's the biggest decision that's been made in a single meeting where like huge, huge amounts of capital or time or commitment were done? I think one of them was certainly the decision to go with a model that would be at least 1.5 times faster than a human for macrohard. It's looking significantly faster than that. 8x maybe, maybe more. The, like, for other human emulator type attempts in the other labs, the approach has been let's do more reasoning and build a bigger model.
38:21we've like that decision put us in totally the opposite track of what everyone else is doing and everything that we're doing really is downstream of that like well not everything but pretty much everything um it impacted and it was very early on uh that this was decided it was sort of expected also um that this is the move especially given the analog to full self-driving um no one's going to wait around uh 10 minutes for the computer to do something that I could have done in five, but if it can be done in 10 seconds, well, I'd be happy to pay whatever amount of money for that. Um, it's just obvious really.
39:00So normally like us engineers would, you know, if it's, would push back and say, Oh, you know, here's the 20 different reasons, uh, that it needs to be this way. Um, but if the decision is made and you work backwards, life finds a way. I remember Elon saying, uh, I think it was at the like Y Combinator, he was doing a like q a with gary tan and gary talked about like ai researchers and he was like no they're just all ai engineers now yeah we did this was someone said that in one of the meetings um we did with him uh talking about recruiting like he was here's the job description or something like that and like for 10 minutes he just goes on the engineers just engineers doesn't matter good engineers engineers just someone who's fundamentally a problem solver it doesn't matter if they did like this you know x y thing and this infrastructure or this you know particular architecture or whatever engineers why is it so important why is that definition so important um it keeps things broad it means that people can come in to us from like a extremely wide variety of places and this has been the case i mean there is i think less so in the ai world but i think there's a lot of spacex stories where people came in from strange walks of life that would not have otherwise seemed to be the case and then ended up doing huge things at spacex in the engineering world as a result so keeping it broad means that those people can have a path to us and and help us accelerate for you personally what's the most like fun thing about working there day today no one tells me no and tells you no no one tells me no um yeah if i like have a good idea i can usually go and implement it that same day and show it off and we'll see if uh if it makes sense we'll we'll we'll run whatever eval or um show it to a customer or show it to elon or whoever and we'll get an answer usually that same day uh as to whether or not that was the right move there's no deliberation there's no waiting for any bureaucracy uh i like that a lot i was expecting to sacrifice some amount of this coming from extremely small startups to a larger company.
41:16Like, I guess joining at 100 people, I mean, to me, it was like a 10x leap of anywhere else I've been. But I guess relatively to Elon companies, it's pretty small. And it does feel very small. There's not a lot of overhead in anything. Did you have any other big assumptions going in that proved completely wrong? I thought there would be more top down. And there's some, but not really that much. especially because of how many there's basically only three layers of management there's um like ICs there's the co-founders and some of the new managers and then Elon and that's it and so because there's so many reports to the managers now um nothing really comes from them top down like we'll usually come up with a solution they'll okay Elon okays we're good if there's feedback than we update, but it's a lot more bottom-up than I expected.
42:09Is it like trying to be designed so that everyone is building things and there is fewer manager-managers and more just builders? Yeah, when I joined, I think every manager also wrote code. And I think largely today they still do. Not as much now that some of them have 100-plus people reporting to them, but everyone's an engineer. I remember actually on my first week, I sat down for dinner and this guy sits next to me and I say, you know, like what team are you on? How are you going? How are you doing? I just joined. And he tells me, oh, I'm on sales and like enterprise deals. And I was like, oh, I don't want to talk to this guy.
42:48He's a sales guy. And then he starts telling me about this model he's training. He's an engineer too. The sales team is an engineer. They're all engineers. Everyone is an engineer. I think at the time, there was probably less than like eight people who were not engineers at the company. in some capacity. And even then, like, yeah, it was really cool. Everyone contributes to the machine. Is it a little bit more like you have a single person working on some project and they just, you know, if you're an engineer and you're working on the thing, you can have a much closer relationship to the customer and like understanding their problem and then rapidly like implementing solutions and stuff.
43:25Yeah, yeah. The less layers you have, the less information is lost. There's less compression, basically. Because you have to communicate less times and language is lossy compared to what's going on in your brain. So if you have to go from customer's brain to words to, you know, salesperson's brain to words to manager, every layer you're losing. It's like telephone. A huge amount of information, yeah. And so if you can cut as many layers as possible, then you've only got one compression step of the customer telling you what to do or what they want. or what their experience is or whatever, and US Danger can solve it directly.
44:06Is there anything specific that you've never heard of or seen at any other company that XAI does that allows things to just happen way faster? The fuzziness definitely between teams and what everyone is responsible for is definitely not what I expected and I don't think exists nearly as much in any large company or even remotely similarly sized company. like if I need to fix something on our VM infrastructure I will do it I will show it to the guy who owns that and they will be like okay and it's merged immediately and deployed like there's not a lot of strict regiment like everyone is allowed to update everything and there's some checks for dangerous things but largely you're trusted to do the right thing and do it right which is really cool I remember when Elon was still like really working on Doge, there was at one point, I think they deleted like Ebola prevention or something.
45:09And then they rapidly like re-put that back in. What things have been like deleted because of this rapid process of trying to figure out, you know, what doesn't need to be done and then like re-implemented. There's very rarely anything like irreversibly destructive. I'm actually not really aware of anything where something was irreversibly destroyed. but like I said, yeah, frequently something will be deleted or removed or something like that and someone will be like, hey, I needed that, you know, an hour or two and then you go and roll back. Or, you know, sometimes it can be months where, you know, someone's building this project and they're depending on, I don't know, some piece of infrastructure, something like that and it turns out we rebuilt that thing three times by the time you go and deploy and need it and so you update and go that way.
45:57Do you think it's helpful to have so few people working there on the engineering team? Yeah, definitely. The more people you have doing it, I definitely say a job for one person done by two will take twice as long. And it applies for every skill, I think. And especially now that you don't need to physically write as much code as you did previously. You can be more of the decision maker and the architect. Everyone can be an architect. You just don't need as many hands. So one brain can do a lot more. You tried starting multiple companies, and you were doing a whole bunch of different projects prior to this.
46:38What about working here? And what about the mission, the culture resonated? Why did you decide to work on this? I've definitely always been very Elon-filled. He's been a big personal hero of mine, especially growing up. I've seen the Falcon landings, the first ones. I went out to launch five of Starship, which was so worth it. It was the first one they caught. It was really cool. It was definitely the coolest thing I've ever seen. So being part of anything remotely related to that sounds awesome to me. Is there a reason why you chose this company instead of SpaceX or Tesla? Yeah, so I'm definitely an entrepreneur by heart.
47:18And XAI is definitely the smallest company, the newest of all of them. I think my assumption is, and this is largely proven true, I think, is where you can have the most leverage and change as an individual person because proportionately you're a much larger percentage of the company than you would be at these other companies. Not to say that they're not doing cool things or everyone's not as important, but just the proportional change of life. The leverage to decision is way higher. Yeah. Yeah. not even the decision but to implementation to seeing the results like it's very quick and I guess another assumption that I thought would be the case but it's wrong that I had is that I would be faster on my own you know to build XYZ thing or try XYZ experiment I'm actually usually faster at XAI just because I have a groundwork and a team who's probably already done a lot of the steps that I would otherwise have to do by hand.
48:22And there's no one saying no. You mentioned like it's kind of a fuzzy blurred line between people working on different people working on different things. Has there been any ability for you to kind of go to other people in the organization and just ask for help? All the time. What does that look like? I walk up to their desk and I say, hey, here's my question. What are you working on right now? Can I support any of that? And can you help me with this? That's it. Everyone's in the same building. Actually, funny enough, we started testing some of our human emulators internally within the companies as employees.
49:04And in some cases, like we didn't really tell anyone about this. And so in some cases, there'll be someone like someone doing some work. and someone is like, hey, can you help me with this thing? Or like, can you do this thing? And the virtual employee is like, yeah, sure. Come to this desk. Come to my desk. And they go there and there's nothing there. It's like the quad situation where it's like we're going to show up. And I think when they first rolled out their vending machine, it was like, I'm going to see you tomorrow. And then it's obviously like a piece of code. Yeah, exactly. And so multiple times I've gotten a ping saying like, hey, this guy on the org chart reports to you.
49:39Is he like not in today or something? it's just like an emulation and it's uh it's an ai it's a virtual employee um but yeah generally we all expect to be like in the same building and reachable to each other um so uh it goes like always and uh i can ask for help i people ask for me for help all the time what have been the biggest like blunders that have happened hmm so um with the human emulators with the customers that we're working with um when we try to understand like we always try to understand what the job that they're doing is and all the facets of it um frequently we'll you know talk to them we'll interview we'll even watch them well actually we'll do the watching the last step so we'll talk to them we'll interview they'll give you go see the write-up or we'll just meet up with them and write notes as to how they do their job and then um like a week later we'll look at the uh mistakes that the virtual employee is making and realize like well it's always making mistakes in these places and in these specific cases what's going on and we go watch the human doing the same thing and there's like 20 different steps that are missing that they just totally left out and we go to them they're like oh yeah we do that like i forgot to tell you my bad it happens all the time um a lot of things people like i guess assume automatically it's all for granted in their head totally an autopilot the same way that you um can like be driving for an hour and not remember a single second of it and not be paying attention can be totally in your own world um this is the same for every thing that a human does uh repeatedly and that's what we're trying to solve basically is all the uh dumb stuff that humans do repetitively right now that they don't need to um try to solve for that case exactly how do you decide like which which thing to go after?
51:26What's like the, in your head when you're thinking about that, what are the biggest things outside of driving that humans do all the time that they just don't need to do? Anything repetitive on a computer. So like customer support is a big one where it's just taking in free form input from arbitrary customer in arbitrary form factor and translating that into a standard workflow that is purpose-built for like uh an AI to take care of that so that human could go and do something more creative and uh use their brain like in a more effective way um totally the same like it's a total parallel to what happened uh in the coding world like okay I don't need to write the same uh you know implementation 20 different times anymore I can describe it in like three words and it's done.
52:19It's a huge compression step. And this is the same thing basically, but for arbitrary digital workflows. On the human emulator side, you run into this problem of humans not existing. And then like someone says, come to my desk and the person doesn't exist. Is there any other thing that's been kind of surprising on rolling that out internally? Surprisingly, we've been able to generalize to more cases than we thought. We test and are pleasantly surprised a lot of times. Like just today, we gave Elon a few cases where we did not train on this task at all, but it did it flawlessly, like perfectly, like way better than we would have expected.
52:57Yeah, the generalization is better than we expected for sure. And we're still at a very early stage, so it's only going to get better. And it's, again, the same parallels to full self-driving where there's stuff not in the training data that the car does react to perfectly due to generalization of an otherwise very, very small model. It's a matter of basically weight efficiency. For the Elon meetings that you've been in, what does that actually look like? They're pretty simple, honestly. I've been lucky that most of the ones I've been in have gone mostly pretty smoothly. What does smooth look like?
53:42Smooth is limited feedback or thumbs up. That means like, okay, you're going in the red direction. Keep going. I'll hear updates next week or whatever it is. If there's feedback or total reversal of direction as a request, then we messed up somewhere. And the question is where. So that's usually we don't even have time to identify that. But that's something you just build up implicitly as a muscle as you go on. And sometimes assumptions also change based on new information. That always happens in every case. So when it comes to the top down, it's a little chaotic. I know with SpaceX, the cost for parts and building things is super, super important because everything basically costs a fuck ton of money and time to do right.
54:33For this sort of thing, I imagine It's a little bit less focused on, you know, he's not like necessarily drilling down on, do you understand every part of every process? What does it look like when he's kind of giving feedback?
54:51Usually it's either at a very high level or at a very low level. It's not really often in between. So either on the high level, it's like a product direction or a customer sense. focus on this segment exclusively or don't do this thing at all or whatever. And then at a low level, especially when it comes to compute efficiency or latency, you'll always have a specific suggestion or let's try this. And he's open to being like proven wrong, but it has to be proof. It has to be like, let's try it and see what the results are. It can't be just someone's opinion. there has to be an experiment done, which has led to some surprising results sometimes, and we go with it.
55:36What have been those? So the compute efficiency of going with the small model has led to, well, a lot of improvements that we wouldn't have otherwise thought. Some of them are secondary, some of them are primary. The obvious ones are, well, obvious, being able to go much, much faster than human, but also as a result, and Tesla found this too, with full self-driving, going with the smaller model, they're able to iterate much, much faster. So not only does the model react to situations faster and can be more, I guess, tolerant of timeframes, you can also just deploy iterations much faster. If it was four weeks before, maybe it's one week now.
56:19So that actually goes back to the experimentations, why we can have 20 different ones going in parallels was a result of that particular decision early on in the chain. And was the initial idea like go just do big, large models and then it changed? Sort of. We definitely wanted to go faster than everyone else. But the question of how much faster was, well, the answer to that was amplified, basically, multiplied by a lot. There's this like a lot of bias and stuff in Wikipedia. And Elon has been like focused on kind of creating an alternate version that's just kind of like, you know, more truthful in effect.
56:55How do you go about basically cleaning up the internet in that way to figure out what is truth? It's a really hard problem. It's very hard, especially because the internet is not usually the ground truth for whatever thing it is. So wherever we can, we try to drill down to the fundamentals, which is very hard. like I don't know what is the fundamentals like in physics of the constitution that's not really a question I think I can answer or anyone could really faithfully answer very well but you try to do something like that um drill down as close as you can and then build up from that which is hard too because there's not actually a big body of like writing that does that um one of the few probably examples is like James Burke um with his connection series is where he'll take two totally seeming unrelated concepts and then connect them um through physics and inventions um it's really cool and we're trying to do the same but uh it's fairly novel how do you find better data data is not the only thing that goes into the results yeah like how you actually train on that data.
58:11And I know it's a pretty broad term, but it is true. Like how you actually evaluate against that data and train against it and your different methods for updating the weights do matter a lot. You can try to faithfully recreate the input or the output given any arbitrary input. And well, you can create basically a horrible copy-paste mechanism if you want, which is a classic problem in ML. There's a bit of an art to it to avoid that problem. But I guess we're a few steps removed from that at this point. We're not measuring the fitness to any particular data set. At this point, we're trying to measure to an arbitrary output.
58:56So it matters a lot how you construct your evals, which is really hard for truthfulness because then you need to know the truth, which isn't always... Well, I mean, that's really the problem we're trying to solve, right? So it's kind of chicken egg. Yeah, there's, like, a lot of different approaches and a bunch of smart people working on it. Yeah, if anyone has suggestions, please send them through. There's, like, a lot of different ways to look at it. There's been, like, moments in time where I've seen Elon on X and someone has said, like, this is obviously not right, and it's, like, some Grok output, and he's like, we're going to fix this.
59:31And then, you know, 12 hours later, 24 hours later, he's like, all right, it's fixed. When that happens, like what happens internally? He shows us what went wrong. And then quickly, whoever is awake at the time will start up a thread to go and solve it. Usually individually and pull in a few others if need be. And then give a post-mortem on what happened. And everyone will understand then what went wrong and how to avoid it in the future, ideally. yeah the like generally making mistakes once is okay but making the same mistake twice is big problem throughout spacex's history there's been a number of and same thing with tesla there's been a bunch of these like surges where randomly elon will like come in at midnight and say you know like everyone that can't come in like send out a company-wide email and say like come in we need to be working that sort of thing um has there any been anything like that it's more for the big models that happens more than anything.
1:00:30For Macroheart specifically, we've been operating in a war room for four months. So we've kind of always been on that push. Do you guys have a sign on the door that says war room? Yeah. Amazing. Actually, well, yeah, we grew the original war room. And so we moved everything out. And I'm told Elon walks into the war room and it's totally empty. And he's like, where is everyone? What? And he walks over to where we are now, which is just the gym, which we cleared out and put everyone in now. And then conducts his impromptu questions of what's going on. That was a long night. What is it like on one of those nights where a lot of things kind of get shaken up and moved forward or like there's one of these surges?
1:01:15What does that feel like? I think actually I saw this from one of the co-founders of XA posted this recently. Igor, who was great to work with. I liked him a lot. It was actually really cool to work with him, side tangent, because his work on StarCraft AI way back, like, I guess 10 years ago now almost, it was one of the first, like, cool ML work that I tried to replicate myself in high school, which was very hard. It was really cool. So it was really cool to work with him. Like, I totally never thought I would get the chance to. But anyway, I saw him post this thing a few days ago where he's like, okay, there are some months where only a few days go by.
1:02:01And then there's some nights where months happen. And that was like one of them for sure. Months might be an exaggeration. I think we would have gotten to the technical result. We would have in a few weeks anyway. But doing it in one night was a huge push. And it was a long night. has there been any moments where the company just didn't leave the office for like five days or like a week yeah the surges for the models usually results in a lot of people staying in overnight um and you mentioned there's like five or six pods that people can sleep in and they like toggle out yeah there's some there's some sleeping pods and we have some bunk beds now too um which are less less nice but they exist um and then when the tent picture came out everyone kept sending that to me and I was like honestly we have tents but I've never seen that many out at once um so yeah I know you worked on a bunch of different projects as a kid and I think I don't know if this was the first one but it was like fidget spinners and and making fidget spinners um I don't think it was in your garage but maybe it was like in your room yeah what kind of stuff like that tinkering mindset how much of that have you kind of taken to this uh quite a bit quite a bit yeah um so I learned programming when I was quite young.
1:03:15Um, my dad got me a book when I was like 11 and I liked it a lot. Um, well, I liked it a bit, but I really started to like it when I realized you could make money from it. And so, um, I met some people online who were basically writing scripts for games as hacks and would sell them online, um, for small amounts of money. But you know, making a couple of hundred bucks online was huge for me. Um, the first time that you like have someone give you money, it's the strangest feeling. It was huge. It was crazy. Yeah. I remember having and I asked my dad for like a PayPal, like custody account or whatever, and getting the money in, and it was like the coolest thing of all time ever for me.
1:03:51Yeah, it was really big. And so I did that for like a couple months and saved up enough money to, at the time I was really interested in added manufacturing, like 3D printers. RepRap was the big thing then, so that was kind of where what kicked off the modern 3D printing revolution. You built your own, right? Yeah, you had to. That was the only option. RepRap is literally just a bunch of university students, basically, who said, like, let's see if we can build a machine that can build almost all the components for itself, which was why it was called RepRap. And they basically built in a variety of universities these rooms where you start with one printer, and then it prints the parts for the next printer, and you go all the way up and you scale up.
1:04:40There's lots of problems with this, as it turns out. And that's what they were solving. And it eventually kicked off the modern 30-per-new resolution. But I was very obsessed with it. And so I took one of their parts lists and bought everything from Alibaba. And a month later, things came in. And I assembled it all one night, which went poorly, actually, when I was unbundling the copper cable for the power supply, which was a very sketchy power supply and did catch fire in the end. All the copper windings came loose and frayed everywhere. And one went two inches into my thumb. um just can't your phone doesn't work or did you go to the hospital or something no so it was a school night and it was like 3 a.m because I wasn't very good at building things at 13 um and I spent like an hour in the bathroom trying to pull it out with tweezers it just wasn't it was like it was bad so I just cut it off and I was like it and so bit by bit over the next few weeks it came out and I would snip it off in the mornings um it was fun um yeah uh but I got the printer assembled.
1:05:40And around that time, yeah, the fidget spinner craze was going off. So I bought a thousand skateboard bearings from China and basically set up a little factory in my bedroom where every two hours at night I would wake up and I would clear the print bed, start a new print of fidget spinners and I would sell them online. And then before school, I had a little assembly line in my garage where I would put in the bearings, spray paint, dry, and then run around to all the other bus stops of the other schools, sell them to my distributors, which were just other kids of other schools, sell all day at school, come back, collect from my distributors, and then sell online, ship, built a little healthy business.
1:06:15And after two months, I ended up getting shut down by the county. Their official quoted reason is that the companies that sell the school food have technically an exclusive license to sell anything in school property. But I think they just didn't like that I was distracting everyone and making money doing it. But it taught me a good, like, healthy disrespect for authority, I think. That has kind of been like a constant theme. What is that actually, how is that materialized in your life with, like, the healthy disrespect for authority? And you even mentioned institutions. Like, you don't necessarily trust institutions.
1:06:52How did you kind of come to that, and what does that look like?
1:06:57I've always known from very young. like I want an unconventional outcome. And so going through a conventional path would pretty much necessarily not lead you to an unconventional outcome. So I grew opposed to any form of convention and institutions necessarily enforce convention. I think creativity and interesting outcomes come mostly from free-spirited individuals in almost every case, if not all of them. I guess it's a bit of a high-minded way of saying it, but yeah, individuals are the most creative you can get. And so staying true to that is the way to go. I do love John Collison's idea of everything is so hard to build and so hard to make, especially put into the real world, that if you look around, it's basically like the world is just filled with people's passion projects.
1:07:53Yeah, it's a total miracle. and there's a story behind every little thing way more than you would think i remember reading about the um i think it was ykk zippers apparently every good zipper like there's two or three companies in the world that make zippers which are actually pretty little little miracles they're very cheap but also mechanically like relatively complicated for how much you pay for them and there's only a few companies that are capable of building or that have set up to build them um And it's, yeah, basically it's one Japanese guy's passion project over 40 years to figure out how to do this properly.
1:08:27And this is the case for pretty much everything. Anything very specific and at scale is probably only done by a few companies or a few people in the world. So, yeah. I mean, you hear about it every so often, right? Like some company in Germany, arbitrary company in Germany shuts down and Volkswagen has to halt all their lines or something like that. Happens all the same. It was a big thing in COVID. Right before we met, you had made a liquid fuel, I think, rocket engine. It was like a very small thing I saw downstairs. But you said we were talking before this that you did it in like 24 hours just on a whim.
1:09:02How did that happen? Yeah. So it was a project over like roughly four weeks. And I started by literally just buying a bunch of textbooks and trying to figure out like what are the design principles behind a rocket engine? Like how do I design it? There's not like, it's totally different from learning software where you can just go in GitHub and download people's code and modify it. There's no file for a rocket engine. You have to learn how to, like, what are the material properties? What's the chemical properties? How do you actually machine it? How do you design the parameters and know what to expect from, in terms of thrust output?
1:09:38And how do you not overpressure the engine and all these kinds of things? How do you design the injector, which is, the injector was very hard. That was probably 50 % of the time. Is that the hardest thing? yeah the injector was very hard and it was like the biggest flaw in the end um so yeah i spent like three four weeks doing this and uh expedited a bunch of parts from china like cnc'd and all that stuff um and uh it was right before thanksgiving i was gonna go fly back to the east coast and visit my family and i was like okay either i fire it like build it and fire it tonight it was all just a bunch of parts at that time uh or i do this in two weeks and i'm like i'm not drank a lot of coffee in the morning and spend the whole day like hacking away at Eliminant extrusions and built out the test frame and then the engine itself and Let it off that night Yeah Which had a lot of um We'll say Concessions made to make it happen that night I Do find it absolutely hilarious that you like you said were you like a couple feet away?
1:10:44Yeah, so I designed it like I wasn't stupid it i designed it so that i could remotely fire it but um i didn't the power supply hadn't come in yet to remotely power the computer that was on board so i had to use a usb cable from my laptop to power the onboard computer and i didn't have a long enough usb cable uh the longest one i had was like six foot so i just stand right next to it and light it off and i was like there's like a 30 chance that this thing explodes or or launches fire everywhere and actually um i don't know if it shows in the video I think it does show in the video but my jacket did catch fire because I wasn't that great at designing the injector and it did create a lot of overpressure events which meant there was a lot of basically uh unburnt fuel spewing out which was ethanol and so that's liquid and just landed something that landed on my on my jacket and caught fire um so yeah that's a trophy still the burnt jacket Bye.
From the publisher
How Elon deals w fires, how xAI built their Colossus data center in 122 days, and why it's important to challenge requirements.
