Own or Be Owned: Why Every Company Needs Its Own AI Model (Yash Patil, Co-Founder & CEO of Applied Compute)

23 Jun 2026 · 1 h 8 min · 25 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode argues that companies should “own their intelligence” by training and serving custom AI models on their own data, rather than relying on frontier closed models that may restrict capabilities or degrade performance. It also discusses why AI job loss predictions are overstated, how cost and model routing drive adoption of smaller specialized models, and how post-training (including RL with verifiable rewards) enables domain-specific performance.

Guest

Yash Patil, co-founder & CEO of Applied Compute. Background: ex-OpenAI researcher (joined via Sam Altman’s residency program after dropping out of Stanford sophomore year). Worked on post-training infrastructure; helped develop early ChatGPT-related systems. At 23, runs Applied Compute, a ~$1.3B company training custom open-weight models on customer data for production.

Key claims

relying on third-party AI tools can “atrophy” deep engineering thinking; frontier providers control access; transformation will take decades; “frontier model for every task” is like using a blowtorch; evals are the “scorecard” and should be owned.

Notable examples

Palantir signing large contracts before solving open research questions; DoorDash storefront menus converted from menu images using a specialized model; discussion of OpenAI’s Sam Altman firing/firing week experience.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Impact of Relying on AI Tools

0:00 to 2:08

The discussion centers around the potential downsides of over-relying on AI tools and the importance of critical thinking in engineering.

“There is truly some atrophying happening by relying too much on AI tools.”

Introducing Yash Patil and Applied Compute

2:08 to 2:48

Host introduces Yash Patil, discussing his background and the focus of his company on training custom AI models for businesses.

“The best founders aren't spending their time on expense reports.”

Introducing Yash Patil and Applied Compute

2:53 to 3:18

Host introduces Yash Patil, discussing his background and the focus of his company on training custom AI models for businesses.

The Importance of Owning AI Models

3:51 to 6:00

Yash discusses the significance of businesses owning their own AI models and the implications of dependency on external AI providers.

“And we've been chatting already sort of before we've sort of kicked this off.”

The Challenges of AI Model Control

6:00 to 8:02

Exploration of how control over AI models by providers can impact users and the competitive landscape of AI development.

“It's kind of like a model-less company is sitting on shifting sand.”

Building Custom AI Solutions

8:02 to 12:39

Yash explains how Applied Compute helps companies create custom AI models tailored to their specific needs.

“Like you cannot cover the full surface area of what these models can do just within the walls of Anthropics.”

Yash's Early Influences and Journey

12:39 to 14:00

Yash shares anecdotes from his childhood, including his early interest in coding and the influence of his family.

“So I want to come back to it, but maybe to sort of set the stage a little bit, I'd love to hear about your story a bit.”

Inspiration from Family: The Coding Journey

14:00 to 17:42

Learn about the speaker's journey into coding, inspired by his older brother and father's encouragement.

“Like they really paved the or, you know, older siblings in general paved the way for the younger kids.”

College Experience: From High School to Stanford

17:42 to 19:38

Discover how the transition to college influenced the speaker's approach to learning and projects.

“And I think I've seen you say somewhere that, you know, you're a very good student in middle school and high school.”

Dropping Out for Opportunities: Joining OpenAI

19:38 to 21:51

Explore the decision to drop out of Stanford to join OpenAI and the motivations behind it.

“So I dropped out of Stanford during my sophomore year.”
Show all 25 chapters

Inside OpenAI: Experiences and Growth

21:51 to 26:32

Hear about the speaker's experiences at OpenAI, learning from legends in AI.

“you got the OpenAI job by cold emailing Sam?”

The Dramatic Week at OpenAI: Sam's Return

26:32 to 28:00

Understand the impact of leadership changes at OpenAI and the atmosphere during that week.

“But like it was kind of this feeling of like, oh, my gosh, like things are going so well.”

The Power of Leadership in AI

28:00 to 29:24

Learn about the importance of leadership in making visionary projects possible.

“And so I like came into the office and Sam was actually coming down the stairs.”

Understanding Post-Training in AI Models

29:24 to 33:46

Explore the significance of post-training in developing AI models for specific tasks.

“And that's really how you get the best people is you work on really hard problems because all the best people want to work on hard problems.”

Evolving AI Services for Businesses

35:40 to 42:00

Understand how companies can train their own models for tailored AI solutions.

“What was the moment that you realized there's a need for the sort of service that applied compute offers where you are sort of going into these companies and doing this more hands-on work?”

Scaling AI Model Infrastructure

42:00 to 45:12

Learn about the infrastructure needed to train and serve AI models effectively.

“But really, it's the infrastructure that is repeatable across all of these companies, right?”

Continuous Engagement in AI Deployments

45:12 to 48:36

Explore how ongoing engagement is crucial for improving AI models post-deployment.

“a bit more narrow, a bit more deterministic that like we can actually very inexpensively do this very effectively.”

Case Study: DoorDash's AI Integration

48:36 to 51:58

Discover how a specialized model improved DoorDash's storefront creation process.

“That is difficult, and that is the problem to solve in continual learning.”

The Future of Open vs Closed AI Models

51:58 to 55:18

Understand the implications of open weight models and their growing importance.

“So I think it just naturally like the kind of work just naturally attracts these these types of folks.”

Building Effective ML Infrastructure

55:18 to 56:03

Learn about the challenges and strategies in developing ML infrastructure.

“It's not like RL is completely figured out and it's kind of like a commodity thing.”

Training AI Models: Asynchronous Methods

56:03 to 58:20

Discover the innovative approaches to training AI models asynchronously to optimize performance.

“So, So you have some checkpoint, some weight checkpoint that's like policy step zero or something.”

The Future of Personal AI

58:21 to 1:00:14

Explore the potential rise of personal AI models for individuals and their implications on compute demand.

“like, so there's that side of the house, but then there's also environment creation, right?”

Challenges in AI Deployment

1:00:15 to 1:02:26

Learn about the challenges and time required for AI diffusion into large organizations.

“This kind of ties back to tokenpocalypse.”

Thought Experiments on Technology and Humanity

1:02:27 to 1:05:26

Engage in thought experiments about humanity's potential to recreate technology post-collapse.

“What has surprised you most in terms of that change management, maybe the slowness of some of these enterprises?”

The Importance of Fundamental Skills in AI Age

1:05:27 to 1:07:15

Reflect on the necessity of critical thinking and fundamental skills in an AI-driven world.

“If you could assign a book to everyone on Earth and you knew they would read it and they'd understand it, what would you want to give to people?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00There is truly some atrophying happening by relying too much on AI tools. The one thing I'm worried about people losing is this ability to actually deeply think about how things should be architected and designed. The best engineers right now are the ones who are like, learn how to code before AI. Because at the end of the day, I think humans will be the overlords over all of these AI tools. And you will still need to do critical thinking to know what you want to build. Palantir is one of the few companies that signs massive contracts without knowing if they can solve the problem. They'll sign like huge, you know,$100 million deals before they've started working out the problem.

0:32I think we've been into that category where there's open research questions. And people who like to work on things that are uncertain are just naturally attracted to companies like AC. I know there's a lot of folks who are like, oh, 50 % of jobs are going to go away in the next two years. Personally, I think that is not going to happen. I think just the diffusion of AI into the real world takes time.

1:02Demand for intelligence is outrunning the chips that supply it. Meanwhile, tokens are being sold at subsidized prices that won't last, demand keeps climbing, and companies are discovering that using a frontier model for every task is like cooking with a blowtorch. Spectacular, and the wrong tool for most jobs. Yash Patil lives at the center of that squeeze. The 23-year-old ex-OpenAI researcher runs Applied Compute, a$1.3 billion company that trains custom models on a business's own data. Smaller, cheaper, and built for the work that that company actually does. One year in, Anjash's startup is already serving a mix of tech startups and established corporations, helping them enter the AI era all while owning their own intelligence.

1:51In this episode, we discuss why open-weight models are quietly taking over real workloads, why Yash believes AI's transformation of the economy will take decades rather than years, and what it was like inside OpenAI the weekend the board-fired CEO Sam Altman. I'm Mario, and this is The Generalist. The best founders aren't spending their time on expense reports. They're busy building. Brex is the agentic finance platform that makes that possible. High limit corporate cards, banking, and AI that handles the back office automatically so your team never has to. Expenses get captured, books get closed, and spend stays in policy without anyone chasing it down.

2:32Your team gets their time back to focus on what actually moves the company forward. Vursell, OpenAI, Anthropic, Granola, and Deepgram already run on Brex. One in three startups in the U.S. does too. It's time to get Brex. Go to brex.com slash solutions slash startups.

3:18Every insight, every answer, every recommendation is grounded in verified knowledge, not outdated information or hallucinations. When your teams and your AIs share one trusted foundation, everything moves faster, with fewer redos, fewer blind spots, and more confidence in every decision. Because in the age of AI, truth isn't just power, it's protection. See what Guru is doing for thousands of companies like Spotify, DHL, and Stripe at GetGuru.com. That's getguru.com. Yash, lovely to have you here. And we've been chatting already sort of before we've sort of kicked this off. And I think it makes sense to maybe start with some of the things we were chatting about.

4:02If that sounds good to you. Of course. Yeah. Yesterday, Fable 5 came out. Big day in the world of AI and has implications for all of us, but including on the sort of the business you're running. I know you've got a few thoughts on it. Like, you know, what came to mind for you when you sort of saw it? I think, first of all, the model is amazing. Like, a ton of our team members were using it. It's like, clearly feels like a step function. You know, I think there have been moments in time where we've seen better and better models. And this is definitely like a point in time that should be noted. I think, though, the thing that you're referring to is kind of when we were looking at the system card, and this was kind of blowing up on Twitter and whatnot.

4:45Yes. There were some stipulations about how you could use the model, particularly some of the guardrails where the model might not answer in a super intelligent manner or it might sort of like answer around particular types of questions. And I think in the past, it's been focused on things that are very reasonable, like safety concerns and things like that. Things that kind of our consensus view models helping with this kind of stuff is not great. But yeah, in this case, it was queries related to like AI model development, which really sparked this question of like, hey, how much control do the frontier model providers have and how you use these models?

5:31And, you know, if they do end up sort of holding out capabilities for particular types of work, what does that mean for the downstream users of these models? So the sort of natural conclusion and what the core thesis of our company is, is sort of own or be owned, like having ownership over your own model and your own intelligence is actually really existential for a lot of people. And, you know, if you don't start investing in owning your own models, you might kind of be in a tough position when, you know, these things do change underneath you. It's kind of like a model-less company is sitting on shifting sand.

6:10Yeah. And, you know, the I think the consensus things that people understand, you know, that Anthropoc might not want to help you with is like biochemical weapons and such. But, you know, folks that are relying on these models to do more research, discovering that that is sort of maybe in some places subtly being degraded is totally, you know, difficult to to adapt to. Totally. And there are stated reasons for why they're not doing this, particularly around safety in terms of model development, where it's unclear whether having everyone have access to these tools is beneficial for humanity. But the big thing is who's going to make that call.

6:54And the people who make the models are going to make that call. So and then this not so charitable view of this, right, is Anthropic is doing very well and making some extremely powerful models. And this kind of prevents other people from being in that race as well if they want to use Anthropic models to build up their model development stack. So it's a little anti-competitive. Yeah. Yeah. My view would be that I think they really believe it. I think they sincerely worry about this stuff. Everything they put out into the world is fairly consistent around that. I don't know. Do you find it to be more of an anti-competitive behavior?

7:34No, I don't. I actually don't think it's anti-competitive. I think it does come from a really good place, which is like, you know, I have friends at Anthropic, and I really do respect the leadership. They handled the whole Project Glasswing and deployment of Mythos really well. Like, you know, they do get some flack for it. But I think if it if you do have a powerful new technology, it makes a lot of sense, you know, for you to take things slowly and roll it out to preferred partners first. Also, you know, with this new Fable model, there's like some stuff around data retention, which makes sense.

8:07Like you cannot cover the full surface area of what these models can do just within the walls of Anthropics. So you actually need to see how people use it. I think the important point is, and these things can be both true at the same time, I think, you know, Anthropic is acting in really good faith, but also at the same time, it does raise the concerns of like, hey, if you don't have control of your model, if Anthropic ever does want to hold out a capability for competitive reasons, you know, what happens to the users of those models? Because, you know, I think Anthropic is an amazing business.

8:40They're like, I mean, they're obviously doing incredibly well and they're moving into the application layer and things like that. So there are products that rely on Anthropic models that are going to be competitive with whatever Anthropic puts out into the application layer. So I guess the question is what happens when, you know, you're relying on those and they want to push their products higher up the stack. Yeah, there was always the phrase, I think it was around Apple, when they would release something that sort of immediately threatened something within their developer ecosystem. It's Sherlocked, right?

9:14Yeah. People would say that. Yeah, yeah. Anthropic, every time one of these models comes out, it's like immediately someone is depositioned potentially in these different ways. To come back to the idea of owning one's intelligence, maybe we should talk about it in the context of your company Applied Compute. What does that mean? Why is that so important? And why is that the challenge that you're spending your time on? Yeah. So what we do at Applied Compute, maybe I'll give like a little bit of an overview, is we help companies own their intelligence by training custom models on their data, being able to serve these in production at scale, and then continuing to improve them the more they're used.

9:53In practice, what that looks like is we benefit from the open source community, basically taking open-weight models that are extremely powerful, actually. I think, you know, if you sort of track the development of open source models compared to closed source models, you know, they've kept up and a lot could argue that they've sort of been closing the gap from when the first open weight models started to come out. But yeah, we basically are able to harness these open weight models, combine it with customer data to create task specific or specialized models. So the idea really comes from the fact that we don't think all AI workloads are homogenous.

10:33There are places where you want to use the right model for the right task, and you want to optimize for cost, latency, performance, modality, deployment, flexibility, that kind of stuff. So our vision is that AI is actually going to be very decentralized in the sense that many companies will own their own models rather than there being a sort of single model that does everything. Yes. And so to put it just in probably a too dumb analogy for folks, it would be sort of a Goldman Sachs of the world has all of this useful data. I think you've even used maybe Goldman as an example before. That's top of mind.

11:13It has all this useful data context. And just unleashing a general model into that loses a lot of what the value of this place is. And you also don't want to necessarily give that data away. Totally. Yeah. I mean, the way these companies operate on the inside is like a lot of what makes them special and differentiated. So, you know, I had Brendan, a good friend of mine and CEO of Mercosur. He was the one talking about the Goldman stuff and he was over at our office last week kind of talking about what they're seeing in enterprises. And, you know, I think the sort of manifestation of what differentiates a company when it comes to AI is kind of coming up in this talk around evals, right?

11:56So, you know, companies have their own definition of what good looks like, how they, you know, do certain tasks and procedures inside of their company, and they're sort of codifying it within evals. and there's no way they would ever give those evals away to the model providers because that's a lot of their secret sauce. There's some things that they might want to share to improve the capabilities of the models, but anything that's proprietary and unique to them, that's what they're going to want to hold close. So many of these topics are, I feel like, so live right now, like the question of the value of routing and cost.

12:39Pokenpocalypse. Yes. So I want to come back to it, but maybe to sort of set the stage a little bit, I'd love to hear about your story a bit. I was seeing that when you announced Applied Compute for the first time, there was a tweet from someone you went to middle school with, which was very sweet. And it described you in such an interesting way that it made me very curious. It describes you as the middle schooler hanging out with high schoolers, And someone who had a very unusual sort of unconventional mind. Yeah. What were you like as a kid? Were you just building things from the very beginning?

13:13Yeah. So that's actually really funny. I know exactly what you're talking about. It's a friend of mine. Her name is Sarah. We went to high school together. But she was a few years above me. So, you know, on a little tangent, there is this club called Science Olympiad. And we had like there was like this middle school where a lot of people went to high school like this specific high school afterwards. She was there in high school. I was there in middle school. And those those science Olympiad teams often like hung out together. So that's how I got to know her before I was there. But yeah, in terms of in terms of childhood, like I can tell you a little bit about my family.

13:50So, you know, a small family of four. Mom is like a pediatrician, runs her own sort of medical practice. my dad's like an electrical engineer works on chips and and things like that very relevant now um and i have an older brother uh who who uh is in tech and uh is like software engineer now works at a really amazing company called chai chai discovery plug for chai i love it chai discovery um but growing up yeah like my brother was actually the person who got me into into coding so So I feel like older brothers always like get way less credit than they deserve. Like they really paved the or, you know, older siblings in general paved the way for the younger kids.

14:34And yeah, he was he was the first to sort of start fiddling around with like building apps and games and things like that. And one of the things he did sort of when he joined high school is he built this grade checking app. So the grade checking system in our school district was terrible. Everyone hated it. So he just built this app and then got a bunch of his friends to use it. Eventually, the whole school was using it. Eventually, a bunch of schools in the whole school district were using it. And I just thought it was so cool that he, as a single person, was able to build this cool thing. And now everybody I know is using it.

15:13So that really got me into this mode of like, oh, wow, I want to do that too. Basically, classic copy of the older sibling. So my dad actually encouraged me a lot to build these kind of like little fun, unique apps or platforms and stuff. But he doesn't know how to code. So he would give me the ideas and then he'd say, go, go build it. I'd go and build it and show him and he'd be like, oh, you should do this or that or something. So like, you know, for example, so my dad grew up in India, like a village and is very communal type type living. and the way it would work is each family would like cook a bunch of food of like the thing that they were really good at cooking yes and then they would like go and like sort of trade food with like other families and things like that that's sweet yeah so so he was like oh you should just you know your mom makes really amazing burritos and then we know this neighbor makes really good barbecue so he was like you should go and like build something that helps us trade food or we're not built it uh it was really fun no one ever used it but it was uh in america exactly but it was a lot of fun to just kind of like iterate on this thing and see something come to life and that that project actually got me my first internship oh wow uh yeah my my neighbor he's in tech and he was working at this uh small fintech startup that was like four people and my dad was like in classic dad fashion was like in the yard talking to him about oh you know yasha's working on this thing like yada yada yada and he was like oh do you want to do you want to come and work at our startup for the summer and I was I'd never done anything like that so I was like oh that's super cool this was like kind of freshman year of high school so um yeah he took me to their office which was just an apartment uh building um like one unit there and I thought it was super cool to see like all of these just normal people like working like it's such a small group of people working on this this sort of like hackathon project yes you know it was it's very different than when i visit my dad at work and he's he's in the cubicles and like yeah like hundreds of people and things like that yes it was like i was like cool this is like a bunch of people who are doing exactly what i wanted to do in my free time with my dad yes but as a job doing it as a job and so i was like this is super cool so uh interned there and a couple other places that's how i kind kind of got into startups and stuff.

17:39But yeah, it was kind of cool. Amazing. And then you go to Stanford. And I think I've seen you say somewhere that, you know, you're a very good student in middle school and high school. And when you came to college, it was sort of like, you know, I've arrived and I'm going to basically watch the lectures and build stuff. No, no, exactly. So I mean, in middle school and high school, grades were super important. I think it's like a common thing and a lot of immigrant families as well um but yeah when i went to to college it was kind of like a give me an inch and i'll take it so i was like oh like you know classes are online that's cool i'll just like watch the lectures online i spend most of my time just like hacking on like basically uh manifestations of what i was working on like childhood like middle middle school and high school so but for like the university setting so nothing i want to be very clear none of this was impressive or like interesting or anything like that.

18:33I was like working on like every single generic college campus coding project you can think of like social calendar for like me and all my friends. You have to do these things or like you get these reps. Exactly. Or like a dining hall voting app or something like that. Because, you know, some food is better than others and things like that. But yeah, just through that, like found other people who were kind of doing the same thing. And I think Stanford's a really cool place because there's these sort of natural communities that form and um you know there were like little informal clubs of like oh let's get together and hack on a project or something like that and then sort of the main organization i was affiliated with at sanford was tree hacks which was like the school's hackathon um that we throw for for like everyone across the nation and even some folks internationally so yeah spend most of my time just like kind of working on projects and stuff rather than uh uh attending class you'll have to correct me if I'm wrong, but you're a Teal fellow, but you also seem to have gotten a degree.

19:35How can this happen? So I did not get a degree. Yeah. Yeah. So I dropped out of Stanford during my sophomore year. Okay. About halfway through my sophomore year. So yeah, was basically there for a year and a half. Dropped out to join OpenAI, though. And then after OpenAI started to apply. And OpenAI, I think, you know, the way you've talked about it before is like you were just so mesmerized by this technology. Like, can you, yeah, tell me about that. I mean, I looked super fondly upon OpenAI because, yeah, obviously the technology was amazing, but the people were like even more amazing. So I actually didn't really take any AI classes when I was at Stanford.

20:22so joining open ai is how i learned a lot of this stuff and there's there are like were incredible people on the post-training side that kind of took them under their wing and like taught me a lot of stuff so you know uh luke metz barrett's off uh a bunch of folks on the the post-training side were like they were my managers and i was just like kind of look like looked up to them you know they're they're legends in like the language modeling space they like made kind of made the first versions of ChatGPT with John Shulman. It was actually really funny. So Barrett, who used to be my skip, told me on his last day when he was leaving OpenAI, he was like, Yash, by the way, I don't know if you know, but when you're joining, you were originally supposed to be on the ChatGPT engineering team.

21:09And we just really needed more people on post-training infrastructure, which is on the research side. So it was like last minute we pulled you in. And I didn't know because you don't know your team before you join and it kind of you know that's basically the reason why i got into all this ai and research stuff is completely by luck yeah it's had a few things in motion yeah it said a ton of things in motion i feel super grateful for like the people that were there and then the fact that like they were willing to like teach me a bunch of stuff and or take me under their wing and yeah i mean like this this guy ian osmond who's now at deep mind through my 21st birthday party.

21:45It was a lot of fun. I was lucky to be surrounded by some really, really awesome people. Am I right in recalling that you got the OpenAI job by cold emailing Sam? Yeah. The backstory there is rewind to freshman year. We had been hacking on some projects and stuff and my friend and I both had internships lined up for the summer. We were like, oh, let's just keep working on cool things and like you know keep keep hacking on stuff um so we stanford is a it's a very unique place because i think there are a lot of people who encourage folks to to do this like to go and just like build stuff so we were you know there were a few venture friends that were like hey like here take a angel check or something like that and go and spend the summer building something and then if you know if you if you end up not wanting to continue it's totally fine or whatever um but growing up like I knew that nothing comes free.

22:41So prior to this, I had been invited to Sam's place for a dinner. He puts on some dinners sometimes for people who are building stuff on campus, whatnot. There were, I think, a couple dozen folks who went. And Sam is someone I deeply, deeply admire. I think he's amazing. And so I had watched all his YC videos and read all his blogs and stuff like that. couldn't work up courage to actually talk to him at that dinner. But a couple weeks later, when we were like, oh, should we do our internships or not? Shot him an email and was like, hey, like there's some people who are offering to sponsor us for the summer.

23:21We don't really know who they are. We're not even sure if we feel comfortable. Like, what do you think we should do? Do you know these people? And he was like, you should totally do it. If you're this is back when Sam was kind of talking about like, oh, people should spend more time building stuff and take time off of school to like try and explore things he was like if you want i can basically give you a small stipend um to just like pay for rent and food to go and work on it and i'll sponsor you for the summer which is like deeply generous of him and and we were like wow this is cool like and so so um so yeah so we we just did that for for the summer um worked pretty closely with someone who was like running uh sort of his his family office at the time who became a friend and then ended up shutting that down though.

24:09It was like an e-commerce company, but ended up shutting it down and coming back to school and like returning the money and things like that. But yeah, that's how I got to know him. And then when ChatGVT came out, I was like, wow, this is the coolest thing I've ever seen. I have to work on it. So then I was like, oh, hey, Sam, I don't know if you remember me, but like, you know, I think ChatGVT is super cool. I'd love to work on it. like are there any ways i can apply for a job and then he was like oh yeah like we have this residency program you should apply for it so he introduced me to the residency folks talked to them they were like oh you have to drop out and i was like i'm fine with that i can i can take time off and stuff and they were like no you really have to drop out like you have to like commit to never coming back burn the boats yeah and i was like um i don't know if i could fully convinced my parents of that.

25:00And then so I was like, Sam, like, you know, should I should I do? And he's like, let me chat with them. The residency is like a six month thing. So you'll have the option if you want. Then I interviewed, was very lucky to get the residency position and then joined like basically like a week after. They needed you to drop out like almost as a proof of conviction somehow. I think at the time it was just it was actually just much smaller company. Right. So like, I think they're there. It's not the, you know, massive company that it is today. This was actually something that I don't think they had done before.

25:35Like they hadn't had people do the residency that had like dropped out of college just yet. But after after I did, I think there were a couple other folks who became more mainstream. But yeah, I guess I think they just wanted like commitment and stuff. But I I already kind of knew at the time I was like, I would be so lucky to be able to work here. You know, I totally would. If it was up to me, It was not up to me. My parents had to have a say, right? Clearly the right choice. Yeah. You ended up living through one of the most interesting periods of technological history ever, and also one of the most dramatic weeks of technological history ever.

26:10Yeah. I'm sure there's lots of that you don't want to talk about, but you were there when Sam was ousted, came back. What was that experience like for you as someone having what was, I think, their first proper job? Like that's a very dramatic thing to go through. It really sucked. So I remember that week pretty well. And I don't want to be overdramatic or anything like, you know, things happen and stuff. But like it was kind of this feeling of like, oh, my gosh, like things are going so well. We're making so much progress. Everything is, you know, everything's happening. You know, why are we halting the progress?

26:48Like, let's just go. Like things are going so well. and so i just remember there were like a bunch of basically that whole week was actually like you know we couldn't go into the office all our laptops were shut down for security reasons and whatnot so it was like every day there'd be someone who'd be who's like oh i'm like hosting at my place just come hang out um because we know that there's no work so um so yeah i just remember like everyone being down on the fact that it wasn't even about like the situation like you know everyone wanted Sam to come back and not no one more than me because I think he's really awesome but everyone was kind of sad about the mission just being on pause they were like this is this is weird like what we were doing so well so and then it was a roller coaster too because you know Sam was uh talking with the board and things like that and you know we we all thought he was coming back and stuff and then there is an announcement that you know he wasn't and there's someone else stepping in.

27:46But eventually, right, I think we were super excited that Sam came back as CEO. And it's actually funny. I remember, so after that was announced, everyone was going to meet back up at the office. And I lived, so I lived across the street from the office in this, like the building, the Madelon. And so I like came into the office and Sam was actually coming down the stairs. So he gave me a big hug. And I'll definitely remember that for a while. But that was like probably one of the coolest, coolest days was just like, it's back on track. Let's go. Right. So it sounds like, you know, there's obviously been, I'm thinking of all the stories that have been out over this year at the books.

28:25Sounds like you still are basically, yeah, a big supporter of Sam and think he's a good executive in that respect. Oh yeah. I think, I think, yeah, a little side tangent on Sam. I think he, he's an amazing CEO and his superpower is he's able to make impossible things sound extremely possible like you know if I were to tell you hey we're gonna spend you know hundreds of billions of dollars to build the biggest supercomputer in the world and train a model that can like you know do like you know 80 percent of all knowledge work you'd think I was crazy but somehow he's able to go up in front of hundreds of the smartest people in the field and be like, guys, this is possible.

29:09We can do it. And I think that extends like the whole like the whole leadership team as well. Like Elio is really, really good at this, too. He had his, you know, the feel the AGI stuff like. Yeah, I really admire Sam's ability to make impossible sounding things sound possible. And that's really how you get the best people is you work on really hard problems because all the best people want to work on hard problems. No one who's really, really smart wants to work on things that are easy. I think that's a, yeah, I mean, that is a skill. I won't make you answer for Sam's journey. In terms of your time at Opening Eye, you mentioned that you ended up spending it on post-training and this sort of had such an impact on what you're building today.

Read the full transcript

29:54For folks that I'm sure most of the audience or a lot of the audience will be really intimately acquainted with this, but for folks that maybe aren't, can you explain a little bit about why post-training matters and how it sort of led to this work you're doing with Applied at the moment? Totally, yeah. So kind of a high-level overview of training models in general is there's a couple phases of training. A lot of people talk about the big, expensive, high-capex phase, which is pre-training, which is let's go and take internet scale data, use this architecture called the transformer to sort of learn language patterns.

30:32And out of that falls intelligence, right? So you get this model that can do next token prediction. And that's kind of like how we got the first sort of GPT models where they just got really good at predicting the next token. And it felt like these models could actually reason and answer questions and things like that. After you do pre-training. You basically have a unaligned model that will just predict the next token until you tell it to stop. To make it useful, that's when you move to things in the post-training realm. So there's a couple of different types of post-training. When we were first training these chat models, where it's like a user sends a message and an assistant sends a message, there's a variety of techniques of supervised fine-tuning, as well as reinforcement learning with human feedback to sort of steer the models to give responses that are sort of coherent, you know, in the format and tone that a human would like and whatnot.

31:31But you're sort of just basically doing behavior cloning on the supervised fine-tuning side. So you're saying, here's a bunch of examples for good answer. If someone inputs this prompt, these are examples of what good completions look like. And then you just train the model to reduce the loss on those like target tokens. Now, what people have been talking about when they say post-training today is something quite different than supervised fine-tuning or reinforcement learning with human feedback. It's this idea of RLVR. So reinforcement learning with verifiable rewards. It turns out if you take a really small, like a relatively small amount of high quality data as compared to like when you were doing pre-training or SFD or something like that, and you actually have a model try to attempt the same problem a bunch of different times and you have a verifiable way of checking if the model gets the answer right, you can use this new reinforcement learning algorithm to sort of show the model, hey, these are examples of what good looks like.

32:36These are examples of what bad looks like. Yes. And sort of nudge the weights to do more of the good reasoning versus bad reasoning. This thing that people talk about when they talk about chain of thought, So all these tokens before the model outputs an answer is actually entirely an emergent property. Turns out like models learn to steer themselves, learn to think, and that's how they get more accurate and get, you know, basically like hill climb any sort of task that you give it. So what the first domain that this started with was math. part of it is because first of all it's like got a really a lot of really good properties which it's like easily verifiable you don't need that many complex tools or things like that to be able to do math and the second thing is just like everyone on the research team at open ai was like competitive math people or things like that so it was like a domain that they deeply understood and also were very excited about and there's also a lot of other physical philosophical reasons about like why math is kind of the root of a lot of other things.

33:41So basically, the first experiments were actually hill climbing on a variety of different math problems ranging from basic arithmetic to harder proofs, algebra, calculus proofs, whatnot. And it turns out that the model is able to learn on easier problems and internalize that and be able to unlock harder and harder problems. So that's this hill climbing that you've seen. And there was kind of like a couple of scaling laws that came out. So obviously you have the pre-training scaling laws. Now you have this like post-training scaling law. And now you have this test time compute scaling law, which is like if you actually allow the models to reason for longer when you're using them in production, they give you better answers.

34:26So that's this like reason when you pick a reasoning level in Chachipati or something like that. That's it. Just reasoning for longer, expending more test time compute before giving a final answer. This episode is brought to you by Persona, the B2B identity platform helping businesses verify users, fight fraud, and build trust. Fraudsters are already using AI to spoof faces, voices, and documents, so your defenses need to adapt just as fast. Persona helps secure some of the Internet's largest and most trusted platforms with identity verification. If you're building a product where trust matters, identity should be a priority.

35:01You've probably already experienced Persona without realizing it. verifying your LinkedIn profile, signing up for Etsy, or renting a scooter with Lime. Trusted by leading companies like Square, Brex, and Twilio, Persona gives you the building blocks to create identity flows that adapt to your customers, risk tolerance, and locales you operate in, whether you're verifying age, onboarding businesses, or automating KYC. It's fully configurable, so you can launch in days, not quarters. Want to see for yourself? Generalist listeners get a free year of the starter plan. Head to withpersona.com slash generalist and check it out.

35:39You're doing this work at OpenAI. What was the moment that you realized there's a need for the sort of service that applied compute offers where you are sort of going into these companies and doing this more hands-on work? So I think Microsoft AI, when they put out their thinking models recently, you know, the title of their paper is like the hill climbing machine or something like that. I think the way of thinking about RL is it really is that it's like if you define the hill to climb well the algorithm works so the way we were training these models was actually just like on the post-training side was hey let's go and make a bunch of high quality data sets that target the capabilities that we want the models to get really good at like basically we'd start with an eval and be like okay if this is the what the eval looks like let's just like go make training data that looks like the eval.

36:29Yes. And that was the way that we were approving these models. So the natural conclusion was kind of like, hey, you know, everyone is going to have their own evals and the things that they want these models to be really good at. And evals extend beyond just, you know, the intelligence of the model. It's like, oh, it needs to, you know, be this intelligent in this cost and latency regime or this modality or something like that. So if everyone's going to have their own evals, everyone should also be able to train their own models that do well on those evals right so so our sort of natural conclusion was like hey maybe since we have this hill climbing machine the best way to sort of manifest it in the world obviously the frontier models are amazing you know these will be like workhorse models for a lot of things yes but if people want to be very specific about the task that they're trying to optimize for they should have the ability to go and push drain their own models on their own data.

37:26Yeah. Evals are, you know, the scorecard for all of these models. And that's so general. It's the new PRD. Yeah. Yeah, exactly. And the reality is that every organization has a different scorecard for what they consider valuable. And that's, you know, sort of what you're saying is like, you know, if you have a different set of things that you're optimizing for, it starts to make sense that you have a machine that is made to hit those optimizations. Exactly. And if we go back to the Goldman example, right? Like if a lot of those evals are going to say private, people are going to want to own their intelligence, own the post-training capability to go and create those models that are good at those evals.

38:02Right. Yeah. So I think that was the primary motivator. And then I think the main sort of driver for why people are thinking about this right now is cost. The token apocalypse has really shown, hey, you need to use the right model for the right task, even beyond just like optimizing for evals. rate. Yeah, that's I'm curious, like how you think about it today. Clearly, that's top of mind. I think in the past you've talked about how general models sort of set the floor and, you know, specialized models can sort of raise the ceiling in some sense. Yeah. In some ways, it feels like, you know, the distance between those two might be getting smaller, but the cost at which you are accessing these different offerings is like, has it become more of a cost story in your head?

38:46cost has become like a a the main driver i think right now but i think in principle the argument still stands which like take kirkland and alice right they caused waves in the twitter community and i'm sure broadly as well about how they're going to spend like half a billion dollars to i forgot what the exact quote was but the essentially the idea was to create ai tools that their competitors can't also use, right? So I think the idea is, and this extends beyond just the model, to be clear, right? The model is part of the moat, but also training that model inside of your harness with your context, that's actually how you're going to get differentiation.

39:29So I think a lot of people are thinking about, hey, if I'm using the same stuff as my competitors, how am I different? And that's also another driver for why they should train their own models on their own data. And I think it's been just almost a little over a year, right? And you've gone from starting the company to a last valuation, 1.3 billion. Been a good year. Been a good year, yeah. Can you tell me a little bit of how you've started to work with folks and what that process looks like with these companies? Totally, yeah. So we work with a variety of companies across AI natives, digital natives, and large enterprises, and hyperscalers.

40:08So basically, if you look at how money is spent on these models right there's some amount spent on like training and it's like a non-trivial amount but like most of the money and most of the usage is on on inference right so when we were thinking about what the future of open models looks like or models that people are going to inference a lot for different things our conclusion was kind of if you are going to train or if you are going to use open weight models and specialize that model you're going to do some sort of training on top of it. So the actual wedge into inference is getting really good at training.

40:47So what we do with our customers is we have a post-training platform where they can integrate all of their data. And we have a team of applied researchers that actually deeply embed with the customer, work with their ML and research teams to train these custom models. And we really want to make sure that the economics are in their favor. So we try to have most of the stuff that they pay for come on the inference side, right? So it's like, once we have a really good model and you guys want to use it, then we can go and host it for you. Yes. The sort of forward deployed engineer model is something that I think people obviously associate with Palantir as a successful example of it.

41:29With Palantir, I think the way they've made it work is that clearly they found a way to get good leverage on it such that the 50th version of it of them taking on a project is very different than the fifth. What's the version of that for you that stops it from being sort of just pure consulting leverage? There is always some component of this that is going to be unique per company. That's kind of the whole thesis, right? It's like each company is going to have different data and want a different model, But really, it's the infrastructure that is repeatable across all of these companies, right?

42:05It's like the scaled infrastructure for training these models, the scaled infrastructure for serving them, and then the sort of techniques for capturing production usage, turning that into more data that we can go and continually train. So our equivalent of Foundry, which is Palantir's platform that they use to give to all their FDs to go and move faster. Yes. We have our own post-training platform where all of our applied researchers and researchers from the companies we work with can both collaborate to train different models. And really, I think the way of thinking about it is it's the infrastructure to make training performant and easy.

42:45Mm hmm. Yeah. Does it end up becoming more of an ongoing engagement than, you know, people might realize from the outside? I could imagine that, you know, one version of it would people would think, oh, you go in, you're sort of doing an AI transformation and then, you know, you're out the door. But I would imagine like the demand for, you know, improving this, deepening it is like leads to quite long engagements. Totally. Yeah. I mean, so I think there's always a desire to continually improve these things, which is why, you know, when we deploy a model in production, a lot of our product is built around, hey, how do we take the usage of this model and turn it into more data that we can continue to continue to train on or build context around the model so that, you know, it's learning actually the more you use it.

43:34So, yeah, I think the idea is that the engagements don't end. Obviously, like the way we generate revenues by serving the models and we need to make sure that those models are always performant. And maybe even to get even more sort of tangible, like can you maybe talk us through an example or two of like what taking this more specialized contextual approach unlocks for a customer, for instance? Yeah. So I think one of the case studies we have on our website is this engagement we did with DoorDash, which is, you know, they have more than 100 ,000 merchants that come to their platform every year and they want to create a new DoorDash storefront.

44:16Right. So when they when they do this, they come with a bunch of unstructured information about their business, you know, what what food they serve, things like that. And one of the things is like a picture of their their menu. So it turns out this was actually an example of where we were able to exceed the frontier in terms of quality because we trained a small specialized model that was able to take pictures of menus and turn them into high-fidelity DoorDash storefronts. Basically, the process was create an eval for what good looks like, create the training data for us to go and do RL training on top of, train the model, and then deploy that and host that.

44:58The trend feels like, you know, if applied compute continues to grow and this is part of sort of a broader movement is like sort of trending towards more of these small models where you do sort of say, here's a set of tasks that like are maybe a bit more narrow, a bit more deterministic that like we can actually very inexpensively do this very effectively. Is that roughly how you see things going? Yeah. And so I think the motto is better, cheaper, faster. what people are starting to realize is that these open weight models provide a ton of flexibility around ownership and they're actually good.

45:37Like you can actually go and use them to perform optimally on your task. So I think you're going to see just a lot of people for, you know, reasons related to cost. I think cost is the big driver right now, but there's just so many other downstream benefits you get by owning your own model. When you sort of move towards the world of you know, a company having multiple of its own models for different tasks, does that create like some sort of maintenance debt for them that they have to sort of stay on top of? And yeah, I mean, there's probably still cases where that's very much worth it for them.

46:13But is that something you've noticed come up? Yeah, I mean, so this is part of like the continuous improvement part is there's obviously like model drift and you know, the distribution of tasks might be changing slightly, but over a long period of time, that really, really matters. Our view is just like training and inference loops. They're like this right now, but they're starting to move closer and closer together. Right now, the form is like people are just getting a lot faster at making new data and training the model again. But eventually, these things are going to be quite intertwined. I think there are some cool examples of this that are kind of like maybe the first sort of beginnings of it.

46:55Like if you saw Cursors, Composers, Real Time RL work, I mean, that was them doing like online training by like taking a bunch of user trajectories, you know, getting, extracting the implicit reward, taking a training step and then deploying a new model in production. I think like that is just kind of an indication of where things are going to go. It's like you're actually going to be doing more and more continuous training because as your model is being used. It's kind of like an exploration technique, right? Yeah. And so what you're describing there is almost like a continuous learning model where it's like every time I happen to do this task, it's actually training this model to be better the next time around.

47:38Totally. And the name of the game is learning from sparse rewards, right? So if you think about the different phases of training, pre-training is like super data inefficient, right? You need to like learn the whole, or you need to train on like the whole internet in order to like learn the language. Then we had this SFT stuff, which was kind of behavior cloning. So you say, hey, here. I'm not sure about that. Yeah, yeah. So SFT is described as behavior cloning because you're giving it examples of what the correct output should look like. And so you give it a bunch of examples. And then for a new task or a new prompt, the model is sort of interpolating, like, oh, I should say these things.

48:16And then we kind of move to RL, which is a lot more data efficient, which is I'm going going to take one data point, one math problem, generate a thousand trajectories, all with different thinking tokens or whatnot, all with different rewards. And then I'm going to do backprop on those. And there the model will kind of learn what good reasoning looks like and bad reasoning looks like. The ultimate thing, right, is being able to do super, super sparse learning, which is like you go and roll out a trajectory, you have some indication of whether it was good or bad or something, and you're able to directly learn from that.

48:53That is difficult, and that is the problem to solve in continual learning. On the open weights models, you were mentioning earlier that you think maybe the gap between open weight and closed models has sort of narrowed over the past year. In some ways, if I look at the hardest benchmarks, I almost feel like it's widened, but actually what we're learning is that just beyond a certain threshold, it almost doesn't matter. Yeah. Like, I don't know. Do you think that's wrong? Yeah, yeah. So I think what we were talking about before is Frontier models are amazing and you should still use them for your hardest tasks.

49:29But if you have like$100 of spend, maybe 20 of those dollars go to the Frontier models and the$80 can be actually$20 that go to better, cheaper, faster models that you train on this task, right? So I think like like the frontier models are obviously very, very, very good. But, you know, you don't need to use them for everything. There's all these memes on Twitter, right? Like where like people are using some big sword to cut like a blowtorch. Exactly. So I think like, you know, there's a lot of truth to that. And that's what you know, you're talking about model routing and things like that. I think there's going to be pretty interesting innovations there.

50:07We're working on some of that as well. The other thing I think about with open wage models is like China has really dominated that so far. Yeah. You know, it feels like America is is clearly leading with some of the frontier models. But yeah, it's like how do you think about the fact that an implication of this could be that a lot of American companies become like actually quite dependent on Chinese models that the rules could get changed up on them? of those could go closed at some point. Is there a robust enough open weight system right now? So I think that you're right. It all boils down to incentives.

50:46So I'm extremely excited that companies like NVIDIA are investing a ton in open source models. And NVIDIA has a real reason to do it, right? The dream for NVIDIA is that every company has their own model and is running it on Nvidia chips. So I think that there are real, there are companies like Nvidia, Nvidia being the big one, that have true incentives to keep models open. And I think there is enough of investment and a lot of smart people working on open models that I think it's going to be a pretty safe and big bet over a period of time. One of the things that I think is really interesting about applied compute sort of outside of the technology piece is that you seem to have managed to attract a huge amount of ex-founders, which is always a very bullish sign, I think, for an investor, you know, to see that happen.

51:43How have you managed to do that? It's something like two thirds, right? Yeah. So it's almost part of like the business model is ever because we're doing this like sort of forward deployed approach where we're embedding deeply with with customers. it requires people who are extremely high agency and like ambitious and move fast. So I think it just naturally like the kind of work just naturally attracts these these types of folks. And then, yeah, I think I think like what we we in our culture, we sort of decided, hey, we want to hire ex-founders who are who fit this archetype. And they just themselves attract other like minded folks.

52:26so yeah I think it hasn't been anything really special like we're not trying like recruiting from some very specific pipeline or something like that but it's just this kind of natural snowball effect where you know the type of work that we do plus the initial team that we got together just attracts more people of the same mentality at a higher level like what are the things that you think are maybe particularly distinctive about applied computes culture Sure. Like, what do you care about that, you know, other other startups, you know, or other places you've worked maybe wouldn't care about as much?

52:59So there's a lot of this like nine, nine, six stuff. Yes. We're not a shop like that. Like we and I deeply respect people who who choose to operate like that. I think the type of culture that we want to embody is people are just very excited about their work. and we have no requirements on how long people work or things like that, but it's a very bullish sign when people are so excited about what they're doing that they want to come in on weekends and sort of spend a bunch of time in the office. So yeah, I think the things that we prioritize or what we look for when people are joining the team is like being deeply mission aligned.

53:40So, you know, we want people who want to come and work on hard problems, but also really believe that, you know, something like this should exist in the world and that this is actually very important for basically every company out there. And then take a lot of pride in working on hard problems that don't have like the thing that's very unique about our company. Right. And it's actually somewhat related to a Palantir. I remember having this unique moment where I was like, Palantir is one of the few companies that signs massive contracts without knowing if they can solve the problem. Right. So interesting.

54:17That's true. Right. Like, you know, they'll sign like huge, you know, a hundred million dollar deals where this is like before they've started working out the problem. Right. So. So I think there are there there's like a unique set of companies that that do that. And and obviously, like, I think we fit into that category where there's open research questions about like what we're trying to go and get really good at and solve. And people who like to work on things that are uncertain and there's a lot of greenfield research to be done are just naturally attracted to companies like AC. When you talk about, you know, wanting people who want to do the hard thing, like what's been the hardest, thorniest technical piece of this so far for you?

54:56Yeah. So I think like one of the really cool places of work in the company right now is basically building the best ML infrastructure for training open width models. So, you know, We're obviously in this compute crunch. The more performantly you can train models and the more accurate they can be, that's true differentiation. There's actual work. It's not like RL is completely figured out and it's kind of like a commodity thing. There's real research happening at the ML infralayer. So we've been hiring a bunch of these really, really smart researchers that are actually able to go deep into the stack, work very close to the hardware to actually make training more performant.

55:39And that's one of the ways we're differentiating amongst like other things that are out there. Can you, you know, maybe part of this is too secret sauce, but like, can you talk a little bit about what that training stack is, like what that looks like? Yeah. So, so basically the, you know, the, we, we procure our own GPUs and then we, when you're actually going and doing training, right, the way training works is you're alternating between like doing inference. So, So you have some checkpoint, some weight checkpoint that's like policy step zero or something. You go and try a bunch of attempts at a problem, right?

56:13And that problem is actually in some sort of replayable environment, right? So say we're training a coding model. You're like spinning up a sandbox that has a code base on it and potentially even more complex than that. Like you have, you know, services that are running that the model need to interact with, yada, yada, yada. All needs to be replayable though, right? because you're repeating this like a thousand times. Yes. You go and do the inference on that, get a bunch of model attempts, assign reward to them based off of like some verifier and then use some sort of objective to be able to go and take the next step on the model and get your policy step one.

56:50And so one of the really interesting problems is like, okay, you're doing this sort of alternating of sample, train, sample, train, sample, train, sample, train. Yes. it's like always slower to do things in sequentially right so what do you actually want to do is you want to do things asynchronously where you can actually be sampling on your next set of problems while training is happening oh wow and you can actually there's like really cool science questions around this of like how far can you push that how asynchronous can you be how off policy can you be can i be sampling problems using policy you know step one or something or sorry policy some some earlier policy while i'm training like the last policy so this is where a bunch of the ml dynamics come in or wait how do you make these off policy corrections where you're trading off performance and speed and throughput versus like ml dynamics and and stuff like that so these are the kinds of like infra problems that a lot of our team is working on it's like how do we make this stuff really performant and and it really matters for longer and longer horizon tasks, right?

57:57Like, you know, if you're running your model for like an hour, your total training runs can be super, super long if you're doing things synchronously, right? Train inference, train inference. You really do want to do things super asynchronous in order to like cut down on your total GPU time. You're sort of optimizing around your layering GPU time because that's like the most constrained. And so this is the other cool thing, right? It's like, so there's that side of the house, but then there's also environment creation, right? Why do people care so much about performance sandboxes and things like that?

58:29The reason is because if you're waiting on CPU workloads, if you're idling your GPUs for CPU workloads, you're probably wasting a lot of money. So there's like a ton of performance engineering on the environment side of the house to make training faster as well. You know, as you were talking about the sort of imperative for every company to own their own intelligence. It sort of occurred to me, maybe this is a very naive thought, like, isn't the sort of logical extension of that, that at some point it'll also trickle down to the human level and all of us will need to have our version of that. Like this, you know, does AC in five years, you know, start to become, you don't need a forward deploy engineer, but you have a version of it that's like more self-serve and, you know, each of us has our own model, our own set of models.

59:17Yeah, I mean, so far we've been pretty enterprise facing, but maybe this is market number two. Excellent. No, I think you're right. I think a lot of people have been very excited about this idea of personal AI. And I think maybe a bit of a jump to a broader point, but if you think about personal AI, that's something that deeply understands you, how you work, and maybe running all the time. right so i think um maybe a broader point is just like as these things start to trickle down the demand for compute and inference is gonna like balloon like crazy you know that that that's something that's like kind of like a broader maybe i actually don't even know if the market is fully pricing in how much compute that we'll need yeah all of this stuff like if everyone has their own personal models and that are constantly being trained and inference and things like that like we just need a ton of compute which i don't know it seems likely that will happen to me maybe that's yeah maybe that's too maybe i'm no no i think i think people totally want it to happen the problem is i mean i think they're just like right they're like physical limits right like uh to how you know how much the supply chain can accommodate and like not even like on the chip like you know there There are ASML machines, like lithography machines that take a really long time to build.

1:00:37So that's why I think we're in for... This kind of ties back to tokenpocalypse. We're going to be in a compute crunch and inference demand is going to crank like crazy. The ability to train Pareto optimal cost performance models is actually going to matter a lot. And I think we're seeing the early beginnings of that. And it's just going to get worse until people really start to invest in this. what's your sort of uh you know not ai 2027 but your view on the next couple years like what do you think yeah the world is looking like from your perch you know i know there's a lot of folks who are like oh 50 of jobs are going to go away in the next two years and things like i think like personally i i think that is is not going to happen um and it it probably is inspiring a lot more fear that maybe doesn't is maybe misplaced yeah i and i and i think it's coming from a good place, which is just like we need to be prepared for this sort of stuff.

1:01:34But kind of being on the front lines, like doing this for deployed work and whatnot, I think just the diffusion of AI into the real world take time. There's a lot of change management that needs to happen. Companies are like large institutions that take time to adopt new things. And we were talking about this before, right? Like OpenAI is this, you know, it's OpenAI, the research and deployment company. And, you know, deployment is so hard that they created a company called Deployco, the OpenAI Deployco, right? And same with Anthropic and stuff like that. So I think like people know that it is going to take some time to go and deploy this stuff.

1:02:16So I don't think overnight everyone's going to lose their jobs. There will be like areas of the economy that I think will experience more change than others just due to the nature of the job. But to me, this is like a multi-decade long rollout and diffusion of AI into the real world. Wow. What has surprised you most in terms of that change management, maybe the slowness of some of these enterprises? What has maybe you couldn't have counted on? I mean, in some ways we knew this would happen, but just seeing it up close makes a lot of sense. There are just established procedures inside of companies, and one of the difficult things is like you know to change those procedures you often have to touch multiple layers of management and and the actual people who are doing the work and it is just a difficult coordination problem right and and i think the other thing that is also obvious but we've just you know michael chan on our on our go-to-market team put out a really good like field notes guide to what forward deployed engineering has been like is like no one is data ready.

1:03:19Absolutely no one is data ready. Things are like, obviously, always very messy and whatnot. But if data is kind of everything, even just the state of where it is, and even the best companies is it's not where it needs to be to actually go and make any sort of deployment extremely smooth right now. So I think there are just these natural sort of roadblocks that will slow things down when you're trying to roll out AI and like big organizations. Well, I always like to end with a couple sort of thought experiment questions. If you were given the chance to conduct any experiment with no operational constraints and no budget constraints, what's an experiment you'd like to run?

1:04:01Oh, this is a fun one we talked about in the office, which was if we all reverted to cavemen and like lost all modern technology, how long would it take us to reproduce the iPhone? So knowing what we know now, how long would it take us to get all the infrastructure in place? Like no fabs, none of this stuff. Man, that's really - Just sand. I used to think about that for like a day at least. Yeah. I don't know. What did you come up with? Well, so, and I guess the darker side of this is you have to factor in everything. So like people going to war against each other and things like that. Yeah, yeah, true.

1:04:46If we were just lost all modern technology. 100%. I conclude that we just never would. You think we would never have anything? Yeah. Like, first of all, most of the population, this is getting a little weird, but I think most of the population would not just like know how to survive the first like 60 to 90 days. And then, you know, we'd have to, there's the ability to go and build all this infrastructure takes a ton of time. If you sort of are thrown into chaos and dystopia, people are obviously going to get mad at each other and create factions and fight. So I think we'd spend most of our time fighting them than actually building a thing.

1:05:23Wow, I'm going to be thinking about that for a while. That's a good one. Okay, sort of last question then, I suppose. If you could assign a book to everyone on Earth and you knew they would read it and they'd understand it, what would you want to give to people? I really like this book called The Design of Everyday Things. It's actually one of the first tech books that my brother gave me. And I think it really set a good framework for how to think about solving problems in non-complex ways. And here's like a crazy segue for you, right? I think in today's age, when you have AI tools, particularly coding, you can actually, by going faster, you can actually be going a lot slower.

1:06:07So like taking the time to, like the one thing I'm worried about people losing is this ability to actually deeply think about how things should be architected and designed. Yes. because if we rely too much on AI tools, and I think every company has seen this in some capacity with this AI vibe coding or whatnot. We're all for AI coding. My main thing that I worked on at OpenAI was codecs, but I think there is truly some atrophying happening by relying too much on AI tools. So back to the basics, back to the roots. Yeah, I mean, things like judgment, taste, all of that stuff is very easy to stop actually building when you have these tools in front of you, right?

1:06:49Totally. I think the best engineers right now are the ones who are people who learned how to code before AI because they actually deeply understand systems thinking and things like that. And that would be my worry as a new grad engineer or something like that is getting too reliant on AI tools. Because at the end of the day, I think humans will be the overlords over all of these AI tools and you will still need to do, you know, critical thinking to know what you want to build. They might help you build it, but you need to know what you want to build. Well, I think that's a that's a nice optimistic place to end.

1:07:26Thank you so much. I really enjoyed it. Oh, thank you. This was awesome. Yeah. Thanks for having me. That's it. Thank you for listening to this episode of The Generalist Podcast. Please subscribe on Apple Podcasts, Spotify, or your preferred podcast app. Ratings and reviews help others discover these discussions, so if you enjoyed the conversation, I'd be grateful if you could take a moment to leave one. For all past episodes and more, visit us at thegeneralist.substack.com. See you next time as we continue to explore the future.

From the publisher

Yash Patil is the 23-year-old founder and CEO of Applied Compute, a $1.3 billion company helping businesses train custom AI models on their own data: smaller, cheaper, and purpose-built for the work they actually do. Before founding the company, Yash dropped out of Stanford and spent two years at OpenAI working on post-training infrastructure and Codex. He left with one core conviction: every company that runs its critical workflows on someone else’s model is building on shifting sand. Applied Compute is his answer to that problem, already serving customers including DoorDash, Cognition, and Mercor.


In our conversation, we explore:

  • Why “own or be owned” is becoming existential for any company that relies on frontier AI models
  • What it was like inside OpenAI the weekend the board fired, and then reinstated, its CEO
  • Why post-training is where competitive advantage is now being built, and what reinforcement learning with verifiable rewards actually is
  • Why evals have become the new production environment, and why companies will never share them with frontier providers
  • How a specialized model built for DoorDash outperformed frontier models on a narrow, high-value task
  • Why cost, not capability, is now the primary driver pushing companies toward custom models
  • Why Yash believes AI’s transformation of the economy will unfold over decades, and why near-term fears about mass job displacement are misplaced

—

Thank you to the partners who make this possible

Brex: The intelligent finance platform.

Guru: The AI source of truth for work.

Persona: Trusted identity verification for any use case.

—

Transcript: https://www.generalist.com/p/own-or-be-owned-why-every-company

—

Timestamps

(00:00) Introduction

(03:50) Fable 5 and the case for owning your own models

(09:22) Why Applied Compute is betting on custom AI models

(12:30) Yash's early influences and first projects

(17:42) His brief time building at Stanford

(19:29) Leaving Stanford for OpenAI

(25:58) Inside OpenAI during Sam Altman's firing

(28:18) What Yash admires about Sam Altman

(29:43) Teaching models to reason

(35:39) The core insight behind Applied Compute

(39:40) How Applied Compute works with its customers

(45:55) Why model training never ends

(48:56) Why not every task needs a frontier model

(51:25) The culture and people of Applied Compute

(54:50) Applied Compute's training infrastructure

(58:43) The coming compute crunch and other predictions

(1:03:48) Final meditations

—

Follow Yash Patil

X: https://x.com/ypatil125

Website: https://yashpatil.me

LinkedIn: https://www.linkedin.com/in/yash-s-patil

—

Resources and episode mentions: https://www.generalist.com/p/own-or-be-owned-why-every-company⁠

—

Production and marketing by penname.co. For inquiries about sponsoring the podcast, email jordan@penname.co.

More from The Generalist

All 50 episodes
Own or Be Owned: Why Every Company Needs Its Own AI Model (Yash Patil, Co-Founder & CEO of Applied Compute)The Generalist · 1 h 8 min
Listen in VO