Why AI Models Aren’t the Product Any More | TWiAI Ep 18

18 Jun 2026 · 1 h 21 min · 27 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI models aren’t the product anymore; the “product” is the agentic application layer (agents, evals, harness, UI) plus proprietary learning loops. The episode also covers SpaceX’s planned $60B stock purchase of Cursor and what that signals for compute, monetization, and model ownership.

Guests and backgrounds

  • Ali Ansari, founder/CEO of Micro One; recruits and manages human experts to help train/fine-tune frontier models with compliance/HR handled by Micro One. Previously built an AI recruiter/screener and a pre-vetted engineer marketplace; later pivoted into the “data space.” Claims Micro One has ~300M ARR (as of April 2026).
  • Ryan Daniels, founder/CEO of Crosby Legal; runs an AI-first law firm that combines AI systems with human legal expertise; charges pricing by the deal (flat rate), not billable hours.

Key claims

  • “Model is no longer the product”; agents + evaluations + application layer are where value accrues.
  • Enterprises will spend most data/effort on agents/application layer (possibly ~100% of data spend).
  • Proprietary evals/feedback loops are required for subjective domains like law.
  • Application companies should avoid giving proprietary knowledge to OpenAI/Anthropic/Gemini; own the intelligence layer.

Notable examples

  • SpaceX buying Cursor (Cursor previously 40–50% of Anthropic revenue; Cursor coding tool launched 2022; $1B+ annualized revenue cited).
  • Crosby: summarizing client review letters with trustable outputs; lawyers become more valuable as models improve.
  • Satya Nadella’s “cognitive loop” framing: learning loops compound human + token capital.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

SpaceX Acquires Cursor

0:00 to 0:46

Learn about SpaceX's plans to acquire Cursor for $60 billion and its implications.

“SpaceX formally announced plans to buy Cursor for$60 billion in stock.”

Introducing Guests and AI Law Firm

3:20 to 4:40

Meet Ryan Daniels and learn about Crosby Legal, an AI-first law firm.

“So our second guest, Ryan Daniels of Crosby Legal.”

Innovations in Legal Services

4:40 to 7:20

Explore how Crosby Legal provides legal services using AI technology.

“They're ending the billable hour forever.”

Ali Ansari's Insights on SpaceX and Cursor

7:20 to 10:30

Ali discusses the implications of SpaceX's acquisition of Cursor.

“They always keep that periphery vision open.”

Challenges in AI Model Development

10:30 to 13:40

Delve into the difficulties SpaceX faces in developing AI models and how Cursor fits into this.

“mainly from the product usage that they have from their customers actually using Cursor, which is one of the most important kind of data flows that you need to actually build a model.”

The Impact of Anthropic on Cursor

13:40 to 14:01

Examine how Cursor's relationship with Anthropic influenced its development and competition.

“There was supposedly like a profile of Cursor CEO and co-founder Michael Truwell, but it sort of delves into their history with Anthropic.”

The Dynamics of AI Model Development

14:01 to 20:08

Learn about the competitive landscape and key insights in AI model development.

“Anthropik executives apparently specifically told Cursor that Cloud Code was always going to be more of a research effort than a major commercial push.”

Understanding AI Model Distillation

20:08 to 22:51

Explore the concept of model distillation and its implications for AI development.

“And if compute moves to space, another advantage for the Cursor team.”

The Future of AI Models and Applications

22:51 to 28:07

Discuss the future landscape of AI models and how companies will adapt.

“I mean, one thing I'd add is I think a researcher recently outlined his vision of the future where there are just two types of companies that matter.”

The Value of Intelligence Layers

28:07 to 29:29

Explore how companies can differentiate themselves by owning their intelligence layer while utilizing closed-source models.

“But you can still own your own intelligence layer while building on top of the closed source models.”
Show all 27 chapters

The Role of Models vs. Products

29:30 to 31:13

Discuss the importance of harnessing models and building products that generate revenue rather than focusing solely on the models themselves.

“You have this very friendly collaboration with the labs where we do a lot of applied evals with both of the open an anthropic, and yet we never know.”

Cognitive Loops and Learning Systems

31:21 to 32:53

Learn about the concept of cognitive loops in AI and how they contribute to enterprise value by combining human and token capital.

“wrote a widely shared X article over the weekend titled A Frontier Without an Ecosystem.”

The Evolution of Legal Services with AI

32:54 to 34:18

Understand how AI is changing the landscape of legal services and enhancing the value that lawyers bring as models improve.

“And I think the hard thing is that feedback loop, the reinforcement learning loop, which is the key, again, like going back to what I mentioned with Ali, you need to be able to eval an output to say if it's good or bad.”

Customer-Centric Solutions in AI

34:19 to 36:00

Examine how companies should focus on solving customer problems effectively rather than fitting into traditional VC paradigms.

“a few months ahead of most other companies.”

Frontier Models and Customer Interaction

36:01 to 37:58

Explore the potential of frontier models in enhancing customer interactions and how this impacts the future of service delivery.

“like Ryan are just going to say, you know what?”

AI in Summarizing Legal Documents

37:59 to 40:04

Learn about the challenges and solutions for using AI to summarize legal documents effectively for different audiences.

“last bit done, not the first draft, but the final ready to file.”

Scaling Models with Real World Data

40:05 to 42:09

Delve into how services companies can enhance AI models by using real-world data and the implications for service delivery.

“This is where Ali, the person who is closest to the customer gets to win.”

The Role of Humans in AI-Driven Services

42:09 to 45:30

Explore how human expertise is essential in AI applications, especially in niche service companies.

“Now, if you think about, again, a services company, there's an abundance of those tasks because it is what you're doing and what you're serving to your customers.”

Challenges of AI in Court Reporting

45:31 to 49:56

Discuss the limitations of AI in capturing the nuances of court reporting and legal transcription.

“There was an interesting piece of the Wall Street Journal over the weekend about how court reporting seems to be a prime target for AI takeover.”

The Future of AI in Legal Systems

49:56 to 56:01

Analyze the potential of AI in legal settings and the implications for access to justice.

“When will a person be able to be their own counsel for, let's call it, not small claims court, but whatever the next level up is?”

Human Involvement in Model Training

56:01 to 57:54

Discusses the importance of human experts in maintaining AI model accuracy.

“It's like a very simple like user interface changes on the proctoring side of things make the model drift in its capabilities.”

Data Collection for AI Training

57:55 to 1:00:06

Explores the concept of collecting operational data from various industries for training AI models.

“And of course, it's not us training models, it's for our customers.”

The Value of Failed Companies' Data

1:00:07 to 1:02:52

Evaluates the potential usefulness of data from failed startups for AI training despite the failure.

“And then the third component, which is really important, is what is the seeded data that it has?”

The Dark Data Pool of Slack

1:02:53 to 1:05:20

Discusses the implications of using Slack data for AI training and the challenges involved.

“They do a better job examining Slack with AI than Slack does with their tools.”

Building a Legal Benchmark for AI

1:05:21 to 1:07:51

Introduces a new benchmark for negotiating SaaS contracts using AI, focusing on multi-turn redlining.

“doing this, they're improving their agents and they're effectively training a model on customer data.”

Collaboration in Legal AI

1:07:52 to 1:10:00

Explains how collaboration among lawyers aids in improving AI's understanding of legal negotiations.

“How do the two of you work together and collaborate?”

Episode Discussion

1:10:00 to 1:21:16
“Like basically the lawyers had a lot of consensus on the first review.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00SpaceX formally announced plans to buy Cursor for$60 billion in stock. Cursor once accounted for 40 to 50 % of Anthropics total revenue. It is Game of Thrones in terms of talent, in terms of territory, in terms of weapons in AI. The model is no longer the product. The agent evaluations that sort of come on top, the harness, the user interface and so forth. You don't want to give your knowledge to Claude. You don't want to give it to OpenAI or Gemini. You want that for yourself. The majority of the data spend, in fact, probably close to 100 % of it, will be on the agent and the application layer versus the model companies.

0:36There's just going to be orders of magnitude more agents built than obviously models. Cursor will have the number one model at this time next year. Thanks to our friends at PayPal, the exclusive sponsor for This Week in AI. Try the payment and growth platform that's trusted by millions of customers worldwide.

0:57All right, everybody. Welcome back to the greatest AI show in the universe. Yes, I've launched another roundtable show. This one is called This Week in AI. I'm running it just like I do my venture roundtable on This Week in Startups or how we do all in, which is to say we have a group of experts and we talk about the news of the week and we'll even tip you off into where the future is going to be. So if you listen to This Week in AI, you're going to be six months ahead of everybody else. Why? Because we pick people for the show who are in the arena, who have deep expertise. But in this four-quadrant chart, they also have to be willing to share it candidly.

1:37So in the top right-hand corner, expertise, building the future, but they have to be willing to talk about it. If they're experts and they're quiet, nope, can't come on the show. So that's what we're optimizing for. This is episode 18. You can get all the information you need about this podcast at thisweekina.ai. I'm going to have our editorial director here at this weekend, Lon Harris, moderate this week, because I want to shoot the ball along with the number one, the number one CEO, just in terms of performance, in our fourth fund, Ali Ansari. He is my favorite, the golden child, or as I call him internally, Travis 2.0.

2:18You know Travis Kalanick, Lon? Of course. You know how he plows through brick walls? He's super pumped. He's super pumped, that guy. He is super pumped. And you know what he did? A lot of bricks went flying. And you know what my job was? Had to catch some of the bricks sometimes or explain what we were doing at Uber, smooth it out a little bit. Yeah, we're watching the troll from the Battle of Pelennor Fields doing the same. Okay, great. Great nerd reference or juggernaut. So I asked Ali, one of the things I like to do, This is a little producing trick. I like to ask the great guests on the program, hey, who else is like really crushing it?

2:57Ollie came back with an incredible guest. And now we have another person crushing it in AI. I think the app layer in AI is going to be a huge win. And Ollie's in partnership with us. So Lon, you have the reins. I am trusting you. With the comms. What is it on the Star Trek? The comms. You have the comm. You have the comm. You're in charge. Get started. All right. Well, yeah. So our second guest, Ryan Daniels of Crosby Legal. He's the founder, CEO. They are, Jason, an AI-first law firm combining AI systems with human expertise to provide a scalable legal service. So rather than selling, you know, like an AI model that's good at legal questions to law firms, Crosby built their own in-house team.

3:40And then these lawyers developed and used the AI product in-house. So you're contracting an AI-first law firm when you go to Crosby, which is an interesting - So Ryan, if I'm doing my Series A, you're saying instead of using traditional Silicon Valley law firm, I could use yours? We'll do it faster. We'll be more reliable. We'll do all your sponsorship agreements. Well, those are easy. Sponsorship agreements, we've got that already dialed in with AI. But in all seriousness, closing a Series A is, I think,$50 ,000 in Silicon Valley now. closing a seed round, maybe 10K if you're using notes. Is that about right in terms of what the cost structure is?

4:20Yeah, that's probably right. And what would you provide those similar services for? I'm curious. So we're getting to financings. We started mostly with commercial agreements, the most recurring thing that our clients do. We can build the best data models around. Our main thing is pricing by the deal, not by the hour. Oh, flat rate? Flat rate, yeah. And so our incentives are really aligned. Yeah. That is my dream. And then the billable hour. Exactly. What did you say, Ali? They're ending the billable hour forever. Well, Ali and I know about the shocking legal bills that we get in sometimes. And oh, my Lord, the sticker shock sometimes.

4:55And if you ask a lawyer, hey, can I get a flat rate? Nine times out of 10, they say we're not the right law firm for you. So, Ryan, there is a clear lane for you here if you can solve this specific problem. Hey, Lon, give a proper introduction to Ali. I gave my seed investor, this guy's my guy. I'm Professor X in this model, right? I run the school for gifted mutants. Right. He's like my Wolverine. So, but explain what he does. Sure. He's the founder and CEO of Micro One. They recruit and manage human experts who help train AI. So think of it as sort of an expertise marketplace, AI labs that have models that they need to fine tune in sort of high level, you know, tasks, the things that you still need a human in the loop to help them figure out, and they team them with pre-vetted domain experts like a full-service operator handling the modeling, compliance payments, basically all the HR stuff you don't want to worry about.

5:52And then you have senior engineers, PhDs, and others providing this critical bottleneck for improving frontier models, getting some human expertise in there. Most recent numbers we have, about 300 million in ARR as of April 2026. Whoa, are we allowed to say that, Ali? Are we allowed to say that? That's a bit of an old number. Oh, there you go. Well, leave it at what it is. I'll tell you, it's the fastest revenue ramp in like two years that I've seen since, I guess, I wonder if it's faster than Uber's first two years. It might be, in fact. It might be. So incredible revenue ramp. And Ali, that was not the business I invested in when we met.

6:34The first business, if I remember correctly, three years ago was you had built an incredible AI to sort through developers and give them tests. Yeah, am I correct? Yeah, exactly. The V1 and Micro One was we built an AI screener. And this was actually for, it was an internal tool that I built for one of my previous companies. And then I decided, okay, this is probably something I could spin out and make a product on its own. So the first version was an AI recruiter, which companies use to essentially source some talent. talent. And then they also, we ended up making this kind of like next generation marketplace of pre-veted engineers specifically, which startups would hire from.

7:11So that was V1. But then we realized very quickly that there's this data space and we sort of pivoted entirely into it. Which, Lon, is what the great entrepreneurs do. They always keep that periphery vision open. So if you ever hear my talk about peripheral vision, when you're a great entrepreneur, You start with a vision, you start building. And then sometimes out of the corner of your eye, you see, a second, is that a diamond mine? I see something glistening. Is that a perfect wave to surf? And I'll give you a great example. If you look at Uber once again, or Airbnb, which we weren't investors in, Airbnb was renting a room in somebody's apartment.

7:50And then they saw, hey, wait a second, what if somebody just rented the whole damn apartment? Incredible micro pivot changed the whole nature of the company. expanded the TAM by 100X. In Uber, it was linking town cars. It was your own private chauffeur. It was for the elites, Uber Black. But in the corner of their eye, they saw a now-defunct company, Sidecar, allowing you to take rides, ride-sharing, which Lyft then copied. This ride-sharing thing, hey, use your personal car and give a personal ride. And they went after that. Obviously, that became UberX, Lyft, and Sidecar. The originator of the idea failed to capture that opportunity.

8:26Let's get started with the docket. Sure. So our first topic, huge news this morning, although sort of the culmination of a story we've been following for a while. On Tuesday, SpaceX formally announced plans to buy Cursor for$60 billion in stock. The announcement, of course, came just days after SpaceX's debut on the NASDAQ in the largest IPO ever. Cursor now will become a wholly owned subsidiary of SpaceX. Of course, everybody knows about Cursor's popular AI coding tool, which launched in 2022. They've apparently crossed$1 billion in annualized revenue as of November of last year. Those are the most recent numbers we have.

9:03SpaceX shot up 60 % on Tuesday. They now have a market cap of$2.88 trillion, Jason. That makes them the fourth most valuable company in the U.S., surpassing Amazon and Microsoft. Of course, this is all subject to regulatory approval. I do have one quote here from Lekker, a cap CIO, Quinn Thompson, that I thought was good on X. This is brilliant corporate finance. Use your newly printed low float retail inflated currency to acquire real businesses ahead of the lockup expiring. Probably the most creative, accretive way to sell as much equity as possible into an IPO pump. I wonder what acquisition is next.

9:42So that brings me to my sort of dual question. What would you acquire next if you were SpaceX in this wonderful vaunted position they have? And do you think this was a done deal all along when we first announced the like, maybe we're going to buy Cursor, maybe we're just going to work on them? And they were just waiting to make the announcement post IPO. Ali, what do you think about has this been a done deal all along and they were just waiting until the post IPO moment to make the big announcement? And if you were running SpaceX, what are some other companies you might be looking to scoop up?

10:11Yeah, so I think it was probably a done deal, I would say. I'm sure there was some sort of regulatory concerns that resulted in them announcing it later. I think this is an incredible deal for both sides. And for Cursor, obviously, they have immense amount of potential to build their own model, mainly from the product usage that they have from their customers actually using Cursor, which is one of the most important kind of data flows that you need to actually build a model. but they you know sounded like they had some compute constraints with which obviously you know spacex had a ton build and so i think that that sort of puzzle matched uh really nicely um i would say the what what spacex has has built in terms of this sort of three core business lines is quite incredible to see and and it's you know they've done it in very quickly it seems like in the last like maybe six months or so obviously the starlink had been there for a while but the The compute business, which I heard somewhere has more gross profits than Anthropic.

11:16That might be true, maybe not, but that's an incredible statistic. And of course, now they have the opportunity to not only serve these frontier model companies, but also to serve their own frontier model, which I think buying Cursor is the best way to potentially do so. Because once you build a model that can navigate the computer in the best way possible, which is the coding capabilities, then you can have a lot of immersion capabilities that come from that, which is in many other domains. So I'm not sure if I would buy any other companies. I think that these three business lines are really lucrative, and perhaps SpaceX should focus on that.

11:51Sure. I mean, I guess, Ryan, to sort of follow up, I mean, why do you think XAI and SpaceX needed Cursor with so much compute, with so much CapEx at their fingertips? I mean, is it that hard to train your own model? and like what did Cursor, what can Cursor's team do that the SpaceX XAI team could not? I'll start by saying Cursor is a very big client of ours. So everything I'm saying has no relation to that. Fair enough. This is all speculative. Great, good disclaimer, good disclaimer. Yeah, good to know. I am a lawyer after all. Look, I think like long short is you really need two things to build a model.

12:25And one is, you know, a ton of compute. And so there's a lot of money. And the other is it's actually extraordinarily hard as a research problem. And my sense is that, you know, I think for physical world problems, Elon is just, you know, bar none the best to solve. Right. And, you know, even the speed with which he set up Colossus was astonishing. But like for the really hard computer science problems, they really struggled. And, you know, they've gone through several waves of researchers at X and still haven't been able to come up with a base model as good as the two labs. my sense just in kind of the you know the folks I know at Cursor they probably are one of the best research teams and so you know if you couple that it starts to make some sense that they could build you know a premier frontier model I guess the only other one that's relevant you know like I guess so you know between Google and Meta they still haven't you know I guess Google has but Meta still struggled so like clearly it's very hard even with great researchers and and I think there's a miraculous story about Cursor as you know a few-year-old startup that's been able to build that kind of research capability.

13:31Jason, I want to ask you about sort of the timing of all this and that quote that I read. But also building on what I just said, there was an interesting article in Business Insider this week. There was supposedly like a profile of Cursor CEO and co-founder Michael Truwell, but it sort of delves into their history with Anthropic. And it's sort of, it almost seems like it's Anthropic's fault that Composer exists. Cursor once accounted for 40 to 50 percent of Anthropik's total revenue. And then prior to launching Cloud Code, Anthropik executives apparently specifically told Cursor that Cloud Code was always going to be more of a research effort than a major commercial push.

14:09Like, we're not going to be a rival to you. And of course, that didn't necessarily hold. So in January 2026, Truel called a red alert and said, OK, we need to design our own AI model. We can't be totally dependent on Anthropik anymore. Do you think how much of that tension, like how much of this is Anthropics fault? Did Anthropic push cursor to make a model and now they have this massive competitor in SpaceX plus cursor? Yeah. Let me give you three really important observations here for entrepreneurs. And this is the most dynamic space I've ever seen in 30 years in the technology business. It is Game of Thrones in terms of talent, in terms of territory, in terms of weapons in AI.

14:51And if you look at cursor cursor had the lead and their partner was in fact claude but all these frontier models eventually are going to need to put up revenue numbers to backstop their trillion dollar valuations right open ai and anthropic are right around that you know 800 billion maybe trading in the secondary market to a trillion dollar when they go public market caps when you look at that That means Sam Altman and Dario have to examine what their platforms are being used for and then compete with their customers. And I've made this point three times, four times in my career. I made this point when people went to Facebook and started working with them.

15:34I made this point when Apple started releasing apps that competed with the best apps in the app store. And, of course, Microsoft is famous for doing this. Microsoft originally came out with Windows. They studied Lotus 1-2-3, eventually launched Excel. They bought PowerPoint. So when you build an application level on another person's platform, they have perfect insight into how you're using their product, and they study you. So Windows knew how many people were using Excel. They understood that, and they understood if they bundled it with the operating system, they would kill Lotus 1, 2, 3, which is what they did.

16:08Facebook did the same thing. They started coming out with games. They looked at their social graph, and they killed Zynga. And Mark, my friend Mark Pincus thought, oh, my relationship with Zuckerberg is so deep and so strong, he would never shiv me. He got shivved in the middle of the night, literally woke up in a pool of blood. Same thing happened to Lotus 1-2-3. Now we look at this, same thing's happening. Now, Cursor did the most important, this is point two, pivot in their career. They said, we have to make language model. That's a lot of effort. They were stuck. And this is point two. Elon stood up Colossus faster than anybody.

16:45That's their data centers. They bought all these GPUs. And Elon had a very famous tweet, buy GPUs and stand them up. Step three, monetize it. Step two, question mark. What happens in between, he didn't know. But he knew that having all that computer footprint would work. So when he met with Cursor, as this story goes that you're reading about in the public, He gave them a massive chunk of compute, which then let them catch up and then even perhaps beat Codex and Cloud Code. So you have this three or four horse race going on for coding co-pilots, agents, et cetera. And peanut butter chocolate, all of a sudden, he makes this option now in 0.3.

17:30Elon is exceptional at buying companies. People don't know this because they haven't been paying attention. But if you look at all the public data, obviously, SpaceX, and when we look at the valuation, $3 billion in revenue approximately per year,$60 billion valuation, right? 20 times top line revenue. SpaceX today, in total, including cursor's revenue, probably is around a$30 billion run rate, but trading at close to$3 trillion. So they have a 100x, 80, 90, 100x multiple. Which means when you make an acquisition like that, the seller, cursor, gets to be in a bigger, more stable, more diversified equity pool, the SpaceX equity pool.

18:17And then SpaceX gets to use that incredible valuation to make intelligent purchases. And if you look at Elon's career of buying these, I just did a quick research project before we got on air. SpaceX has bought a half dozen companies, the most important of which is probably swarm technologies that some of my friends were investors in. They were the IoT small satellite startup. They were doing incredibly important research. And now they are the Starlink direct to sell. And they had a lot of great innovations there. And they bought them for a half billion dollars, 524 million. Now you look at how savvy that was.

18:54He also bought Twitter. He created XAI, put it into SpaceX. He bought Twitter and Cursor, put that into SpaceX. And now you have two very important key assets, X, which has all that real-time data from Twitter. And then you have that one. And then you look at Tesla. Tesla's made, I think, close to 10 really important acquisitions, most of them in batteries and battery technology and wireless charging. But of course, SolarCity was in there. And SolarCity became Tesla Energy. I think that was a$2.6 billion acquisition. That acquisition is the entire or a large portion, let's say, of the energy business of Tesla, which is extraordinary.

19:35Eventually, these two companies are going to come together, as everybody has been speculating, and you'll buy ticker symbol E-L-O-N. Elon, yeah. Ticker symbol Elon. The true Elon industries like Stark Industries. We'll finally get it. Yeah, exactly. So this is just brilliant across the boards. And I think Cursor will have the number one model at this time next year in terms of revenue, in terms of market share. And they can play the long game because they have so much compute. Everybody else has to build all this compute. And if compute moves to space, another advantage for the Cursor team. And the Cursor team is exceptional.

20:14This is an exceptional team. Yeah. One thing that was interesting, I think, Kursar built off of Moonshot's AI. They're admitting that that composer was built initially distilled from Moonshot, Kimmy, and then they sort of built it. But now they say it's 85 percent their own model. So I feel like there's a lot of people don't really understand how distillation works. We think of it as like ripping off someone's model. But it's really like almost entirely their own product at this point. Ali, explain distillation for folks. And you have a very unique position in the market because you help frontier models and all the application layers fill in their data.

20:53And data on the open web is questionable in terms of quality. And it's a commodity because everybody's stolen every little nook and cranny. And I know these frontier models, they need your resource to get that data in there. So maybe explain that a bit. Yeah, so distillation has a bit of a negative connotation because the word is when you distill a closed model where you essentially prompt it at a large scale and you try to create a massive SFD, what's called an SFD data set, which you can think of as prompt response pairs that you get from frontier closed models to then use as pre-training data or some sort of post-training data.

21:36That obviously is not a good thing to do. And we've heard Chinese companies do it to closed frontier models in the US. But I think what I would call what Cursor did is I actually wouldn't call that distilling a model. It's just you're building on top of a baseline reasoning model that is open source. And that is like the perfect thing to do. That is the right thing to do, which is if you already have a really good baseline reasoning model that has been pre-trained and it's already a massive model that has pretty good capabilities and it's like literally open source, then you shouldn't reinvent the wheel, at least for your first model.

22:13And you should just post-train it in ways that improves the model for coding, which is what, you know, Cursor did. And the post-training efforts, they are so large that they sort of like resemble a pre-training effort. So that's why it sort of requires a lot of compute as well. And if you do a really massive post-training effort, you're changing the weights in so many ways where the model is just yours. And that's why I haven't heard the 80 % number, but that sounds right, which is like you're essentially changing 80 % of the weights, quote unquote. And so it is your model. I mean, you fundamentally changed the model, but you've kept some of this sort of baseline knowledge and baseline reasoning that came from the internet scale pre-training?

22:55I mean, one thing I'd add is I think a researcher recently outlined his vision of the future where there are just two types of companies that matter. One type of company makes models and the other type of company does not. And there will be maybe five or six companies that can make their models. And that's like kind of a hard cap. And it's basically what we talked about before. it's a function of computer into capital. And, you know, each model is now, you know, pre-training starts at about a billion dollars, but probably looking more like four or five to make a frontier model, maybe even more.

23:28And then the research capabilities, which keep getting more and more difficult. And so if you accept that view of the world, this acquisition makes a ton of sense, right? Like SpaceX needs to be one. It has some of the most capital availability, especially now with its share price. And Elon, you know, it's like inevitable for Elon to have his own model. It's just been taking longer than I think I would have anticipated. Probably could have been open AI. And so when you sort of understand his ambitions there, it makes sense that Elon will be one of the companies that have crossed that, you know, kind of frontier, cross that plane, that's binary to be one of the companies that is able to make models.

24:04And if we believe that models can build the next models, which is already happening, then it's just like a recursive improvement for three or four companies that become bigger and bigger and bigger. So that argument's compelling. And I think it kind of frames where this goes and why cursor thought to sell um given that they have the capabilities but they need to compute and they need the capital infrastructure yeah it's undisclosed how much revenue uh open ai's codex revenue that's not disclosed publicly but claude code there was a rumor of 2.5 billion in revenue or run rate in february 26 and cursor in june reported i think a$4 billion run rate.

24:41And so this is a big space. Just code in and of itself is going to be $100 billion in revenue in the next year or two, just from those top three players. And you're correct, Ryan. XAI had a bunch of turnover with the founders or the founding team. We all saw that in the press. And Elon rebooted everything. And I do think he will catch up. My question, Ali, and I'm sure you have something to add here, is our open source models, what percentage of a frontier model are open source models today, Ali, in terms of capabilities, et cetera? And when do you think, if ever, it will flip where the open source models will exceed frontier models or be within plus or minus 10, 15 percent?

25:27In other words, it's not discernible to the majority of users that there's a difference between an open source model and a frontier model. I want to sort of define frontier model for a second. And I totally agree with Ryan's take of like, there's probably, we're probably going to converge to like four to five really massive model companies, which I would sort of define them as baseline reasoning model companies versus like frontier models. And the reason for it is because I would actually define frontier models as sort of what's built it's actually in the product layer where, and this relates to Satya's post a bit, which I'm sure we're going to get into.

26:06But essentially when you take a baseline resilient model, whether it's closed source or open source, and you build a product on top of it, the best way for you to make sure that the product is actually reliable is to, in a way, build your own model. And you're not, of course, you're not doing like full-on pre-training and doing these sort of post-training runs that require a lot of compute, at least for most companies. But when you build a probabilistic software with evaluations as the sort of core of the full product lifecycle and the full product build out, you're very much doing what model building looks like.

Read the full transcript

26:45And this goes back to this notion of there is, and Greg from OpeningEye says this as well, he publicly tweeted it, which is the model is no longer the product that what comes on top of the model, the agents, the agent evaluations that come on top, the harness, the user interface, and so forth, that is the product. And I think those are areas where frontier models are actually built. And I don't think there's an infinite set of domains that you can take this final mile in. And I don't think the baseline reasoning model companies, which I think will be still the biggest companies in AI, will take the last mile, which we sort of call this like infinite last mile in every single one of those domains.

27:34So there's going to be countless companies that are really like today, we'd call them AI application companies that end up building their own models in some ways. And of course, we're seeing that with Cursor. And actually, I think those are where the frontier models are going to be, not the open source versus closed source that are these like massive models. And I actually don't think the percentage of open source versus closed really matters too much. I think in a lot of domains where the last mile is taken, an open source model will be totally okay to sort of fine tune and evaluate your way into what I would call frontier in that product.

28:06And a lot of cases, the closed source models is where you'll sort of build on. But I think the real final sort of value will come from all the evaluations that results in these product companies owning their own intelligence layer, which doesn't mean they're not building on top of these closed source models. But you can still own your own intelligence layer while building on top of the closed source models. So that's what I would define as frontier. Ryan, in terms of Crosby, you probably looking at the cursor example are thinking, hey, are my model, frontier model folks going to partner with me or eventually try and take our business?

28:46Now, you have a services component, obviously, but how do you think as an entrepreneur about being either headless or not dependent on any one model? And to Ollie's point and Greg Brockman's, the model itself isn't the product, the harness, the instructions, the skills, memory. There's a lot of other components to this. Yeah. I mean, I think there's just so many ways I can go. I think this is the fundamental question that anybody doing, but I guess we'd call application companies previously, and now it's becoming, I think, even more opaque, kind of where we differentiate in terms of, you know, our own type of intelligence in one specific domain.

29:22We're all like, you know, there's like this, I think we can oftentimes fall into nihilism where we're like, nothing matters, don't build anything, the labs will do it. You have this very friendly collaboration with the labs where we do a lot of applied evals with both of the open an anthropic, and yet we never know. like what are they gonna you know and and i think it actually so a couple of things i think one is it makes companies like like always like micro one much more interesting because for companies like us we have this vertically integrated solution we have tons of unique data that is very very hard to create and we built it ourselves with our own law firm and that is a huge edge um but you know if you don't know how to build your right evals and kind of like tune models to work well for that like it's not super useful i think it's one thing i think the other is it's very domain specific, I think the non-self-verifiable domain.

30:09So like things that are really taste-based, obviously law is one of them. It's super subjective. You know, as lawyers get more senior, they disagree on what quality looks like. You know, code is very self-verifiable, right? And so you can see why that's something easy for the lab to get really good at. But the more taste-based you get, the more I think the value can accrue more easily to these application layer companies that are just closer to the work product. And that's our bet. And I think it might just mean that we get more specialized over time. But that's kind of what we just need to focus on.

30:39Whereas I think the labs will, they already have, both OpenAI and Anthropic, have legal point solutions that are, I think, very base level and not super specialized. Yeah, that makes sense. Also, just one thing to add here in terms of the product, or the model not being the product, is OpenAI and Anthropica are the best example of this, where they're coding product, not modeled, but the product is what is generating a very large portion of the revenue. And I think, you know, Cloud Code obviously is the perfect example where they had the best coding model. They obviously still do. But when they came out with the desktop application, which was an actual product, that's when the revenue, you know, skyrocketed.

31:21Satya Nadella, the Microsoft CEO, wrote a widely shared X article over the weekend titled A Frontier Without an Ecosystem. He argues that the future is a, quote, cognitive loop connecting people and systems. So this means enterprises will be driven by humans building learning loops on top of AI models and turning workflows into more like agentic systems that improve themselves through reinforcement learning. And then those loops are the company's IP where all of the value rests. I pulled a key quote. This means the real opportunity is not in picking the best model, but instead building a learning loop on top of models where human capital and token capital compound.

32:03You can offload a task or even a job, but you can never offload your learning. The future of the firm is the ability to compound that learning across people and AI. The Post, by the way, staggeringly popular, 63 million views, 12 ,000 RTs, a genuine banger, even from the CEO statement genre. So I guess my first question would be to Ryan. This reminds me a lot of what Crosby has sort of designed, its own kind of in-house proprietary feedback loop, much like The Post describes. So can you walk us through like how that works and how the combination of human and token capital is really producing these like higher level results?

32:44This post was so validating in so many levels. You know, we've hired now about 50 attorneys that we've hired it and we have a captive law firm that we've built. And, you know, it's not easy when you're, you know, raising venture to explain the value of a law firm that looks very much superficial like a services business. it's not easy to explain to lawyers that are really anxious about models getting better like why are we just trying to have them train themselves out of a job and our fundamental belief and i used to be an attorney but our fundamental belief is that the lawyers get more valuable as the models get better because they can do the things that only humans do and and the you know cost per dollar of that time actually go up over time and all the stuff that you're paying somebody a thousand dollars an hour for that you can definitely off you know hand off to a lawyer you no longer have to pay for it and we kind of know that i think people get frustrated with their lawyers because we all kind of know that.

33:31And I think the hard thing is that feedback loop, the reinforcement learning loop, which is the key, again, like going back to what I mentioned with Ali, you need to be able to eval an output to say if it's good or bad. If you can't do that, the models can't reinforcement learn on it. And for the subjective domains like poetry, like forms of accounting, like literature, obviously like legal, where there's a variety of opinions for what's a good output, this is where this becomes a very proprietary thing to work on. but like i think josh koshner wrote a post last week that i felt like i was like why don't we think of that where he says we're long humans here at thrive holdings like this is just like we've just been so banging this drum of we're long humans we don't want to build an application that can replace lawyers for this domain that's very subjective and where the quality of judgment matters so much the more we can kind of create these loops of learning on top of the models um in a vertically integrated system is so powerful so like and and i think we're probably a few months ahead of most other companies.

34:25But I think every law firm will have to figure this out or they won't exist in four or five years. Yeah, this is, I think, going to be the trend of 2027, Lon. You're going to see open source, on-prem servers, building your own language model, or the harness around it becomes the trend. And when you read Satya's post, you immediately think about a company like Ryan's because this concept of there's a software layer, there's a services layer, there's an intelligence layer, all of that is going to be attracted into what I call the problem being solved layer. The problem for your customer being solved.

35:08Ryan's customers do not care how the problem is solved. They care the problem is solved faster, better, cheaper. Some combination of those things. So if you fast forward to 2027, you don't want to give your knowledge in terms of Ryan's company or many other companies, accounting companies, we have tax GPT in our portfolio. You don't want to give that to Claude. You don't want to give it to OpenAI or Gemini. You want that for yourself. And so we as venture capitalists investing in these companies, we never wanted to do services. We always wanted to do SaaS, higher margin. Please, if you have any services revenue, hide it from the venture and investment community.

35:51None of that matters. The whole paradigm has been abstracted into customer solution. Money changes hands. That's it. And so if you take this to its end state, I think savvy founders like Ryan are just going to say, you know what? I don't care what's underneath the hood here. I care about my customers and I don't care about VCs and their lack of understanding or being able to put me in a box. And that's what Kushner is, I think, also saying in his missive. And everybody's kind of coming to the same conclusion. Just don't worry about what box this startup fits in. Are they delighting customers with a better, cheaper, faster solution?

36:30Period. Full stop. And Ollie's company has the same. MicroOne has the same confusion that VCs have had, right, Ollie? You had this like three months ago, people saying like, these companies that are doing reinforcement learning, their services, they're not software. It doesn't matter what a VC thinks. VC's opinions do not matter. The customer's opinion is what matters. So we just have to get this VC paradigm biases off the table and just look at what customers, Ollie's customers, like Ryan, love what he does for them. Ryan's customers love what they're doing. So how does the frontier model, now the challenge is on the frontier model.

37:09How do they get to play with the people who are touching the customers. And I think that's what could be radically changed. Don't be surprised if everybody working at Ryan's company has a workstation running a local model, studying their behavior. And at some point, I think, Ryan, you could just have everybody on a Mac studio or with a terabyte of RAM running local models that you own, and you never touch a frontier model. Is that a possibility, Ryan? There is always a gap between what the frontier models can do and like great work output. And then the question is like, what domains do you really just pay the extra premium for that last 5 % to get to the good work output?

37:49And what things are you kind of okay with cutting the corners or doing yourself? And for domains where there's sensitivity, especially like legal, medicine's another one, the value all the cruiser power law, the premium all goes to like getting that last bit done, not the first draft, but the final ready to file. you know, ready to sign signature kind of contract. And so for us, like we've been in the last year in particular, dogmatic about the codifying the judgment of our lawyers who are all fairly skilled, you know, five plus years, they're expensive, but like this is where the value started to accrue.

38:23And, and so, yeah, we're becoming much more thoughtful about what are we doing with that data? You know, like, you know, should we be posting our models with it instead of like sharing anything with, with frontier models? And. And I think, you know, Jason, the point you made is like spot on. Like, I think this, you know, hand-wringing about is it services and software? It's just like when human labor can be done by a model for the first time ever, and every six months, like even more and more labor, it's just such an anachronism to ask the question. It doesn't matter anymore. And we're always surprised by how many things we're like, lawyers will always have to do that.

38:56And then one day, like they no longer have to, and we just watch it with our evals. So examples of in the lawyer stack things that have already become automated versus things that still require human in the loop and then things that you think are going to be the last castles to fall. I mean, we'll get to this when we talk about the benchmark. I think like. I'll tell you one thing that took us a long time was summarizing a review for a client. Sounds really trivial, right? We make a bunch of changes. We send them a cover note saying, here's what we did. There's an enormous amount of judgment that goes into that, right?

39:27If you're talking to a lawyer, you give a very in-depth summary. You really want them to get in the weeds. And the way we evaluate ourselves is we don't want our country to ever have to open a document that we did because that means that we haven't described it well and they don't trust it. So the real time saving with AI, you have to read the document every time, look at the output and verify it. With us, you should just trust it. That's the key. For a non-lawyer, they want a very cursory summary and they don't want to know the 10 changes you made. They want to know the one thing that they really might want to know about.

39:53And getting models to be thought, and we could just create all these hard-coded rules. here are the five things you're summarizing these instances and it was too brittle and then we just put in a bunch of really good instructions on like hey given these are the parameters of the deal and it's a really high velocity deal it's not a big client and it's a salesperson not a lawyer give a cursory summary and here's what kind of cursory looks like and i think with 4.6 it started to kind of come together where it was like making like you would read it like yeah that's a pretty reasonable note but for the longest time we're sending these like huge cover notes our clients were like, what am I supposed to do with this?

40:22Like, this is crazy. And yeah. This is where Ali, the person who is closest to the customer gets to win. Yeah. Yeah. I think this is actually a really important point, which is that there is one of the main, of course, as we've said, one of the main ingredients to train models is obviously data, but there's two sort of subcategories of that, which I almost want to call like the two kind of new scaling laws that are specifically in this data category. The first is the horizon of tasks that you get access to and you're able to kind of keep increasing the horizon of. And the second is the real worldness of the tasks and of the data points.

40:59And if you think about what a services company does, is they're delivering real world economic value to their customers, which means it is as real world as it gets. It is literally the real world. There's not a proxy to the real world. And it's also, as you kind of keep improving the capabilities of the sort of, in this case, the AI native law firm, you're able to deliver longer and longer horizon tasks to your customers and eventually deliver as long as possible. And those two dimensions can be very much scaled up to then improve the intelligence layer that you have at the application company or the service company, whatever you want to call it.

41:41And this is something that labs are very much looking for. A lot of the work that we do with the Frontier Labs, they're always trying to look for ways to make the data longer horizon, but also make it more realistic. And so one of the things that we've actually done recently is we've started to partner with companies that are using some of these applications that are built by these AI labs in their real workflows versus these kind of like fake sandboxes to then send back trajectories of the real tasks that they do with those tasks evaluated, which is structured and evaluated, which means like how well did these applications actually do on these tasks for these model companies to then train on.

42:21Now, if you think about, again, a services company, there's an abundance of those tasks because it is what you're doing and what you're serving to your customers. So this goes back to the point of, I think, you know, the services companies that have a very niche focus and sort of one area that they want to kind of use intelligence plus humans on is really those are the companies that are going to build the frontiers of these domains. And also we are, I love that everyone's sort of going this route of like long humans and humans first. We've been saying this for a long time, which is human brilliance is needed more than ever.

42:58Like when you start to improve model capabilities in a certain area, the expertise from humans, the judgment to continue training in that area, plus the judgment required to actually deliver value in those areas increases and the demand for it increases. And coding is actually a perfect example where a lot of folks think that like as coding capabilities get really good, where like the model can basically write almost 100 % of the code. There's this notion that like, oh, humans are actually not needed, but it's not true. I mean, there are more engineers than ever. there's also more engineers training the models than ever.

43:31Like there's the pipelines that we kick off right now at MicroOne for training model capabilities, a very large portion of them are coding. And the demand is the highest it's ever been. And of course, the capabilities is also the highest it's ever been. So we will likely see this pattern follow along with every other domain as well. Ali, do you think that like your next client base in 2027, 28 will be like accounting firms, consulting firms, trying to have these learning loops done locally? 100%. In fact, it is a current client base as well. We are starting to see enterprises that are, let's call it less AI native than others that are coming to us that are saying, we have this pilot of an agent that we've built, but every time we test it, there's just random people within the company that go and do these like anecdotal QA on this probabilistic software.

44:27So we don't know if it's actually going to work or not in real life setting or like in production. And that's exactly where the evals come in. And that's where like them owning their own intelligence comes in. So we actually believe that in the long run, the majority of the data spend, in fact, probably close to 100 % of it will be on the agents and the application layer versus the model companies. And this doesn't mean the model companies like reduce their spend. I think they're going to continue increasing their spend in major ways. But there's just going to be orders of magnitude more agents built than obviously models.

45:00And it's not like 100x or 1 ,000x, maybe like a millionx more agents built than models. And it's sort of a similar argument to where compute is used, which is obviously compute was mainly used for training a while ago. Now it's largely inference, and it's probably at some point going to be essentially almost 100 % inference. I think that's sort of like the equivalent trajectory that we're going to see on data, which is the majority of data needs will be on the application layer from all types of companies, even the ones that are like mom and pop shops that are still owning their intelligence layer.

45:30I thought this was interesting, especially talk about the nuances of the law that's sometimes hard to bake into these models. There was an interesting piece of the Wall Street Journal over the weekend about how court reporting seems to be a prime target for AI takeover. There's a shortage of trained humans to do it. And you're just listening in. It seems like a speech to text sort of thing. Of course, a court reporter, they're the ones typing the real time transcripts of everything that's spoken aloud during a legal proceeding like trials, depositions, hearings or whatever. But the thing about what they're typing is that transcript becomes its own legal document.

46:06So say if there's an appeal, future lawyers and judges are relying on that transcript as an accurate verbatim copy of what was said. So you would think speech to text has come such a long way, but the real work of court reporting includes a lot of stuff AI isn't great at yet. Capturing nonverbal cues, filtering out ambient noise, essentially making it certified that we can rely on it, you know, as a verbatim transcript of what happened. The National Court Reporters Association, maybe they're a little biased, but they argue AI transcription remains too error prone for use in courtrooms. Some states are already trying it, though.

46:44North Dakota has eliminated stenographers and already switched over its courts to electronic recordings. And, of course, this is a difficult thing for humans to do because you have to type Texas for its court stenographers requires 95 percent accuracy at 225 words per minute for five minutes straight. I can type pretty fast, but I'm not I can't do that much. That's like some people compare it to learning a foreign language or a musical instrument or something like that. So my question here is for Ali, is this a matter of time? No, maybe for Ryan since he's in the legal space. Sure, for Ryan. Is this a matter of time before you think AI could do this as well as any human?

47:22Or is this always going to be something we need a human in the loop for? And, you know, for Ali too, is this something we could ever train or fine tune with the use of experts to make AI better at it? I think the obvious answer, I actually think AI can do this already, honestly. Like, I think some of the voice models, like Whisper, it's kind of astonishing. Whisper's crazy. And so I think it's solved. I don't think that's the question. Like, the business that I would have started, if not this one, is an AI private dispute resolution company. And there's like AAA, the American Arbitration Association, and JAM.

48:03These are two huge arbitration societies. and businesses typically internationally for jams and in the U.S. for AAA agree when they sign a contract ex ante, we're going to like surrender our jurisdiction to these to these arbitrators because they're smarter than judges. They're faster. They're smarter for our need for business. But here's the key. The outputs of those tribunals are enforceable by a court. So it's not just like whatever they say we might go with. We consent to a court binding us what a decision of this tribunal is. But the key is that we have such a need for dispute resolution as humans.

48:35That's just in the US, that's how we solve our problems. And the US court system, especially the federal system, you need an enormous amount of political capital to appoint any new judge. It's pretty deadlocked. And so any change is really, really hard to make. I think that AI lawyers will overwhelm courts far faster than courts will use AI to keep up. There will just be way more lawsuits filed. It's already happening. And so, and I think one of the key things to get past is there is a collective myth that we surrender to as Americans that the courts are fair. And like, we just know they're not like, we just know that juries are not fair.

49:11They have a lot of bias. And so we're judges are much more likely to grant bail after lunch than before lunch. There's a great famous study on that. Um, because, because, so if they're hangry, they, yeah, super famous, super famous study, It's a hangry index? Wow. Yeah, big time. Super famous study from like seven, eight years ago by these Israeli researchers. So like, we just know. However, we just believe, we surrender this collective belief in the fairness because it's done by humans. And so I think courts, like, it is inevitable an access to justice increases. And about 70 % of Americans don't have access to a court or a lawyer when they need it.

49:45And it will because of AI. And I just don't know if the public judicial system will be able to keep up with the reality of technology. And this is just another example of that. But they're like endless examples. When will a person be able to be their own counsel for, let's call it, not small claims court, but whatever the next level up is? Like a, let's call it a$100 ,000 dispute. I think small claims court is under - Like under 25, 10, depends on the state. Yeah. So let's say$100 ,000 dispute about, I don't know, I was building my dream home and I'm fighting with my architect and the construction company.

50:21I don't want to hire a lawyer. it's not worth it for this 100k and damages but if i can use ai to represent myself maybe i'll take a swing at it when will that moment happen do you think you have to be in a position to cut someone open because it's dangerous if you're not licensed and so but you know like let's be real i think some things in court require that level of training for sure you really need to know how to argue the right motion and um your client will be at a significant disadvantage if they represent themselves. I think by late 27, models will be on evals, which are really hard to verify, as good as many lawyers, right?

50:58We have like almost 3 million barred attorneys in the US, about 800 ,000 practicing. So let's index close to a million. And there's an astonishing variety in the quality of the lawyer getting, right? Like, I mean, there's just like some pretty crummy lawyers. And yet people don't have access. And then I think the hard question for courts are grapples when models are verifiably as good as the mean of a lawyer, how can you ethically not give people access to those who represent themselves? And that's coming. Well, we saw a study, I think it was out of Harvard, that consumers were preferring the bedside manner and the fidelity of answers from a general practitioner.

51:39And that's already happened in healthcare. Now, people might not know that or might not be doing it, but the number of health searches you're seeing in these large language models and people going to them first is obvious. I apparently have a mosquito bite allergy. It's a silly, simple thing, but I went to an emergency care because these like three mosquito bites were just, you know, getting inflamed. And I just said, hey, you know, can I record this emergency person at the urgent care? I was like, sure, no problem. I took it, I dump it into the transcript and the audio file into one of the large language models.

52:12And it gave me an absolutely spectacular answer in real time that I then had a Socratic dialogue with the person. And she's like, yeah, that's correct. That's correct. That's correct. Literally everything she told me was then verified and then some. And then it gave me other opportunities. And then I shared it with my wife and said, hey, by the way, I think one of our daughters also has this mosquito allergy thing. And here's the protocol. We're done. like amazing right i didn't have to go to the urgent care if they allowed me to get this certain steroid to put the inflammation down i would have been on my own i could have solved the problem without going to urgent care and that to me was like hmm okay we're already here in my mind i mean ali like do you think if you were to create evals somehow that could show some of these really subjective domains were as good as a licensed professional do you think there would be would there be even need for them?

53:06Or does the licensing regime just make it uninteresting? Yeah. So I think part of it is the licensing regime, but I think another part of it is, this is not a favorable statement towards what we do at Micro One, but evals, no matter how representative you try to make them and how much you deem that this is verifiably better than humans, you can't actually get to this sort of like 100 % representative state where you determine that this legal benchmark is now better than a general practitioner or this medical benchmark very much proves that models are better than all urgent care folks that practice.

53:54I think in most cases, there's certainly going to be some tasks that are like fully automated away. But I think in most cases, even in the long run, the value of AI will be ultimately delivered by humans in some way. And I'll use, I mean, coding is the obvious example where, again, like there's sort of a renaissance of software happening, but it's the software engineers and the many new folks that are now calling themselves software engineers that are actually delivering that value. But that's sort of the obvious example. I'll use another example that's very niche and it's actually like a simpler problem, which you can like determine solved, which is we built a proctoring model at MicroOne.

54:35It's like, you know, part of what we do is obviously the AI recruiter that we built. And there's a bunch of like different models that we have that kind of make this work to source and vet experts. One model that we built is specifically the proctoring model, which takes in, it's a very simple model. It's not LLM. It's like a pre-trained model from scratch, which takes in video embedding of someone doing an interview. And then it sort of spits out a probability that they've cheated. And we've trained this on lots and lots of data. And it's sort of like a fairly simple model. You can kind of like determine this solved.

55:07However, in this case, humans are still delivering the value of this proctoring model in two ways. One is the recruiters that are checking the proctoring scores are in almost every case and pretty much every case, doing a very quick sanity check if this proctoring score hallucinated or not before they determine whether or not this person should be marked as cheating or should be not marked as cheating. And that can't really change no matter how well we say the model does. That's sort of the first way. The second way is this model, we've trained it with a pretty good sort of architecture and There's a nice data pipeline that has a lot of like self-improvement, you know, capabilities built in.

55:52But there's like a massive amount of drift that happens, even though there's no like law changing here. There's no like societal implications that result in this drift. It's like a very simple like user interface changes on the proctoring side of things make the model drift in its capabilities. So that simple example on its own results in a bunch of human experts that we have, which I think now is like it used to be about 10 people. Now we actually increase it as like 20 or 30 people that are constantly working on labeling which videos they think is cheating or not to not allow for this model drift to happen.

56:30So this very simple example of something that is very much sort of solved is still delivering value through humans, both on the side that actually makes the ultimate decision, but also on the side that continues to train the model because of the drift. And I certainly don't see this changing in areas that are way more complex than this proxering model, such as law, such as medicine. I think you should, Ali, get an urgent care facility and just buy it. you should buy one of these Texas-based urgent care facilities. You know, they're like franchises kind of, or they're like local businesses, like local restaurants.

57:08Some are definitely franchises for sure, yeah. They feel like it, I don't know exactly how it works, but if you bought one of those, Ollie, and you just used it for data collection, and you just said, there's going to be a microphone in every room, we're going to anonymize the patient data, HIPAA, whatever, and you just recorded every session, and then every time people came in with Cedar Fever, you know a dislocated pinky whatever it happens to be mosquito bite allergies and you just had that data set i think you could run the business at a break even and then have the best data set in the world i wonder if somebody's done that yet they're always smiling not not to i i don't want to make this promotional but this is actually like literally my number one focus not not on urgent cares there's like a lot of sensitivities on medical data and so forth but what you just described, Jason, which is the real operational workflows of companies is the most valuable thing for model training.

58:02And of course, it's not us training models, it's for our customers. But we're going out right now. And I think you guys will find this interesting because the magnitudes of money paid in this space is quite high, which is we are paying millions of dollars to companies that are oftentimes not even that large of companies, like 30 employees, 40 employees to get access to anonymized PII sort of transformed versions of their entire corpus of company data to seed for training environments. So it's very similar to the clinician example. Yeah. Do you remember, Jason, there was that story a few months ago about companies like startups that have gone belly up and they're selling their like Slack histories for AI data?

58:46No, I never saw that. Yeah, it was like failed. So, well, but that was a lot of the like snark on social media was, but then aren't you training your new data to be a failed startup? Like, why would you trade it on companies that went belly up? So do the opposite of what you read in this slack. Right. Well, but I mean, that's my question for Ali now is like, would that data still be valuable even though the company didn't work out? Yeah, there you go. So even though that company didn't work out, is there still value to be gained from just the workflows, what people were saying, the conversations that were happening behind the scenes?

59:21Yeah, there is. So we're not going necessarily too much after failed companies because, like, obviously you can argue that there's some bad data as well. But even for those, it's still valuable because if you think about what an RL environment is trying to accomplish, it essentially gives an agent, which is some frontier model, a task. And the task could be, you know, go read this document. And based on the scope of work that you see in this document and the constraints that you see in these Slack conversations that may change the scope of work, build me this front end of this web application, just as a random example.

59:58That task has sort of a few components, right? One is the actual prompt, like what is it that you're trying to do? The other is the verifiers, which we'll touch on what that is when we talk about the Crosby thing. And then the third component, which is really important, is what is the seeded data that it has? Which is like, what documents are you telling it to actually review to then create this web app? Or what is the Slack conversations that you're telling it to review to look at the constraints of this web app? And if you create synthetic data for that seeded environment, you get really not the best results because it's too simple.

1:00:32There's no noise in the synthetic data. The Slack conversation is very sort of structured and the scope of work document is probably not the most complex that it's probably too well-defined. So even if the company has failed, you can still take all of these different documents as seeded data for these RL environments to refer to as you build these tasks. That's kind of one of the main use cases for these company purchases? The ultimate dark data pool is Slack. They are not allowed to take that pool of data and train AI on it. That cannot be possible in the terms of service. However, if there was a way for Slack to make it free, but we get to train our LLM on what you're saying, that would be the ultimate product in the world because all human knowledge work winds up in some way inside of Slack eventually.

1:01:27Salesforce ascended. I mean, but they're not allowed to. So now you start thinking about that as a possibility. What if somebody took an open source version of Slack and then said, hey, you can use this for free forever, but we just want to be able to train on what's in there. That would be the ultimate crazy bargain. Don't take that bargain. But what I want to do, and this is where Slack is so frustrating for me, is getting access to your entire corpus, DMs, et cetera, requires you to keep like paying up to like 30 or$40 a month for a person. Like if you want that version that has compliance as a FinTech company, let's say we're monitoring everything.

1:02:07And I think we pay for that. So if somebody on our team sends a DM to another person with like an insider trading, you know, idea. We're all aware that you're reading our Slack DMs. The word is actually reading it. No, I told everybody up front, Like we're recording all this, just so you know, never say anything derogatory about a founder, never do anything stupid, you know, in the corporate Slack. But what I really want to do is take the entire Slack, export it every night or in real time and train my own internal venture model. That's my ultimate goal. And I'm kind of doing that with Notion.

1:02:41Now, Notion's AI is doing a version of that. And we've authenticated Notion to look at our Slack. So now when I ask Notion AI a question, Notion is just such a brilliant product. We have to have them on this week in AI. They pull Slack in and they do a better job. Sorry, Mark Benioff. They do a better job examining Slack with AI than Slack does with their tools. And that to me is like just an incredible future if we could get all that data. And now when I ask, hey, give me the history of our investment in Micro One. Um, if I do it in Slack, result is like a two out of 10. When I do it in Notion, it's like a seven out of 10.

1:03:21Yeah. If I had my own Hermes open claw version of this, or maybe even using clawed code eventually to build something, I think we could get to like a nine or 10 out of 10 where now I'm asking questions, Ali, like, you know, when did we, how did we meet Ali? How did we, what's our investment history? It's starting to make me a dossier of the entire relationship, which is really amazing. And what companies did we miss? We have their interviews and their notes. To me as a venture capitalist, this takes it from, hey, I'm the world's greatest angel investor picker to, hey, tell me where I effed up the most.

1:03:59like it's the difference between uh the nba players not wearing a whoop not studying their data with orico or whatever it's called that we're investors in i've got the name of it um there's like a company that studies not only the um game data but the health data they have and then they put them together so here's jalen brunson's sleep patterns plus his performance the next day. Here's Steph Curry's three-point percentage versus his nutrition. All kinds of weird stuff are coming out of the NBA now where they can make these players really moneyball them with AI. It's just a really bright, interesting future.

1:04:38But the dark data pool of Slack is the ultimate data pool for me. What do you think, Ali? By the way, are we sure Slack is not training on the data? I mean, I just asked ChatGPT, I mean, of course, we're not sure if this is true, but it says it is explicitly prohibited to train foundational models on the data, but it is not explicitly prohibited using interaction data, feedback results, and outputs for product evaluations and improvements of SLAT AI itself. And again, going back to the point we made earlier, which is the intelligence layer, you don't need to have this distinction of like training models.

1:05:16You can just eval your way to quote unquote training models. And if they're doing this, they're improving their agents and they're effectively training a model on customer data. By the way, I think Slack is, I'm not sure if this is true. Of course, this is just based on strategy thing that just came out. But if it is, then I think that does mean they're effectively training customer data. And I also think Slack, by the way, needs some competition. We use Slack, but they are very aggressive. They have a very aggressive sales team. And I think competition for Slack will be a very good thing for the world.

1:05:47I would give a million dollar seed funding to a team to make an AI first Slack that charged a flat rate for up to 500 users. That was like... Wasn't that Glue? Wasn't that Saks' AI Glue project? Saks had an AI version of Glue, yeah. I think there needs to be another one that's not priced off of this. And I don't know what happened with Glue. I don't hear much about it. I don't hear about it anymore. I think this might be an example of being too early. There's some, what would you, Ryan, if you were going to make a disruptive version of Slack, let's game play it here. And then I want to end with your partnership with Ali, but just let's imagine what that would look like.

1:06:33AI first, Slack, disruptive. How do you disrupt Slack? So there's this company called Ando. It's Sarah Dew's company, still in stealth. this is and and and she's great and uh this is exactly a n d o and oh yeah and um it is an agent first slack and i'm dying to get our company access to it and it's i think our team's too big and the and it resembles slack in many ways from a product perspective but it's built so that agents run natively right now i still struggle to use agents in slack and it doesn't make any sense it is where we do all of our work but the idea that like i could have all of my agents that work on my behalf and run through Slack and talk to my teammates agents and actually get things done and then give me the summaries, we will obviously be there a year.

1:07:17The fact that we're not there now when the models can probably do it is just a product limitation. So it's coming. So if Slack can't figure it out, I'm dying to switch us to endo. Yeah. Wow. It says the product thesis is the current platforms like Slack, Teams, and Discord were built for human to human collaboration, not for humans collaborating with AI agents and agents collaborating with one another. So that That does end it. I just submitted myself for endo. Yeah. A-N-D-O dot S-O. We'll put it in the show notes and we'll book that founder here on the show. All right. I wanted to end with you guys talking about your partnership, Ali, because this is a really unique one.

1:07:54How do the two of you work together and collaborate? Yeah, I'll give the brief. Ryan, feel free to add on to this. But essentially, we are building a benchmark. And I think by the time this episode comes out, it will be live. We are building a benchmark that is around multi-turn redlining for SaaS contracts and other types of contracts. And we are having about 20 to 30 top lawyers that are building these verifiers based on contracts that are real negotiations that happen between the lawyers, where a lawyer submits one form of redlining based on some real documents that are, in a lot of cases, public documents.

1:08:39And then there's another lawyer that redlines back. And then there's the sort of four or five turns that happen. And there's about four to five categories that each turn sort of cares about in terms of legal accuracy, the verboseness of the red lines, negotiation leverage, and a few other things. And our goal is to simulate real world redlining scenarios with this benchmark where it's first of its kind. and also our goal is to completely open source the data set that is created here, not just the full reports and sort of a sample subset of this, but the entire data set for folks to analyze on their own and sort of check the accuracy of.

1:09:20But yeah, Ryan, if you want to add on to that. I think the key here is we came to Ali, we've mentioned it for a while now, and I said we have this really hard problem that I would have thought would be quite simple, which is these routine commercial contracts, like any tech company, anytime they want to close, let's say anything over 30k 50k has to close a contract and it takes sometimes months usually at least weeks and that to me was crazy like i was a gc of a startup it was crazy that was the bane of my existence and like i just thought this would be solved by ai by now and it's not and it's far from it and so we we like came to all of this like we did it's like it's for some reason the models are just like they can't seem to figure out the judgment calls that lawyers are making so so we put these lawyers to negotiate against each other and each had the same instructions but we'd have five lawyers doing the exact same tasks to see where they agreed and didn't and the results were kind of amazing.

1:10:06Like basically the lawyers had a lot of consensus on the first review. They would all do the same thing. Models were all over the place. And then for subsequent reviews, there's a really easy way to close a deal, which is just to say yes, right? And that's not legally protective. That's being a bad lawyer. That's what a salesperson wants to do, but not what a lawyer would do. And we found is that models are really likely to just say yes to everything and not be legally protective. And so we found all these interesting nuances. It's the first benchmark that shows an entire negotiation from start to multiple terms.

1:10:33Most of these benchmarks just show one task. And the goal here is that agents should just negotiate against each other and close a deal. And lawyers can weigh in just at the end. That's what we're working towards. And so this is the first step in publicly showing what would it take for agents to just run the deal. They know what I care about. They know what Ali cares about. And they can just simulate us all the way to the end. Yeah, I think one other nuance here is that, of course, legal is a very subjective field in a lot of ways. And one of the things that you sort of need to do is you need to simulate a really good debate, essentially, that happens before you get to this final structured judgment, whether it's red lines or any sort of outcome of a case.

1:11:15And so in this case, the way we've designed it is instead of taking the score of one lawyer that determines the verifiers for one set of red linings, we actually take an average of, I believe, four to five lawyers that do the same redlining, which in other words, they create the rubrics that define what a good redlining would be, which give a score to the model once the model actually attempts that same redlining. And then the sort of average is what ends up being the model score. And there is – so we sort of remove the noise due to subjectivity and we try to converge a little bit more towards the truth here.

1:11:55Based on the current benchmarks and the results that you're seeing, like how good are the top models doing like Fable 5 or whatever? Like how close to a human lawyer's job are they able to do right now? So Fable, we're still just given everything that happened over the weekend. We're having to play with some of those numbers because we lost access before we were done. Fair enough. Not your fault. State of the art is like, you know, between 10 and 20 percent of the way a lawyer would get there. It's like there's still so much that goes into it on so many dimensions. Wow. We're talking about like it's really a chess game between two lawyers.

1:12:30Like how does each person negotiate? Are they really forceful? Are they really gentle? What's the right tact here given who your counterparty is? and the board is constantly evolving with each step. And I am positive models will be able to do this as we've staked our business on it. But we're focusing a lot of our efforts now, more than we would have anticipated on research to be able to try to figure this out first because every transaction really runs on some sort of contract. Yeah. A follow-up question here just because I'm curious. Is there a future where there may be multiple legal models available and you're picking like, I want the shark model.

1:13:06versus I like, you know, we're having a friendly divorce. I don't need the shark model. I'll take the more mild-mannered, reasonable one. So there's a paper that came out of Stanford about a month ago where they just had a very simple negotiation over just like some fixed sum of money and had models negotiate against each other. And 4.6 was just by far the most aggressive and refused to say no. So like, I could find the paper. And so like, there may be a case where you're like, end up paying up more in high stakes negotiations. but again like the the what's the right result is just what your canter party accepts and so there's just so much game theory about like when you use models when you bring a lawyer in and this is the bane of our existence and and uh this is what you pay a good lawyer for right like they get it done and they get it done for you like better than than you could have done yourself do you guys have strong opinions on the mythos getting pulled not not strong i i think the mythos thing obviously without commenting too much on the specifics is i think there actually is like a very simple way for the government here to do a good job regulating models.

1:14:09And that is very specifically create a data set that redlines model capabilities before they come out and make sure that this data set has a very wide range of coverage, whether it's cybersecurity, nuclear weapons, chemistry, physicists, and just have a very large set of expertise that sort of create these tasks that may be harmful and run this data set each time, which takes, you know, it can be very quick each time a frontier model is before release and dedicate sort of a recurring budget to updating this data set. Because at some point the models will, you know, saturate this benchmark essentially, just like they do every other benchmark.

1:14:56and if you just continue to update this and this becomes a source of truth kind of best redlining for final safety issues, this can become a really good way to allow frontier models to just not be delayed in releases because it could just be a very quick inference fraud. And that's all. I think we sort of keep it very simple at just this data set. This is the great idea. Now, if we're going to regulate these models, I think self-regulation is the exact way to do it. This list of tests should be done by AISafety.org, a consortium of the top 20 companies who each put a little bit of money in to fund this organization.

1:15:39This organization is responsible with AI safety folks to create this third-party certification, bioweapons, cyber hacking, harmful, what is the term for kids and porn and all this other stuff? CSAM. CSAM test, just all these different tests, revenge porn tests, whatever. And then when you have yours, it does it and it is responsible for doing it quickly and it gives you the certification. Hey, our 4.1 has been done. Okay, we now have 4.7. We need everybody's model to go through it. Boom. Everybody goes through the same safety testing. Everybody gets the same certification. And the government has access to that.

1:16:21It can talk to AISafety.org. Just like the MPAA exists in movies to rate them, that is not a government agency. The government agencies will F it up. No. They'll family show it up. It's run by the industry. The MPAA is the industry self-policing. That's exactly right. And if they screw up, Ryan, then it's on them. And then you could say, hey, the MPAA gave PG-13 to this movie, but kids went to it. And parents agree. Common Sense Media is another organization that kind of does this. So between MPAA, the independent common sense media company, you can kind of get to some ground truth. And there could be multiple organizations that do this, and you could subscribe to two or three different ones.

1:17:05Yeah, Ryan? Yeah, I think given just how fast things are evolving, you need experts to regulate something like this. And they will never be at a government agency in the same way, just given the speed of change. there's this very you know shaky prospect of if a model gets big enough and a company gets big enough is it going to be nationalized and i think like just given where things are geopolitically like the u.s government never wants to get anywhere close to looking like it's taking control or overseeing the way a model functions even as these intelligence capabilities get greater and so and i think like even something like this starts towing the line too closely so you know like self-regulating and you're totally right.

1:17:46If it trips the line, then the government get involved. But, you know, I think there's a great incentive for them not to trip a line. Yeah. And I think just one last point on this, I think it's a great idea, Jason, in terms of an organization that actually kind of involves a lot of these different frontier model companies, because the argument that you can make against this is each company is going to have obviously a sort of bias against their competitors and a bias towards their model is actually doing well on these data sets. And I think in this case, it's actually okay to have that because of course they can say that they don't, but I think there's just that innate bias that isn't actually going to exist.

1:18:20But that's okay in this design because if there are enough companies that are organized and they are sort of adversarial to their competitors, that actually is a good design because you're trying to create literally what's called adversarial tasks to get the models to fail and be jailbroken. And so if the competitors are, you know, by design being adversarial towards each other, that's actually the point of the organization. So that's, I think that allows for this organization to exist in kind of a good state. And maybe there's a government overseeing it in some way. But I think if there's enough companies that dedicate equivalent budgets in some way towards this, it could be a very nice self-regulating entity.

1:19:03Awesome. Well, thank you so much to both of our guests. An amazing show. Ali Ansari, it's micro1.ai is the website and micro.ai slash careers. If you want to go work for Micro One. Is there any jobs you're particularly looking for at the moment, Ali? We should search for it. We are hiring researchers in all three labs that we have, Realm, Robotics, and Cortex. So if you're a frontier researcher wanting to focus on the data stack and only the data stack, which we believe is the most important ingredient, then come join. And I can tell you, getting some equity as an equity shareholder, having equity in Ali, that's a pretty good bet.

1:19:45He's going to run through the walls. I would buy stock in Ali. With his juggernaut helmet. Ticker Simber, A-L-I-A. Yeah. Also, Ryan, Daniels, thank you so much for being here. Crosby, legal is the company. Crosby.ai is the website. And Crosby.ai slash careers, if you'd like to go work for Crosby, what positions are you looking for right now? Great lawyers, all big law firms is our bias. And MLO engineers who are interested in this problem of automating a highly subjective, high impact career profession. Another fantastic. No plug for me? No plug for me, Lon? I get no plugs on my own show. Here's the plug for me.

1:20:24Jason, what's your website? Who are you hiring for right now? aside from a new co-host, aside from a new AI co-host. Yes. No, you're a great broadcaster, Lon. I just told you in the group chat, great broadcasting. I appreciate that. We have an associates in training program we run every summer, AIT. We don't have a landing page for it. We do it every June at graduation. We select them in April and May. And if you wanted to join the launch team, we do have some open positions there on the website. What is the website? Careers.launch.co. Thank you, Ali, for plugging me. You can see them there, but we're going to make an AIT.

1:21:01So ask them to create launch.co slash AIT to make an associate in training. And, you know, then people could sign up at any time for that. So hopefully we'll have that landing page open before this episode comes out. I'm sure. Great job, everybody. We'll see you next time. Bye bye. Bye, everybody.

From the publisher

SpaceX bought Cursor for $60 billion. Satya Nadella says companies need to stop relying on third party AI models and build their own “token capital.” The focus is shifting from LLMs to the application layer sitting on top of them. Here to unpack what that means are guest experts Ali Ansari (Micro1) and Ryan Daniels (Crosby).Plus we get a sneak peek at their new contract redlining benchmark, a crucial eval for how well LLMs don’t just answer questions about the law, but demonstrate actual legal reasoning.Timestamps:0:00 SpaceX acquires Cursor7:36 Distillation vs. building your own model19:43 Nadella's "Frontier Without an Ecosystem"30:21 AI in the courtroom32:15 Ando: the intriguing new workplace tool1:05:38 Inside Micro1 and Crosby's new benchmark1:07:14Guests:Ali Ansari: https://x.com/aliansarinikMicro1: https://www.micro1.ai/Ryan Daniels: https://x.com/ryanjdanielsCrosby: https://crosby.ai/Relevant Links:Bloomberg: “SpaceX acquires Cursor for $60B”: https://www.bloomberg.com/news/articles/2026-06-16/spacex-cements-60-billion-deal-to-take-over-ai-startup-cursorQuinn Thompson “brilliant corporate finance” post: https://x.com/qthomp/status/2066859672749977988Business Insider: “Inside Cursor’s Wild Rise”: https://www.businessinsider.com/cursor-ceo-michael-truell-spacex-elon-musk-anthropic-2026-6Satya Nadella: “A frontier without an ecosystem is not stable”: https://x.com/satyanadella/article/2066182223213293753Joshua Browder’s Do Not Pay: https://donotpay.com/Harvard Magazine: “AI Outperforms Doctors in Emergency Room Tasks”: https://www.harvardmagazine.com/ai/ai-outperforms-doctors-diagnosis-harvard-studyJoshua Kushner “Long Humans” post: https://x.com/JoshuaKushner/status/2065093542809092465Ando: https://ando.so/Nim Ravid on X: https://x.com/Nim_Ravid1Subscribe to the TWiST500 newsletter: https://ticker.thisweekinstartups.comCheck out the TWIST500: https://www.twist500.comSubscribe to This Week in Startups on Apple: https://rb.gy/v19fcpFollow Lon:X: https://x.com/lonsFollow Alex:X: https://x.com/alexLinkedIn: ⁠https://www.linkedin.com/in/alexwilhelmFollow Jason:X: https://twitter.com/JasonLinkedIn: https://www.linkedin.com/in/jasoncalacanisCheck out all our partner offers: https://partners.launch.co/Great TWIST interviews: Will Guidara, Eoghan McCabe, Steve Huffman, Brian Chesky, Bob Moesta, Aaron Levie, Sophia Amoruso, Reid Hoffman, Frank Slootman, Billy McFarlandCheck out Jason’s suite of newsletters: https://substack.com/@calacanisFollow TWiST:Twitter: https://twitter.com/TWiStartupsYouTube: https://www.youtube.com/thisweekinInstagram: https://www.instagram.com/thisweekinstartupsTikTok: https://www.tiktok.com/@thisweekinstartupsSubstack: https://twistartups.substack.com

More from This Week in AI

All 34 episodes
Why AI Models Aren’t the Product Any MoreThis Week in AI · 1 h 21 min
Listen in VO