OpenRouter: from Seed to Stripe — with OpenRouter’s Alex Atallah & AMP’s Anjney Midha

25 Sep 2026 · 1 h 21 min · 31 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

OpenRouter’s approach to model “pub-sub” and routing—why a marketplace-like layer is needed to make many LLMs usable, discoverable, and swappable for both developers and agents.

Guests (backgrounds)

  • Alex Atallah: OpenRouter co-founder; early community contributor; previously worked in systems/flywheel-building roles (mentions OpenSea/marketplace context).
  • Anjney Midha: investor and operator; previously head of platform at Discord (crypto/NFT launch security, phishing/social engineering/DDoS defense); later led series into Mistral; enterprise deployment and safety/security experience.

Key claims

  • LLM consumption is continuous and SKU-changing, so inference needs a pub-sub-style model subscription plus routing/orchestration, not just static SDKs.
  • “Monopoly” concerns are overstated; scaling creates many model companies, but labs often lack developer-focused distribution/plumbing.
  • OpenRouter adds strategic value beyond “a wrapper”: model discovery, versioning, provisioning, routing, and production reliability.
  • Enterprises need control/management because closed-model guardrails can refuse valid use cases.

Notable examples

  • Alpaca (synthetic-data fine-tuning of LLaMA) as the “$600” proof that models can be monetized as services.
  • Discord AI bots: Clyde setup help and moderation; OpenAI refusals drove demand for open models and orchestration.
  • MidJourney/Stable Diffusion communities on Discord as a template for feedback loops and “home base” UX.
  • Mistral “Mixtral” price/performance competition as an early marketplace working example.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Exploring OpenRouter and Product Principles

0:45 to 2:12

Discussion on the PubSub principle and its application to OpenRouter.

“And marketplaces are an easy, easy, easy example of this.”

The Rise of Alpaca and Model Development

2:12 to 3:45

The impact of the Alpaca model and its implications for new business models in AI.

“This is something that was more consensus now, but not consensus when you guys started, which was that there is such a demand for swapping models and changing things out and that people would not use the native SDKs.”

Creating a Marketplace for AI Models

3:45 to 6:04

The discussion centers around the need for a marketplace to manage AI models.

“one, we have a whole new way of monetizing data for the first time.”

Personal Journey: From Stanford to OpenRouter

6:04 to 8:08

Anjney's recollections of his early interactions with Alex and their shared history.

“And then Anj, no stranger to wanting more model diversity, at the time, you're a couple of years into your anthropic journey, which we've covered in the previous podcast as well.”

Challenges in Content Moderation and Model Control

8:08 to 11:52

The complexities of moderating content on Discord and the need for open models.

“It was at Old Union, if I remember correctly.”

Initial Struggles and VC Perception of OpenRouter

11:52 to 14:01

Discussion on the initial struggles faced while raising funds for OpenRouter.

“because if you outsourced it to the labs and they controlled the guardrails and their guardrails, their safety policies forbid the model from responding to your prompts, that was quite catastrophic.”

Exploring the Limitations of LLMs

14:01 to 14:54

Discussion on the refusal of LLMs to assist in certain narratives and the potential for model variety.

“the LLMs would just refuse to help with that part of the story and then these would be like okay this is not like structurally inherent to LLMs.”

Challenges in the Early Days of OpenRouter

14:55 to 16:04

Recollections of challenges faced while raising capital and the skepticism from VCs about scaling laws.

“Well, I was going to say that the biggest objection we got is big model win, which is all the value.”

Understanding the Innovation Landscape in AI

16:05 to 17:40

Discussion on the innovation across different models and the importance of competitor diversity in AI.

“So it didn't seem like, you know, would be a really crazy outcome if that happened.”

The Developer Experience and Distribution Challenges

17:41 to 19:48

Insights into the lack of distribution strategies for AI models and the need for improved developer experiences.

“the research teams were fantastic at figuring out how to reason about new capabilities.”
Show all 31 chapters

Marketing and User Engagement for AI Models

19:49 to 21:40

Exploring the marketing strategies for AI models and the importance of user engagement in their success.

“With Google, DeepLine is done training a new checkpoint and then they push a button and it gets blasted out across all their surfaces from Google Docs to, you know.”

The Misunderstandings of VCs About OpenRouter

21:41 to 24:18

Discussion on the misconceptions some VCs had regarding the value of OpenRouter as merely a marketplace.

“We are a neutral layer looking at this market like it's a big, dark room with all the corners completely obscure to users.”

Building a Sustainable Ecosystem for AI Models

24:19 to 28:00

The conversation shifts to the strategies for building a robust ecosystem around AI models and the importance of operational experience.

“at the scale the OpenRouter team had started just doesn't happen by default.”

The Axie Infinity Community and User Experience

28:00 to 32:58

Learn how building community engagement was crucial for Axie Infinity and OpenSea.

“But we initially connected when you were at Discord and we talked about the Axie Infinity server.”

Lessons from Early AI Applications

32:58 to 40:00

Explore the evolution of AI applications and their community-centric development.

“that very few other marketplaces have been able to achieve over the last five years.”

The Role of Open Router in AI Development

40:00 to 42:00

Understand the need for OpenRouter in providing APIs for AI creativity.

“And of course, there was the crazy distribution that you enabled for a lot of these developers.”

Emergence of OpenRouter and Developer API Needs

42:00 to 43:08

Learn how the launch of Stable Diffusion sparked the need for OpenRouter, enabling developers to create their own applications.

“And shortly thereafter, Stable Diffusion launched.”

Mistral's Impact on AI Pricing and Performance

43:08 to 45:10

Discover how Mistral's introduction affected AI pricing and performance perceptions in the industry.

“The shape of Open Router enabled is that.”

Comparisons Between Arena and OpenRouter

45:10 to 49:54

Understand the distinctions in mission and target customers between Arena and OpenRouter, despite external perceptions of similarity.

“a few days after Mixtral came out at NeurIPS.”

The Importance of Focus in AI Development

49:54 to 52:44

Explore why maintaining focus is crucial for success in AI businesses and how it influences brand perception.

“Whereas the highest expectation customer, from my perspective, that Alex really understood and the mission was to serve was like a developer, right?”

AI Pair Programming and Initial Goals

52:44 to 55:09

Learn about the early goals and initial focus of Anthropic on AI pair programming and how it shaped the company’s trajectory.

“like people think that the early days of Anthropic were like super easy because they were on the GPT-3 guys who left.”

Exploring Unlaunched Ideas and User Needs

55:09 to 56:00

Delve into the ideas that were prototyped but ultimately not launched, and the ongoing challenges of meeting user needs.

“Were there other ideas that you wanted to pursue that you turned down?”

Exploring AI Use Cases and Mentorship

56:00 to 58:08

Discuss various AI applications including personal practice and mentorship scenarios.

“And it was like growing and we were building more conviction in it over time.”

Fine-Tuning Models and Market Positioning

58:08 to 1:00:44

Delve into the importance of fine-tuning AI models and positioning within the market.

“I mean, we decided, really we like leaned into our focus and figured that like there are, like we just saw the ecosystem develop over time.”

The Evolution of Fusion Technology in AI

1:00:44 to 1:02:40

Examine the challenges and advancements in AI model fusion technology over time.

“And in our case, the technology was a little too early.”

OpenRouter's Growth Milestones and Ecosystem Dynamics

1:02:40 to 1:05:29

Outline key milestones in OpenRouter's growth and the dynamics of its ecosystem.

“Yeah, and it came on your fable so you were like, this is fable level.”

Fraud Prevention and Trust in AI Ecosystems

1:05:29 to 1:10:01

Discuss strategies for preventing fraud in AI applications and working with partners like Stripe.

“It was like a swing action where like model labs would come up with some sort of frontier innovation.”

Understanding Fraud in AI Inference Market

1:10:01 to 1:12:09

Explore the challenges in the AI inference market related to fraud and safety.

“And we help them like regain control and detect it.”

The Growth of Token Economy and Fraud Risks

1:12:10 to 1:14:59

Learn about the exponential growth of the token economy and the subsequent risks of fraud.

“And ones that like are not focused on just adding a markup on top of inference.”

The Need for Security in the Token Economy

1:15:00 to 1:18:36

Understand the necessity of new security measures to combat fraud in the token economy.

“Where we are today is roughly there on tokens, but over the next even five years, we're expecting the token economy to get to like roughly$5 trillion.”

OpenRouter and Stripe Partnership Insights

1:18:37 to 1:19:51

Discover how the partnership between OpenRouter and Stripe aims to enhance trust and safety.

“is like a huge percentage of that I think is going to be fraud, abuse.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:02Okay, we are here in Anj's house, which is where all great startups in San Francisco start. Howdy. And congrats on Cursor, Mistral. God knows what else, you've got so much stuff going on. There's a lot going on. Well, OpenRouter has probably been the most, I would say, one I'm excited about recently. Yeah, and we have Alex, first time on the pod, but you've been in the EIE a few times. I appreciate every time you've shown up for the community. Congrats. I just like, what a journey. When I was looking back at your past posts, one of the earliest principles that I saw you write as a sort of product person is PubSub as a product principle.

0:42And I wanted you to maybe explain how you think about what should exist in the world. Yeah, the PubSub piece, which was early 2023, I didn't think about it until we talked like 10 minutes ago, is about how there is like a way of thinking about products as an intersection between subscribing to data and publishing data. And marketplaces are an easy, easy, easy example of this. You have suppliers that are publishing some kind of product to a skew. And the skew is kind of like a pub-sub topic that a consumer is subscribing to and just going to consume whenever they want. And humans consume in a very, like, discrete, ad hoc way.

1:28It's not very scalable. You know, all their attention is on the topic when they're buying the thing, and their attention is nowhere else when that happens. Agents and consumers of inference don't act like that. They're consuming continuously, and they're changing the SKUs that they consume from all the time. So open router is sort of like a blend between a normal API experience and a marketplace where we create model slug. We have the auto router. We have all kinds of like product SKUs that you can subscribe to. And then you can like continuously add like derived value and make decisions based on those consumers.

2:11Yeah. This is something that was more consensus now, but not consensus when you guys started, which was that there is such a demand for swapping models and changing things out and that people would not use the native SDKs. I guess for each of you, what was your sort of realization moment that this would be it? You've given a talk at AIE about Alpaca as one of your inspiring moments. Alpaca, I can like rehash the Alpaca moment for a sec. like the very beginning at the end of 2022, OpenAI was the only game in town. There was like OpenAI, Cohere, and then a smattering of like early attempts at open weight models.

2:55When WAMA came out in January of 2023, it was like, whoa, really exciting. This is really big. It outperforms GPT-3 on one or two benchmarks. But you can't chat with it. It wasn't like an, it wasn't actually an engaging model. But it seemed like someone just needed to fix a couple of things and do some RLHF on it to get it all the way there. And Alpaca was the first model that I saw that did that. It only took$600 to do. A team at Stanford generated a bunch of synthetic data, fine-tuned Llama, and made Alpaca, 7 billion parameter model. Maybe it was 13 billion parameters. And it was so good. Like, I was just, like, on an airplane using it.

3:37But in many cases, I could not discern a chat GPT versus an alpaca result. And I figured if it was this easy to make a model, one, we have a whole new way of monetizing data for the first time. You can just take really valuable data and turn it into a service in$600. And that cost will probably go down over time. When you say monetizing your data as what eventually will become an MCP endpoint, or as a training data for a model? Yeah, training data for a model. An abstract way of saying, hey, I have this data. Compress it into a model. It makes sense for me in my product, but I could repackage it in the form of a model and sell it.

4:24And so it's just a whole new business model for the economy. It also, of course, provides a way of following what Frontier Labs are doing, but in a way that a single developer or a small team of developers can roll on their own. And so whenever you have an example of that, like a breakout app that's doing really well, and then some kind of framework for imitating it in your own flavor, you have an immediate ecosystem. of like an immediate ecosystem like should arise because there's just a huge gap between the like decisions that the single company is making and all of the variations in those decisions that like a wider ecosystem can create themselves.

5:13And so then, you know, you need a marketplace to like discover all of those services and all of those products. There wasn't any place on the internet that like was like a home base for LLMs in terms of seeing how much they were being used and seeing who was using them and why. The closest would be Hugging Face. Hugging Face was the closest. They just started Hug like a few years ago before that. Yeah, and Hugging Face also didn't have the closed source models. Yeah. And they didn't, you couldn't use the models at the time. And there wasn't data about who was using that. There are like a bunch of differences between Open Router and Hugging Face.

5:49And those differences felt really critical to me, especially when I was just trying to learn about LLMs. and why people are choosing these different little ones that are emerging over time. Got it. And then Anj, no stranger to wanting more model diversity, at the time, you're a couple of years into your anthropic journey, which we've covered in the previous podcast as well. What was your introduction to Alex? Well, the introduction was, I think, 13 years before that. Oh. But the open router handshake actually happened right over there, if you remember. Yeah. which was Alex and I met I believe a sophomore now if I remember at a Stanford Review meeting for the first time I think so yeah so Stanford Review was the libertarian newspaper on campus at Stanford that Peter Thiel started back in the day and for whatever for whatever reason Alex and I both showed up to one of the meetings and I remember the editor-in-chief was a mutual friend of Boris Lisa was really a really great editor-in-chief.

6:54Part of an editor-in-chief's job is to assign responsibilities to people and make sure the work gets done. And I may be misremembering the details, but I remember wanting to, it was kind of surprising to me that at the time there was no dedicated technology section in the newspaper. Because it's political, right? It originally started as a political... States and all those things. To take us back in time, you may remember this, but there was this technology kind of legislation that was being debated called the Net Neutrality Act. And net neutrality is like inherently this political concept, right?

7:33It's about the regulation of internet broadband access. And so there was a community of us who were kind of technologists, but also debating the politics of the technology. And I thought the review would be a great place to like write about that. And I was working on, I think I had a net neutrality article and I remember proposing, well, maybe you should start a technology kind of section. And Alex was one of the only people who said, yes, that would be cool. And I forget whether we ended up writing stuff together, but that's when we first met was 2011 or 12. I forget which year it was. It was one of those.

8:08Yeah. It was at Old Union, if I remember correctly. That's where we used to meet. But along the way, Alex and I have had a chance to hang out often. And probably the time when we had the most professional overlap was when I was running the platform at Discord and it had become this explosive kind of platform for crypto and NFTs in the middle of the pandemic. Which also, by the way, you were in charge of safety and security as well, right? I was the head of platform, which meant all of the crypto, DAO and NFT launch security debugging fell on me. And the phishing. The phishing, the social engineering attacks, like a ton of DDoS that we were getting hit by.

8:53It's around the time I started teaching security at Stanford, CS153. And Alex was at OpenSea at the time, and I was trying to figure out how we could defend against all these attacks that we were, like, and at peak, I forget if you remember how much NFT volume was running through Discord, but it was like a meaningful amount of like, it was like several billion dollars in NFT volume of GMV, so to speak, were running through the platform. and it was all coming from OpenSea. It was these like buy, sell, trade servers. I mean, the D in DAO is Discord. Yes. So that's when I think we had hung out professionally.

9:28But a year after that, OpenAI gave Discord early access to GPT. Sorry, GPT-3. No, it was GPT-3.5, actually. GPT-3.5, which is the RL version of GPT-3. And that's around the time we made a Discord bot with OpenAI for internal deployment. and that's when I realized we would need, like since I was part of the deployment team. What was the use case? There were two that were, and there's actually a post now called Discord is your place for AI with friends that somebody sent me recently that I wrote and published in 2023. But there were two use cases. One was Clyde, which was like a first party friend inside of Discord that could help you set up your Discord server and talk to you about onboarding and get your friends to hang out more.

10:16And then there was content moderation. And one of the realizations we had with content moderation was it would refuse to moderate. Like, we just refuse our prompts because the RL, the post-training was, we were very early in the post-training era and it would just, our prompts would trigger it. It's like guardrails. And we told OpenAI, hey guys, we need access to the weights because if we're going to be doing content moderation at scale, we have 250 million monthly active users, we need more reliability that the model will do what we need it to. And they said, well, sorry, guys, that's not how this works.

10:50We're a closed source company. And so that was my first realization that we needed open models and the enterprises would need more control over capabilities and then ultimately would need some kind of control plane or management system to orchestrate these open models. But there were no good open alternatives until maybe six months later when Llama came out. And six months after that, I led the series into Mistral, which was started by Guillaume and the Llama team. And around that time is when I remember hearing about Alex launching OpenRouter and going, these worlds are going to collide. And I don't know when it'll make sense to team up, but Alex was so early and could see, I think he was totally right about this ecosystem starting with Llama that then needed like an easy layer to manage for it, especially for, I was approaching it from the enterprise perspective because I had been that, like as the VP of platform at Discord, it was my job to ensure that when we deployed models to like 250 million users, they did what we wanted them to.

11:51And that was very hard because if you outsourced it to the labs and they controlled the guardrails and their guardrails, their safety policies forbid the model from responding to your prompts, that was quite catastrophic. Yeah, but a moderation is a thing that they want to support. And obviously, beyond that, OpenAI would work with you, presumably, to give you a moderation endpoint, which they offer for free. It was an interesting use case. So they did give us a moderation endpoint. However, as you guys know, every Discord server is like a mini deployment of itself. And so the use case was, instead of having human moderators that have to interpret the norms of the community, you just give the, they often, like every subreddit, Discord servers, public ones, have their own rules that the user, the users create.

12:41Oh yeah, you run the links based Discord, yeah. And then humans used to read those norms and then enforce it every day manually, like observing each message in these communities. And these communities have like millions of users. So we had a 5 ,000 plus person team globally on the Discord content moderation team. These are outsourced contractors who had a really tough job. And so the idea was instead, if you could give the norms of that server to the LLM, then the LLM would do custom moderation for that server. It's almost like a, like in context moderation for that server. And many of those servers norms just violated OpenAI's rules.

13:19And so that, it was like, we had our own custom eval. So each server had its own custom eval, but at the time, OpenAI's evals, we were all so primitive in our thinking about how to deploy these LLMs that often the post-training prompts were super heavy-handed. It said, oh, anything about Harry Potter, anything that has trademarked content, you know, don't refuse. And if it was a fan, Harry Potter fan community, this is a real use case that had a content moderation, the LLM would just refuse. and that was just not precise enough. Another one that we heard was like if someone was trying to write like a detective story and there's one chapter with a lot of violence like maybe someone kills someone the LLMs would just refuse to help with that part of the story and then these would be like okay this is not like structurally inherent to LLMs.

14:13There must be like some choice out there so that I can like switch to another model when I'm getting a refusal or a bad result from the main one that I have. And that tension also drove me for a marketplace. Yeah. I think that is well accepted now. What was it like back then when you were raising or starting this? Did people get it? What was some of the struggles? Basically, I like getting stories out of him about how other VCs don't get it. So anything you want to talk about, now that the early journey of OpenRouter is done. You can obviously talk about some of the early day stuff. Well, I was going to say that the biggest objection we got is big model win, which is all the value.

15:04Scaling laws. Yeah, scaling laws and natural network effects are just going to kind of accrue to one company, which will be like, it'll be a Google-style monopoly, just like how Google won the search market by a large, large margin. And you'll just be fighting for scraps at the end, basically. That was probably the biggest objection we got. You know, it is interesting that Google won the search engine race with such a huge margin. You know, I think like, had there been more interesting benchmarks or had like search engines been, you know, had people like seen them a little bit more like LLMs where there are services that you can build companies on top of, that might not have been the case.

15:50But LLMs don't merely have a user interface. They're also like ways of building entirely new businesses. And, you know, a Google-level monopoly would be like the Dutch East India Company times a quadrillion in magnitude because the whole economy ends up like depending on the one monopoly as well. So it didn't seem like, you know, would be a really crazy outcome if that happened. And it's also less likely because the economics of creating good competitors are much more decentralizable. Everything Alex said is true. And I came at it from a completely different perspective, which is... Yes, this is why we're here.

16:32The scaling laws were never like, in my mind, were always a feature, not a bug, for why OpenRouter would be very valuable. Because I was one of the first investors in Anthropic and it was obvious to me that other researchers in our friends group, I went to grad school for machine learning and I just had a lot of friends in the ML community who it was very obvious to us that the bitter lesson holds. And so I was like, oh, fantastic. Now we have at least two proof points that compute scaling works. It was OpenAI and Anthropic. And by the time I think we decided to team up on OpenRouter, I'd already invested in Mistral and Black Forest Labs and Luma.

17:09So there was multiple model companies and teams that I was working with. But you did other modalities, whereas this is literally text. Exactly. And it was so obvious to me that an ecosystem of different kinds of models were being created. And that this whole narrative of like only one company will dominate, like Google, was like maybe true. But one, I don't believe that. been too, there was so much extraordinary innovation happening across several different research teams. But the shared problem I was noticing across all of them was often, you know, the research teams were fantastic at figuring out how to reason about new capabilities.

17:48They think in terms of capabilities, but never, like are not developer mindset oriented. Like what happens after the training is done and the checkpoint comes out, like you'd be shocked how, how like similar to their early pre-chaining teams at Anthropic, BFL, Mistral, were in their default approach to taking their research out of the lab and scaling their impact, which was often, oh, the checkpoint is done, put it out as an API, done, and then there'd be crickets. In the case of Claude, the first Claude checkpoint was actually done a year before they released it internally. And then ChatGPT came out and we decided, okay, yes, it's a good idea to release a Cloud version externally.

18:33And they had no plan. Like no plan for how to get developers to actually try it out. And so if you go to the Cloud One blog post, you'll notice there are like three kind of developer examples for users of the API. And one is a Discord bot, and the second is Vivian, my wife's startup called Juni Learning. And then there was like Notion, because these were all friends of like the Anthropic team, because that's how like last minute the planning was around, hey, once the model's done training, how do you get it out to the world? There was no distribution platform that understood what developers needed, all the key management, provisioning, like simple endpoint management, versioning control, like all these things that scientists and researchers go, I mean, that's plumbing, I don't really think about it.

19:16And instead, Alex came at it from that perspective. And so, it was so obvious to me that every single lab I was funding would spend literally sometimes billions of dollars into training and then a checkpoint would be done and there'd be crickets during early access because they were like, oh, that's right. It's hard to use a checkpoint to make anything. You actually need a whole bunch of plumbing around it to make it usable by a developer. And so by the time, I think we think it was so obvious to me that a distribution platform like Open Router was critical to have in the ecosystem if we wanted there to be competition to Google.

19:51With Google, DeepLine is done training a new checkpoint and then they push a button and it gets blasted out across all their surfaces from Google Docs to, you know. Everywhere, even if I don't know. Everywhere, like on Android. Like overnight, they can deploy a new checkpoint to like a billion devices, right? And that invisible infra advantage, distribution advantage, most people don't realize, but until Open Router showed up, you had to think about all of that yourself as a model lab. And it was very daunting. You know, at Anthropic, I think it took more than 12 months to get to our first 10 million in revenue.

20:23and in contrast with Blackforce Labs, I remember the early days, you guys had a conversation with the BFL team and it was so simple for OpenRider to say, oh, no problem. Like the day you launch, we can send a million developers to you. You know, that was crazy. That was like a step function change in like power. Is that a real number, million? I think today it's like four million. How many developers are on OpenRider today? Over 10 million, but it's hard to... Yeah, I don't know how to... We do a lot of account deduping work, but no one... If you could get 1 ,000 developers, just to put it in context, if you get 1 ,000 developers who actually try the model on day one after you release it and just do inference and give you feedback, that's 1 ,000 more developers than they knew how to get to on their own.

21:19Well, BFL had a reputation, but yes. They had one with stable diffusion. And with Mistral, I don't know if you guys remember, but the first checkpoint they released was like torrents. It was like torrent weights. Yeah, they just put up a magnet link. There was no API. Because they weren't in for people. It's like, okay, download these weights and you guys go through. He has a story on his side, yeah. Yeah, I mean, in addition to the building a really good developer experience around it, The marketing that we do for different models is totally different and perceived totally differently from the marketing that a model lab does for itself.

21:56Yes, 1 ,000%. We are a neutral layer looking at this market like it's a big, dark room with all the corners completely obscure to users. And users are walking into the room and feeling around and trying to figure out what objects to grab off the tables and build into their companies. It's an insane way of working. Models are not products where you can just enumerate all their features onto a web page. They're all black boxes, including the open-weight ones. So you need to shine lights on all corners of this room so that people can see what makes this model good. And you need the company shining that light to be a neutral third party, which is what we specialize in.

22:41So in addition to developer experience, there's also a very important marketing and product packaging component and a way of routing and discovering models becomes critical to your go-to-market as a provider or a model lab or a server tool and more in the future. And this value, to your earlier point, about how many VCs just don't... One of my biggest frustrations is that But venture capitalists, many of them just don't have any operating experience in the field. So unlike a traditional investor who's just maybe come up through the ranks as an associate working on financial modeling or maybe hasn't been a real operator in the field for more than 10 years, which is a big part of the industry now, I had just arrived at A16Z a year after running the platform.

23:33And so I knew what the challenges were of building a real great developer experience and actually being able to create a working piece of software with a model. And there were a few, I won't name names, but there were investors who were looking at Open Router and felt at the time, when I would compare notes with people, that it was just, I quote unquote, just a marketplace. Yeah, just a thin layer, just a proxy. a wrapper or whatever on other people's APIs. And I was like, you have no idea how strategic the value that OpenRouter has created by being able to orchestrate even three APIs in production.

24:12The amount of both engineering work and community design that goes into getting that actually live and running in production at the scale the OpenRouter team had started just doesn't happen by default. And that was one of the things that stood out to me about Alex from the earliest days. He just understood, from a systems perspective, how do you get these flywheels going? That stood out to me with OpenSea when we were working together on the NFT integration at Discord. Alex had a level of systems thinking on how you get these flywheels going that most scientists and machine learning people just don't think of.

24:49We often think in terms of pre-training, mid-training, post-training. It's a linear stage. It's this linear pipeline. It's no loop, yeah. It wasn't until much later that the modern context feedback loop cycle really got standardized in the industry. But at the time, if you remember, machine learning was like, mostly we did a lot of ML like when I was in grad school on a laptop. So you just like download a data set, ran some ablations, and you looked at the loss curves and you're like, great, I made AI. And the idea that you have to like deploy those capabilities, collect feedback trajectories, then like put those into a continuous loop came much, much, much later.

25:23And it was very counterintuitive to the traditional AI mindset. I do remember doing the investment phase for Open Router. I just didn't try and re-educate a bunch of other VCs on why it was not just a marketplace. I was like, you know what? I'm just going to invest. And I'm going to take the opportunity to partner with Alex. And if no other VCs get it, that's totally fine. Because at the time, it was not obvious, I think, to several other investors that Open Router was not more than just a wrapper around APIs. That infuriated me. I was like, you know, I don't have time to debate you. I'm just, we're going to invest.

25:59And then I think like a month later, Matt Murphy marked it up by 10X. I forget what the exact post money was and so on. But to his credit, Menlo Ventures realized, okay, there's actually much more strategic value here as well. Maybe you didn't hear all these conversations behind the scenes. But that frustrated me a lot. There was a lot of this opining about rappers. And if you're like, oh, an app is just a rapper on a model, then like, and OpenRouter is like this wrapper on top of other APIs. This is the most stupid reductive framework. So it's clearly somebody who has no experience deploying products.

26:34It's the thing you dismiss other things with. Like everyone's a wrapper on everything, right? Like if there's some point, some wrappers have value. I mean, investors are wrappers on LPs, right? Like metric capitalists. So, I mean, yeah, it's all wrappers all down to bare metal, I guess. When I started the whole AI engineer, I guess the coining in 3.23, that was the number one pushback, is that this is no value, you should actually just train models. And yeah, I mean, obviously this is like, you guys are one of the testaments to the fact that you can actually build very valuable wrappers, but also very valuable model companies.

27:05It's so hard to be, like the day a model launches, the fact that you have an open router endpoint for that model frequently at the top of Hacker News on day one, people don't realize the amount of work that goes into accomplishing that. And OpenRouter used, like, that would happen over and over again. And I remember going, people have no idea how hard that is. You know, that's not... Yeah, we've covered some of the inference engineering that goes behind some of the... Yes. With Base 10 and all those. Well, today you have, you know, all those, like, cool code name things where people guess what OxyAlpha is and all those things.

27:42But I guess one of the things that you're teasing is how do you get that initial flywheel going, right? Because today you have your scale and your reputation, all these things. So obviously you drive immense distribution. But when you're early on, when it's mostly... The bootstrap, yeah. What is the bootstrap like? I mean, to bring it back to early Discord days, I think we initially connected with... This is an OpenSea story, technically. But we initially connected when you were at Discord and we talked about the Axie Infinity server. Oh, yes, yes. This server was like the biggest server at the time at Discord.

Read the full transcript

28:18That's right. And you were kind of like constantly bumping up the limits. The limits on the server. Oh my God. For those who don't know, like 10 % of the Philippines was on that server. It was like a meaningful contributor to the GDP of the country. It was an NFT like crypto game. It was like a Pokemon breeding thing. Yeah. Similar. Yeah. There was battling. There was breeding. And then there was like a marketplace for trading. Play to earn as well. Yeah, play to earn. And like the graphics were really cute and fun. And you kind of like, you know, you get kind of emotional about your Axie that you make.

28:54So to like start a community like that, which we had to do many times at OpenSea with basically every early project for us to create a marketplace for it, we need to make sure that like the community actually wants it. And it's kind of like building something that people want and going and telling them about it. Like you can do that on a one-on-one basis, but as a way higher leverage to do that in a community where everyone can talk to you at the same time. So we spent a lot of time like building things that the community really wanted. We did the same thing for Open Router. And, you know, like the Axie community was one of like a zillion communities we did that with.

29:34And Ange like saw us doing it. And because you could just see people sharing OpenSea links constantly in that Discord. Like users sharing links is a really clear indicator that like something important is going on. So we spent, you know, a lot of time, like first figuring out what the gap is in the technology that people care about. Like what was the actual problem that needs to be solved? You know, in early LLM days, It was, you know, OpenAI refusing to finish the prompt or to complete the task. It was also, you know, inability to customize models. And so there are communities that are just completely blocked on that issue.

30:21And those are the communities that are most useful to sort of learn about and dive into and explore. Something that really struck me at that time, as I was just hearing your talk, I remember noting how, you may not remember this, but we were, we had these like working Zoom calls that we were doing a sprint around for like this OpenSea integration with Discord. And, you know, we'd get, it was myself, my engineering team, I think you were there. And I remember, you know, Alex in the middle of one of those calls, just like, there was like silence. You know, we were, we were all like, oh yeah, this totally makes sense.

30:59Let's do this. And then there's some, everybody aligned. And Alex was like, no, this makes no sense to me. And everyone's like, I remember going, what? Like it works. Like you click on a link and then it bounces you out to like OpenSea. And he was like, it's not a good user experience. Yeah, we should not do this. And I remember going, you know, he was the only person out of all of us to actually raise his hand and go, yes, it made sense from a technical implementation perspective. Like we are bouncing the user out into OpenSea. and so it kind of checked the box of the product manager's requirements on both sides but Alex went one step further and was like you know what would be better guys if we just embedded the experience right here inside of Discord so the link opened up as an embedded iframe and you can just check out right there and not one person on the call there were like seven of us who had met like you know week after week and it's the guy who doesn't work for Discord and it's the guy who doesn't work for Discord like technically you benefit if they bounce exactly and that was like adversarial to keep the user inside of Discord would be adversarial to OpenSea.

32:04And yet Alex put that user experience first. And I was like, that's special. Because it's very hard to have somebody who's technical like Alex and understands the developer flow, but also understands the best user experience and wants to prioritize that. And that's two sides of the flywheel that you can get spinning, like is often hard to stop. And you just reminded me, like that one was one of those moments where I go, I realized I got to be better at user experience because I should have been the one who came up with that. And I didn't. And I learned from you. and I think that went into one of our case studies for the PM training program at this point.

32:35I don't know if it's there. You need an Alex, this is the conclusion. Yeah, you need an Alex. And this is why nobody should be surprised why Stripe decided they had to buy OpenRouter because it's a really rare combination of people who understand the machine learning community, the developer experience, and the end user experience. And putting all that together has resulted in this extraordinary scale that very few other marketplaces have been able to achieve over the last five years. Yeah, well, we should talk about the other reasons for acquisitions, which you've written about. I want to sort of proceed somewhat chronologically as well.

33:11So there is a point that, you know, one of the questions that Dave from HF0 sent in was when did you know it really started to work? And you brought up Mixtraw. I don't know if you want to bring up that story. Oh, yeah. Which obviously you overlap with, so. Yeah, the MOE was, I don't know when I mean there's no like one moment where I was like oh this is you know officially starting to work it was like moments of kind of increasing connection like really super early on oh yeah but well yeah so before Open Router I wanted to like explore a bring your own model experiment and which anyone familiar with crypto is like you know Phantom and all these things yeah yeah so it felt like doing a MetaMask analogy for AI would be kind of a fun way of exploring that.

33:58And at the time, there were no AI apps. There were probably as many AI apps that were like hitting an LLM via an API call as there were like games just doing it in JavaScript. Basically, like there was a moment in time where it could have been the case that web apps call LLMs through the browser, like through some kind of desktop managed app that is controlled by the user. And of course, there are like, I think many reasons that that did not happen. But back when the days were that primordial, I built a Chrome extension called Window AI. And - With Plasma. With Plasma. I had come across early on and I was like, who's going to actually use this?

34:48You did. Plasma had a couple, like, I think Phantom was using it. There were some other like real companies. There was like a shim, basically. React for Chrome extension. It compiles to all these. Yeah, kind of like Next.js for Chrome. Next.js, Next.js. Okay. And yeah, built Window.ai on top of it. The creator of Plasma like started contributing code to Window.ai in GitHub. And that turned out to be Louis Vichy, who is the co-founder of Open Router. You have told me this is how you met Louis. Yeah. Okay. So that allowed users to kind of like configure which model they wanted to use for a web page in their browser.

35:27And then like the app would just call out to that model when it needed to do things. You know, not the right form factor for LLMs. But, you know, it's like fun experiment. You learn a lot. And like, I, you know, open sourced it. And, you know, the main learning is like, okay, this has to be an API. And it has to look a little bit like Like there has to be more of a developer experience here and more of a discovery experience as well. Like I don't know where to use these models. And a little Chrome extension is not going to help me discover. It's not enough real estate. I need more space. I need visuals.

36:01I need graphs. I need, you know, examples. I need images. I need to like, I need to be able to like explore both as a human and as an agent. So that's kind of how OpenRider came to be. You know, a meta point that I think is underappreciated, but Alex is reminding me, is that we were quite lucky that we were so, we were like adjacent to the crypto community in those days. Because in hindsight, crypto ended up being kind of like a dress rehearsal for generative models, right? If you think about the Axie experience, you know, Alex is totally right. There were not that many AI apps at the time. and while I was dealing, my job was to be the head of platform at Discord, which meant to be a general purpose place for communities and friends to create, for developers to create apps and bots and other services that could be deployed across Discord.

36:54And while 80 % of the attention at the time was being spent on crypto, because that's where all the NFT volume was, there was like 20 % of my time I was spending with a friend who would get hotpot with me and asked me for, we'd play Magic the Gathering on weekends and he was working on a little Discord bot that could take a text input and turn it into an image. And it was called MidJourney. Was that David? That was David Holtz. He was a good friend. And David and I have both been sort of failed AR, VR founders before that. And I remember this, MidJourney was one of the fastest growing communities we had after Axie Infinity started to peter off.

37:31And many of the abstractions and the infrastructure decisions we made to scale Axie happened just in time because they actually did this and then fell off a cliff. And then as MidJourney was taking out, we like explicitly decided to help David make the server, the MidJourney server the primary place for interaction with the model because it was very hard for people to understand how to use the model if they couldn't see other people using it and copy them. And so the single player MidJourney web app on its own like midjourney.com had like terrible retention because people would show up, they'd see this empty field.

38:06So it's kind of like Dali too. And they would type in like cat or dog. And it was like paralyzing for them to have this blank canvas that they had to fill because they never used an AI model before. But instead in a Discord server, you could see other people using it and riff off of their prompts and the engagement was off the charts. And so scaling, you know, mid-journey from zero to like 10 million monthly actives was a much smoother approach post-AXIE Infinity. And so - Don't forget the best of four pictures. Which is the feedback loop. The RLHF feedback loop, which by the way, separately, like Tom Brown, David and I used to play Magic the Gathering on weekends.

38:40And so like it was one group of friends who would hang out and we'd like, these concepts were all being discussed all the time. But, you know, there was, I think there were few of us who bridged both the crypto worlds and the AI worlds and compared to crypto where it was all, the question was always, what's the use case, you know, for this technology? There was never any need to ask that for AI because the use case was so visceral. It was like, I can create now anything I can imagine. and I can write novels, I can code. And the infrastructure that those of us who believed in the distributed systems, like value of crypto, like the censorship resistance part, found this use case that was explosive.

39:16And I think between MidJourney, you know, Claude was a Discord bot pre-launch, you know, that we were using internally as an LLM. 11 Labs had a TTS model that we had on Discord as well. Like Discord became this petri dish for like early apps to innovate. And I don't think it's a coincidence that they found a home there before OpenRouter gave the world like a public home store or like a, you know, storefront. Discord was this like almost kind of petri dish storefront that was kind of like piggybacked on the infra we'd built for crypto communities. And then I think Alex was one of the first people to realize, wait a minute, like these apps need their own home on the internet.

39:58And then OpenRouter to me was a continuation of that community's needs. And of course, there was the crazy distribution that you enabled for a lot of these developers. So then my question is, how come you were, my perception is OpenRouter is not that Discord-centric, right? You have a Discord. Yeah. And you use it to engage your community. But it's not like mid-journey where like, no, that is like the primary way people experience OpenRouter. Yeah, mid-journey, like, it really helps us see visually really quickly how people are using the model and how to prompt it. And I think that is partly why the server was so critical.

40:33It's like, it is the user experience. It actually adds a ton. Yes. And you can go the whole mile with just like prompting via mid-journey, like via the mid-journey Discord server, getting your images and then sharing them and having fun. For OpenRouter, for LLMs, like you need a lot of user experience around LLMs to make them like really usable. Charts, chart point. And yeah, like seeing the examples of other people is also not as useful because it's a lot of stuff to read. It takes a long, long time. You need like code-based integration, not possible to do in a Discord server. You need, or technically it's possible.

41:11I shouldn't say that. It's just not a great developer experience. You need like, you need governance for, at the point where you got code-based integration, now you need governance for managing the LLMs that have access to it, the data policies, which teams, all that stuff needs a lot more than a Discord server can provide. So it's just like it's not the right. Well, in addition, you're not wrong, but also there's the very important distinction that, you know, Midjourney was an end-user application. Right. And, you know, that's why Discord, which has 250 million monthly end consumers, you know, it made sense for Discord to kind of be a host for that application experience.

41:54What I knew was going to happen soon after MidJourney found explosive product market fit, because I think when MidJourney launched, from launch to 100 million revenue run rate, it was less than eight months. And shortly thereafter, Stable Diffusion launched. And all of us used to hang out in the Discord server. It was the... The Stability Discord? It was the Lion. Yeah, Lion. The Lion Discord server. The image community that spawned Stable Diffusion. Yeah. And so when Stable Diffusion came out, I realized, oh, now other people can build their own MidJourney. Because until then, MidJourney did not have an API.

42:30So they were a full stack company. They were training their own models and they were deploying them as an application. But if you want to build your own MidJourney, there was no API of that quality. And I think DALI, too, was still quite primitive. Like MidJourney actually had great quality. And then when Stable Diffusion came out, suddenly there was this new person who could, There was this new capability in the world, which is a developer could create their own mid-journey. And that, I think, created the need for something like Open Router, because then you need an API to... If you had the kind of creativity of David Holtz and you had stable diffusion as the model and you wanted to put these things together, how could you do that without having to figure out how to host the weights?

43:06And what Open Router... The shape of Open Router enabled is that. When you have open models, alternatives to closed sort of applications, Open Router's value in the world becomes extraordinary because now any developer can just show up and use the API. Did you just say the shape of Open Router? Oh, no. It's the real Ange. I'm misaligned now. I've been overtrained. I've been using Claude way too much, haven't I? Claude-ish is what people are saying. Claude-ish. Oh, God, I got to untrain myself. Okay, and I just want to cap off the mistrial side. My TLDR is there was a mistrial price war is what I called it, right?

43:44Like, roundabout in Europe since 2023. or four. December. They launched Mistral 8x7B. And the price went down like 80%. To me, that's very positive because it's like the first real competition to host Mistral. Is there more? Yeah, that was, I'm like trying to remember it, all the things that happened. Like we saw that model come out and immediately saw people say that it was the best model in the world. Yes. Like, this was, to my knowledge, the first time an open weights model was called that in real seriousness. It's hype, right? Is it? You know? It was hype. It was hype. It was also, like, hype from AI influencers at the time.

44:31And there were many examples where it was, like, outperforming GPT-4. So, people really wanted to try it out and see, is this going to be true for me, too? And if so, at what price? and the inference landscape was really messy. Yes. We cleaned it up. It allowed providers to compete on price so we could give users the best price in one spot. And so it was, I think, the first clear example of a provider marketplace working in a way that adds value to developers. Sean, you may not remember this, but I think we met for the first time a few days after Mixtral came out at NeurIPS. at a luncheon. Yeah, that's where I also met BFL as well.

45:18And Guillaume was there. I was at New York at that time. You were there too. And we had just announced the Mistral investment and Guillaume was over there and I remember turning to Guillaume and asking him, like, how are you feeling after the launch of Mistral and 7B? And him in his typical French fashion was like, I mean, it's an okay model. It's not that good. And I was like, it was so in contrast. But I remember... Him also saying that part of the reason he felt a lot of people thought that it was better than GPT-4 was because of the speed. It was an MOE model that they had absolutely figured out how to make super efficient.

45:57It was on the period of frontier. And this is an important thing about LLMs. Sometimes when they're faster, you think they're smarter. Even though if you did these common evals that you do seven tries, and I don't actually remember. I think we should go back and figure out what the data says. But I wouldn't be surprised if it turns out on an end of seven attempts, GPT-4 was smarter on evals, but the perception on correctness would be smarter or more accurate. But people from a human preference perspective felt that it was smarter because it was so fast. And actually most queries do not take that level of intelligence.

46:40So this is the start of humans as router, which then eventually becomes open router as router of like the auto mode. Oh, that's interesting to think about it. Because humans are the routing mechanism. Like I will ask the fast model first and then if like, oh, not good enough, I'm going to upgrade manually. Yes, yes, yes. But then he's going to auto it. I hadn't thought of it that way, but I mean, that makes sense. Which then there's a lot more techniques like fusion. Fusion is the thing that we should talk about. Before I move on to those things, I just want to close off the sort of early years.

47:06One thing that I observe, which you are also an investor in Arena. Right. And we talked about mid-journey having that feedback loop of ABCD and choosing that being very important. And you understand the flywheel. So how come you didn't build Arena and how come Arena didn't build OpenRouter? Well, Arena started before OpenRouter, right? They had the school project. Yeah, Ellum says. And then they became a company. Ellum Arena, yeah. And I know you had some Arena experiences, like the heads-up comparison type things. Yeah. But you never really went as hard as Arena did. and doing heads of experiences.

47:42And LMS actually did have a router project based on Elemarina Elo's, which they never commercialized. It's hard to do a company that does both because one company is taking data and selling it and the other company really can't by default. So, you know, I think there is like a branding reason that there are two companies here. Like when you set up OpenRouter, There's no training, no prompts, aside from what your provider policies set. Like OpenRouter can't see your prompts or completions. If you want to see that as an org, you have to opt into it and enable it. And so we're pretty conservative and careful about data policy and security and privacy.

48:29And Elmarino, their business model is oriented around the labs. Because they give it for free, right? You don't give it for free, they give it for free. Yeah. I mean, we do give some, we like have free endpoints too, but like those free endpoints, I think we're not collecting any prompts. We're not like monetizing the data unless you, you know, opt into it for some reason. This compares, I mean, you're not the first person to ask me this and Alex knows this, but I, you know, I was the first CEO of Arena for the first five months when we were helping Anastasius and Waylon kind of spin out of Berkeley.

49:03And I did invest in that before Open Router, But it was very strange to me, the comparisons that outside folks would make between the two projects, because the missions were completely different. The founding entity for Arena, we called it the AI Reliability Institute, because it was actually there as an eval service. Like the data, so to speak, that they were originally kind of offering the labs was, how do you make the evaluation of models more reliable than kind of like the state of the art at the time, which is like really just finger in the wind. That's kind of what Anastasios and Wayland's PhD work was as scientists at Berkeley, was on statistical methodologies for sort of correcting, you know, eval estimates based on like intrinsic biases and how you collected the data.

49:54Style control. Style control and stuff like that, which is very much like a, hey, if you're a scientist and you're trying to kind of, the highest expectation customer for Arena was always like a post-training and like a research at a lab. Whereas the highest expectation customer, from my perspective, that Alex really understood and the mission was to serve was like a developer, right? Who then takes the result of the research and then produces an application that's deployed to the world. It was actually a completely different problem and person that these two teams were focused on. And so from the outside in, actually, I don't know if you remember this, but I have a distinct memory of a few weeks before we did the term sheet together for Open Router, I'd given you a call because we were trying to get a pooled data set together from Open Router and from Arena to create like an open source repository of prompts.

50:51I mean, these projects were so kind of different in their goals that it was totally normal to me to be like, oh yeah, let's call Alex and see if you'd want to team up on pooling data because they're so different. We actually don't have that kind of data at all. We didn't have API prompts. We didn't have what developers want to do with the models, which is very different from what researchers inside a model lab want to do before releasing the model. Does that make sense? And so to this day, I think you see that this difference, even though at a 30 ,000 foot level, I guess you could kind of conclude that Arena and OpenRouter are adjacent, but the roadmaps, the missions, and so on at the time, at least, were in very different sort of directions.

51:36That ideal customer, I totally get that. Yes, yes, yes. As a founder, I want to own everything. That's possible. This is clearly the adjacency, and I'm going to explore that. Own everything, meaning you don't know what to do yet, so you want to make sure you catch PMT. No, I think what he says is you want to own the entire infrastructure space, and so you kind of expand to whatever demand you can capture. Yeah, I think that's hard, you know, in reality, because serving multiple customers is difficult. Clearly, you know, this is when it's a focus, right? Yeah, I still think even in the age of AI, like focus is underrated and critical, not just because you end up with a better product by focusing your humans on it, but also because the world knows what your focus is.

52:221 ,000%. The world can map like, oh, I have this issue. Which brand out there is going to help me with that issue? This is the brand that's known for that focus. So like if I want real attention on this issue, like this really matters to me, I should go with the brand that cares the most about it. To underscore Alex's point about how important focus is, in the early days of Anthropic, it was not easy to, like people think that the early days of Anthropic were like super easy because they were on the GPT-3 guys who left. But it was actually very competitive. The company was starting$10 billion behind OpenAI, right?

52:58And so to get to the frontier, like the big question was, what do we want to be known for? What's the mission? And the mission was AGI repair programming. And so to the exclusion of all kinds of other things that were really shiny at the time, like image models and video models that were getting lots of momentum, the Anthropic team was like, we just got to focus on coding. Like that is the core capability that we're focused on. Today, you can see the results, right? It's a trillion-dollar company within five years. And that focus, I think, like the focus on who your highest expectation customer is and how you exceed their expectations, because exceeding anyone's expectations is hard, and doing it for multiple customers is so even more difficult, is part of the reason why OpenRider succeeded and Anthropica as well.

53:39Was the focus on coding that early, though, or did it come later? Literally from day one, it was AI pair programming is responsibly commercialize an AI pair programmer was the seed memo. That was when I invested, right? We actually kind of like refined that memo a lot. Well, you got to ask Dario and Tom for permission on that. But it's an extraordinary piece of writing that they had put together. And AI, you know, commercializing, responsibly commercializing an AI pair programmer was the mission, you know, from day one. And I would say there was maybe like a couple moments in the company's history where like they did experiments to kind of see if like little detours made sense like a general chat bot, like Cloud AI when ChatGPT was really taking off.

54:18But at the end of the day, especially once they got their really significant pre-training computer online, I think all the main evals of the company, for example, have always been coding evals, long horizon, agentic programming. I mean, from day one, that was always the plan. Because when Cloud Instant came out and Cloud 2 came out, I remember the marketing mostly being focused on pros. Like this model could write better and long context. It was the first 100K. Long context. This directly affected me because I know something like that. What did you make? Small developer, which was my Devin before Devin.

54:54Oh yeah, yes. That's small. Yes. And, you know, so I think there's all that really like good like focus is another thing. That is a question that people do want to ask. You know, you could have built any other things like, and obviously OpenRod was working, working, working. Were there other ideas that you wanted to pursue that you turned down? You know, just the path, road's not taken. We made a couple of prototypes for things that we didn't launch. One was a fine-tuning model as a service. Lots of that, open pipe and all those things. But it kind of was in a very consumery form factor where you would give us a YouTube video or two or three.

55:32We would then extract all the transcripts from it and try to fine-tune a model to talk like the person in the YouTube video or the people in the videos that you sent. So like a really, really easy way of creating a fine-tuned model based on like some kind of videos that you like. That would be so useful. We made it too. It was like kind of a – Nobody used it. It was – we didn't actually like test it with that many people because the model marketplace was our main focus. And it was like growing and we were building more conviction in it over time. Just as a creator. Yes, yes. I have 500 hours of recorded voice of myself, make a thing of you, charge access to it.

56:20It works for OnlyFans, doesn't work for us as regular people. I think mostly it's just a glorified rag bot. Whether it's in the weights or it's outside the weights, it doesn't really matter. You're just doing rag on the videos. And people ultimately always just want to find the source video that directly answers it. My use case was mostly to practice with myself because I often like to see what, like the way I practice for a job interview or like I'm hiring a candidate or public speaking or whatever, is I wish there was like a good mini-me that I could like critique because it's kind of hard to pull yourself out.

56:51I would never offer it to other people in service. Like pick your top five mentors that didn't talk to them instead of talk to yourself. That'd be cool too, yeah. That was, that's the character AI, that's a replica. And that was the use case we were aiming at. I see. It was like, you want to create an experience. Like AI Steve Jobs. And AI Steve Jobs was the initial use case. Even though it's not allowed. That's a common prototype, yeah. Talking about adjacencies, fine-tuning as a service, as part of the router service, is something that I would typically think about as well, right? Like, why don't you do that?

57:23Because if people are running already, their inference through you, store everything, log everything, fine-tune to a smaller model that is cheaper, faster, all these things that's within your control, right? You didn't do that, but other people would have pitched that in the general state of an infra startup. I think you were just maybe a little bit early because today that's an extraordinarily fast-growing segment like from Astral where they do a lot of enterprise deployments. It's often fine-tuning custom models for ASML or whatever. But not as a router. They're just like, I come to you because I like your Astral models.

57:52I want custom Astral model, right? It is not I want to run all my OpenAI prompts, store all my results, and then just move off of OpenAI. They're not doing that. As a way to export off of dependency on a Frontier Lab, I have not seen that yet. Yeah. Which was your kind of your... Key position to do. I mean, we decided, really we like leaned into our focus and figured that like there are, like we just saw the ecosystem develop over time. All these inference providers that do want to help companies do that. Like it makes sense for us to partner with them and to like give users lots of choice and to like, you know, figure out what makes them, what gives them competitive advantages.

58:36It's a whole new business, basically. And there's value in being a neutral marketplace that just kind of like works with those companies. Could you share a little bit, to Sean's point, like how you prioritized, what are some ways you prioritize features? Because you've always done it so elegantly that it never just happens. And you make all the right decisions that always have product market fit from the outside looking in. But consistently, you seem to have prioritized a lot of hit features that worked. And maybe I have a sample set bias or whatever. Can you list what you think hit features worked?

59:09Oh, the leaderboards. Leaderboard, okay. Yeah, you know, like from day one. Charting, BYOK. But he had like plugins, you know, he had like, and I think there was a whole thing I want to get into about like completions versus check completions versus completions. And then also let's call it like the rise of the reasoning models and how you deal with those, multimodality, all those things. BYOK, that was a huge one. There was one, like, I think it was in early 2024. Very early 2024, we thought it might be interesting to fuse the results of multiple models together. And we launched a prototype called MOM, mixture of models, that let you, like, pick a couple models.

59:55We'd pick them for you, and then it would fuse the results together at the end. and it would show you all the intermediate results in this like big Kanban board looking product. What does the fusion at the end? Another model? Another model. The smartest of the set. Of the three, of the set. So this is like a council idea? It was like a very early LLM council. This is a multi-agent swarm as like they would call it at one of the frontier labs in the early days, you know? Yeah, like some of those ideas are like going in the right direction, but the devil's in the details. There's a lot of like product refinement needed to make them really work.

1:00:31They take your focus away from, you know, whatever else you have going on. And there's a lot of like community building and learning that you need to do. And the technology might be too early. So there are like all kinds of reasons they might go wrong. And in our case, the technology was a little too early. In other words, the fused result was a little bit worse, sometimes the same as the best model that was being used to fuse because the best model was so far ahead of options two and three at the time. You know, over time, the top three or four LLMs have gotten closer together, still neurodivergent, but like all capable of inserting like pretty interesting ideas.

1:01:14Like RL has basically like expanded the surface area of creativity for machine learning researchers within each lab. And so they can, you know, diversify the reasoning power of different models more effectively. At least that's my theory for why Fusion works better than it used to early 2024. So the technology was a little bit too primitive. The form factor was not right. And so we would have had to go through a couple more iterations. And so we decided to just delete all the code. And then years later, middle of 2026 or early 2026, we're like, let's bring it back. Like the research is looking kind of promising for Fusion.

1:01:58The models now have like two, three, four top frontier models that are all really good. And like I'm frequently trying to like consult multiple models to get the best results. And then I, you know, I ran a little personal experiment where I was like, I'm going to like do an architecture plan for a code change. I'm going to give it to all the models. I'm going to fuse the result. and I'm going to ask all the models if the fused result is better than the individual result each model came up with. And they all said yes, that the fused result was better. And this happened a couple of times and I was like, okay, spot check, pretty good.

1:02:36We should like benchmark this and that's how we built Fusion. Yeah, and it came on your fable so you were like, this is fable level. Yeah, yeah. Let's start leading up to this year which we haven't gotten to this year. Can you mark out the main milestones in the journey? I think it seems like your promise was routing. You decided the business model very early. You take a cut. And what are the major milestones that inflect the growth? You're growing like 9 % week on week now? Is that the official number? In terms of token volume, I think that sounds about right. Yeah. So can you mark out the sort of brief history of OpenRouter up to the acquisition?

1:03:20Let's call it, we're just talking about, you know, people are, you have sort of your birth moment with the Michelle stuff where people are really competing. You have your state of AI thing where it's very cute. You have a hundred trillion tokens, ha ha ha. Because now you're doing 10 a week, you know. We're doing 10 a day. 10 a day now? Yeah, more. So yeah, you do this in 10 days. Like what are the major points there? You know, I just want to like, there's a smooth curve, but you feel the infections. A lot of this is kind of oriented around model launches. We had a huge focus on pros all the way up through May of 2024 because coding was just not there and no apps were able to build much on top of it.

1:04:11So a diversity in models, but not a wide diversity and not a wide diversity in use cases. Dream Tavern was one of our top apps at the time. The creator of Dream Tavern now runs product at Cognition, Devon.

1:04:30Then in the middle of 2024, we saw Claude Sonnet 3.5. That came out, incredible leap forward in coding. And we saw the dynamics of apps building on top of us change. we saw a huge surge in volume in like users using open router. And this is when I think people started to look at the like money that they were spending and get a little bit like, whoa, what's going on? I might need to like think about like more cost efficient but equivalent models. And shortly after that, I think it was after Sonic 3.5, Mixtrol 8X7B came out. And everyone was like, whoa, this is the model. Like the open weights community delivered.

1:05:14And so it was really good timing from the strong. Basically all of Anja's portfolios are just helping you out. It takes an ecosystem to grow an open router, you know? Yeah, that was the, yeah, it was. Like it was this early ecosystem. It was like a swing action where like model labs would come up with some sort of frontier innovation. Like usage would surge. Then users look at their invoices 30 days later and are like, whoa, what's going on here? And then open-weight models would deliver cost-effective options two, three months later. We saw that happen several times. One thing you also did with the coding agents was that you broke out which are the top coding agents, and they love that.

1:06:00They love that leaderboard. The Klein versus the root code versus the what have you. Yeah, yeah. Like Klein was like the top of our leaderboard at the time, we then at the end of, and I'll skip forward a little bit, the end of 2025, there were quite a few coding apps on the leaderboard, but they were all IDs or terminal-based agents. And at the end of 2025, we saw OpenClaw appear. And OpenClaw was like particularly interesting because one, it was like a new form factor that like brought in a new type of user, not just a developer, but like a productivity or sort of like an internet creator came to AI for the first time.

1:06:48And it also had an interesting architecture where it was like calling your chosen model for these heartbeats to see if it was still alive in addition to actually using the model for real tasks. And the heartbeats are like, they're kind of, you don't want to pay a lot of money to do a heartbeat. Yeah. So the auto router that we provided was really, really useful to this like wide range of users all of a sudden. And so we just saw it rocket exponentially. And then we saw, you know, like OpenCloud just blow up and a couple other apps lean into that new paradigm and do something similar. Hermes came out and really leaned into things like the auto router and built like a really good community and leaned into like basically skill management and making it really easy and effective for people to like set their memory in the agent and build really good skills.

1:07:45Which another thing you never did, memory skills, sandboxes, all these like adjacent things you could have done. Could have, but like, I think like - It's hard to bet. They're also very, there are things that developer, that really matter for like the developer use cases that were coming out at the time. Like developers wanted to architect those things. Those are kind of critical to building a good user experience. It's been hard for companies to find abstractions that work for all developers on the memory layer. It is, you know, there are some, like Mastra has done a pretty good job, for example.

1:08:21But like developers have like lots of varied preferences for them. and then we saw, you know, the way our leaderboard has changed over time is kind of like a movie of how the AI space has changed over time. If you just sort of like go to the Wayback Machine and look at the rankings leaderboard and the apps leaderboard over time, it sort of shows you like what's happened in AI over the last couple of years. To me, the coming of each moment was Andre Carpathia was like, I no longer read Local Llama because I just go to Open Router's leaderboard. which I remember that I think he probably like said like sorry guys I'm gonna send a bunch of traffic to you so I was gonna bring it into the Stripe uh thing how does that kind of conversation start we had this long-standing relationship with Stripe though from you know like many different projects that we had worked on with them we invest you know a lot of effort in countering abuse and token fraud.

1:09:26Can you give some numbers just so people understand? I think I posted about this. We blocked 10x as much dollar volume last month as the month before. And the types of token fraud are diversifying quite a bit. There are fraudsters going after typical stolen credit cards, but there are also people trying to resell traffic against the terms of service. There's like hacked accounts. There's people who just lose, you know, like their whole company is compromised and they don't even realize it. And we help them like regain control and detect it. There are accounts that are like reselling inference on the side.

1:10:09There are accounts that are dealing with, you know, like an accidental runaway agent and they don't realize it. Not a hack, but it's something that blows up and the company doesn't want it. And so our trust and safety team works a lot on all of these categories of problems and helps block it and detect it. And so we've built these, we have models around them. We worked closely with Stripe for a while on this. And I think it's going to become a huge problem in the ecosystem. We're already seeing a lot of companies start to see these fraudsters like spread and look for other ways other you know other than open router to other fraud vectors and if you're making a gateway or selling like generalized inference you are a target for fraud if you're selling very discreet like intelligence products intelligence products that are like doing something pretty specific but not like you know, just reselling inference with some added capability, then you're way less likely to get these fraudsters.

1:11:20So I think we'll see companies also move away from just reselling inference with some sort of like added capability and move towards sort of like discrete tasks and charging for those tasks and charging for those enhancements and letting people bring their own inference like in a third-party way. Whoa. Okay. And yeah, obviously you would power that. But do people pay for outcomes or per task? I think people will pay, you know, I think like the data dog pricing page is a good look at like the future to come. It's like companies, like infrastructure companies will like charge for different types of events that they're providing.

1:12:06And there'll be lots of like continuous pricing models that look like that. And of course, there will be like, if you go down towards consumer apps, you know, simpler pricing, more subscriptions, you know, fewer events to worry about. And ones that like are not focused on just adding a markup on top of inference. Not just because fraud is hard, but also because the pressure from the labs and from good inference providers to do a commit and then bring your inference elsewhere is going to be very high. Any comments? Two. One, I think Alex has done a very eloquent job of describing something. you know, counterintuitively, I knew would be a thing at scale like four years ago because of Discord.

1:12:56And the particular experience that taught me this was, you know, as we started scaling MidJourney, you know, one of the primary ways that we used to give away or like get people to try MidJourney early on to get to their first 10 generations because, you know, 10 generations of, 10 images generated was roughly the magic moment activation point we found. Like once you've done 10, you were like, this is extraordinary. But for that, so we had a free trial with MidJourney. And one day I woke up because they had a platform and had to monitor, I had all these dashboards. You know, I had like three missed calls from David and it turns out like there had been this flood of new users overnight.

1:13:36And we were like, this is great. And he was like, no, actually we shut down the free trial. And I was like, why is that? And he said, Ange, look at the geolocation IP addresses. And basically somebody in China had started to resell MidJourney free, subscriptions with the free trial as a way to like basically, you know, it was fraud abuse, right? Even for a specialized model like MidJourney. Yeah, and that was actually an application. So this idea that, I think the big picture realization I had back then was, hey, there's a new type of unit of value that's being streamed across the internet called a token.

1:14:11And over the next 10 years, the entire internet value chain was going to have to deal with the fact that like, the more valuable tokens got, the more bad actors were going to try to get their hands on those tokens. And anytime you scale something and the payload gets more and more valuable, more bad people try to get access to that value. And so it was very obvious to me back then. And so, look, to this day, I don't think there's a free turn. I don't think Majority has ever actually turned on the free trial since then because it was really not an easy problem to solve in terms of trust and safety.

1:14:49That's why I started teaching the class Security at Scale at Stanford. Like one of the, that and the anthropic learnings, to me it was clear that the need for security at scale was going to be enormous a few years from then. Because if you just do the math, right, think about if where, you know, online payments, you know, started roughly in the 80s and 90s, right, and grew to over a trillion dollars over the next 10 years, and we needed to build entirely new payment solutions to deal with online fraud. Where we are today is roughly there on tokens, but over the next even five years, we're expecting the token economy to get to like roughly$5 trillion.

1:15:29And over the next 10 years, I'd be shocked if we went to$10 trillion of token flow. And so if we were starting to see such aggressive abuse and fraud at subscale, mid-journey, remember mid-journey at this point was like less than 300 million revenue run rate a year. I just realized we were going to need entirely new systems to deal with the fraud that was going to happen for trying to get into the token flow. And so I forget the board meeting it was when you brought up that Stripe wanted to partner up. And it made so much sense to me because Stripe Radar, when I was a Kleiner 10 years ago, we invested in Stripe.

1:16:04and the whole pitch that Patrick and John communicate so eloquently was like, hey, unlike traditional payment tools like Braintree that do a seven-day verification, like KYC and AML to get the fraud out of the way, we actually just bite the fraud cost up front as a customer acquisition cost and tell a developer, just use five lines of code and we start accepting your payments in five minutes. And what'll happen is over time, we'll collect all this data on the developers. Cloudflare model. Is the Cloudflare model, right? And they did. Five years later, they launched Stripe Radar. And Stripe really today is a security company.

1:16:37That's the real. People think it's a payments company. No, the reason, there's lots of other payments providers today that give you like cheaper payments transmission. But the reason Stripe keeps, you know, being the dominant one here in Adyen and Europe is because they have extraordinary fraud detection that they've built, you know, over the years. This is the same story with Elon and Max Levchin. And Affirm, yeah. So, you know, I think the story shows up over and over again, where every time you have value streamed across the world in large amounts, You need new protection and security infrastructure to fight to keep the bad guys out and allow the good people to have their transactions happen really fast.

1:17:11And so I think this is why, from my perspective, the Stripe and open router story is a security story for the internet ecosystem, for the frontier AI ecosystem. Without a partnership like that, it becomes very hard to defend the quality of experience and the speed and all the good stuff without letting the bad guys get in the way. The second is that, you know, there's this underappreciated thing about like the fact that you need to, like all the bad things that Alex described as being perpetuated by humans right now is going to be perpetuated by AI agents over the next 10 years, right? So think about the like recursive scale we're about to see of bad actors.

1:17:51It's not just bad human beings. It's all the bad agents that are going to be attacking the token flow. And it's very hard if you're a researcher or an AI lab to reason about that problem because the only data you have is how agents you're training are going rogue. But that's just a fraction of all the bad behavior on the internet that we're going to see. And so what you need is defenders, new sheriffs in town, which cowboy has, that can see all the bad behavior from AI agents across the ecosystem from different model labs and different post-trained deployments and different developers and take all of that data and say we're going to build a shield for the entire token economy because without that, you know, the amount of fraud we're going to see of this$10 trillion in GMV and global GDP growth is like a huge percentage of that I think is going to be fraud, abuse.

1:18:41And we might never get there if people just don't trust tokens, right? And I don't think this infrastructure exists. So you have your work cut out for you at Stripe, but I don't think people have realized the scale at which agentic fraud, like bad behavior perpetuated by AI agents is about to hit us like a tsunami. Yeah, I mean, there's a lot to dig into there. I want to give you the last word. We do have to wrap. What can people expect from Open Router and Stripe? I mean, I think this is a really good way for us to accelerate go-to-market and to go up market more quickly. It's also, you know, as Ange eloquently described, there's a really clear better together story here when it comes to improving trust and safety and making it really easy to accept tokens and let people bring their own inference to your app and to help developers just build on top of inference going forward.

1:19:36We have a really strong brand with Open Router, and we're keeping the brand. So Open Router, as a product and the roadmap and the name and the brand, is staying the same. And so what you should expect in the next six months is that most things will be like what we would have done had we been independent, except everything will be moving faster. And that's kind of like our near-term goal, longer term. Hopefully I can comment on it soon, but I can't now. Okay, well, we'll hopefully do a follow-up at some point. But thank you for being so generous with your time and congrats on the partnership. I mean, this is one of the most beautiful bromances I've seen in AI.

1:20:22Just starting out. Starting from Stanford to here. Lots more to do. Lots of sheriff policing to do of the token economy. We need new sheriffs for sure. Yeah. Awesome. Thank you. Thank you.

From the publisher

From the earliest days of open-weight models to becoming the neutral routing layer for more than 10 million developers, OpenRouter is one of the clearest bets that the future of AI will be multi-model. In this episode, OpenRouter co-founder & CEO Alex Atallah, with AMP’s Anjney Midha returning with swyx to unpack how OpenRouter emerged from the first wave of Llama, Alpaca, Mistral, and Midjourney, why model diversity mattered before it was consensus, and how a company dismissed as “just a wrapper” became critical infrastructure for the AI ecosystem.

We go deep on the product and distribution lessons behind OpenRouter: why model labs can spend billions training a checkpoint and still struggle to get it into developers’ hands, how Mistral helped prove the value of a competitive inference marketplace, why OpenRouter chose focus over expanding into fine-tuning, memory, and other adjacent products, and how its rankings became a real-time map of how AI usage was changing. Alex also explains OpenRouter’s early experiments with model fusion, why they deleted the first version and brought it back years later, and how the platform grew to more than 10 trillion tokens per day.

Finally, Anjney explains why Stripe and OpenRouter fit together, why token fraud may become one of the defining security problems of the AI economy, and why the next wave of fraud won’t just come from humans but from autonomous agents attacking increasingly valuable token flows.

We discuss:

* Why OpenRouter bet early that no single AI model would win everything

* Alpaca, Llama, and open models becoming impossible to ignore

* Why Discord’s early AI deployments exposed the limitations of closed models

* Why model labs can spend billions on training and still fail at distribution

* How OpenRouter became a neutral distribution layer for model developers

* Why VCs dismissed OpenRouter as “just a marketplace” or “just a wrapper”

* The Mistral price war and the first real proof of an inference marketplace

* How Midjourney scaled through Discord and what it taught the AI ecosystem

* Why crypto infrastructure became a dress rehearsal for generative AI

* OpenRouter vs. LM Arena and why their missions are fundamentally different

* Why focus became one of OpenRouter’s biggest strategic advantages

* Anthropic’s early focus on AI pair programming and coding

* The OpenRouter products that were prototyped but never launched

* MOM, OpenRouter’s early Mixture of Models experiment

* Why model fusion failed in 2024 — and why it works much better now

* How OpenRouter’s leaderboard became a live map of the AI industry

* OpenClaw, auto-routing, and agents reshaping AI usage

* How OpenRouter reached 10+ trillion tokens per day

* Why inference gateways are increasingly becoming targets for fraud

* Why Stripe’s fraud infrastructure is strategically important to OpenRouter

* The coming rise of agentic fraud and attacks on the token economy

* What changes and what stays the same as OpenRouter joins Stripe

Alex Atallah

* LinkedIn: https://www.linkedin.com/in/alexatallah/

* X: https://x.com/alexatallah

* Website: https://alexatallah.com

Anjney Midha

* LinkedIn: https://www.linkedin.com/in/anjney/

* X: https://x.com/AnjneyMidha

* AMP: https://www.amppublic.com/

Timestamps

00:00:00 Introduction

00:02:12 Alpaca, Llama, and the Multi-Model Bet

00:06:04 Discord, Open Models, and OpenRouter’s Origins

00:14:28 Why “One Model Wins” Was the Wrong Bet

00:17:27 Why Model Labs Struggle With Distribution

00:23:04 “Just a Wrapper”: Why VCs Misunderstood OpenRouter

00:27:58 Bootstrapping OpenRouter Through Community

00:36:16 Crypto, Midjourney, and the Early Generative AI Ecosystem

00:43:38 Mistral and the Birth of the Inference Marketplace

00:47:10 OpenRouter vs. LM Arena

00:52:08 Focus, Anthropic, and Roads Not Taken

00:59:34 Mixture of Models and Model Fusion

01:02:44 Sonnet, OpenClaw, and OpenRouter’s Explosive Growth

01:09:03 Why Stripe Acquired OpenRouter

01:12:45 Fraud and the Emerging Token Economy

01:17:47 The Coming Wave of Agentic Fraud

01:19:07 What’s Next for OpenRouter at Stripe

Transcript

Introduction: OpenRouter, Marketplaces, and Pub-Sub as a Product Principle

Swyx [00:00:00]: Okay, we are here in Anja’s house, which is where all big startups in San Francisco start.

Anjney Midha [00:00:08]: Howdy.

Swyx [00:00:08]: And, congrats on Cursor, Mistral. I don’- God knows what else. You got so much stuff going on.

Anjney Midha [00:00:17]: There’s, there’s a lot going on. Well, OpenRouter is probably the - has been the most, I would say, like, one I’m excited about recently.

Swyx [00:00:24]: Yeah. And we have Alex, first time on the pod, but,

Anjney Midha [00:00:27]: Thanks for having me.

Swyx [00:00:27]: You’ve been in the IE a few times. I appreciate every time you’ve shown up, for the community. Congrats. I just, like, what a journey. When I was looking back at your past posts, one of the earliest principles that I saw you write as a product person is sub as a product principle. And I wanted - you to maybe explain how you think about what should exist in the world.

Anjney Midha [00:00:49]: Yeah. The sub piece, which was early 2023, I didn’t think about it until we talked like 10 minutes ago, is about how there is like a way of thinking about products as an intersection between subscribing to data and publishing data. And marketplaces are an easy example of this. You have suppliers that are publishing some product to a SKU. And the SKU is like a sub topic that a consumer is subscribing to and just going to, like, consume whenever they want. And humans consume in a very, like, discreet, ad hoc way. It’s not very scalable. all their attention is on the topic when they’re buying the thing, and their attention is nowhere else when that happens. agents and consumers of inference don’t act like that. They’re consuming continuously, and they’re changing the SKUs that they consume from all the time. So OpenRouter is like a blend between a normal API experience and a marketplace where we create model slug. We have the auto router. We have all kinds of, like, product SKUs that you can subscribe to. And then you can, like, continuously add, like, derive value and make decisions based on those consumers.

Alpaca, Llama, and the Multi-Model Bet

Swyx [00:02:11]: Yeah. This is something that was more consensus now, but not consensus when you guys started, which was that there is such a demand for swapping models and changing things out and, that people would not use the native SDKs. I guess, for each of you, what was your realization moment that this would be it? I, - You’ve, you’ve given a talk at EIE about Alpaca as,

Anjney Midha [00:02:33]: Yeah.

Swyx [00:02:33]: One of your inspiring moments.

Anjney Midha [00:02:35]: Alpaca, I can, like, rehash the Alpaca moment for a sec. Like, the very beginning, at the end of 2022, OpenAI was the only game in town. There was, like, OpenAI, Cohere,

Swyx [00:02:47]: Yes.

Anjney Midha [00:02:48]: And then a smattering of, like, early attempts at open weight models.

Swyx [00:02:54]: Yeah.

Anjney Midha [00:02:54]: When Llama came out in January of 2023, it was like, “Wow, really exciting. This is really big.” It outperforms 3 on, one or two benchmarks. but you can’t chat with it. It wasn’t like - It wasn’t an engaging model, but it seemed like someone just needed to fix a couple things and do some RLHF on it to get it all the way there. And Alpaca was the first model that I saw that did that. It only took $600 to do. A team at Stanford generated a bunch of synthetic data, tuned Llama, and made Alpaca, billion parameter model. Or was - Maybe it was thirteen billion parameters. And it was so good. Like, I was just, like, on an airplane using it. I, - in many cases, I, like, you could not discern a ChatGPT versus an Alpaca result. And I figured if it was this easy to make a model, one, we have a whole new way of monetizing data for the first time. you can just, like, take really valuable data and turn it into a service in $600. and that cost will probably go down over time.

Swyx [00:04:03]: When you - So sorry. when you say monetizing your data as, what eventually will become an MCP endpoint or as a training data for a model?

Anjney Midha [00:04:12]: Yeah, training data for a model.

Swyx [00:04:13]: Awesome.

Anjney Midha [00:04:13]: Like, an abstract way of saying like, “Hey, I have this data.”

Swyx [00:04:15]: Compress it into a model.

Anjney Midha [00:04:16]: Like, it makes sense for me in my product, but, like, I could repackage it in the form of a model and sell it. And so it’s just a whole new business model for the economy. It also, of course, provides, like, a way of following what Frontier Labs are doing, but in a way that, like, a single developer or a small team of developers can roll on their own. And so - Whenever you have an example of that, like a breakout app that’s doing really well, and then some framework for imitating it with - in your own flavor, you have an immediate ecosystem of, like an immediate ecosystem, like, should arise because there’s just a huge gap between the, like, decisions that the single company is making and all of the variations in those decisions that, like, a wider ecosystem can create themselves. And so then, you need a marketplace to, like, discover all of those, services and all of those products. There wasn’t any place on the internet that, like, was like a home base for LLMs in terms of seeing how much they were being used and seeing who was using them and why.

Swyx [00:05:29]: The closest would be Hugging Face.

Anjney Midha [00:05:30]: Hugging Face was the closest at the time, yeah.

Swyx [00:05:31]: They just started Hugging, like, a few years ago before that.

Anjney Midha [00:05:34]: Yeah, and Hugging Face also didn’t have the closed-source models.

Swyx [00:05:37]: Yeah.

Anjney Midha [00:05:38]: And they didn’- you couldn’t use the models at the time. and there wasn’t data about who was using them. There were, like, a bunch of differences between OpenRouter and Hugging Face, and those differences felt really critical to me, especially when I was just trying to learn about LLMs and, like, why people are choosing, like, Different little ones that are emerging over time.

Discord, Open Models, and the Origins of OpenRouter

Swyx [00:06:03]: Got it. And then, Ansh, no stranger to wanting more model diversity, at the time, you’re a couple of years into your Anthropic journey, which we covered in the previous podcast as well. What was your introduction to Alex?

Alex Atallah [00:06:16]: Well, the introduction was, I think, thirteen years before that.

Swyx [00:06:20]: Oh.

Alex Atallah [00:06:20]: But the OpenRouter handshake happened right over there, if you remember.

Anjney Midha [00:06:23]: Yeah.

Alex Atallah [00:06:24]: Which - So Alex and I, met, I believe as sophomores now, if I remember at the Stanford Review,

Anjney Midha [00:06:32]: That’s right

Alex Atallah [00:06:32]: Meeting for the first time.

Anjney Midha [00:06:33]: I think so, yeah.

Alex Atallah [00:06:35]: Yeah.

Anjney Midha [00:06:35]: Yeah.

Alex Atallah [00:06:35]: So Stanford Review was the libertarian newspaper on campus at Stanford that Peter Thiel started back in the day. And, whatever-- for whatever reason, I, Alex and I both showed up to one of the meetings, and I remember, the editor-chief was a mutual friend of ours. Lisa was really a really great editor-chief, where, part of an editor-chief’s job is to assign responsibilities to people and make sure the work gets done. and I, I may be misremembering the details, but I remember wanting to. It was surprising to me that at the time there was no dedicated technology section in the newspaper.

Alex Atallah [00:07:11]: You

Swyx [00:07:13]: Because it’s political, right?

Alex Atallah [00:07:14]: It is primarily

Swyx [00:07:14]: Like, it’s talking

Alex Atallah [00:07:15]: It originally started as like a

Anjney Midha [00:07:16]: Yes.

Swyx [00:07:17]: Yeah, states and all those things.

Alex Atallah [00:07:17]: Correct.

Swyx [00:07:18]: Yeah.

Alex Atallah [00:07:18]: But it, - To take us back in time, you may remember this, but, there was this technology, legislation that was being debated called, the Net Neutrality Act. And net neutrality is, like, inherently this political concept, right? It’s, it’s about the regulation of - internet broadband access. And so there was a community of us who were technologists, but also debating the politics of the technology. And I thought the Review would be a great place - to, like, write about that. And I was working on, I think, a net neutrality article, and I remember proposing, “Well, maybe we should start a technology section.” And Alex was one of the only people who said, “Yes, that would be cool.” And said. I forget whether we ended up writing stuff together, but - that’s when we first met,

Alex Atallah [00:08:03]: Was 2011 or twelve. I forget which year it was. It was one of those.

Anjney Midha [00:08:09]: Yeah.

Alex Atallah [00:08:09]: It was at Old Union, if I remember correctly.

Alex Atallah [00:08:11]: That’s where we used to meet. But, along the way, Alex and I have had a chance to, To hang out often. And probably the time when we had the most professional overlap was when I was running the platform at Discord, and it had become this explosive platform for crypto

Swyx [00:08:32]: Yeah

Alex Atallah [00:08:32]: And NFTs in the middle of the pandemic.

Swyx [00:08:35]: Which also, by the way, you were in charge of safety and security as well, right?

Alex Atallah [00:08:38]: I was the head of platform, which meant all of the crypto - the DAO and NFT launch security debugging fell on

Swyx [00:08:45]: And their phishing and.

Alex Atallah [00:08:47]: The phishing, the social engineering attacks, the katana DDoS that we were getting hit by. but it’s around the time I first started teaching security at scale at Stanford, CS 153. And Alex was on the, - at OpenSea at the time, and I was trying to figure out how we could defend against all these attacks that we were. Like, and at peak, I forget, if you remember how much NFT volume was running through

Swyx [00:09:10]: Discord

Alex Atallah [00:09:10]: Discord, but it was, like, a meaningful amount of, like, it was, like, several billion dollars in NFT volume of GMV, so to speak, were running through the platform, and it was all coming from OpenSea. It was these, like, buy, sell,

Swyx [00:09:20]: The

Alex Atallah [00:09:21]: Servers

Swyx [00:09:21]: The D in DAO is Discord.

Alex Atallah [00:09:25]: Yes. And so that’s when I think we had hung out professionally. But a year after that, OpenAI gave Discord early access to GPT. Sorry, three. No, it was five. Yeah, five, which is the RL version of three. And that’s around the time we made a Discord bot with, OpenAI for internal deployment, and that’s when I realized we would need. Like, since I was part of the deployment team.

Anjney Midha [00:09:50]: What was the use case?

Alex Atallah [00:09:51]: There were two that were. And there’s, there’s a post now called “Discord is Your Place for AI with Friends” that somebody sent me recently that I wrote, and published in twenty-three. But There were two use cases. One was Clyde, which was the - like, a party friend inside of Discord that could help you set up your Discord server and talk to you about onboarding and get your friends to hang out more. and then there was content moderation. And one of the realizations we had with content moderation was - it would refuse to moderate. Like, it would just refuse our prompts because the The training was. We were very early in the training era, and it would just. Our prompts would trigger it, its, like, guardrails. And we told OpenAI, “Hey, guys, we need access to the weights because if we’re gonna be doing content moderation at scale, we had 250 million monthly active users, we need more reliability that the model will do what we need it to.” And they said, “Well, sorry, guys, that’s not how this works. We’re a closed-source company.” And so that was my first realization that we needed open models, and the enterprises would need more control over capabilities, and then ultimately would need some control plane or management system to orchestrate these open models. But there weren’t no good - there were no good open alternatives until maybe

Alex Atallah [00:11:10]: Six months later when Llama came out. And six months after that, I led the series A into Mistral, which was started by Guillaume and the Llama team. And - That, - Around that time is when I remember hearing about Alex launching OpenRouter and going, “These worlds are gonna collide, and I don’t know when it’ll make sense to team up.” But Alex was so early and could see. I think he was totally right about this ecosystem starting with Llama that then needed, like, a, an easy layer to manage for, especially for. I was approaching it from the enterprise perspective because I had been that, like, the. As the VP of platform at Discord, it was my job to ensure that when we deployed models to, like, 250 million users, they did what we wanted them to. And that was very hard, because if you outsourced it to the labs and they controlled the guardrails and their guardrails are their safety policies. Forbid the model from responding to your prompts. That was quite catastrophic.

Swyx [00:12:05]: Yeah. But what, a moderation is the thing that they want to support. And obviously, beyond that, they would - OpenAI would work with you, presumably to give you a moderation endpoint, which they offer for free.

Alex Atallah [00:12:16]: It was an interesting use case, that - So they did give us a moderation endpoint. However, as you guys know, every Discord server is like a mini deployment of itself. And so the use case was instead of having human moderators that have to interpret the norms of the community, you just give the, - Often, like every, subreddit, Discord servers, public ones have their own rules that the user, the users create.

Swyx [00:12:41]: Oh, yeah. We run the LinkedIn Discord in. Yeah.

Alex Atallah [00:12:43]: And then humans used to read those norms and then enforce it every day manually, like observing each message in these communities. And these communities have like millions of users. So we had a 5,000+ person team globally in the, on the Discord content moderation team. These are outsourced contractors who had a really tough job. And so the idea was instead, if you could give the norms of that server To the LLM, then the LLM would do custom moderation for that server. It’s almost like a, like context moderation for that server. And many of those servers’ norms just violated OpenAI’s rules. And so - It was like we had our own custom eval. So each server had its own custom eval. But Discord-- at the time, OpenAI’s evals, we were all so

Alex Atallah [00:13:28]: Primitive in our thinking about how to deploy these LLMs that often the training prompts were super handed. It said, “Oh, anything about Harry Potter, anything that has trademarked content, don’- refuse.” And if it was a fan - Harry Potter fan community, this is a real use case, that had content moderation, the LLM would just refuse.

Swyx [00:13:48]: Yeah.

Alex Atallah [00:13:49]: And that was just not precise enough.

Anjney Midha [00:13:52]: Another one that we heard was like if someone was trying to write like a detective story, and there’s one chapter with a lot of violence, like maybe someone

Alex Atallah [00:14:01]: Right

Anjney Midha [00:14:01]: Like kills someone, the LLMs would just refuse to, like, help with that part of the story.

Alex Atallah [00:14:07]: Yeah.

Anjney Midha [00:14:07]: And then - like, we used to be like, okay, this is not like structurally inherent to LLMs. There must be, like, some choice out there so that I can, like, switch to another model, when I’m getting, like, a refusal or a bad result from the main one that I have. And that, like, tension also drove me for a marketplace.

Why “One Model Wins” Was the Wrong Bet

Swyx [00:14:28]: Yeah. I think that is well accepted now. What was it like back then when you were raising or, starting this? did people get it? what was the, some of the struggles? I like getting stories out of him about how other VCs don’t get it. So like anything you wanna, talk about, now - Let’s, let’s call it, that the early journey of OpenRouter is done, right? You can obviously talk about some of the early days stuff.

Anjney Midha [00:14:54]: Well, I was gonna say that, like, the biggest objection we got is big model win, which is - all of the

Swyx [00:15:03]: Scaling laws.

Anjney Midha [00:15:04]: Huh?

Swyx [00:15:04]: Scaling laws.

Anjney Midha [00:15:05]: Yeah, scaling laws, and natural network effects are just gonna accrue to one company, which will be - It’ll be a Google-style monopoly, just like how Google won the search market, by a large margin, and you’ll just be fighting for scraps at the end. That was probably the biggest objection we got. it is interesting that Google won the search engine race with such a huge margin. I think, like, had there been more interesting benchmarks or had, like, search engines been, - had people, like, seen them a little bit more like LLMs where they’re services that you can build companies on top of, that might not have been the case. but LLMs don’t merely have a user interface. They’re also, like, ways of building entirely new businesses. And, a Google-level monopoly would be like the Dutch East India Company times, quadrillion in magnitude because the whole economy ends up, like, depending on the one monopoly as well. So it didn’t seem like would be a really crazy outcome if that happened. And it’s also less likely because the economics of, like, creating good competitors are much, like, much more decentralizable.

Alex Atallah [00:16:25]: Everything Alex said is true, And I came at it from a completely different perspective, which

Swyx [00:16:31]: Yes, this is why we’re here.

Alex Atallah [00:16:32]: The scaling laws were never - In my mind, were always a feature, not a bug for why OpenRouter would be very valuable. Because, I was one of the first investors in Anthropic, and it was obvious to me that other researchers in our friends - I went to grad school for machine learning, and I just had a lot of friends in the ML community who it was very obvious to us that the bitter lesson holds. And so I was like, “Oh, fantastic. Now we have at least two proof points that compute scaling works.” It was OpenAI and Anthropic. and by the time I think we decided to team up on OpenRouter, I had already invested in Mistral and Black Forest Labs and Luma. So there was multiple model companies and teams that I was, working with.

Why Model Labs Struggle With Distribution

Swyx [00:17:14]: But you did other modalities, whereas this is literally

Alex Atallah [00:17:16]: Across different modalities, yes

Swyx [00:17:17]: Text.

Alex Atallah [00:17:18]: Exactly. And it was so obvious to me that an ecosystem of different kinds of models were being created, and that this whole narrative of, like, Only one company will dominate like Google was, well, like maybe true, but one, I don’t believe that. But two, there was so much extraordinary innovation happening across several different research teams. But the shared problem I was noticing across all of them was often, the research teams were fantastic at figuring out how to reason about new capabilities. They think in terms of capabilities, but never - like, are not developer mindset-oriented. Like, what happens after the training is done and the checkpoint comes out? Like, you’d be shocked how, like, similar the early training teams at OpenAI, sorry, Anthropic, BFL, Mistral, were in their, like, default approach to. Taking their research out of the, lab and scaling their impact, which is often, oh, the checkpoint is done, put it out as an API, done, and then there’d be crickets. in the case of Claude, the first Claude checkpoint was done a year before they released it internally. And then ChatGPT came out, and we decided, okay, yes, it’s a good idea to release a Claude version externally.

Alex Atallah [00:18:34]: And they had no plan, like no plan for how to get developers to try it out. And so if you go to the Claude one blog post, you’ll notice there are, like, three developer examples for users of the API, and one is a Discord bot, and the second is Vivian, my wife’s startup called Juny Learning, ‘- And then there was, like, Notion, because these were all friends of, like, the Anthropic Because that’s how - like, last minute the planning was around, hey, once the model’s done training, how do you get it out to the world? There was no distribution platform that understood what developers needed, all the key management, provisioning, like, simple, like, endpoint management, versioning control. Like, all these things that the scientists and researchers go, “ that’s plumbing. I don’t really think about it.”

Swyx [00:19:15]: Implementation detail.

Alex Atallah [00:19:16]: Right. And instead, Alex came at it from that perspective. And so, it was so obvious to me that, like, every single lab I was funding would spend - like, literally sometimes billions of dollars into training, and then a checkpoint would be done, and there’d be crickets, like, during early access because they’re like, “Oh, that’s right.”

Alex Atallah [00:19:35]: It’s hard to use a checkpoint to make anything. You need a whole bunch of plumbing around it to make it usable by a developer. And so by the - I think - it was so obvious to me that a distribution platform like OpenRouter was critical to have in the ecosystem if we wanted there to be competition to Google. Like, unless-- ‘cause with Google, DeepMind is done training a new checkpoint, and then they push a button, and it gets blasted out across all their surfaces from Google Docs to,

Swyx [00:20:01]: Everywhere, even if I don’t want it.

Alex Atallah [00:20:02]: Everywhere. You wanna know about, like, on Android, like, overnight, they can deploy a new checkpoint to, like, a billion devices, right? And that invisible infra advantage, distribution advantage, most people don’t realize, but until OpenRouter showed up, - you had to think about all of that yourself as a model lab. And it was very daunting. at Anthropic, I think it took, well, more than twelve months to get to our first 10 million in revenue. And in contrast with Black Forest Labs, I remember the early days, you guys had a conversation with the BFL team, and, it was so simple for OpenRouter to say, “Oh, no problem. Like, the day you launch, we can send 1 million developers to you.” that was crazy. That was like a step function change in, like, an hour.

Swyx [00:20:46]: Is that a real number, a million?

Alex Atallah [00:20:47]: I,

Swyx [00:20:48]: Okay. All right.

Alex Atallah [00:20:48]: I think today it’s, like, 4 million. How many developers are on OpenRouter today?

Anjney Midha [00:20:52]: Over ten,

Alex Atallah [00:20:54]: Yeah.

Anjney Midha [00:20:54]: Over 10 million, but, like, it’s, it’s hard to, you

Alex Atallah [00:20:59]: I, yeah, I don’t know how to. Yeah.

Anjney Midha [00:21:00]: We do a lot of, like, account duping work, but, no

Alex Atallah [00:21:04]: If you could get 1,000 developers, just to put in context If you get 1,000 developers who try the model on day one after you release it and just, like, do inference and give you feedback, that’s a thousand

Anjney Midha [00:21:15]: That’s huge

Alex Atallah [00:21:16]: More developers than they knew how to get to on their own.

Swyx [00:21:19]: Well, BFL had a reputation, but yes.

Alex Atallah [00:21:21]: They had one in Stable Diffusion.

Swyx [00:21:22]: Yeah.

Alex Atallah [00:21:23]: And with Mistral, I don’t know if you guys remember, but the first checkpoint they released was, like, torrents. It was, like, torrent weights.

Swyx [00:21:31]: Yeah, they just put up a magnet link.

Alex Atallah [00:21:33]: Yeah, there was no API.

Anjney Midha [00:21:34]: Yeah.

Alex Atallah [00:21:34]: Because they didn’- they weren’t infra people.

Alex Atallah [00:21:37]: ? Like, it’s like, okay, download these weights, and you guys go figure out how to host it.

Swyx [00:21:39]: Well, he has a story on his side, yeah.

Anjney Midha [00:21:41]: Yeah, in addition to the, like, building a really good developer experience around it, the marketing that we do on, like, for different models is totally different and perceived totally differently

Alex Atallah [00:21:54]: Right

Anjney Midha [00:21:54]: From the marketing that a model lab does for itself.

Alex Atallah [00:21:56]: Yes, 1,000%.

Anjney Midha [00:21:57]: Right? We are like a, neutral layer looking at this market like it’s a big dark room with all the corners completely obscure to users, and users are walking into the room and, like, feeling around

Alex Atallah [00:22:09]: Yeah

Anjney Midha [00:22:09]: And trying to figure out what objects to grab off the tables and, like, build into, their companies. And it’s just an insane way of working. Like, models are not products where you can just enumerate all their features onto a web page. They’re all black boxes, including the open weight ones. So you need to, like, shine lights on all corners of this room, so that people can see what makes this model good, and you need the company shining that light to be a neutral third party, which is what we specialize in. So the, like. It’- In addition to developer experience, there’s also, like, a very important, like, marketing and product packaging component

Alex Atallah [00:22:50]: Yeah

Anjney Midha [00:22:50]: And a way of, like, routing and discovering models becomes, like, critical to your market as a provider or a model lab or a server tool and more in the future.

“Just a Wrapper”: Why VCs Misunderstood OpenRouter

Alex Atallah [00:23:03]: And this value, to your earlier point about how many VCs, like, just don’t. One of my biggest frustrations is that venture capitalists, many of them, like, just don’t have any operating experience in the field. so unlike a traditional investor who’s just maybe come up through the ranks as, like, a associate working on financial modeling or maybe hasn’t been a real operator in the field for, like, more than ten years, which is a big part of the industry now, I had just arrived at a16z, like, a year after running the platform. And so I knew what the challenges were of, like, building a real - great developer experience and like, being able to create a working piece of software with a model. And there were a few, I won’t name names, but there were investors who were looking at OpenRouter, and, felt at the time, like, when I would compare notes with people, that it was just, I quote unquote, “just a marketplace.”

Swyx [00:23:59]: Yeah, just a thin layer, just a

Alex Atallah [00:24:00]: Correct

Swyx [00:24:00]: Just

Alex Atallah [00:24:01]: A wrapper or whatever on other people’s APIs. And I was like, “You have no idea how strategic the value that OpenRouter has created by being able to orchestrate even three.” APIs in production. The amount of both engineering work and community design that goes into getting that live and running in production at the scale the OpenRouter team had started just doesn’t happen by default. And that was one of the things that stood out to me about Alex from the earliest days. Like, he just understood, like, - from a systems perspective, like, how do you get these flywheels going? Like, that stood out to me with OpenSea when we were working together on the NFT integration at Discord. Like, Alex had a level of community-- like, systems thinking on how you get these flywheels going that most scientists and machine learning people just don’t

Alex Atallah [00:24:48]: Think of. Like, we often think in terms of training.

Swyx [00:24:52]: It’s a linear stage.

Alex Atallah [00:24:53]: It’s this linear pipeline.

Swyx [00:24:53]: There’s no loop yet.

Alex Atallah [00:24:54]: Yeah. It wasn’t until much later that the modern context feedback loop cycle really got standardized in the industry. But at the time, if you remember, machine learning was like. Like, mostly we did a lot of ML, like, when I was in grad school on a laptop. So you just, like, download a dataset, ran some ablations, and you looked at the loss curves, and you’re like, “Great, I made AI.” And the idea that you have to, like, deploy those capabilities, collect feedback trajectories, then, like, put those into a continuous loop, like, came much later. And it was very counterintuitive to the - like, the traditional AI mindset. I do remember doing the investment phase for, OpenRouter, I just didn’t try and educate a bunch of other VCs on why it was not just a marketplace. I was like, “ what? I’m just gonna invest.”

Anjney Midha [00:25:41]: Yeah.

Alex Atallah [00:25:41]: And I’m going to, like, take the opportunity to partner with Alex, and if - no other VCs get it, that’s totally fine. ‘Cause at the time, - it was not obvious, I think, to several of the investors that, like, OpenRouter was not more than just a wrapper around APIs. And - that infuriated me. And I was like, “ what? I don’t have time to debate you. I’m - we’re gonna, we’re gonna invest.” And then I think, like, a month later, Matt Murphy marked it up by 10x. Like, - I think. I forget what the exact money was and so on, but, to his credit, Menlo Ventures realized, “Okay, there’s much more strategic value here as well.” Maybe you didn’t hear all these conversations behind the scenes But that frustrated me a lot. there’s a lot of this, like, opining about wrappers. and if you’re like, “Oh, an app is just a wrapper on a model,” then, like. And, OpenRouter is, like, this wrapper on top of other APIs, and this is the most stupid, reductive framework.

Alex Atallah [00:26:31]: And so it’s clearly somebody who has no experience deploying product at scale.

Swyx [00:26:34]: It’s the thing you dismiss other things with. Like, you’re a - everyone’s a wrapper on everything, right? Like, and there’s, there’s some Some wrappers have value.

Alex Atallah [00:26:40]: Investors are wrappers and LPs, right?

Alex Atallah [00:26:42]: Like venture capitalists. So, yeah, it’s all wrappers down, all down to bare metal, I guess, and like energy.

Swyx [00:26:46]: Yeah, there - When I started the whole AI engineer, I guess, the coining, in 2023, like, that was, like, the number one pushback is that this is no value. You should just train models.

Anjney Midha [00:26:56]: Right.

Swyx [00:26:57]: And, yeah, obviously this is, like. you guys are one of the testaments to the fact that you can build very valuable wrappers, but also very valuable model companies.

Alex Atallah [00:27:06]: It’s so, hard to be. Like, the day a model launches, the fact that you have an OpenRouter, endpoint for that model frequently at the top of Hacker News on day one, people don’t realize the amount of work that goes into accomplishing that. And OpenRouter used. Like, that would happen over and over again, and I remember going, “People have no idea how hard that is.”

Alex Atallah [00:27:30]: That’s not.

Swyx [00:27:31]: Yeah, we’ve covered some of the inference engineering that goes behind,

Alex Atallah [00:27:34]: Yes

Swyx [00:27:34]: Some of - with Base Ten and all those. Well, today you have, all those, like, cool code name things that people guess what Oxy Alpha is and all those things. But, like, I guess one of the things that you’re teasing is, how do you get that initial flywheel going, right? Because today you have your scale and your reputation, all these things, so obviously you - you’re driving immense distribution. But when you were early on, when it’s mostly

Bootstrapping OpenRouter Through Community

Alex Atallah [00:27:55]: The bootstrap, yeah.

Swyx [00:27:56]: Yeah.

Alex Atallah [00:27:56]: What was the bootstrap like?

Anjney Midha [00:27:58]: To bring it back to early Discord days, I think we, like, initially connected with. This is an OpenSea story, technically. But, and we initially connected when you were at Discord, and we talked about, like, - the Axie Infinity server.

Alex Atallah [00:28:13]: Oh, yes. Yes.

Anjney Midha [00:28:14]: This server was, like, the biggest server at the

Alex Atallah [00:28:17]: Yeah

Anjney Midha [00:28:17]: At Discord.

Alex Atallah [00:28:18]: That’s right.

Anjney Midha [00:28:19]: And you were like, constantly bumping up the

Alex Atallah [00:28:22]: The limits on the server. Oh, my God

Anjney Midha [00:28:24]: Of how many people could be in the server.

Swyx [00:28:24]: For those who don’t know, like, 10% of Philippines was Axie.

Alex Atallah [00:28:29]: Was on that server. That’s a big hit.

Swyx [00:28:31]: It was, like, a meaningful contributor to the GDP of the country.

Alex Atallah [00:28:33]: It was an NFT, like, crypto game, but it

Swyx [00:28:35]: It was like a Pokémon breeding thing.

Anjney Midha [00:28:36]: Yeah.

Alex Atallah [00:28:36]: Yeah. Similar. Yeah. There was battling, there was breeding, and then there was, like, a marketplace for trading.

Swyx [00:28:43]: Earn as well.

Alex Atallah [00:28:45]: Yeah, earn. And, like, the graphics were really cute and fun, and you like, you get emotional about your Axie that you make. So to, like, start a community like that, which we had to do many times at OpenSea with every early project, for us to create a marketplace for it, we need to make sure that the, like, the community wants it.

Anjney Midha [00:29:09]: Right.

Alex Atallah [00:29:09]: And it’s like building something that people want and going and telling them about it. Like, you can do that on a one basis, but there’s way higher leverage to do that in a community where everyone can talk to you at the same time. So we spent a lot of time, like, building things that the community really wanted. We did the same thing for OpenRouter. And, like, the Axie community was one of, like, a zillion communities we did that with. And Anj, like, saw us doing it and. ‘Cause you could just see people sharing OpenSea links constantly in that Discord. Like, users sharing links is a really clear indicator that, like, something important is going on. So we spent, a lot of time, like, first figuring out what the gap is in the technology that people care about. Like, what was the actual problem that needs to be solved? in early LLM days, it was, OpenAI refusing to finish the prompt or,

Anjney Midha [00:30:09]: Yeah

Alex Atallah [00:30:10]: To, like, complete the task. It was also.

Anjney Midha [00:30:13]: Inability to customize models. and so there are communities that, like are just completely blocked on that issue, and those are the communities that are most useful to learn about and dive into and explore.

Alex Atallah [00:30:28]: Something that really struck me at that time, - as I was just hearing your talk, I remember noting - you may not remember this, but we - we had these, like working, Zoom calls that we were doing a sprint around for, like this OpenSea integration with Discord. and, we’d, we’d - it was myself, my engineering team. I think you were there. And I remember, Alex, in the middle of one of those calls, just like there was like silence. we were all like, “Oh, yeah, this totally makes sense. Let’s do this.” And then there’s - every, like everybody aligned. And Alex was like, “No, this makes no sense to me.” And everyone’s - I remember going, “What? Like, it works. Like, you click on a link and this, then it bounces you out to, like, OpenSea.” And he was like, “It’s not a good user experience. Yeah, we should not do this.” And I remember going, he was the only one person out of all of us to raise his hand and go, yes, it made sense from a technical implementation perspective. Like, we were bouncing the user out into the, into OpenSea. And so it kinda checked the box of the product manager’s requirements on both sides. But Alex went one step further and was like, “ what would be better, guys? If we just embedded the experience right here inside of Discord so the link opened up as an embedded iframe, and you can just check out right there.”

Alex Atallah [00:31:47]: And not one person on the call, and there’s like seven of us who had met, like, week after week.

Swyx [00:31:52]: And it’s the guy who doesn’t work for Discord.

Alex Atallah [00:31:53]: And it’s the guy who doesn’t work for Discord.

Swyx [00:31:55]: Like, technically, you benefit if they bounce.

Alex Atallah [00:31:57]: Exactly. And that was, like, adversarial. To keep the user inside of Discord would be adversarial to OpenSea. And yet Alex put that user experience first. And I was like, “That’s special.”

Swyx [00:32:08]: Wow.

Alex Atallah [00:32:08]: Because it’s very hard to have somebody who’s technical like Alex and understands the developer flow, but also understands the best user experience and wants to prioritize that. And that’s two sides of the flywheel that if you can get spinning, like is often hard to stop. And you just reminded me, like that one was one of those moments where I go, I - I realized I gotta be better at user experience because I should have been the one who came up with that, and I didn’t. And I learned from you. And, I think that went into one of our case studies for the PM training program at Discord.

Swyx [00:32:34]: Whoa.

Alex Atallah [00:32:36]: I don’t know if it there is Because of

Swyx [00:32:38]: You need an Alex is the conclusion.

Alex Atallah [00:32:40]: Yeah. You need an Alex. And this is why I’m not, nobody should be surprised why Stripe decided like they had to buy OpenRouter because it’s a really rare combination of people who understand the machine learning community, the developer experience, and the user experience. And putting all that together has resulted in this extraordinary scale that very few other marketplaces have been able to achieve

Window AI, BYOM, and Finding the Right Form Factor

Swyx [00:33:02]: Yeah.

Alex Atallah [00:33:02]: Over the last, five years.

Swyx [00:33:04]: Yeah. Well, we should talk about the other reasons for acquisitions, which

Alex Atallah [00:33:07]: Yes, we should.

Swyx [00:33:07]: You’ve written about. I wanna proceed somewhat chronologically as well. So - there is a point that, one of the questions that, Dave from H of Zero sent in was, when did it - really started to work? And you brought up Mixtral. I don’t know if you wanna bring up that story.

Alex Atallah [00:33:22]: Oh, yeah.

Swyx [00:33:23]: Which obviously you overlap with, so.

Anjney Midha [00:33:26]: Yeah, the MoE was. I don’t know when. there’s no like one moment where I was like, “Oh, this is, officially starting to work.” It was

Swyx [00:33:36]: The moment where you had a Chrome extension, like, really super early on.

Anjney Midha [00:33:39]: Oh, yeah. But, well, - yeah. So before OpenRouter, I wanted to, like, explore a bring-your-own-model experiment. And,

Swyx [00:33:47]: Which anyone familiar with crypto is like, yeah, Phantom and all these things.

Anjney Midha [00:33:50]: Yeah. So it felt like doing a MetaMask analogy for AI would be a fun way of exploring that. And at the time, there were no AI apps. There were probably as many AI apps that were, like, hitting AI - like, hitting an LLM via an API call as there were, like, games just doing it in JavaScript. like there was a, there was a moment in time where it could have been the case that web apps call LLMs through the browser, like through some desktop

Alex Atallah [00:34:27]: Yes.

Anjney Midha [00:34:27]: Managed app that is controlled by the user. and of course, there are like, I think, many reasons that did not happen. But back when the days were that primordial, I built a Chrome extension called Window AI

Swyx [00:34:43]: With Plasmo.

Anjney Midha [00:34:44]: With Plasmo.

Swyx [00:34:45]: I had come across early on, and I was like, “Who’s gonna use this?” You did.

Anjney Midha [00:34:49]: Plasmo had a couple, like, I think Phantom was using it. there were some other, like real companies using it.

Alex Atallah [00:34:56]: It was like a shim.

Swyx [00:34:57]: React for Chrome extension. It compiles to all

Anjney Midha [00:35:00]: Yeah.

Alex Atallah [00:35:00]: I see.

Anjney Midha [00:35:00]: Like Next.js for Chrome extensions.

Swyx [00:35:01]: Next.js, Next.js.

Alex Atallah [00:35:02]: Okay.

Anjney Midha [00:35:03]: And yeah, built Window AI on top of it. The creator of Plasmo, like started contributing code to Window AI, in GitHub, and that turned out to be Louis Vicchi

Alex Atallah [00:35:15]: Oh, you’

Anjney Midha [00:35:15]: Who is the founder of OpenRouter.

Alex Atallah [00:35:17]: That’s right. You have told me this is how you met Louis. Yes.

Anjney Midha [00:35:19]: Yeah.

Alex Atallah [00:35:19]: Okay.

Anjney Midha [00:35:20]: So, that allowed users to like configure which model they wanted to use for a web page in their browser, and then, like the app would just call out to that model when it needed to do things. not the right form factor for LLMs, but, it’s like fun experiment. You learn a lot, and like I open sourced it. And the main learning is like, okay, this has to be an API, and it has to look a little bit - like, there has to be more of a developer experience here and more of a discovery experience as well. Like, I don’t know where to use these models, and a little Chrome extension is not gonna help me discover. It’s not enough real estate. I need more space. I need visuals. I need graphs. I need, examples. I need images. I need to, like, I need to be able to, like explore both as a human and as an agent.

Crypto, Midjourney, and the Early Generative AI Ecosystem

Alex Atallah [00:36:10]: Yeah.

Anjney Midha [00:36:10]: So that’s how OpenRouter came to be.

Alex Atallah [00:36:13]: A meta point that.

Alex Atallah [00:36:16]: I think is underappreciated, but Alex is reminding me, is that we were quite lucky that we were so. we were, like, adjacent to the crypto community in those days. Because in hindsight, crypto ended up being like a dress rehearsal for generative models, right? If you think about the Axie experience, Alex is totally right, there were not that many AI apps at the time. And while I was dealing-- my job was to be the head of platform at Discord, which meant to be a general purpose place for communities and friends to create-- for developers to create apps and bots and, other services that could be deployed across Discord. And while 80% of the attention at the time was being spent on crypto, because that’s where all the NFT volume was, there was, like, twenty percent of my time I was spending with a friend, who would get hotbot with me and ask me for. We would play Magic: The Gathering on weekends, and he was working on a little Discord bot that could take a text input and turn it into an image, and it was called Midjourney. You

Swyx [00:37:15]: Is that David?

Alex Atallah [00:37:15]: It was David Holz.

Alex Atallah [00:37:16]: He was a good friend. And David and I have both been failed ARVR founders, in the before that. And, I remember this. Midjourney was one of the fastest-growing communities we had after Axie Infinity started to peter off. And many of the, like, the abstractions and the infrastructure decisions we made to scale Axie happened just in time because they. Axie did this and then fell off a cliff. And then as Midjourney was taking off, we, like, explicitly decided to help David make the server, the Midjourney server, as the primary place for interaction with the model, because it was very hard for people to understand how to use the model if they couldn’t see other people using it and copy them. And so the single-player Midjourney web app on its own, like midjourney.com, had, like, terrible retention because people would show up, they’d see this empty field. It’s like E 2, and they would type in, like, cat or dog. And it was, like, paralyzing for them to have this blank canvas that they had to fill because they’d never used an AI model before. But instead, in a Discord server, you could see other people using it and riff off of their prompt, and the engagement was off the charts. And so scaling, Midjourney from zero to, like, 10 million monthly actives was a much smoother approach Axie Infinity. And so,

Swyx [00:38:29]: Don’t forget the best of four pictures, and you choose one.

Alex Atallah [00:38:31]: The best, yeah, and then the other, we

Swyx [00:38:32]: Which is the feedback loop.

Alex Atallah [00:38:33]: The RLHF feedback loop, which, by the way, separately, like, Tom Brown, David and I used to play Magic: The Gathering on weekends. And so, like, it was one group of friends would hang out, and we’d. Like, these concepts were all being discussed all the time. But, there was.

Alex Atallah [00:38:47]: I think there were few of us who bridged both the crypto worlds and the AI worlds. And compared to crypto, where it was - the question was always, what’s the use case, for this technology? There was never any need to ask that for AI because it’s, like, the use case was so visceral. It was like, I can create now anything at - I can imagine. I can write novels, I can code. And the infrastructure that those of us who believed in the distributed systems, like, value of crypto, like the censorship resistance part, found this use case that was explosive. And I think between Midjourney, the, Claude was a Discord bot launch, that we were using internally as an LLM. ElevenLabs had a TTS model that we had on Discord as well. Like, Discord became this petri dish for, like, early apps to innovate. And I don’t think it’s a coincidence that they found a home there before OpenRouter gave the world, like, a public home store or, like, a, storefront. Discord was this, like, almost petri dish storefront that - had, like, piggybacked on the infra we’d built for crypto communities. And then I think Alex was one of the first people to realize, wait a minute, like, these apps need their own home, on the internet. And then OpenRouter, to me, was a continuation of that community’s needs. And of course, there was the crazy distribution that you enabled for a lot of these developers.

Why OpenRouter Couldn’t Just Live Inside Discord

Swyx [00:40:07]: So then my question is, how come you were. My perception is OpenRouter is not that Discord-centric, right? You have a Discord.

Anjney Midha [00:40:14]: Yeah.

Swyx [00:40:14]: And you use it to engage your community, but it’s not like Midjourney where, like, no, that is like the primary way people experience OpenRouter.

Anjney Midha [00:40:21]: Yeah, Midjourney, like, it really helps to see visually really quickly how people are using the model and how to prompt it.

Swyx [00:40:29]: Yeah.

Anjney Midha [00:40:29]: And I think that is partly why the server was so critical. It’s like it is the user experience. It adds a ton.

Swyx [00:40:36]: Yes.

Anjney Midha [00:40:37]: And you can go the whole mile with just, like, prompting via Midjourney, like, the, via the Midjourney Discord server, getting your images and then sharing them and having fun. For OpenRouter, for LLMs, like, you need a lot of user experience around LLMs to make them, like, really usable.

Swyx [00:40:54]: Charge point.

Anjney Midha [00:40:55]: And yeah.

Anjney Midha [00:40:57]: The, like, seeing the examples of other people is also not as useful because it’s a lot of stuff to read. It takes a long time.

Swyx [00:41:03]: Yeah.

Anjney Midha [00:41:04]: You need, like, based integration. Not possible to do in a Discord server. You need, Or technic- it’s possible. I shouldn’t say that. It’s just not a great developer experience. you need, like, - you need governance for. At the point where you got based integration, now you need governance for managing the LLMs that have access to it, the data policies, which teams. All that stuff needs a lot more than a Discord server can provide. So it’s just

Swyx [00:41:30]: Yeah

Anjney Midha [00:41:30]: It’s not the right.

Alex Atallah [00:41:32]: Well, in addition, you’re not wrong, but also there’s the very important distinction that, Midjourney was an end user application.

Swyx [00:41:40]: Right.

Alex Atallah [00:41:40]: And, that’s why Discord, which has 250 million monthly end consumers, made, it made sense for Discord to be a host for that application experience. What I knew was gonna happen soon after Midjourney found explosive product-market fit, because we. I think when Midjourney launched, from launch to $100 million revenue run rate, it was less than eight months. And shortly thereafter, Stable Diffusion launched. And, all of us used to hang out in the Discord server. There, I think it was the,

Swyx [00:42:13]: The Stability Discord?

Alex Atallah [00:42:14]: It was the

Swyx [00:42:16]: Yeah, LAION.

Alex Atallah [00:42:16]: Yeah, the LAION Discord server.

Swyx [00:42:17]: The image community that spawned Stable Diffusion.

Alex Atallah [00:42:19]: The image community. Yeah. And so when Stable Diffusion came out, I realized- Oh, now other people can build their own Midjourney.

Alex Atallah [00:42:27]: Because until then, Midjourney did not have an API, so they were a stack company, right? They were training their own models, and they were deploying them as an application. But if you wanted to build your own Midjourney, there was no API of that quality. and I think E two was still quite primitive. Like, Midjourney had great quality. And then when Stable Diffusion came out, suddenly there was this new person who - there was - this new capability in the world, which is a developer could create their own Midjourney. And that, I think, created the need for something like OpenRouter, because then you need an API to. If you - if you had the creativity of David Holz and you had Stable Diffusion as the model and you wanted to put these things together, how could you do that without having to figure out how to host the weights? And what OpenRouter, - the shape of OpenRouter enabled is that. Right? When you have open model alternatives to closed applications, OpenRouter’s value in the world becomes extraordinary because now any developer can just show up and use the

Stable Diffusion and the Need for a Model API Layer

Swyx [00:43:20]: You just love model diversity.

Anjney Midha [00:43:21]: Did you just say the shape of OpenRouter?

Alex Atallah [00:43:23]: Oh, no.

Anjney Midha [00:43:25]: Were you in cloud? What is this the real Han?

Alex Atallah [00:43:26]: I’ve been, I’ve been - I’m, I’m misaligned now. I’ve been overtrained. I’ve been using Cloud way too much, haven’t I?

Swyx [00:43:34]: Claude-ish is what people would say.

Alex Atallah [00:43:35]: Claude-ish. Oh, God, I gotta untrain myself.

Swyx [00:43:38]: Okay. - And I just wanna cap off the Mistral side. my TLDR is there was a Mistral price war, is what they called it, right? Like, round about NeurIPS is twenty-three or twenty-four.

Mistral and the Birth of the Inference Marketplace

Anjney Midha [00:43:47]: Yes. December

Swyx [00:43:48]: They launched, the Mistral 8x7B, and like the price went down like 80%.

Anjney Midha [00:43:54]: Yeah.

Swyx [00:43:54]: To me, that’s very positive because it’s like the first, like, real competition to host Mistral. Is there more?

Anjney Midha [00:44:01]: Yeah, that was. I’m, like, trying to remember it, all the things that happened. It. Like, we saw that model come out and immediately saw people say that it was the best model in the world.

Alex Atallah [00:44:15]: Yes.

Anjney Midha [00:44:15]: Like, this was, to my knowledge, the first time an open weights model was called that in real seriousness.

Swyx [00:44:22]: It’s hype, right? Is it?

Anjney Midha [00:44:25]: It was hype. It was hype. It was also, like, hype from AI influencers at the time. And there were many examples where it was, like, outperforming four. So people really wanted to try it out and see, is this gonna be true for me too? And if so, at what price? And, the, like, inference landscape was really messy.

Alex Atallah [00:44:49]: Yes.

Anjney Midha [00:44:50]: We cleaned it up. - it allowed, like, providers to compete on price, so we could give you just the best price in one spot. And so it was, I think, the first clear example of, like, a provider marketplace working in a way that adds value to end developers.

Alex Atallah [00:45:08]: Sean, you may not remember this, but I think we met for the first time a few days after Mistral came out at NeurIPS

Anjney Midha [00:45:15]: Yeah.

Alex Atallah [00:45:15]: At a luncheon.

Swyx [00:45:16]: Yeah. That’s where I also met BFL as well. Yeah.

Alex Atallah [00:45:18]: And Guillaume was there.

Swyx [00:45:19]: Yeah.

Anjney Midha [00:45:19]: I was at NeurIPS at that time.

Alex Atallah [00:45:20]: You were there too. And, we had just announced the Mistral investment, and I remember Guillaume was over there, and I remember turning to Guillaume and asking him, Like, “Is it is all the. Like, how are you feeling after the launch of Mistral and seven B?” And, him in his typical French fashion was like, “ it’s a, it’s an okay model. It’s not that good.” And I was like. It was so, in contrast. But I remember him also saying that part of the reason he felt a lot of people Thought that it was better than four was because of the speed. - it was an MoE model that they had, like, absolutely figured out how to make super efficient. It was on the Pareto frontier. And this is an important thing about LLMs, right? Sometimes when they’re faster, you think they’re smarter, even though, like, if you did, N of, these common, like, evals that are - you do seven tries, and I don’t remember. I think we should go back and figure out what the data says, but I wouldn’t be surprised if it turns out, oh, on an N of seven attempts, four was smarter on evals, but the perception of on, like, or correctness would be smarter or more accurate. But, people, like, from a human preference perspective felt that it was faster because it - or smarter because it’s so fast.

Swyx [00:46:36]: Yeah. And most queries do not take that level

Alex Atallah [00:46:39]: Don’t take that. That’s true.

Swyx [00:46:40]: Right? So this is the start of humans as router

Alex Atallah [00:46:42]: Yes.

Swyx [00:46:42]: Which then eventually becomes OpenRouter as router of like the

Alex Atallah [00:46:45]: Oh, that’s interesting way to think about it. Yeah.

Swyx [00:46:47]: Like, because humans are the routing mechanism. Like, I will ask the fast model first, and then if, like, oh, not good enough, I’m gonna upgrade manually.

Alex Atallah [00:46:52]: Yes.

Swyx [00:46:53]: But then he’s gonna auto it.

Alex Atallah [00:46:54]: I didn’t, I hadn’t thought of it that way, but that makes sense.

Swyx [00:46:57]: Which then there’s, there’s a lot more techniques, like fusion. Fusion is the thing that we should talk about. Before I move on to those things, I just want to close off the early years. one thing that I observe, which you are also an investor in Arena.

OpenRouter vs. LM Arena

Alex Atallah [00:47:10]: Right.

Swyx [00:47:10]: And we talked about Midjourney having that feedback loop of, A, B, C, D, and choosing that very. being very important. And you understand the flywheel. So how come you didn’t build Arena, and how come Arena didn’t build OpenRouter?

Anjney Midha [00:47:23]: Well, Arena started before OpenRouter, right?

Swyx [00:47:27]: They had the school project

Anjney Midha [00:47:29]: Yeah, LM

Swyx [00:47:29]: And then it became a company.

Anjney Midha [00:47:31]: LM Arena, yeah.

Swyx [00:47:32]: So, but, and I know you had some Arena experiences, like the up comparison type things.

Anjney Midha [00:47:37]: Yeah.

Swyx [00:47:37]: But you never really went as hard as Arena did.

Swyx [00:47:40]: And,

Anjney Midha [00:47:40]: In doing up experiences?

Swyx [00:47:42]: Yes. And LM Arena did have a router project based on LM Arena ELOs, which they never commercialized.

Anjney Midha [00:47:48]: It’s hard to do a company that does both because one company is taking data and selling it, and the other company really can’t by default. So, I think there is, like, a branding reason that there are two companies here. like, when you set up OpenRouter, there’s no training, there are no prompts, right, aside from what your provider policy set. Like, OpenRou- like, OpenRouter can’t see your prompts or completions. If you want to see that as an org, you have to opt into it and enable it. And so we’re, like, pretty conservative and careful about data policy and security. And privacy. And LM Arena is like, their business model is like oriented around the labs and,

Swyx [00:48:34]: Because they give it for free, right? You don’t give it for free to give it for free.

Anjney Midha [00:48:37]: Yeah.

Anjney Midha [00:48:38]: But we do give some. We like have free endpoints too, but like those free endpoints, we, I think we’re not collecting any prompts. We’re not like monetizing the data unless you, opt into it for some reason.

Alex Atallah [00:48:48]: This comparison. you’re not the first person to ask me this, and Alex knows this, but I was the interim, like the founder, like first CEO of Arena for the first five months when, and we were helping Anastasios and Waylin spin out of Berkeley. And, I did invest in that before, OpenRouter, but it was very strange to me the comparisons that outside, folks would make between the two projects because the missions were completely different. The founding entity for Arena, we called it the AI Reliability Institute because it was there as an eval service. Like the data, so to speak, that they were originally, offering the labs was how do you make the evaluation of models more reliable than like the state of the art at the time, which was like really just finger in the wind.

Alex Atallah [00:49:38]: That’s what Anastasios and Waylin’s PhD work was as scientists at Berkeley, was on statistical methodologies for correcting, eval estimates, based on like intrinsic biases and how you collected the data.

Swyx [00:49:54]: Yes.

Alex Atallah [00:49:54]: And

Swyx [00:49:54]: Style control.

Alex Atallah [00:49:55]: Style control and stuff like that. And which is very much like a, hey, how. If you’re a scientist and you’re trying to. the highest expectation customer for Arena was always like a training and, like a researcher at a lab. Whereas the highest expectation customer from my perspective that Alex like really understood and was the mission was to serve was like a developer, right? Who then takes the result of the research and then produces an application that’s deployed to the world. It was a completely different problem and person that these two teams were focused on. And so from the outside in. I don’t know if you remember this, but I have a distinct memory of a few weeks before we did the term sheet, together for OpenRouter, I’d given you a call because we were trying to get a pooled data set together from OpenRouter and from Arena to, create like an open source repository of prompts. these projects were so different in their goals that it was totally normal to me to be like, “Oh, yeah, let’s call Alex and see if he’d want to team up on pooling data,” because they’re so different. We need. We don’t have that data at all. We. Like, we didn’t have API prompts. We didn’t, we didn’t have like what developers want to do with the models, which is very different from what researchers inside a model lab want to do before releasing the model.

Swyx [00:51:15]: Yeah.

Alex Atallah [00:51:15]: Does that make sense? And so to this day, I think you see that this difference, even though at a 30,000-foot level you could. I guess you could conclude that Arena and OpenRouter are adjacent, but, the roadmaps, the missions and so on at the time at least were like in very different directions.

Swyx [00:51:36]: That ideal customer, I get. I totally get that.

Alex Atallah [00:51:39]: Yes.

Swyx [00:51:39]: As a founder, I want to own everything, right?

Alex Atallah [00:51:41]: That’s possible.

Swyx [00:51:42]: Like this is clearly an adjacency that I’m like gonna explore that.

Anjney Midha [00:51:45]: Own everything meaning like you don’t know what to do yet, so you wanna like make sure you catch PM

Focus, Anthropic, and Roads Not Taken

Alex Atallah [00:51:51]: No, I think what he

Anjney Midha [00:51:52]: As quickly as possible.

Alex Atallah [00:51:53]: You want to own the entire infrastructure space, and so you expand to whatever demand you can capture.

Swyx [00:51:58]: You want to have a play in each end.

Alex Atallah [00:51:59]: Yeah, I think that’s, that’s hard, in reality, because serving multiple customers is difficult.

Swyx [00:52:05]: Clearly, this is the one focus, right?

Alex Atallah [00:52:08]: Yeah.

Anjney Midha [00:52:08]: Yeah. I still think even in the age of AI, like focus is,

Alex Atallah [00:52:12]: Is critical

Anjney Midha [00:52:13]: Underrated and critical, not just because you end up with a better product by focusing your humans on it, but also because the world knows what your focus is.

Alex Atallah [00:52:22]: One thousand percent.

Anjney Midha [00:52:23]: The world can map like, “Oh, I have this issue. Which brand out there is going to help me with that issue? This is the brand that’s known for that focus.”

Alex Atallah [00:52:31]: Yes.

Anjney Midha [00:52:32]: So like if I want real attention on this issue, like this really matters to me, I should go with the brand that cares the most about it.

Alex Atallah [00:52:39]: To underscore Alex’s point about how important focus is, in the early days of Anthropic, it was not easy to. Like people think that the early days of Anthropic were like super easy because they were on their 3 guys who left, but it was very competitive. The company was starting 10 billion dollars behind OpenAI, right? And so to get to the frontier, like the big question was, what do we want to be known for? What’s the mission? And the mission was AGI pair programming. And so to the, exclusion of all kinds of other things that were really shiny at the time, like image models and video models that were getting lots of, momentum, the Anthropic team was like, “We just got to focus on coding.” Like that is the core capability that we’re focused. And today you can see the results, right? It’s a trillion-dollar company within five years. And that focus, I think, like the high. The focus on who your highest expectation customer is and how you exceed their expectations, because exceeding anyone’s expectations is hard, and doing it for multiple like customers is so even more difficult, is part of the reason why OpenRouter succeeded and Anthropic as well.

Anjney Midha [00:53:39]: Was the focus on coding that early, though, or did it come later?

Alex Atallah [00:53:42]: Literally from day one it was AI pair programming is. Responsibly commercialize an AI pair programmer was the seed memo. That was when I invested, right? We like refined that memo a lot. Well, you got to ask Dario and Tom for permission on that.

Alex Atallah [00:53:57]: But it’s an extraordinary piece of writing that they had put together. And AI, commercializing it. Responsibly commercializing an AI pair program was the mission, from day one. And I would say there were maybe like a couple moments in the company’s history where like they did experiments to see if like little detours made sense, like a general chatbot, like Claude.ai when ChatGPT was really taking off. But, at the end of the day, but especially once, they got their like significant training compute online, I think like the. All the main evals at the company, for example, have always Coding evals, long horizon agentic programming. from day one, that was always the plan.

Anjney Midha [00:54:34]: Because when, like, Claude Instant came out and Claude 2 came

Alex Atallah [00:54:38]: Yes

Anjney Midha [00:54:39]: I remember the marketing mostly being focused on pros. Like, this

Alex Atallah [00:54:43]: Yeah

Anjney Midha [00:54:43]: Could write better

Swyx [00:54:44]: Yeah Long context. It was the first of its kind.

Anjney Midha [00:54:47]: Long context,

Swyx [00:54:49]: This directly affected me ‘cause I built something on that. Yeah.

Alex Atallah [00:54:51]: What did you make?

Swyx [00:54:52]: A small developer, which was my Devin before Devin.

Alex Atallah [00:54:54]: Oh, yeah. Yes.

Anjney Midha [00:54:55]: Yes.

Alex Atallah [00:54:55]: Small.

Swyx [00:54:56]: Yes. and, so I think, like, there’s, there’s all that really, like, good, like, focus is another thing - That is a question that people do wanna ask. you could have built any other things. Like, and obviously OpenRouter was working. were there other ideas that you wanted to pursue that you turned down? just the paths, roads not taken.

Anjney Midha [00:55:16]: We made a couple prototypes for things that we didn’t launch. One was a tuning model as a service.

Swyx [00:55:23]: Yeah. Lots of that with OpenPipe and, all those things.

Anjney Midha [00:55:25]: But it - It was in a very consumery form factor, where you would give us a YouTube video or two or three. We would then extract all the transcripts from it and try to tune a model to talk like the person in the YouTube

Alex Atallah [00:55:40]: Yeah

Anjney Midha [00:55:40]: Or the people in the videos that you sent. So, like, a really easy way of creating a tuned model based on, like, some videos that you like.

Alex Atallah [00:55:48]: That would be so useful.

Anjney Midha [00:55:50]: We,

Alex Atallah [00:55:51]: No

Anjney Midha [00:55:51]: We made it too. It was

Alex Atallah [00:55:53]: You don’t think so?

Anjney Midha [00:55:54]: It was, it

Alex Atallah [00:55:55]: And nobody used it?

Anjney Midha [00:55:55]: It - We didn’t like, test it with that many people because the model marketplace was our main focus, and it was, like, growing, and we were building more conviction in it over time.

Swyx [00:56:09]: Just, you

Alex Atallah [00:56:10]: Yeah. Why,

Swyx [00:56:10]: As a creator

Alex Atallah [00:56:11]: Yes. I’m a creator.

Swyx [00:56:11]: Have you been pitched many, like, - I have five hundred hours of recorded voice of myself.

Alex Atallah [00:56:17]: Right.

Swyx [00:56:17]: Make a thing of you, charge access to it. it works for OnlyFans, doesn’t work for

Alex Atallah [00:56:23]: I see

Swyx [00:56:23]: As regular people. I think - this is mostly, - It’s just a glorified RAG bot.

Alex Atallah [00:56:28]: Right.

Swyx [00:56:29]: Whether it’s in the weights or it’s outside the weights, doesn’t really matter. You’re just doing RAG on the videos, and people ultimately always just wanna find the source video, that directly answers it.

Alex Atallah [00:56:36]: Oh. my use case was mostly to practice - - with myself ‘cause I often like to see what. Like, the way I practice for a job interview or if I’m hiring a candidate or public speaking or whatever is I wish there was, like, a good

Swyx [00:56:48]: Yeah

Alex Atallah [00:56:48]: That I could, like, critique ‘cause it’s kinda hard to pull yourself out. I would never get. I would never offer it to other people as a service.

Swyx [00:56:54]: I wish there were, like, pick your top five mentors that, then talk to them instead of talking to yourself.

Alex Atallah [00:56:57]: That’d be cool too, yeah.

Anjney Midha [00:56:58]: That was, that’

Swyx [00:56:59]: That’s the creator AI. That’s a replica.

Anjney Midha [00:57:01]: And that was the use case we were aiming at.

Alex Atallah [00:57:02]: I see.

Anjney Midha [00:57:03]: Is like, you wanna create an experience

Swyx [00:57:06]: Like AI Steve Jobs and.

Anjney Midha [00:57:07]: And AI Steve Jobs was the initial use case.

Alex Atallah [00:57:11]: That’s a,

Anjney Midha [00:57:12]: Even though it’s not allowed.

Alex Atallah [00:57:14]: That’s a, that’s a common prototype, yeah.

Swyx [00:57:15]: Talking about adjacencies, tuning as a service, as part of the router service is something that I would typically think about as well, right? Like, why don’t you do that? ‘Cause if people are running already their inference through you, store everything, log everything, tune to a smaller model that is cheaper, faster, all these things that’s within your control, right? you didn’t do that, but, like, other people would have pitched that in the general state of a infra startup.

Anjney Midha [00:57:37]: Yeah. Yeah.

Alex Atallah [00:57:37]: I think you were just maybe a little bit early ‘cause today that’s an extraordinarily growing segment. Like, from Mistral, where they do a lot of enterprise deployments

Fine-Tuning as a Service and Infrastructure Adjacencies

Anjney Midha [00:57:44]: Right

Alex Atallah [00:57:44]: And stuff and tuning as, custom models for ASML or whatever. And often

Swyx [00:57:48]: But not as a router. They’re, they’re just like, “I come to you because I like your Mistral models. I want custom Mistral model,” right? It is not, “I want, to run all my OpenAI prompts, - store all my results, and then just move off of OpenAI.” Right? They’re not doing that.

Alex Atallah [00:58:01]: As a, as like a way to export off of dependency on a Frontier lab, I have not seen that yet. Yeah.

Swyx [00:58:08]: Right.

Alex Atallah [00:58:08]: Which was your vision.

Swyx [00:58:09]: Is efficient to do.

Anjney Midha [00:58:10]: We decided. Really, we, like, leaned into our focus and figured that, like, there aren’t. Like, we just saw the ecosystem develop over time. All these inference providers that do wanna help companies do that, - Like, it makes sense for us to partner with them and to, like, give users lots of choice and to, like, figure out what makes them, what gives them competitive advantages. It’s, it’s a whole new business and there’s, there’s value in being a neutral marketplace that just like, works with those companies.

Alex Atallah [00:58:45]: Could you share a little bit, to Sean’s point, like, how you prioritized. What are some ways you prioritize features? ‘Cause you’ve always done it so elegantly that I never. it just happens, and you make all the right decisions that always have product-market fit from the outside looking in. But consistently, you seem to have prioritized, a lot of hit features that worked. And maybe I have a sample set bias or whatever, but Sean’s question

Swyx [00:59:06]: Can you list what you think hit features worked?

Alex Atallah [00:59:09]: Oh, the leaderboards.

Swyx [00:59:10]: Leaderboard, okay.

Alex Atallah [00:59:10]: Yeah. like, from day

Swyx [00:59:13]: That’s charting, right? That’s the feedback loop.

Alex Atallah [00:59:14]: Charting, BYOK.

Swyx [00:59:15]: But, like, he had, like, ins. he had, like, And I think there was a whole thing I wanna get into about, like, completions versus

How OpenRouter Prioritizes Product

Alex Atallah [00:59:22]: Yes.

Swyx [00:59:23]: Check completions versus completions. And then also, let’s call it, like, the rise of the reasoning models and how you deal with those, multimodality, all those things, right?

Alex Atallah [00:59:31]: Yeah. BYOK.

Swyx [00:59:32]: BYOK, yeah.

Alex Atallah [00:59:32]: That was a huge one.

Anjney Midha [00:59:34]: There’s one I. Like, I think it was in early 2024, very early 2024, we thought it might be interesting to fuse the results of multiple models together, and we launched a prototype called MOM, Mixture of Models, that let you, like, pick a couple models, or we’d pick them for you, and then it would fuse the results together at the end, and it would show you all the intermediate results in this, like, big Kanban looking product.

Mixture of Models and Model Fusion

Swyx [01:00:05]: What does the fusion at the end, another model?

Anjney Midha [01:00:07]: Another model. The,

Swyx [01:00:08]: The smartest of

Anjney Midha [01:00:09]: The smartest

Swyx [01:00:10]: Of the set

Anjney Midha [01:00:10]: Of the three, of the set.

Swyx [01:00:12]: Okay. So this is like a council idea?

Anjney Midha [01:00:13]: Yeah. It was a model. It was like a very early LLM council.

Alex Atallah [01:00:16]: This is a agent swarm as, like, they would call it at one of the Frontier Labs, in the early days?

Anjney Midha [01:00:23]: Yeah, like some of those ideas are, like, going the right direction, but the devil’s in the details.

Swyx [01:00:27]: Yeah.

Anjney Midha [01:00:27]: There’s a lot of, like, product refinement needed to make them really work. they take your focus away

Swyx [01:00:34]: Right

Anjney Midha [01:00:34]: Whatever else you have going on. And there’s a lot of, like, community building and learning that you need to do. And the technology might be too early. So there are - like, all kinds of reasons they might go wrong. And in our case, the technology was a little too early. In other words, the fused result was a little bit

Swyx [01:00:53]: Right. Like a Frankenstein

Anjney Midha [01:00:54]: Sometimes the same as the best model that was being used to fuse because the best model was so far ahead of options two and three at the time. over time, the top three or four LLMs have gotten closer together, still neurodivergent, but, like, all capable of inserting, like, pretty interesting ideas. Like, RL has like, expanded the surface area of creativity for machine learning researchers within each lab, and so they can, diversify the reasoning power of different models more effectively. At least that’s my theory for

Swyx [01:01:29]: Yeah

Anjney Midha [01:01:30]: Fusion - it, like, works better than it used to, but early twenty-twenty-four. And, so the technology was a little bit too primitive. The form factor was not right, and so we would have had to go through a couple more iterations. And so we decided to just delete all the code. And, then years later, middle of twenty-twenty-six, or early twenty-twenty-six, we’re like, “Let’s bring it back.” Like, the research is looking kinda promising for fusion. The models now have, like, two, three, four top frontier models that are all really good and, like, I’m, I’m frequently trying to, like, consult multiple models to get the best results. Like, and then I ran a little personal experiment where I was like, “I’m gonna, like, do a, an architecture plan for a code change. I’m gonna give it to all the models. I’m gonna fuse the result, and then I’m gonna ask all the models if the fused result is better than the individual result each model came up with.” And they all said yes, that the fused result was better. And this happened a couple times, and I was like, “Okay, spot check, pretty good. We should, like, benchmark this.” And that’s how we built fusion.

Revisiting Fusion as Frontier Models Converge

Swyx [01:02:40]: Yeah. And it came on your Fable, so you were like, “This is Fable level.”

Anjney Midha [01:02:43]: Yeah.

Swyx [01:02:44]: Let’s start leading up to this year, which we haven’t gone to this year. can you mark out the main milestones in the journey? I think, it seems like your promise, was, routing. You decided the business model very early.

Swyx [01:02:59]: You take a cut. And, like, what are the major milestones that, inflect the growth, right? Like, you’re, you’re growing, like, 9% week on week now? Is this the official number?

Anjney Midha [01:03:10]: In terms of token volume, I think that sounds about right, yeah.

Swyx [01:03:13]: Yeah. So just, like, can you mark out, like, the brief history of OpenRouter up to, the acquisition? Let’s, let’s call we’re, we’re just, we’re just, talking about, people are, - you have a your birth moment with, the Mistral stuff where people are really competing. You have your state of AI thing where,

Anjney Midha [01:03:32]: Yeah.

Swyx [01:03:32]: It’s very cute. You have a hundred trillion tokens, ha, ‘cause now you’re doing ten a week, .

Anjney Midha [01:03:39]: Yeah. We’re doing ten a day.

OpenRouter’s Growth Inflections

Swyx [01:03:41]: Ten a day now?

Anjney Midha [01:03:42]: Yeah. More.

Swyx [01:03:43]: So yeah, you do this in ten days.

Swyx [01:03:45]: Like, what are the major end points there? I just wanna. Like, there’s a smooth curve, but, like, you feel the inflections.

Anjney Midha [01:03:51]: A lot of this is oriented around model launches. we had, a huge focus on pros all the way up through May of twenty-twenty-four, because coding was just not there, and no apps were able to build much on top of it. So, a diversity in models, but not a wide diversity and not a wide diversity in use cases. Dream Tavern was one of our top apps at the time. The creator of Dream Tavern now runs product at Cognition, Devon. - Then - In the middle of twenty-twenty-four, we saw Claude 3.5 Sonnet. That came out, incredible leap forward in coding, and we saw the dynamics of, like, apps building on top of us change. we saw a huge surge in volume in, like, users, using OpenRouter. And this is when I think people started to look at the, like, money that they were spending and get a little bit like, “Whoa, what’s going on? I might need to, like, think about, like, more efficient but equivalent models.” And shortly after that, I think it was after Sonnet three five, Mixtral 8x7B came out, and everyone was like, “What? This is the model.” Like, the OpenWeights community delivered. And so it was really good timing from Mistral.

Swyx [01:05:17]: All of Anja’s portcos are just helping you out.

Alex Atallah [01:05:21]: It takes an ecosystem to grow an OpenRouter?

Anjney Midha [01:05:24]: Yeah, that was the. Yeah, it was. It like, it was the, like, this early ecosystem, it was like a swing action where, like, model labs would come up with some frontier innovation. Like, usage would surge. Then users, look at their invoices 30 days later and like, “Whoa, what’s going on here?” And then OpenWeight models would deliver, like, a, like, effective options two, three months later. We saw that happen several times.

Swyx [01:05:54]: By the way, one

Anjney Midha [01:05:55]: Yeah

Swyx [01:05:55]: One thing you also did with the coding agents was that you broke out which are the top coding agents, and they love that. They love that leaderboard. The Klein versus the Rue code versus the what have you.

Anjney Midha [01:06:04]: Yeah. Like, Klein was, like, the top of our leaderboard at the time. We, We then, at the end of. And I’ll skip forward a little bit. The end of twenty-twenty-five, there were quite a few coding apps on the leaderboard, but they were all IDs or, terminal-Agents. And at the end of twenty-five, we saw OpenClaw appear. And OpenClaw was, like, particularly interesting because, one, it was like a new form factor that, like, brought in a new type of user, not just a developer, but like a productivity or a, like an internet creator came to AI for the first time. And it also had an interesting architecture where it was, like, calling your chosen model for these heartbeats to see if it was still alive in addition to using the model for real tasks. And the heartbeats are like, they’re kind

OpenClaw, Hermes, and the Auto Router

Swyx [01:07:02]: Fréquence.

Anjney Midha [01:07:02]: You don’t wanna pay a lot of

Swyx [01:07:03]: Every thirty minutes

Anjney Midha [01:07:04]: To do a heartbeat.

Swyx [01:07:05]: Yeah.

Anjney Midha [01:07:05]: So, the auto router that we provided was really useful to this, like, wide range of users all of a sudden. And so we just saw it rocket exponentially, and then we saw, like OpenClaw just blow up and a couple other, apps lean into that new paradigm and do something similar. Hermes came out and really leaned into things like the auto router and built, like, a really good community and leaned into, like, skill management and making it really easy and effective for people to, like, set their memory in the agent

Swyx [01:07:44]: Yeah.

Anjney Midha [01:07:44]: And build really good skills.

Swyx [01:07:45]: Which another thing you never did, memory skills, sandboxes, all these, like, adjacent things you could have done.

Anjney Midha [01:07:52]: Could have, but It’- I think,

Swyx [01:07:54]: It’s hard to bet.

Anjney Midha [01:07:55]: They’re also - There are things that developer-- that really matter for, like, the developer use cases that were coming out at the time. Like, developers wanted to architect those things.

Swyx [01:08:05]: Right.

Anjney Midha [01:08:05]: Those were kinda critical to building a good user experience. It’s really-- It was, like, - It’s been hard for companies to find abstractions that work for all developers on the memory layer. It is, it - Yeah, there are some, like Mastra has done a pretty good job, for example. But, like, developers have, like, lots of varied preferences for them. And then we - - the way our leaderboard has changed over time is like a movie of how the AI space has changed over time. If you just like, go to the Wayback Machine and look at the rankings leaderboard and the apps leaderboard over time, it shows you, like, what’s happened in AI over the last couple of years.

Swyx [01:08:48]: To me, the coming of age moment was, Andrej Karpathy was like, “I no longer read Local Llama ‘cause, like, I just go to OpenClaw-- OpenRouter’s leaderboard.”

Leaderboards as a Map of the AI Ecosystem

Swyx [01:08:57]: Which I remember that. Yeah. I think he probably, like, said, like, “Sorry, guys, I’m gonna send a bunch of traffic to you.”

Swyx [01:09:03]: So I also wanna bring it into the Stripe, thing.

Why Stripe Acquired OpenRouter

Swyx [01:09:07]: How does that conversation start?

Anjney Midha [01:09:09]: We had this longstanding relationship with Stripe, though, from, like, many different projects that we had worked on with them. We invest, a lot of effort in countering abuse,

Swyx [01:09:24]: Token fraud.

Anjney Midha [01:09:24]: And token fraud.

Swyx [01:09:26]: Can you give some numbers just - so people understand?

Anjney Midha [01:09:29]: I think I, like, I posted about this. We blocked 10x as much dollar volume last month as the month before. And the types of token fraud are diversifying quite a bit. there are, like, fraudsters going after typical stolen credit cards, but there are also, people trying to resell traffic against the terms of service. There’s, like, hacked accounts. There’s people who just lose - like, their whole company is compromised, and they don’t even realize it, and we help them, like, regain control and detect it. There’- There are accounts that are, like, reselling inference on the side. There’- There are accounts that are dealing with, a, like, an accidental runaway agent, and they don’t realize it. Not a hack, but it’s something that blows up and the company doesn’t want it. And so our trust and safety team, like, works a lot on all of these, like, categories of problems and helps block it and detect it. And so we’ve built these. we have models around them. We - We worked closely with Stripe for a while on this, and I think it’s gonna become a huge problem in the ecosystem. Like, we’re already seeing a lot of companies start to see these fraudsters, like, spread and look for other ways other than OpenRouter to other fraud vectors. And if you’re making a gateway or selling, like, generalized inference, you are a target for fraud. If you’re selling very discreet, like, intelligence products that are, like, doing something pretty specific, but not, like, just reselling inference with some added capability, then you’re way less likely to get these fraudsters. So - I think we’ll see companies also move away from just reselling inference with some like, added capability and move towards like, discreet tasks and charging for those tasks and charging for those enhancements and letting people bring their own inference, like, in a party way.

Fraud, Abuse, and the Emerging Token Economy

Swyx [01:11:39]: Whoa. Okay. and yeah, obviously you would power that.

Anjney Midha [01:11:44]: Right.

Swyx [01:11:44]: But you - People pay, for outcomes Or per task?

Anjney Midha [01:11:48]: I think people will pay. I think, like, the Datadog pricing page is a good look at, like, the future to come. It’s like companies, like infrastructure companies will, like, charge for different types of events that they’re providing, and there’ll be lots of, like, continuous pricing models that look like that. And of course, there will be, like, if you go down, towards consumer apps, simpler pricing, more subscriptions, fewer events to worry about, and ones that, like, are not. Focus on just adding a markup on top of inference.

The Token Economy and Security at Scale

Swyx [01:12:28]: Yeah.

Anjney Midha [01:12:28]: Not just because fraud is hard, but also because the pressure from the labs and from - like, good inference providers to, like, do a commit and then bring your inference elsewhere is gonna be very high.

Swyx [01:12:44]: Any comments?

Alex Atallah [01:12:45]: Two. One, I think Alex has done a very eloquent job of describing something, counterintuitively I knew would be a thing at scale, like four years ago because of Discord. And the particular experience that taught me this was, as we started scaling Midjourney, - one of the primary ways that we used to give away or, like, get people to try Midjourney early on to get to their first ten generations. Because, ten generations - ten images generated was roughly the magic moment activation point we found. Like, once you’d done ten, you were like, “This is extraordinary.” but for that week, so we had a free trial with Midjourney. And one day I woke up, because I was the head of platform and had to monitor, I had all these dashboards, and I had, like, three missed calls from David. And it turns out, like, there had been this flood of new users overnight. And we were like, “This is great.” And he was like, “No, we shut down the free trial.” And I was like, “Why is that?” and he said, “I want you to look at the geolocation IP addresses.” And somebody in China had started to resell Midjourney free, subscriptions with the free trial as a way to, like, you - It was fraud abuse, right?

Swyx [01:13:54]: Even for a specialized model like Midjourney.

Alex Atallah [01:13:56]: Yeah. And that was an application. So this idea - I think the big picture realization I had back then was, hey, there’s a new type of unit of value that’s being streamed across the internet called a token.

Alex Atallah [01:14:11]: And over the next ten years, the entire internet value chain was going to have to deal with the fact that, like, the more valuable tokens got, The more bad actors are gonna go to try to get their hands on those tokens. And anytime you scale something and the payload gets more and more valuable, More bad things, people try to get access to that value. And so it was very obvious to me back then. And so, look, to this day, I don’t think there’s a free turn. Like, I don’t think Midjourney’s ever turned on the free trial since then, because it was really not an easy problem to solve in terms of trust and safety. that’s why I - started teaching the class Security at Scale at Stanford. Like, it was like one of - that and the Anthropic learnings, to me, it was clear that the need for security at scale is gonna be enormous a few years from then. Because if you just do the math, right, think about, like, if we’re. online payments, has started roughly in the eighties and nineties, right, and grew to over a trillion dollars over the next ten years, and we needed to build entirely new payment solutions to deal with online fraud. where we are today is roughly there on tokens, but over the next even five years, we’re expecting the token economy to get to, like, roughly 5 trillion dollars. And over the next ten years, I’d be shocked if we weren’t at 10 trillion dollars of token flow. And so if we were starting to see such aggressive abuse and fraud at subscale, Midjourney, remember Midjourney at this point was, like, less than three $100 million revenue run rate a year.

Alex Atallah [01:15:44]: I just realized we were gonna need, like, entirely new, Like, systems to deal with the fraud that was gonna happen for trying to get into the token flow. And so, - I, - I forget the board meeting it was when you brought up that, Stripe wanted to partner up, and it made so much sense to me because Stripe Radar. When I was at Kleiner ten years ago, we invested in Stripe, and the whole pitch that, Patrick and John communicate so eloquently was like, “Hey, unlike traditional payment tools like Braintree that do a day verification, like KYC and AML to get the fraud out of the way, we just bite the fraud cost upfront as customer acquisition cost and - tell a developer, like, just use five lines of code, and we start accepting your payments in five minutes. And what’ll happen is over time, we collect all this data on the developers.”

Swyx [01:16:31]: Cloudflare model.

Alex Atallah [01:16:32]: Is the Cloudflare model, right? And they did. Five years later, they launched Stripe Radar, and Stripe really today is a security company. That’s the real. People think it’s a payments company. No, the reason. There’s lots of other payments providers today that give you, like, cheaper payments transmission. But the reason Stripe keeps, being the dominant one here and Adyen and Europe is because they have extraordinary fraud detection that they’ve built, - over the years.

Swyx [01:16:52]: It’s the same story with Elon and Max Levchin

Alex Atallah [01:16:55]: And affirm, yeah.

Swyx [01:16:56]: Yeah.

Alex Atallah [01:16:57]: So, I think the story shows up over and over again, where every time you have value streamed across the world in large amounts, you need new protection and security infrastructure to fight, to keep the bad guys out and allow the good people to, like, have their transactions happen really fast. And so I think, - this is why - from my perspective, like, the Stripe and OpenRouter story is a security story for the internet ecosystem, for the frontier AI ecosystem. Without a partnership like that, it becomes very hard to defend the quality of experience and the speed and all the good stuff without letting the bad guys get in the way. the second is that, there’s this underappreciated thing about, like, the fact that you need to. Like, - all the bad things that Alex described as being perpetuated by humans right now is going to be perpetuated by AI agents over the next ten years.

Swyx [01:17:46]: Oof.

Alex Atallah [01:17:47]: Right? So think about the, like, recursive scale we’re about to see of bad actors. It’s not just bad human beings, it’s, it’s all the bad agents that are gonna be attacking the token flow. And there’s. It’s very hard if you’re a researcher and at an AI lab to reason about that problem because the only data you have is how agents you’re training are going rogue. But that’s just a fraction of all the bad behavior on the internet that we’re gonna see. And so what you need is defenders, new sheriffs in town, which cowboy hats, that can see all the bad behavior from AI agents across the ecosystem, from different model labs and different trained deployments and different developers, and take all of that data and say, “We’re gonna build a shield for the entire token economy.” Because without that, the amount of fraud we’re gonna see of this 10 trillion dollars in GMV and global GDP growth is, like, a huge percentage of that, I think, is going to be fraud, abuse. And we might never get there if people just don’t trust. Tokens, right? and I don’t think this infrastructure exists. So you have your work cut out for you with, at Stripe, but I don’t think people have realized the scale at which agents, agent, agentic fraud, like bad behavior perpetuated by AI agents is about to hit us like a tsunami.

OpenRouter + Stripe: What Changes Next

Swyx [01:18:58]: Yeah. there’s a lot to dig into there. I wanna give you the last word. We do have to wrap. what can people expect from OpenRouter and Stripe?

Anjney Midha [01:19:07]: I think this is a really good way for us to accelerate market and, to go upmarket more quickly. It’s also, as Ansh eloquently described, this is, there’s a really clear better together story here when it comes to improving trust and safety and making it really easy to, like, accept tokens and let people bring their own inference to your app and to help developers just, like, build on top of inference, going forward. We have a really strong brand with OpenRouter, and we’re keeping the brand. So, like, OpenRouter, like, as a product and the roadmap and the name and the brand, like, is staying the same. And so what, like, you should expect, in the next six months is that most things will be like what we would have done had we been independent, except everything will be moving faster. And that’s like our, term goal. Longer term, hopefully I can comment on it soon, but I can’

Closing: Building the Infrastructure for the Token Economy

Anjney Midha [01:20:11]: Now.

Swyx [01:20:11]: Okay. Well, we’ll hopefully do a follow-up at some point, but thank you for being so generous with your time, and, congrats on the partnership. this is one of the most beautiful bromances I’ve seen in AI.

Alex Atallah [01:20:22]: Just starting out.

Swyx [01:20:23]: Starting from Stanford

Alex Atallah [01:20:24]: Just starting.

Swyx [01:20:24]: To here.

Alex Atallah [01:20:24]: Yeah. Lots more to do.

Anjney Midha [01:20:26]: Yeah.

Alex Atallah [01:20:26]: Lots of sheriff, policing to do of the, of

Swyx [01:20:29]: Yes. The cowboys in town.

Alex Atallah [01:20:30]: Of the token economy. We need We need new sheriffs for sure.

Swyx [01:20:33]: Yeah. Awesome. Thank you.

Anjney Midha [01:20:35]: Thank you.



This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

More from Latent Space: The AI Engineer Podcast

All 247 episodes
OpenRouter: from Seed to Stripe — with OpenRouter’s Alex Atallah & AMP’s Anjney MidhaLatent Space: The AI Engineer Podcast · 1 h 21 min
Listen in VO