#259 Anjney Midha: a16z’s Strategy to Turn AI Startups into Unicorns

2 Jun 2025 · 57 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

```markdown Eye On A.I. Episode #259: Anjney Midha - a16z’s Strategy to Turn AI Startups into Unicorns

Podcast Overview

  • Host: Craig S. Smith
  • Guest: Anjney Midha, General Partner at Andreessen Horowitz (a16z)
  • Topic: Transforming cutting-edge AI research into scalable, real-world businesses.
  • Focus: Challenges and strategies in AI infrastructure, model reliability, and evaluation.

---

Key Takeaways

Introduction to AI Startups

  • Discussion on the journey of turning advanced AI research into viable companies.
  • Insight into a16z's approach to investing in AI startups.

Anjney Midha's Background

  • Early involvement in machine learning at Stanford, focusing on bioinformatics.
  • Transitioned to venture capital at Kleiner Perkins.
  • Founded Ubiquity 6, which focused on 3D mapping for applications like robotics and gaming.
  • Joined a16z after selling his company to Discord and became involved with projects like Anthropic.

Notable Projects

  • Anthropic: Early investment and support for its growth; significant focus on compute requirements for foundational models.
  • Mistral and Stable Diffusion: Investments and support in developing open-source language models and image generation technologies.

Current Trends in AI

  • Model Architectures: Discussion on the shift towards hybrid architectures combining various methods (transformers, diffusion models, LSTMs).
  • AI Evaluation: Critique of existing benchmarks and the need for real-world evaluation methodologies.
  • Concern over models being trained to beat static benchmarks rather than solving practical problems.

The Future of AI Agents

  • Definition of agentic systems where models can act autonomously, not just produce text.
  • Excitement about code workflows and how agents are currently effective there.
  • Challenge of reliability and accuracy in non-deterministic agentic applications.

Insights on Investment Strategy

  • Importance of understanding scientific advancements and commercializing them.
  • Need for deep conversations with researchers to help translate technology into marketable products.
  • Emphasis on the collaboration of scientific expertise with product vision to create solutions for customers.

Emerging Challenges

  • Reliability in AI: Focus on making models interpretable and developing better evaluation metrics.
  • Debate over whether to invest in open-source models versus proprietary ones.

---

Episode Highlights

Major Themes Discussed

  • Transition from general models to more specialized last-mile solutions.
  • Importance of AI reliability, especially in mission-critical applications such as healthcare and finance.
  • Future potential of AI agents in automating complex tasks with an emphasis on verification and reliability.

Anjney's Perspective on Investments

  • Focus on high-quality, scalable AI products rather than just models.
  • Acknowledgment that while models are critical, successful products require a comprehensive solution tailored to user needs.

Summary of Insights

  • AI's evolution requires a blend of traditional engineering and deep scientific understanding.
  • The necessity for AI to move toward real-world applications with reliable evaluations.
  • Insights into the future of AI development and its implications for industries.

---

Conclusion The episode provides an in-depth understanding of how leading minds in AI are navigating the complex landscape of technology commercialization. Anjney Midha's insights on evaluating AI models, investing in emerging technologies, and developing reliable systems are crucial for anyone interested in the future of artificial intelligence.

For more updates, follow Craig Smith and Eye on A.I. on X:

  • Craig Smith on X: [@craigss](https://x.com/craigss)
  • Eye on A.I. on X: [@EyeOn_AI](https://x.com/EyeOn_AI)

--- ```

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00We're entering a regime where the model is taking action versus just producing words. we're starting to get into an agentic system. The more autonomy you give that model, the ability to call those tools on its own and self-learn, the more agentic I guess it is. But we are squarely already in the middle of agents. I would say a lot of code workflows is probably where I'm most excited about agents actually working. The last few years have been the unabashed growth of transformer pre-training scaling. And that meant for a long time, people thought that meant that general models would win. And I think actually now we're transitioning to an era where general models with specific last miles are winning.

0:33Building multi-agent software is hard. Agent-to-agent and agent-to-tool communication is still the wild west. How do you achieve accuracy and consistency in non-deterministic agentic apps? That's where agency, A-G-N-T-C-Y, comes in. The agency is an open source collective building the internet of agents. And what's the internet of agents? It's a collaboration layer where AI agents can communicate, discover each other, and work across frameworks. For developers, this means standardized agent discovery tools, seamless protocols for interagent communication, and modular components to compose and scale multi-agent workflows.

1:26Build with other engineers who care about high-quality multi-agent software. Visit agency.org and add your support. That's A-G-N-T-C-Y dot O-R-G. Visit them today to support high-quality multi-agent software. So yeah, Anjani, can you talk about your background before a 16 z and uh and and about the anthropic angel investment that's that's pretty fascinating but obviously you must have done quite a bit before that to be an angel investor uh i i got my start a long time before that in machine learning as a graduate student at stanford i was in the bioinformatics department there and this was when there was an earlier wave of machine learning around 2011 and 12, where everybody in Silicon Valley thought they needed a machine learning strategy.

2:25But I was mostly on my way to a career in academia. I was really fascinated by using some of the new techniques that were becoming possible in deep learning in healthcare. And being at Stanford at the time in the med school meant that I got access to tons and tons of really fantastic, really valuable data, sort of clinical data, healthcare data, patient data. And I thought I was going to go be an academic when I got a chance to spend a summer at a firm called Kleiner Perkins, basically helping their portfolio companies get started with their machine learning pipelines at a time when everybody in Silicon Valley thought they needed it.

3:09So one of the partners there said, look, we need somebody to be our first machine learning engineer. And that's kind of how I first learned about venture capital. I ended up spending four and a half years there at Kleiner as an investor. They ended up asking me to come over from the engineering side to investing, both at the early stage and the growth stages in what now we would call developer infrastructure businesses. But at the time, it was basically companies that were figuring out how to build new products and services on top of GPUs that had only recently become available in the cloud, on places like AWS and so on.

3:47And then around 2017 is when I decided to leave and start my own company called Ubiquity 6. And our goal was to try to use some of these modern deep learning techniques, primarily in computer vision, to allow high precision 3D mapping for applications like indoor robotics and self-driving cars. At the time, there was this explosion in computer vision AI techniques that made it very possible to do high-precision 3D reconstruction on commodity hardware like smartphones. It turns out that there's a little app that took over the world called Pokemon Go that you might remember around that time. And that ended up being one of our primary sources of customer demand was game developers trying to build applications, augmented reality applications in real world spaces.

4:43And we ended up serving millions of users who were walking around the world mapping real world locations for us. And one of the biggest bottlenecks that we had to solve for them was massively multiplayer networking, which meant when you had hundreds of people in one physical location, how do you keep everybody's physical coordinates updated in real time? It's kind of a hard networking problem to update everybody's positions in real time. And so we built a fair amount of interesting technology to allow that. And it turns out when the pandemic happened a couple of years later, and everybody was online, not walking around in the real world, but doing a bunch of multiplayer experiences online, that infrastructure became really valuable because we had built essentially a form of networking called serverless networking, which meant it was really easy for thousands and thousands of people to get together in the same sort of virtual experience very quickly without it costing the developer a ton of capital.

5:42And so I ended up selling that business to Discord at the end of 2020, which had also exploded in the pandemic from being primarily a chat application for gamers to being used by all kinds of groups. And there, my job was to both integrate our backend infrastructure into the company and launch the company's developer platform business. They wanted to allow developers to build third-party applications on top of that infra we'd built. And this was around the time when I got the call from friends who were running research at OpenAI. Tom Brown, who was one of the leads on GPT-3 and had been a longtime personal friend, gave me a call and said, Anj, we're thinking about leaving and starting this new company.

6:26We'd love your help on figuring out how to fundraise and do all the foundry things you've done as a CEO and founder who's sold his company. And so for the first six months of 2021, Dario, Tom, and I just did many, many working sessions on business plan, capital fundraising, how to assemble sort of the commercial side of Anthropic. topic. And then, you know, a few months later is when I was I came on as their angel investor and help them then go raise the first hundred million dollar seed round. And then later, we we got fairly involved in acquiring lots of compute for the company. It was a it was sort of net.

7:09Today, it's well understood that foundation model businesses need hundreds of millions of dollars in compute to get going. But at the time it was a fairly novel thing. And around the end of that year, 2021, is when I started realizing that the scaling laws and language models that we were seeing at Anthropic were going to hold for other modalities like image and video. And I had the chance to, in building the developer platform business at Discord, one of the things you always have to do is figure out what are the early killer apps that you think will drive value right on the platform. And I had the chance to team up with an old friend who was working on a little text to image app called Midjourney.

7:49And we created an app store front page at Discord where we put Midjourney on the front page and helped them launch one of the world's first text to image models. And the Midjourney team did an extraordinary job taking that business from basically zero to many tens of millions of dollars in revenue very quickly. And that kind of gave me the inkling that we were in the middle of a new platform ship fundamentally where entirely new startups could be built in new modalities, whether that was video or audio or code. And I left then Discord in early 2023 to start investing in companies full time and helping researchers and scientists take their research out of the lab and turn them into new foundation model infrastructure businesses.

8:37And since then, you know, I've had the chance to work with some amazing folks like the founders of Mistral and who are working on open source language models. Robin Rombach and team who are the creators of Stable Diffusion and then founder a company called Black Forest Labs that I'm on the board of. And that that that's why I that's what I spend most of my days doing is working with researchers and scientists, turning their research into into businesses. That's why I joined Andreessen Horowitz shortly after I left Discord to do that full time. And that's what I spend most of my days doing now.

9:11Yeah. Wow. Going back to the Pokemon Go, you weren't involved in building that, were you? No, no, no. I was not at Niantic, but they were one of the developers who wanted to use our technology. And early on, actually, John Henke gave me a call and said, this is exactly what we've been looking for. How quickly can you get the service up and running? So I really owe it to him for, you know, gaming was not at all on our roadmap. Indoor robotics, indoor navigation was kind of our primary focus. And I credit Niantic entirely for opening our eyes to the world of location-based gaming. Yeah. And with then from Anthropic, you stayed focused on foundation models.

9:57I mean, at least for a time. And just maybe to jump ahead a little bit, what I'm interested in hearing is how you see, I mean, we're all familiar at this point with foundation models and they raise some issues that I'd like to hear your thoughts on, you know, the open source versus proprietary debate. of and and then you know distillation and small models models for the edge but recently i've been talking to people about you know kind of backwaters in ai research that have been passed by uh during the transformer uh you know phase but when you started you were talking about uh vision right uh you know that was a super supervised learning phase and that's all anybody talked about and everyone was you know drilling down and in these incremental improvements and then transformers came along now everyone's you know working on how to uh you know tweak foundation models to do this or that or uh but there are these other strategies You know, I had on the podcast a bit ago, Sep Hockrider, you know, that's not how the Germans pronounce it, but, you know, the guy who came up with LSTM.

11:36Right. And he's, you know, still pursuing that architecture and doing really interesting things with it and things. There's a lot of promise. reinforcement learning was kind of in the background for a long time and and it's it's being talked about again because it's being used to train foundation models but you talk to the guys up in alberta and and they're doing amazing things and have great optimism about building intelligence using rl alone or at least as a primary strategy. And then I was talking to yesterday this guy, Friston, with the free energy principle and energy-based models. And there's all kinds of stuff going on there.

12:35You know, there's Mamba now, and that's opened up. You know, I had a call with, I'm not going to be a member. I'll cut that out. But anyways, the guys who have productized that in a model called Jamba. Right. So, I mean, are you focused on sort of transformer-based models, or are you scanning the horizon for these other strategies to see which ones might become important in the product space? No, look, I'm constantly looking for new architectures because the idea that a one-size-fits-all architecture just works is deeply counterintuitive, right? Yeah. And this is why I think it's called a bitter lesson, which is that if you get your training as a machine learning, in sort of classical machine learning, what you learn is that there should be specialized architectures.

13:42If you value efficiency, that there should be a more efficient architecture for every different use case or task. That task-specific architectures historically have been more efficient when it comes to production, right? Yeah. And that's just turned out to be unintuitively, frustratingly untrue, right? Like the transformer has just marched ahead. And I think what's going on is that the bitter lesson is just turning out to be true because Moore's law continues being one of the biggest drivers of efficiency of computational workflows. And what we're learning is that, you know, to be honest, these transformer models are so efficient at extracting sort of latent representations of knowledge that in a sense, the architecture, you don't have to get that fancy on architecture if you can process and filter the data correctly.

14:37Now, the big question I think is, the right one you're asking is, where's the limit, right? of how far do transformers get us and I I have been shocked basically by how good RL is at improving performance um let me put it more more clearly what I I do think that we were the last few years have been the unabashed growth of of transformer pre-training scaling Yeah. And that meant for a long time, people thought that meant that general models would win. And I think actually now we're transitioning to an era where general models with specific

15:28last miles are winning. That last mile can often be a combination of an LSTM, an SSM, a completely different non-transformer based approach. And in some cases, they're diffusion models. right so if you if you take the case of multimodal image generation um i think we're heading to an era of transfusion where you take a transformer which is autoregressive and then you combine it with some of the best parts of diffusion so that you get quality um in image generation and that's where we're at we i do think we're headed to an era of hybrid architectures i'm not a transformer maximalist but i do think the network effects of of infrastructure where a lot of the world's like sort of inference serving pipelines are now designed to be for auto aggressive you know work makes it makes the bar higher and higher for a new architecture to to try to jump over before it gets adopted yeah yeah yeah uh but you're hopeful or or you anticipate that there will be a new architecture I mean no one saw transformers coming um you know this is where depending on who you ask if if you ask a traditional machine learning researcher they'll tell you there are no new architectures they've all been invented before you know and a hawk rider will tell you that so will you know the in the entire there's a whole body of of sort of european researchers who will tell you that the transformers is basically just you know sequence to sequence learning which was invented 40 years ago you know with with boltzmann machines right yeah um and so some i think i think there's two sides of the argument one side one school will tell you no way craig there there there are no new architectures we've all invented them before this is just the new kids on the block inventing stuff we've done 40 years ago um and then there'll be the other school of thought which is which is often non-traditional researchers you know the the anthropic guys actually came from physics backgrounds you know dario has a phd in biophysics um jared kaplan who's a chief scientist and you know was was sam did the formative sort of seminal work around scaling laws was a physicist by training not a computer scientist and they'll tell you of course these are new architectures and of course we'll have you know sort of this is fundamentally a new body of work and we might have you know new new um representation learning mechanisms in the future and i to be honest with you my where i fall in the debate is does it even matter right because the thing that continues is sort of being the most important is that they just, they work and they're driving such incredible advancements in the end user's experience that none of us expected that it doesn't really matter to me very much.

18:10Even though my training is in machine learning, I have learned to let go of the ML purist in me and stop worrying about the architecture and asking, you know, do we have enough training data for the models to keep improving? yeah um so what what kind of things you're looking at now that that are sort of beyond their horizon look the the number one thing that keeps me up at night right at night right now is evaluation right how do you actually tell how good these models are we're we're past well past the era of the low-hanging fruit right where when gpt3 first came out and you know daria and tom sent me a screenshot of of their ablations that showed just by increasing the compute by 100x on on a bunch of tasks like um you know they had they had a chart showing how well a human could detect if a piece of if a news if a news article had been written written by an ai model or not and just by 100xing the compute they basically crossed the turing test on on that path yeah right and there were these low-hanging fruits sort of tasks all around us two years ago four years ago actually at this point where in coding for example just by 100xing the compute gpd3 relative to gpd2 was able to solve programming problems that had never seen before in the training data right there was no even if there was no python code in the training data and there's only javascript it was able to still solve python problems during inference time which is what what they'd call in-context learning.

19:45I think now we're headed to the era where those academic benchmarks like MMLU and so on are no longer really, we have moved from the research era to the deployment era. And for the deployment era, you need benchmarks and evaluation to tell you how good these models are in the real world. And to do that, you need a completely different mindset and methodology than I think the previous era, which is you'd have these sort of static benchmarks. benchmarks, you'd kind of train a model to beat those benchmarks, then you present at a research conference. And that was good for the low hanging fruit era.

20:20Now we're in the era where if you want to know if a new model actually solves a physician's problem in the field or a scientist's problem in the field, well, you need these models to be trained in that real world evaluation loop. And that's what I'm looking for right now. Yeah. Yeah, and benchmarks from a layman's point of view, you can train to a benchmark. And I don't necessarily trust these companies to be doing it correctly. Exactly. So everyone comes out with their chart and their model is like the middle finger pointing up there against everybody else. But who knows how they got around to that.

21:07Well, Goodhart's law is real, right? Which is that when a measure starts becoming a target, it ceases to become a good measure, which is exactly what happened with all these static academic benchmarks like MMLU and so on. So now I've become much more partial to crowdsource the wisdom of the crowd, things like LMSIS, right? Where you have real world users show up and rank two side-by-side responses from LLMs. And that's much, much harder to try to sort of, I think it's much harder for that to be overfitted on. Now, you could debate that you can still overfit on that, sure, but it gets harder and harder to do that at scale.

21:47It gets harder and harder to do that in real world tasks. And therefore, I'm much more a fan of sort of that type of evaluation to solve. Look, the big problem to answer here is AI reliability, right? Yeah. The biggest problem holding back AI models from being useful, I would say, in the most mission critical industries of life, you know, defense, health care, financial services is reliability. And I'm just very convinced that the answer there lies in good, robust measurement. right and and that means evaluation with real world with real users in real world context because otherwise the measurements from sort of academic setting are not reliable benchmarks or measures of performance ultimate yeah uh are you looking at uh at any of these you mentioned uh LSTM are you you know the Huck Ryder am I saying it right has a new company I'm not going to be able to remember uh NXAI or something like that where he's trying to productize his the advances he's made on LSTM architecture I I haven't talked to him about the company I'm not familiar with what their particular plan is but i would say there's a number of companies in that vein who believe that specialized non-transformer architectures are more reliable because they're more interpretable in the real world right yeah as opposed to uh trans the transformers a black box right and so there's only two solutions i think either you have an you you have an architectural change that um makes these black boxes transparent right and the body of work they're called broadly speaking interpretability that that i think is fantastic i spend a lot of time paying attention to what's happening in interpretability there's actually a couple of great papers that came out last week from anthropic that show that we're beginning to to to peer into what what's going on in the transformer um so that's either either you've turned the black box into a clear box or the second approach is you start measuring the output of the models to be sufficiently predictable that they're reliable.

24:15And I'm kind of bullish on both. I don't think you need an architectural shift like LSTMs to solve that. I think if you could make transformers more interpretable, you could get the same benefit. Yeah. You're talking about Anthropik's work on I can't remember now, intelligence graphs or what did they call it? But they've developed a tool to kind of like a functional MRI to be able to look in and see which nodes in a network are being activated at different stages of it. Exactly. We call it circuit tracing, right? Right. And the idea is that you can group the parts of a neural network that are activated when you ask it something into these clusters they call features.

25:05And seeing which clusters of these features fire for a particular topic, you can start to trace why a language model says what it does. And you're exactly right. The analogy would be inventing a microscope for what's going on in the LLM. once you have a microscope then you can move on to the part the next step which is okay can we start actually gene editing if you don't have a microscope to look at the genes then you can't actually start editing but once you have a microscope then you can start talking about how to actually do steering and and and control yeah and it seems to me that they they were able to do that in those experiments they the i'm thinking there was uh one of the examples was they started a poem with a word that rhymes with rabbit.

25:56I can't remember. No, that's exactly right. That's right. Yeah, and they could see that the model was looking at habit and, or no, it was carrot, or habit and carrot, and they could sort of direct the model toward one or the other. Exactly. I think what they identified was that the models plan ahead for where they're going to go. And if you can kind of trace the circuits that the model uses to plan ahead, and you can block out or mask certain endpoints of those traces, then you can steer it towards other endpoints. And so the idea was that they could get it to write a different second line of the poem than it was originally planning to by kind of masking, you know, the original plan.

26:46And I think it's very interesting. I think we're starting to get, it's finally starting to get with interpretability going from research to engineering, right? Because it's obviously a toy example, but this is how we do engineering, right? we you first try to prove it in a petri dish and then if it works then you start scaling it up and so i i you know my my understanding is the next step now is to try to scale it up to prove that they can actually steer it in a useful enough fashion outside of those toy examples it's very exciting but up until until that happens and it may be yours i'm a big fan of evaluation uh in real world scenarios like you know lmsys and chat arena i think that move is is is positive i do think

27:34um one one of the core tensions in reliability in in in llms is control over the weights and which is why we're entering an era where geopolitics is obviously pretty important lots of sovereign nations now want their own ai stack they want their own models because ultimately unless the models are fully interpretable there's no other real way to control them unless you have the weights. I guess I would say there's three big solutions to AI reliability. One is mechanistic interpretability or just making the models more interpretable in general. You could have side-by-side evaluation, things like Arena, or you could have open-source, open-weight models that give full control over the inference to the customer running the inference.

28:20I spend my time thinking about all three of those things right now. yeah well what about agents that's the other hot area that that everyone's talking i would think that you guys would be all over that or is that someone else at the firm oh i think that everybody's got to contend with models going from being next word prediction machines to next action prediction machines right and the beauty of the bitter or the the bitter lesson playing out with with transformers is that it's, it doesn't, it's not that hard to fine tune a model to take action on what, you know, would be called tool calling, right?

28:59So the first kind of generation of agents we're seeing right now is just LLMs trained to call APIs when they need help, right? And so you can tell an LLM, hey, I need, tell me the time today. And if you don't know what the time is, then that's fine. Just call out to a clock. Now, some people will tell you that's not really an agent but i i i don't i'm not that dogmatic about these definitions to me basically if if we're entering a regime where the model is taking action versus just producing words we're starting to get into an agentic system um the more autonomy you give that model the ability to call those tools on its own and self-learn the more agentic i guess it is um but we are squarely already in the middle of agents i would say you know a lot of um code workflows is probably where i'm i'm i'm most excited about agents actually working you know the reality is for things like browsing general purpose agents are very brittle right yeah but in workflows like coding uh where the outcome's actually quite easy to verify and score you have we have unit tests right in software so that's why in software engineering and code generation it's quite easy for agents to actually be told whether be given a scorecard.

30:13And I think if your question is, how do you measure AI reliability with agents? I think it's very simple. The general heuristic is exactly how you would score a human being on that task. You write a bunch of tests and you figure out whether it passes those evaluations or not. And the good thing about programming and software is that a lot of those tests are automated by design. Building multi-agent software is hard. Agent-to-agent and agent-to-tool communication is still the Wild West. How do you achieve accuracy and consistency in non-deterministic agentic apps? That's where agency, A-G-N-T-C-Y, comes in.

30:59The agency is an open source collective building the Internet of Agents. And what's the Internet of Agents? It's a collaboration layer where AI agents can communicate, discover each other, and work across frameworks. For developers, this means standardized agent discovery tools, seamless protocols for interagent communication, and modular components to compose and scale multi-agent workflows. Build with other engineers who care about high-quality multi-agent software. Visit agency.org and add your support. That's A-G-N-T-C-Y dot O-R-G. Visit them today to support high-quality multi-agent software.

31:51Yeah. Two questions for someone in your position. One, how do you keep your finger on all of this? Are you like me? You subscribe to a million newsletters. You have a million meetings. uh and you're capturing the conversations uh with some AI like otter I use otter other people use other things and then you're kind of constantly uh you know summarizing and looking through for ideas or or are you relying on on deal flow and what people are coming to you with I think so the answer is yes, all of the above. But there's just no substitute, I find, for deep, long-form conversations with scientists.

32:50Yeah, that's right. And so this was the moment when three, four years ago, when Dario and Tom first gave me that call after GPT-3 and said, you know, we think we're, we've discovered this thing called scaling laws. And we are, we think there's a new lab to be built that prioritizes interpretability first. And I said, sure. And they said, you know, can you help us out as an investor? And I said, sure. How much do you need to get started? And Dario said, I think we can get by with five. And I said, okay, look, 5 million, not a problem. I should be wired over next week. And he said, well, you're off by a couple of zeros there i'm going to need 500 million and i said okay things are going to be a little different but sure walk me through you know usually i find that when when some of the most the leading scientists uh kind of challenge a base assumption you have it's usually a good and you feel uncomfortable it's it's a good i i find it's it's a good practice to dig into that on discomfort and kind of big ask the five whys on from a first principles basis why are they so convicted in that right and um you know in in in his case dario tom and i spent many weeks unpacking why they felt they needed 500 million at the end of that conversation i realized actually that 500 million was extraordinarily efficient to train to build a lab of a capital race to build a lab around because by you know by then open ai had spent over four billion dollars getting to that point yeah and that's when we really got into the idea of compute multipliers and why having compute multipliers allows you to train uh models for six times more efficient uh compute spends than you know the first generation and and i find those long-form conversations you just can't substitute no amount of reading a paper can explain okay can can compensate for those in-person working sessions and that's what i love doing with all the founders you know i work with I often find it most rewarding to be the first call for a scientist or a researcher about commercializing their research before there even is a company.

34:56And that's what I spend my days usually filling. And I find that often means you're quite early. The problem with spending your days with scientists who are living two steps ahead is that most of the world today is not ready for their realizations. And so when I introduced them to, you know, Dario to 22 other investors up and down Sandhill Road, they got 21 no's. And, you know, the reactions were everything from borderline, this is snake oil, to, you know, it just doesn't make sense. There's no way that just throwing more compute will eventually produce models that will be able to solve all intelligence problems under the sun.

35:32It's so absurd. Now, they turned out to be right. And so I often find it's not easy to communicate my excitement about these realizations to other people in the field. So I actually I try to spend as little of my time talking to other investors as I can, because I find by the time other investors have conviction in something, it's already consensus. Right. Yeah. What I'd like to spend most of my days is in long form conversations with with scientists and researchers who are in the lab. and I find the most exciting moment is when science is going from being or some strain of technology is going from being inevitable to being imminent because that's what can be most useful to them because they're often looking for my help to commercialize their work and scale the impact that they can have by turning their research into a product that could be used by millions of people and there's no papers you can really read on that it's often the hard work of sitting down and discussing face-to-face the technology, the capital required, the compute needs, what the industry needs.

36:43And that messy search for product market fit is what I find most exciting. Yeah. And how many was this? Yeah. And I read a lot of papers and I'm a journalist. uh you know and i i'm not a practitioner i'm not an ml person i mean i didn't go to school for that and a lot of times i'll see something and i'll think wow you know you could you could build a product off of that uh and a lot of that research just gets shelved or buried in in archive and no one ever does anything and maybe you know maybe it's the idea is rediscovered 10 years later and someone says oh yeah well so and so wrote that uh 10 years ago is most of your i mean are you investing i mean that's that's that is your role right yeah that's right i'm a general partner i spend my days investing um but i i think the the discipline of investing weirdly has changed so dramatically in the last few years that it actually looks, at least in the world that I spend my time in, which is AI infrastructure, you know, at the earliest stages of taking AI technology out of the lab, it looks much more like the early days of the venture capital industry than, you know, let's say the 2010s, which were, you know, in the earliest days in the 1970s, when you had firms like Kleiner Perkins and Sequoia getting started, these were founded by former semiconductor engineers, right?

38:19You know, Eugene Kleiner and Tom Perkins, obviously from the Fairchild line, right? Genentech, for example, as a company was founded literally in the basement of Kleiner Perkins by a former Kleiner partner and Herb Boyer, right? A scientist at UCSF. And the way these companies were formed is often you'd have a leading scientific mind who felt like their research was ready to leave the lab, who'd often team up with a former operator who understood that capital could be used in really strategic ways to accelerate the commercialization of that science. And when I got the Kleiner, you know, I was 19 at the time, I was really lucky that Brooke Byers, who was the B in KPCB, took me under his wing and kind of, you know, gave me a crash course into how venture capital used to be practiced.

39:03And that was quite different from the way it was being practiced by some of the newer partners there because we had gotten into this era where i think fueled largely by the rise of cloud and mobile um a lot of the the the technology that businesses that were being built did not really have a heavy scientific component right um and so it i i think yeah so you're right i do spend my days investing, but it often looks kind of like, it looks quite different. Even though these businesses that I end up helping these scientists start are often software businesses, you know, Anthropic, Mistral, Black Forest Labs, Sesame is one that we just took out a stealth that is building a conversational voice companion.

Read the full transcript

39:52They look, on the surface, they look like software businesses but under the hood the early days the practice of investing in these companies is often actually looks like uh like building a biotechnology business like building a you're often it's like discovering a new drug right these these foundation models are training these models look more like discovering and developing a new drug than they do building traditional software because frankly you know training a model off of data is is a is actually an a research process it's not really traditional software engineering where you you write up a list of features and you give that to a group of engineers and and then they build that and ship it right um yeah so so yes that's what i spend most of my days doing is uh often the earliest days working with these scientists of investing in them it often means i'm having to do things like procure you know 2000 h 100s and actually set up a cluster before there even is a company for them you know in the case of uh black forest labs which is a open source image and video model lab that i um uh got involved with last year as uh as a founding investor where it was a group of scientists they had trained and created a model called stable diffusion before they wanted to start a new company they were still trying to figure out what that would even look like and they gave me a call and said you know would you invest in us and would you help us build this a commercially viable business around open source image models.

41:19And so I sat down, we spent many weeks kind of planning what it would take to train a frontier image model. I think I even, we like set up the cluster and we purchased the cluster before the company was even formed so that on day one, they could hit the ground running. And so yes, it is investing, but I find it's not sort of the kind of investing that I was seeing lot of venture firms do 10 years ago it it feels like closer to the kind of investing that the 1970s and 80s era uh was common but let me let me ask i mean black force labs and sesame are those are interesting examples um i i've looked at sesame i've played around a little bit with its public interface and it's uh it's good but i've seen a lot of companies that do you know there's a company called speechify is it speechify or speechmatics or somebody speeches in the right british british outfit um you know with these incredible human sounding voices i don't i mean how do you build a business uh 11 labs does a great job as it is i mean how it's a crowded space i just that it amazes me frankly that that they can raise money because what's the differentiator uh in black forest labs uh maybe you can explain it to me but stable diffusion's open source right and um yep yeah i just i mean what what is the new what's new there that that gives you confidence to put money in it or is it that that it's it's not the underlying science it's it's the go-to-market strategy it's it's the how professional the the management how strong the team is and there's going to be a shakeout at some place so you're not going to bet on one horse you're going to bet across the field yeah no this look the the the reality is the answer and this is why i love you know doing what i get to do which is the answer is always different for every different company and so let's you know like we can take a couple of examples to to break this down so it's sesame right you know you brought a company up called uh 11 labs which is an extraordinarily successful company which i'm an angel investor and i had the chance to invest in you know, Maddie and Peter years ago.

43:53And while on the surface, it might seem today like, you know, all audio models are the same. If you actually start looking into the research that Sesame and Eleven Labs are doing, what you'll realize is actually within the field of audio, text-to-speech is completely different from the modality of conversational speech. And this is something I realized early on when I was running the platform team at Discord, where something like 60 % of daily active sessions were spent in voice channels. People just spent time on the platform talking in voice. And we actually tried to use 11 Labs to create a voice companion that you can talk to.

44:34And while 11 Labs is extraordinary for workflows like dubbing of your favorite Netflix show, which is a largely asynchronous text-to-speech pipeline where you give it a ton of text, or you give it a ton of audio and then ask it to convert that into some other audio in batch to be consumed later. Conversational speech, which is real time, can be approached in a completely different way. Right. So you could get to something good enough with Eleven Labs for that use case, but Eleven Labs was not designed for that. They are the world's best at the modality of text to speech. Sesame approaches their problem as a different modality, which is what they call conversational speech.

45:18When you need to train a model to be a two-way companion that can talk to you in real time, that can understand my speech tokens as I'm talking to it, convert that into audio and play that back to me in a way that's both realistic enough that I feel like I'm talking to a human and is fast enough to proxy human speech. the kinds of decisions you end up making to build a product off of that are completely different from what you'd build if you were building a text-to-speech API like 11 Labs for the Netflixes of the world. So I think this is the biggest, one of the biggest changes from four years ago where people were thinking about AI in largely modalities like, you know, language, image, video, and code, right?

46:03And actually within each of those modalities, you have a ton of different sub-modalities that are completely different when you get down to it. And each of these, by the way, I think each of these modalities can be ginormous spaces in and of themselves. So in the case of, so that's one answer for why these are all different. The second is that I don't think models have ever really been products. Models are phenomenal components. You can think of them almost like chips or transistors. They are critical components of bigger products, but ultimately customers don't buy components, they buy solutions.

46:41And so in the case of Mistral, people think of them as a lab that puts out models, but that's not actually what their customers use. their customers use a product called La Platform, which is an end-to-end service that allows a customer to show up and say, I want to use, my problem is that I'm a logistics and shipping company and I'd like an AI agent that can handle the entire workflow of logistics for us. And here's my data, which is because it's regulated, it sits on my own warehouse and I don't want to send that through to somebody else's cloud. And so I have La Platform coming and being deployed on my warehouse with a number of hooks that learn my enterprise's context, how we do business, and then customizes one of their models off the shelf for my company.

47:35That end-to-end solution is what they actually sell. Now, the research community loves them for all the open source models they put out, but those models are just a component of their overall solution. So where I think we're going, and I've always believed this, is the businesses that end up winning are the ones that have the world class scientific and research team to know how to leverage new models and how to improve them, but ultimately have the product vision to then turn it into some end solution for a customer. And that's where I get really excited. In the case of Anthropic, right, early on, their vision was an AI pair programmer.

48:11What they wanted to build, their belief was that, hey, there's about 30 million programmers in the world. If you assume each programmer creates about a hundred thousand dollars of economic value right if we can build a gi that can improve the productivity of those engineers by even 10 percent that's a three that's three trillion dollars of gdp we've created yeah right and their vision was a product it was a gi as expressed through the market of coding now they happened when chat gpt came out later you know about a year later they realized okay this is this is much larger than just coding but early on in the earliest days I get really excited about scientists who also have a very specific view for how their technology can get expressed to the world as a product or solution.

48:55Yeah. Yeah. And certainly they're not in the lead necessarily, but they're at the front of the pack on coding today on generated code with Sonnet. Right. Right. I think Sonnet was an extraordinary model that then allowed a bunch of other new products to be built. Right. Like there's a bunch of AI native code editors that are now possible that got a huge boost through models like Sonnet. So I think a lot of the debate and discourse, because we're technologists, we like to talk about models. But in reality, models are not products. and Sonic can be expressed in completely different ways for different products that can all, I think, be large, massive standalone businesses.

49:44But it kind of sometimes confuses the conversation where people think models are the end product. But they're for them. Let me just, we could talk forever. You're being very generous with your time. I would guess you're seeing a lot of people building DeepSeq-like models that are much cheaper to train than Anthropics or OpenAIs. I would also guess you're seeing a lot of Manus-type multi-agent systems, autonomous multi-agent systems. Rabbit AI just came out with this thing. Rabbit OS intern was one of the things they're calling it, or large action model playground. I don't know if you've looked at that, but it's like Manus.

50:42I mean, you give it a task and it has access to a lot of tools and can go out and do things autonomously on the web. are those two categories like those big multi-agent systems inspired by manis and the uh the more inexpensively trained models like deep seek are you seeing a lot of uh activity in those two spaces yeah so the i think the the answer or or the the the way the space is developing right now is that it seems to be that in any domain where you can write a pretty good verifier for a task whether a task was done well or not in an automated fashion there seems to be no end in sight for reinforcement learning to improve progress there in that domain yeah so as long as a domain can if the task can or or the a sequence of tasks that you would ask a human to perform can be codified in some way in a formal verification process like we are seeing every flavor of entrepreneur attack that that uh flow with with rl and it the the more data you have there that's high quality and verifiable there's there seems to be no diminishing sort of uh the returns to investing in that in that type of research now i think where things are much more brittle is in things like general browsing where it's very hard to verify whether somebody clicked on the right part of a page that you arbitrarily give it right and so i think i'm familiar with rabbit and manis you know i haven't spoken to those teams directly but i i'm very excited that they're pursuing that research because for a long time if the bitter lesson ends up being true then all you'll need to do is give a model a keyboard a mouse and a screen right and say just observe enough human beings browsing their day and and at some point you should be able to run reinforcement learning on on that but we're not there quite yet right now i think the regimes that are working are where there is a workflow where you can write a verifier for and and agents in that are in in that kind of area like customer service and you know coding obviously is the biggest one where if you can if you can uh for let's you know let's take a a place that we're seeing a ton of great progress and which is front-end development front-end web dev it's quite easy to write unit tests for right you can you can write automated verifications on whether uh a a website and a web app compiled successfully and kind of did some basic uh you know hit some basic um unit tests that that you've written for it agents are working phenomenally well there i'm i'm a huge i'm a huge believer that you know within a year you will be able to ask an ai assisted agent to or an ai agent to create a website for you end to end and we're almost already there to be honest um yeah but if you ask an agent to just sort of end-to-end plan a a a a trip for you and and go ahead and book all the you know the reservations and so on which is kind of a a user story that you see often from a number of these companies in practice i find it's still quite brittle and and we're just not there yet and it's it's still it's still faster for me to just go do it myself and i think that bar for general browsing is very very high yeah yeah uh and actually there's another sort of industry that's developing uh that fascinates me because one of the reasons that's difficult is the web is not set up to to allow bots right in but there's this these companies now they're developing uh uh you know agent identify identity verification systems right so that you know your bank or amazon or something when an agent comes to it it's not going to ask for two-factor identification or for it to complete a capture, there'll be some exchange of tokens to verify the identity, the agent, the permissions.

55:12And that'll really unlock a lot of value, I think. But I don't want to keep on going. I just have so much I could talk to you about. Unfortunately, I do need to help. So hopefully you feel like we got, you know, maybe we can wrap with however you end these episodes. but well what i want to know just how many deals do you do you invest in a year i i um i get quite involved with each company at the founding stages and so i find um i i at most i i helped start one company a year yeah and okay and that's roughly been you know the the speed at which i've been able to help founders out for the last few years.

55:54Anthropic was 2021.

56:00Mistral was... Mistral was...

56:06Okay, you can edit this part out, but I'd say roughly one a year is the pace at which I'm able to help founders get their, turn their research into new companies. And I'm usually involved before there even is a company. And the early days, there's just so much to do that can end up consuming the entire year. yeah okay great and it's really fascinating and i do hope we get a chance to talk again i'll be a little more uh focused in my questioning but but this is fascinating

From the publisher

AGNTCY - Unlock agents at scale with an open Internet of Agents. Visit https://agntcy.org/ and add your support.



What does it really take to turn cutting-edge AI research into a successful foundation model company?

 

In this episode of Eye on AI, we sit down with Anjney Midha, General Partner at a16z, to unpack how he helps scientists and researchers transform their breakthroughs into scalable, real-world AI businesses. 

 

From his early backing of Anthropic to launching Mistral and Black Forest Labs, Anjney shares a behind-the-scenes look at how AI infrastructure companies are born.

 

We dive into the critical challenges of model reliability, evaluation beyond academic benchmarks, and the rise of hybrid architectures combining transformers with diffusion and LSTMs.

 

If you're building in AI or investing in it, this is the roadmap to what's next.



Stay Updated:

Craig Smith on X:https://x.com/craigss

Eye on A.I. on X: https://x.com/EyeOn_AI

 

(00:00) Turning AI Research into Real Companies  

(02:08) Anjney’s Journey into Venture Capital  

(05:44) The Birth of Anthropic  

(08:26) Backing Mistral and Stable Diffusion  

(13:16) Are Transformers Really Enough?  

(18:36) Why AI Evaluation Is Broken  

(22:10) Making AI Models More Interpretable  

(28:38) The Real Potential of AI Agents  

(32:43) How a16z Spots AI Breakthroughs  

(37:45) Investing Like It’s the 1970s  

(43:31) What AI Voice Tech Needs Right Now

(46:32) Models vs Products  

(51:17) What’s Holding Back AI Agents  

(55:41) Anjney Startup Investing Strategy 

 

More from Eye On A.I.

All 266 episodes
#259 Anjney Midha: a16z’s Strategy to Turn AI Startups into UnicornsEye On A.I. · 57 min
Listen in VO