[AIE Summit Preview #2] The AI Horcrux — Swyx on Cognitive Revolution

8 Oct 2023 · 1 h 30 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Latent Space: The AI Engineer Podcast - Episode Summary

Episode Title

[AIE Summit Preview #2] The AI Horcrux — Swyx on Cognitive Revolution

Episode Description This episode serves as a preview for the AI Engineer Summit, featuring Swyx, who discusses the state of AI engineering. It covers topics like hiring AI engineers, key tools for AI development, skepticism about AI's longevity, and the future of AI agents. This episode is designed to prepare attendees for the summit's discussions and serves as an introduction to hot topics in AI engineering.

---

Key Themes and Topics Discussed

  1. Definition of AI Engineer
  2. Emergence of Role: The term "AI engineer" is gaining traction as a distinct role from traditional software engineering, focusing on AI applications.
  3. Background: Most AI engineers transition from software engineering, acquiring skills in AI tooling and infrastructure.
  1. Skillset for AI Engineers
  2. Core Skills:
  3. AI UX and coding tools
  4. Familiarity with LLM tooling, AI infrastructure, and inference hardware
  5. Fine-tuning models and working with generative AI frameworks
  1. Tools for AI Engineers
  2. Key Tools:
  3. HumanLoop: A platform for prompt management and fine-tuning.
  4. Guardrails: Output validation framework for generative AI.
  5. Langchain: Application framework for LLM deployment.
  1. Hiring AI Engineers
  2. Criteria for Hiring: Look for candidates who are entrepreneurial, adaptable, and able to navigate uncertain environments.
  3. Skepticism in AI: Some engineers are hesitant about the longevity of AI technologies, stemming from past experiences with tech fads.
  1. Future of AI Agents
  2. Expectations: AI agents are anticipated to evolve significantly, with discussions around their capabilities and practical uses.
  3. Doubts: There is skepticism about whether AI can autonomously perform complex tasks (like PRs) in software development due to the intricacies involved.
  1. Upcoming AI Engineer Summit
  2. Content and Speakers: The summit features speakers from various AI-focused organizations, discussing current trends and future developments in AI engineering.
  3. Goals: Aim to establish a community for AI engineers focusing on practical, technical discussions, rather than policy and regulation.
  1. Competitive Dynamics in AI
  2. OpenAI's Role: The potential for OpenAI to overshadow smaller companies by offering integrated services and competing with them directly.
  3. Market Dynamics: The tension between established companies and new entrants in the AI space, with the need for innovation to remain competitive.

---

Key Takeaways

  • The AI engineer is a rapidly emerging role, crucial for navigating the evolving landscape of AI technologies.
  • There is a growing need for specialized tools and frameworks to streamline AI development processes.
  • Hiring practices should focus on adaptability and entrepreneurship, rather than solely on traditional credentials.
  • AI agents are an exciting prospect but come with skepticism regarding their capabilities and practical application.
  • The AI Engineer Summit aims to unify and empower engineers while addressing current trends and challenges in AI.

Conclusion This episode serves as a comprehensive overview of the state of AI engineering, introducing key themes and discussions that will be central to the upcoming AI Engineer Summit. With insights from Swyx, listeners gain a deeper understanding of the evolving dynamics within the AI landscape.

---

For more details, visit [Latent Space](https://latent.space).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Ladies and gentlemen, it's the Latent Space Weekend Edition. Woohoo! This weekend is a special one, as we are gathering many of our former and upcoming guests and over 10 ,000 of you for our very first AI Engineer Summit, both in San Francisco and on YouTube. We were interviewed by a few of our fellow AI podcasters about the summit and figured we would cross-post them over the weekend to help you prepare, even if you can't join us in person. Now, we'll have a very current episode recorded with Nathan LeBenz of Cognitive Revolution, where we discussed how to hire AI engineers, key tools for AI engineers, skepticism around AI being a fad, the AI engineer conference speaker lineup, and then an hour of AI podcast inside baseball around the future of AI agents, multimodal chat GPT, and AI horcruxes.

0:54While you are listening, there are two things you can do to be part of the AI engineer experience. One, join the AI Engineer Summit Slack. Two, take the State of AI Engineering Survey and help us get to 1 ,000 respondents. Both are linked in the show notes, and we would really love to have you. Now here's Swix's conversation on the cognitive revolution.

1:17Swix, welcome to the cognitive revolution. Thanks. I've been a long-time listener and very excited to be a first-time caller. Well, thank you. Glad to have you here. And I'm also a big fan of your work with the Latent Space podcast, the newsletter, and also looking forward to what you guys are putting together with the AI Engineer Summit, which is coming up in just a couple days. So I'm excited to get into all of that with you. Yeah, happy to dive into that. We did a cross-post, I think, a few months ago, And I really liked your deep dive into the tiny stories stuff. And that's the one that we featured on our feed.

1:52And so I feel like you have the room to go much more in depth than us. So I really appreciate the work that you're doing, going like these two hour things with researchers. It's really impressive. Thank you very much. I really appreciate it. I guess for starters, I thought we'd kind of organize this by taking a little bit of like a broad view survey of your work over the last year or so. As far as I know, you've coined this term AI engineer. And so I guess I wanted to start off by just kind of asking you, like, what is an AI engineer? I'm fascinated in general by these like new AI jobs, right?

2:26We've got the prompt engineer and kind of a few different things have been put forward. It seems like the AI engineer, though, might have more staying power than like the prompt engineer. So how do you think about that new emerging role? Yeah, I definitely think of prompt engineering as like still 2022 and AI engineering as still 2023. And I feel like this is a controversial take a little bit because everyone should be able to use AI. There is no restriction on who does and does not use AI. But I do think that people who are choosing to specialize in the AI stack probably deserve a full-time role that describes what they do.

3:06And out of all the possible names that people have proposed, like cognitive engineer, LLM engineer, probably the one that is going to win is AI engineer. And so I'm not so much coining it as observing that this is a trend that's happening and putting all my chips on red, as they say. So the AI engineer is a software engineer specializing in AI and the emerging AI stack. An ML engineer is not an ML researcher, both of which are much more established roles and much more on the sort of research-oriented and ML ops side of the fence. It is everything to do with what happens after you have a model in production and maybe with a little bit of fine-tuning, which is, as everybody knows, just extra training on top of the vast amount of pre-training that has already been done.

3:50And I think it's basically an emerging category for a few reasons. It's more or less just demand and supply. And I started life as a finance and economics guy. I was a trader and hedge fund guy for quite a few years before I was a developer. And I just think it's just pure demand and supply. There's maybe like 5 ,000 good LLM engineers in the world, LLM researchers in the world, and you cannot hire them as the average company, average startup, whatever. There's no way you'll ever actually be able to build this talent in-house. For better or worse, these models are now available as APIs or as open source models.

4:28There will be a corresponding rise in demand for engineers who are capable of putting them to use, even though they don't necessarily train them on a foundation model research basis. I do think that there will be a rise in this category of the AI engineer. There's definitely a rise in startups serving that category. And so I've pivoted late in space, the newsletter, the podcast, and now the conference, all towards serving this persona, which I identified to be one of them. Yeah, I think I would count myself among the AI engineers as well. I probably, I guess I probably come to it from a somewhat non-standard background, but maybe you can tell me, like, where do you see the AI engineers coming from?

5:06Are they mostly software engineers who've kind of just taken an interest in AI and gone down the rabbit hole, so to speak? Or are there other backgrounds that you see as well? Yeah, they're going to be mostly software engineers taking an interest in getting deeper in AI. In the post that I wrote on Nainspace, the key visual to have in mind is this kind of left to right spectrum of research constraints work versus product and customer constraint work. or basically how close are you to machine learning research or how close are you to applications of that research? And so on the far left is the research scientists and the ML researchers.

5:44A little bit further along is the ML engineers, data scientists as well, I would class within that category. After you've sort of put those models into production, then you get the AI engineers, which are the sort of emergent category. And then finally the full stack generalist engineers who are working on just like the last mile of UI and UX. to deliver to AI products, to people. And so I think just like this emerging category of AI engineers, it's much more people on the right inside of the spectrum. The software engineers moving left into AI. So they're learning this, they're suddenly learning what tokenizers are and what embeddings are, what vector databases are, why prompting a chain of thought makes sense rather than a zero shot generation of stuff.

6:24And then it is much more likely to be people from the right moving left than people from the left moving right. meaning the ML engineers and data scientists moving right into applications, even though you do get those. So, for example, Raza Habib, my most recent guest from Human Loop, has a PhD in probabilistic programming with DPO networks. But he's working on an LLM office solution just because he sees that a lot more people can be served that way. So there are people with all sorts of backgrounds, but I do think it's primarily software engineers. We've had somewhat of a parallel trajectory in the podcasting game over the last few months.

6:58Roz was also an early guest because I've been a customer of Human Loop since early this year and definitely have got a lot of value from the platform. Now starting to use it also to, I think, really getting to the part of the vision that they probably had in mind early on, which is the really accessible fine tuning loop that they've enabled. and we just did an episode kind of going down that rabbit hole a little bit talking about like I think my number one insight there was using GPT-4 reasoning as part of the fine-tuning data set for 3.5 as a way to really in my experience dramatically improve the results and that was a something that human loop made much much easier than it otherwise would have been it's funny because I had already been a customer for months but I hadn't used all the latest stuff.

7:49And then I was like, kind of thinking to myself, like, how am I going to kind of code up my own loop to do this? And then thought, well, I should check human loop and see what all the latest features are. And sure enough, they had done a really nice job of anticipating the need and really streamlining that process. So what do you think are the key skills for the AI engineer? If I were to come to you and say, okay, I've got a background in software development, I've played around with ChatGPT and almost everybody's at least used Copilot or something like that at this point. Where do I go from there to make myself employable as an AI engineer?

8:28What do the employers, the app developers most need? Yeah, totally. I think that's where most people start, which is use the off-the-shelf models that have the most adoption and that is going to be Copilot and ChatGPT. I do think I have been mapping this out. So if you go to the about page on Latent Space, we also have an emerging email course that we call Latent Space University to trigger the Louisiana State University fans. And it's kind of mapping out the curriculum of what I expect to be the baseline competencies of an AI engineer. Like just write the job title. Like what do you expect people to do?

9:06And I think I want to ground this in a very practical manner because it's a demand and supply issue. Companies are trying to hire these people and people are trying to learn to be useful to companies so that they can be part of the AI movement, but then also apply their engineering knowledge towards building useful and interesting things. And so I'll just kind of list through some of them, but we can go into more details as needed. The high level themes are AI UX, AI coding tools, LLM tooling, AI infra and inference hardware. And that includes fine tuning as well. And then finally, the most speculative would be AI agents, which everyone has been talking about, but doesn't have that much practical use.

9:46So I think there's definitely a tradeoff between what people actually use at work versus what people just like to talk about on Twitter and star on GitHub and promise AGI without actually doing anything. So maybe that might be a hot take there. In my course, so we have a seven-day free email course type of thing. And that's meant to be like, I'm never going to make money from that. that's just meant to be like, start here. If you want to learn AI engineering, you're a software engineer. You've never dealt with any of these APIs. So day one would be just try the GPT API for the first time. And just a lot of software engineers just haven't bothered.

10:20And I think you'll be surprised at how much controls, how many options there are available, thinking about running into the context limit for the first time, which try GPT and Copilot kind of hide away for you. And just get familiar with all those basics. I think it's very important. Day two will be prompt tooling and memory. And that's where you get into like the vector databases and the land chains and all that. And those are just kind of frameworks that will hopefully help with the developments. And I do encourage people to kind of build some extractions of these themselves before using something off the shelf, you know, so you get a better understanding.

10:52Day three is code generation. Day four is image generation. Day five is speech to text. Day six, fine tuning and running open source models. And then day seven is building an agent. And I think once you'd finish that little sampler course or like a tour, You have the base capabilities that I expect every AI engineer should have, such that whenever any PM or any CEO comes to you with like, hey, I have an AI products that I would like to build. Is this possible? You can figure out if you can build it or you can just tell them like, hey, it's not possible for X, Y and Z reasons. You know, that way you'll be useful as an AI engineer.

11:23So that's really interesting because this is like, you know, the old XKCD. Somebody's wrong on the Internet, right? I almost have my own version of that with people posting GPT-4 can't do X things when in fact, like it definitely can. And they're either in many cases, like using the wrong model or prompting it wrong or doing like a bunch of things that are wrong. So I actually find that people often come to wrong conclusions about what AI can and can't do to it. Obviously, it's a fast moving frontier. So like your value of knowledge decays pretty quickly. I mean, seven days definitely seems like enough, especially if you really go for it, too.

12:04And this is like a remarkable thing, right? The part of what I love about the AI space in general is just that you can kind of jump right into the frontier. I think that's just extremely cool and super fun. And I would think most technologists would agree with that. So I do think like a seven day intensive is a pretty reasonable amount of time to like get mostly up to speed. I guess if there was one thing I would wonder about, it would be, would folks after that kind of intensive have a well-developed sense for what really is or isn't possible? And how would you coach people maybe on continuing to refine their sense of what is and isn't possible?

12:44Because I see way too many people just giving up too soon. Yeah. This is interesting. So first of all, I actually wouldn't describe the course as intensive. It's just an email course. It's meant to take an hour a day. and then we leave breadcrowns to explore with a lot more details and suggestions for side projects and stuff. I think that if you want to be up to speed, there's a vast gulf between, hey, cover the fundamentals and be cutting edge. Because the people who are cutting edge are the people who are hiding in discords and on Twitter and talking about very niche-y, jargony stuff that you won't even see for a few months out.

13:20So I don't really know if like that's a realistic goal for most people. I think most people want to make sure they cover the fundamentals and then be able to build most projects that you've seen out there that make money. And ultimately, I think that's what most people want. To be cutting edge means you have to go down a lot of rabbit holes. And I'm not exactly sure if I would recommend that for most people because that is a full-time job in and of itself, which is why, by the way, I call it an engineer because I do think that this deserves a separate category or a job title because this is a full-time job keeping up on things.

13:53But I do have a recommendation, which is I have a list of Twitter people to follow. I have a list of Discord communities that I watch. And obviously a list of podcasts and newsletters, which I recommend, which you're definitely on. So all of those are on my GitHub. I have an AI Notes GitHub where I link to all these lists of Twitter, YouTube, podcasts, and newsletter people. By the way, we're also running a survey of who people listen to. So a little bit of a market share competition going on. And if you want to Google state of AI engineering survey, you can get a sense of who people are listening to.

14:31And I think that's important to understand what drives people's attention towards individual projects. That is something that I should keep note of as well. Yeah, interesting. I'll have to definitely subscribe to your Twitter list, which I did not know that you had out there. But I do actually get most of my stuff, I think first from Twitter still, and then definitely have like a ton of discords that I've joined over time as well. So I have this little project where it basically scans discords on a daily basis and then summarizes them into an email. And I'm wondering if I should just release that as like a thing that people can subscribe for.

15:08Because I think it'd be kind of popular, but also very noisy because people discuss all sorts of things in discord. And like, it might not actually make sense. well i mean that's kind of the ai challenge in general right is like how reliable can we make these things i definitely think if you could get it to work well it would be valuable i'm i've got to be in i don't know 50 different discords at this point and it is super noisy in there and i honestly kind of i join them i'll often kind of scan around see what's happening you know see what discussion is there. And then often I don't really come back very much.

15:46So good God, the number of, you know, discord notifications have become unmanageable. So that just kind of gets tuned out. So I actually wonder if there would be a way to intercept the notifications or kind of use that as a particular signal, as opposed to like, you know, having to kind of grab and summarize everything. I do think if it worked well, and it could kind of create an auto GPT style digest of the news from all these different projects of interest, I definitely think that would be really interesting. I might release that as a side project. I don't know how popular it would be, but I think it definitely solves a pain point for myself.

16:24It's funny because I was one of our unreleased episodes that we just recorded. This is Jeremy Howard from Fast.ai. And in his recent post he had a post recently on sort of single one shot learning or like sudden uh drops in the the loss curves trying to figure out like the origins of that and he mentioned the alignment lab discord and i always just find these discords of very elite lm people just to emerge from nowhere like news research is another one i'll just throw these names out there uh for people who want to find it they're all on my list so you can go check it out but like how do you find them as they're emerging because most people found out about Eleuther AI after Eleuther was forming and the core community was already established.

17:07And now it's kind of, people have departed Eleuther into what I've been calling the Eleuther mafia recently. Eleuther is still going, obviously. And I just didn't find like, you want to join these communities when they're early and everyone's still trying to figure it out and building something interesting. And Jeremy said, most of this stuff is also happening in private channels, not public channels. So it's just an extra... wall garden inside of a wall garden. And just joining Discord isn't enough. You have to invest enough to get access into the private channels is what I'm saying. I need to get myself invited to some more private channels.

17:40For what it's worth, the bar for latent space Discord is you show up, you introduce yourself, you're invited in. Well, I appreciate your inclusive approach. A few more just kind of general state of the field questions. Then I want to kind of shift toward the event that you have coming up soon and talk a little bit about kind of where things are going as well. The tooling for AI engineers, if I understood you correctly earlier, it seems like the pattern is mostly identifying the best available tools, and then figuring out how to make those work together. As opposed to, you know, obviously, you kind of distinguish this from like, deep ML research.

18:21But even from like, self hosting, it seems like most of these things are kind of services that people are figuring out how to put in concert and not too often spinning up their own services at this point. Is that accurate? Sure. Even though they absolutely can, right? Like this is the power of an engineer that you can decide to build versus buy at every point in time. And that's always a fun topic of conversation whenever you need to build versus buy. When you think about that build versus buy, like that's such a common debate, obviously. And it's like a tricky one for a lot of folks because the tension that I experience is like, I want to be early to market, if not first to market, right?

19:02Like at my company, Waymark, it's like, I want to be the, we have this video maker. I want to be the first video maker that does like A, B, and C, X, Y, and Z for you. And to do that, you know, you kind of have to build more than you would have to build if you just waited a little while. because if you have a little patience, then sure enough, the market kind of provides solutions for a lot of these things that you need. So I think it's obviously super contextual, but I wonder how if you have like any sort of framework or kind of general high level view of how people should be thinking about build versus buy, especially given how quickly things are coming online for us to be able to buy.

19:43I generally have a pretty vendor-friendly version of this because I have been a developer relations person for five years. But I also am very sympathetic to the choose-boring-technology, try-to-limit-your-independencies-as-much-as-possible mindset. And so basically, typically what I often say is buy first. And then the moment you understand your domain well enough, be in a position to be able to build yourself. And this way, you get the benefits of people setting you on the right path with best practices that they've learned from everyone else. But you're ready at this point to understand that this field is so immature, you may have to rewrite, you may have to rip out something that you started not liking.

20:28So this was the topic of a lot of discussion this summer with Langchain, where a lot of people had this exact journey, right? They started out building your hands with Langchain, and they ran into a lot of problems that maybe Langchain wasn't well designed for. and then they ripped out LankChain and they say LankChain is crap. And I think that's kind of unfair to LankChain just because they're evolving as well. I mean, the company is less than a year old, you know? So I think it's fine. Like everyone, like if you chose to be in this field, you are choosing to be on the cutting edge and sometimes the cutting edge cuts you.

20:57And that's what they do. That's the nature of these things. I would basically buy off the shelf first and understand what people are building out there because you plug into existing communities of people who have encountered these problems for far longer and have dealt with them and thought with them much deeper than you have. And then anything that you don't understand or don't appreciate after working with them for a few weeks, then you can rip them out and build your own. But I think at this point, we have on our show notes, HumanLoop, Guardrails, and Langchain. These are fairly well-established frameworks among the community that you should at least understand as table stakes for what AI engineering is today, just because there's been a few hundred thousand people ahead of you mapping all these things out.

21:40What other tools would you put on that list? I mean, we talked a little bit about Human Loop as kind of, you know, just to set the level there, at least the way I think about it, first is kind of a playground in which to develop prompts. So my workflow is I go there and I workshop a prompt until I get it working reasonably well. Then the reason I do it there as opposed to anywhere else is because I can hit save. And then immediately I have an API that I can call that basically puts all the prompt management stuff on the human loop platform so that all I have to do as a developer from outside is just provide a couple variables, one or more, whatever I set up.

22:22And that makes it super easy to develop against. Then I can also update my prompt without having to do code changes, which is quite nice. And they log everything for me and allow me to kind of come in later and post-process it, evaluate it, export certain subsets of data, which I might use to power fine-tuning based on the most successful data points, whatever. That's really cool. We've had Shreya from Guardrails on the show a while back as well. You want to give your take on Guardrails and how that's used? Yeah, Guardrails is, I would say, in the output validation side of the fence. And, you know, as the name implies, it puts sort of safety rails on what generative AI does, specifically generative text.

23:09And I would say it is competing somewhat with the other orchestration frameworks because everyone, there's only one room for the sort of LLM interface layer and everyone's fighting for it. Human Lube wants to be that interface so they can track everything. And so do all the other LLM ops and prompt ops companies, is what they'd be calling. in my podcast within Raza called it the foundation model ops companies, but they all effectively have the same thing, right? Like we'll track and version manage your prompts and then we'll eval them. And then we'll like, you know, we'll track your costs too and track your latencies and whatever, right?

23:45They all have the same roadmap because it's fundamentally an ops product. And Guardrails and blockchain, slightly different. They're much more in the application framework side of the fence with an open source first and monetization second. And Guardrails would do things like validate that the SQL generated is correct if you're trying to use AI to generate SQL. They used to have sort of output validation. Langchain also has this, by the way, output validation for valid JSON if you're sort of outputting valid JSON. And that went away when obviously OpenAI came out with its own version of that. And so I think the roadmaps with these sort of orchestration or application LLM specific frameworks will evolve.

24:22But Guardrails is much more focused on safety and production readiness, let's say, And Lionchain is much more on the orchestration, even though they collide on some features, because all of them are both of each other's features. Yeah, that's super interesting. Let's come back to that in a second and get a little bit more into the competitive dynamics. But before we do, and this starts to give you an opportunity to talk about the upcoming AI Engineer Summit as well. But, you know, that and and more broadly, how would you suggest that people go about looking for AI engineers today? So, you know, whether you're whatever kind of company you might be, right, software application company, or, you know, I would say, honestly, any company that has kind of a lot of operational overhead could probably make great use of an AI engineer, even if they're not putting a, you know, a public facing or customer facing application out there necessarily.

25:18What do you look for if you're hiring that skill set? You know, it's so new. People are like lost. They, you know, people are kind of just looking for somebody to tell them what to do. But obviously that creates a lot of opportunities for them to hire somebody that doesn't really know, you know, what to do. So if you don't have the skill set yourself, like where do you go to look for it? How do you evaluate it? And of course, you know, the upcoming summit, you know, could be one of those places. I honestly, I really should set up a job board. I've been chatting with a couple of companies on like, hey, I need someone to just set up a job board.

Read the full transcript

25:49You know, I don't even have to make money on these things. I just want to make sure that people who are looking for others, whether it's, you know, you're hiring or you're wanting to be hired, you have a common place to match. Right now it's on Twitter. You know, anyone like actively talking and shipping projects is fair game. Within Discord, it's like in the Latent Space Discord, we have a hiring channel and people post jobs there. And I know people have gotten hired off of there as well. I think just regular channels and communities, posting in the monthly sort of who is hiring job posting boards on Hacker News.

26:21That's typically the way I would recommend these things. And obviously, coming to AI engineering conferences and talking with people who are also attending, that tends to be a very high signal filter for who's very engaged. And I would say, especially if you yourself are not an AI engineer and you want to hire someone who is an AI engineer, The kind of core competencies that we listed out earlier in this podcast, you know, that we have on the About page, that we have on the Latent Space University curriculum, is what I would expect a bare minimum for people to understand for working as an AI engineer.

26:54And I don't think it stops there, because I do think that a core requirement that you cannot really test for, you just have to sort of observe in an AI engineer, is that they're kind of entrepreneurial. real. They're comfortable with things that are not that well-defined because prompt engineering is not that well-defined because the model landscape shifts every single month with the release of a new model or whatever. They have to be able to be on the ball and proactive and not just, hey, I'm a sort of data scientist and now I do LLMs now and now I'm an AI engineer. And I don't think that that is the kind of person that will put your business at the cutting edge of what you can do with generative AI.

27:41They have to be a little bit entrepreneurial. They have to proactively come to you with ideas instead of saying, hey, we'll throw an LLM layer on top of the existing app, which is totally fine. But I do think that, especially for me and especially the people hiring that I talk to, they want someone who is a little bit entrepreneurial. So people who can just kind of like shift their own projects and get attention for them or like use new techniques and put them in practice. And I do have some in sort of the Rise of the Yacht Engineer blog post. There's some sort of role models that I do highlight.

28:15I think the entire Vercel team has been doing an awesome job of showing what it is like to be entrepreneurial. Even though they're not a foundation model team, they can put these things to good use. and in a way that is appealing to enough people that they get millions of users on their free projects. And I think that is something that every company should have. And that is very much the ethos of the kind of AI engineer that I want to encourage. Why are not all software engineers jumping at the chance to be AI engineers? It seems like there's something weird going on there. If somebody says to me today, oh, these AIs, they'll never be really useful.

28:54They're not that useful. There's no good use cases, whatever. The first thing I would say is code. Okay, like you can, you know, question anything, but you really cannot question the utility of a copilot. And you definitely cannot question the utility of coding with GPT-4. It is, you know, for me, I would say a multiple speed up safely. And I maybe am not like the very best software developer in the absence of GPT-4. but like you know i've got plenty of experience and definitely can make stuff work given how powerful it is for developers in their existing workflows and you know and the fact that it's like kind of been integrated into their tools before most other tool sets what am i missing why is it not the case that like everybody is kind of gravitating this way yeah i do think there's some baseline skepticism about whether or not this is a fad.

29:50And a lot of people have been burned by effectively crypto. It's a very fair thing to have some skepticism about anything new. It's a very fair thing to have some self-doubt over your credentials. Like, do you need a PhD to make progress in this field? And very much like why I'm trying to promote the AI engineer is to encourage the idea that no, you actually don't need credentials to make progress in this field because everyone is effectively uncredentialed when Transformers themselves are six years old and when GPT-3 itself is three years old. And many of the techniques and the companies that we talk about in this space are all less than one year old and all making tremendous progress in AI.

30:26And by the way, one of the biggest promoters of the AI engineer concept is Andrei Karpathy, who very kindly supported the concept with the idea that there should be more AI engineers than ML engineers, and they'll be very successful without trading anything. So I think there's there's some mix of skepticism about AI being a fad. And then there's some mix of self skepticism about whether they can contribute in AI. And then there's always like fundamental misalignments or some fundamental doubts about the stochastic parrots argument, right? Whether or not just multiplying a bunch of matrices can actually approach anything regarding simulated intelligence, which we can always talk about in sort of like those kind of dorm room hallway style conversations, like what is consciousness, what is intelligence.

31:11But I mean, you and I know, especially we're coding with AI, these things are actual productivity enhancers. And they've gone from single line autocomplete to function autocomplete to code-based, entire code-based generation. They can help you as tools of thought, with a human in the loop, or they can function without humans in the loop, all of which we're going to see as part of the conference. And that leads us, the more and more you have this sort of AI in the driving seat, that leads you more and more towards autonomous agents. There's a wide spectrum of AI as sort of autocomplete all the way to AI as autonomous agents.

31:44And I do think you just kind of depict your line of like where you think usefulness currently is and understand that that will probably move over time. We have about six orders of magnitude more improvements in terms of scaling according to Nat Friedman until the end of the decade. So take whatever we have today and project forward. Yeah, it's going to get wild. It's pretty hard to imagine what pops out the other end of a million fold more compute than went into GPT-4. So tell me about the speaker lineup. I went through, there's 28 accounted speakers shown on the website. I was proud to have had seven on the podcast as former guests, which was pretty cool.

32:25And it seems like you're basically kind of creating a lineup that covers kind of all the inputs, basically, that an AI engineer would need, right? From foundation models to these kind of frameworks to the quality control measures. And then there's like a number of folks from agent startups as well. Run that down. And particularly on agents, I'd love to hear your perspective. I wonder, is that something that you see as being like the next thing that's going to come on for the AI engineer to tap into? Or is that kind of a different sort of thing where it's like those more are the AI engineers that are building the coolest new stuff on those lower levels?

33:02I basically pulled together from my network all the tall speakers that I thought were building interesting things for AI engineers. specifically for more technical audience because there's other conferences happening all this fall in San Francisco, but they're all very high in general and they spend a lot of time talking about policy, safety, regulation, copyright and all those things. But for builders, I think there wasn't a builder-specific conference until mine came along. And then obviously OpenAI had to top it with Dev Day, which we can also talk about, which I think you and I are also going, which is, I'm very excited.

33:37We should do a live podcast on Dev Day, by the way. So yeah, we have people from AutoGBT is our presenting sponsor with OpenAI, with Microsoft speaking, Notion speaking, Amazon speaking, and GitHub. Like all these are top names, I think, in the AI engineering field. But I also wanted to balance it out with names that you've never heard of. Like Mithun Hunsar from the Rust LLM community is speaking. But we also have the first, I think, one of the world's first demos of Adept, one of the world's first demos, a new computer, and just trying to get a mix of projects that you've never seen before. So Lindy is a very, very hyped agent project that they've never done a public talk here.

34:18And so I think we're trying to be the stage where people launch these new features, new projects for the first time to provoke some thought, and then balance it out with people who talk more about active production issues. So I think one perception of AI is it's all greenfield. It's all 20-something-year-olds building toy projects. is not actually in serious work. And so I want to balance it out with Eugene from Amazon, who is running Amazon Books with language models, you know, and Abhi Arion, who just finished the O 'Reilly book on LLM ops in production. People like that who are actually implementing AI in like large scale production systems.

34:57Hex and Prefects are also speaking with us on how to pivot an existing non-AI company into an AI company. And I think both of them have done a fantastic job of that. So I'm kind of curating this list. I do have very large gaps that I'm very conscious of, right? So I don't have image generation people, and I don't have that much infrastructure based hands and replicates of the world. And I would like to feature them next year. I just had a limited schedule of a two-day single track conference. The underlying thing behind this conference is I'm trying to basically create the ICML for engineers. So ICML, the International Conference of Machine Learning, is like the ultimate, like, if you go to one conference a year, that's the conference you go to if you're a machine learning researcher.

35:39There's no equivalent for the engineer. And so what we're trying to do is like try to offer that. And so this thing is going to grow over time. I kind of see this as like a 10-year commitment towards building the ultimate survey of the field in any given point in time for engineers. I'm jealous that you have a demo from Flo and Lindy among a bunch of other cool stuff as well. I had him on as an early guest and continue to follow their progress from little hints they give out in the public. But I still have not been able to get into that thing even as an alpha user. So I'd be very, very curious to see what they're about to show off.

36:16I'm very excited, too. And honestly, I think you're doing a great job with your reports as well. And I'd love to have you as a speaker in the next iteration. But yeah, maybe I'll spend a bit of time talking about agents, right? Which is a very, very hot topic. So I would say, you know, maybe the surprise now is that AutoGPT is obviously now a company, being one of the fastest going open source projects ever. I think there's a ton of interest in what they're doing. And Torin will be speaking on our stage, I think, for his first time ever as a conference speaker. So I think that will be super exciting.

36:50I do think that there's a range of agent type projects, right? So the open source agents definitely tend to be sort of single use, I would say. Like their goal is really trying to optimize for what's the one most impressive thing that you can do in a single, you know, short video recording or Twitter screenshot. And I do think like that really is pushing the boundaries in terms of like what is possible. And then obviously being open source, people can actually go through the source code and copy ideas off of that. And I think a lot of people have done that with both AutoGPT and BabyHai. And I think the closed source agent type companies like the Adept, like the Lindy's, like the New Computers and the others that are the people working on here, they tend to work on very much more mundane, but daily use type of use cases because they're trying to work towards like, what would you subscribe for?

37:44Like a$20 a month,$100 a month subscription. I think both approaches are valid. And I do think that you need to have representation from both. the amount of human intervention is something that people are trying to get a grip of because the way that auto gpd does it is they basically ask you for a confirmation step before they do anything um and like that's fine but that's not autonomy that's just like you know assisted prompting whatever what you want is to fire and forget like that that is where ultimately these things have to go you know where i sit in san francisco and cruises and waymos are now common modes of transportation and literally i just get in a car i never talk to anyone and it just brings me to a destination.

38:23That's what I want. And we might be permanently for a while, at least five to 10 years away from full self-driving in agents, which is where self-driving was like 10 years ago. So I do think this is one of the most speculative areas, but I have seen some of these demos live because I'm friends with most of the speakers that I'm featuring. And I will say they are useful today. So they may look trivial now, but there's active research going on towards making them more and more substantial. I'd love to hear more about specifically like what use you have been able to get out of any of the agents. I've kind of played around with everything I've been able to get my hands on as well.

39:04And I recently called them broadly as a class, like still just for fun. Yes. Although I don't think it will stay that way. In fact, If anything, I would say maybe I would put my timeline to AI agents, you know, kind of crossing whatever chasm they need to cross. Probably short, if I understood you correctly, probably shorter than, you know, multiple years. I think I'd probably put my money more on like, I guess, kind of the anthropic timeline, if you will. You know, Dario has said in a couple interviews that, you know, they're going to. Two years to AGI. Come on. Yeah. Something along those lines.

39:43I mean, I can't confidently rule anything out at this point. I'm not confident that that will happen. It does seem like probably the biggest question in the whole space right now is like, does the current paradigm get to the ability to do sort of insightful work of the sort that currently humans are kind of the only things that are able to do? You could come short of that, though, and still have like very useful agents, right, that can decompose somewhat complicated tasks and reliably execute on them. I mean, there's a lot of different places on this spectrum of AI progress we could zoom in on.

40:22But I'd love to hear your take on both, why you think AGI is farther off than that, and then coming in a little bit toward AI agents or taking them whatever you want. But also on the AI agent side, I would be thinking like a year from now, they probably will be working well and like reasonably commonplace. But it sounds like you're maybe not expecting it to happen that fast. I think this will happen in gradations. And this is one maybe difference between our two podcasts. I tend to not discuss AGI. I tend to not discuss timelines, even though obviously there's some implicit assumption of them in everything that we do.

40:58Just because it's so hard to predict the future and it's not falsifiable in any way. So it's just a fun dinner topic conversation. So I will say there are some categories which I'm more interested in than others. So the original founding prompt of AutoGPT was I want to increase my net worth. That kind of category of agents, not interested. That's too general for me. That's too AGI, bro sciency. But there are very limited scoped agents, which I think are useful today. I have a piece on agents, which was relatively popular there at Roni Naples. And I said, actually, the most useful agents that I use today is one that has no LLMs in it at all.

41:37And that's the SavvyCal or Calendly agent. Right? Because I send you a link and then you schedule a meeting with me at your own convenience. And if you need to reschedule or cancel, it just happens on your side of the fence. And it just kind of happens autonomously without me. And that is kind of the experience I just want for all my agents. And the fact that we don't have that with LLM-enabled agents is a problem. And we're going to sort of slowly emerge to get there. The second form of agents, which I think have kind of already proven themselves relatively successful is code interpreter or advanced data analysis, as it is now known, because it can generate code, run that code, and then use the output of the code to decide whether or not it needs to fix that code or to stop.

42:22And that is the beginnings of a loop that is basically required for agents to have some level of autonomy in decision-making. And if you sort of break that down even further, you need some ability for planning and prioritization. You need a broader set of tools. You need memory to go with it. And then you need ways to interact with the outside world, whether it's sort of generating text or, you know, manipulating some kind of UI. I think those are sort of like the agent's research. For those who are interested, definitely be Lillian Wang from OpenAI. Her blog on agents, I think, is one of the most comprehensive survey overviews of the research in agents to date.

43:01So I think some categories of those will emerge. So one of the people that I forgot to mention, Deddy from Codium AI, that's an Israeli company that raised an$11 million seed for building coding agents. They only focus on test generation as an agent. To have an agent sort of independently running around generating tests in your code base and then for you to accept reject them. I think that's a relatively skilled problem that I, yeah, I'd be perfectly happy to agree with you. It will probably be useful and commonplace in a year. But things which require a lot more self-driving and a lot more, have a lot more degrees of freedom to fail, those things I think will be waiting for them for quite a while.

43:40My last observation is I think you'll be surprised what Lindy and Notion have to show. I can't say more than that. Yeah, I can't wait. I feel like you have more insider information than I do. And yet I want to take the bull side of a bet, perhaps. Like we need to maybe, you know, think it through to really refine the decision criteria. But I'm trying to zero in on what you think would fail that I think would work at a certain point in the future. Yeah. What do you think I would think would fail? And then maybe we can kind of discuss there. Because I'm positive. I just, I just, I'm much more interested in the engineering challenges than timelines.

44:18I do think people go well past the point of usefulness on their timeline discussions and their, you know, P dooms, if you will. Whenever anybody asks me about my P doom, I always say somewhere between five and 95%. And my quick, you know, add on to that is like, and I'm not sure it's that much worth trying to narrow it down further because 5 % is enough to be very concerned about it in my mind and like five percent you know if it is 95 then the five percent chance of you know surviving is like worth fighting for so kind of anything in there you know to me is like likely enough to be a problem and yeah five percent makes you a doomer basically well the real doomers would go much higher uh i don't i don't know that they would have me in their camp as a you know bona fide doomer with a five percent number because because you're multiplying by by by infinity, right?

45:11So 5 % rounds to one. Well, it certainly, you know, it does make it for me, like the issue of our time. I, again, I don't try to zero in on a specific number much more than that. And I'd say same thing for timelines. Like it seems like things could get really crazy in the next two to three years. It also seems still pretty plausible that we just kind of hit some sort of top out where it's like, hey, this paradigm kind of closes in on like expert performance on a lot of things, but never really achieves that sort of eureka insight capability. And we kind of stay there for a while until, you know, something else happens.

45:49I definitely think both of those are true or not true, but like plausibly true. Going back to the question of like, you know, where our expectations might differ. I mean, I guess the stuff that Lindy has shown would be like the kind of thing that I do expect to work, you know, in the not too distant future. And to, you know, describe those a little bit. It's kind of like text to automation is a lot of what the demos are like. Flow will post something where he'll have like a pretty simple prompt that's like, hey, every time somebody emails me, you know, from this domain, check this other thing, you know, draft me a response, put that in a calendar invite, you know, send it here, whatever.

46:30And he kind of just says all this stuff. And then the platform, you know, just judging from the screenshot that he posts, interprets that, sets up this kind of automation workflow. And in theory, then it's like set up, ready to go, almost like you have a Zapier Zap that you've just conjured, you know, with like two sentences. Some of them that he's shown have been reasonably complicated. And, you know, I assume that it doesn't always work or he'd probably have launched it by now. But I guess I think that that probably will start to work pretty well in the not that distant future. And I guess maybe to refine it a little bit more, it seems like once you can fine tune GPT-4, probably should be able to make a lot of those things happen, right?

47:20If you can just kind of document the reasoning, the breakdown, the planning. I mean, GPT-4 is currently doing everything. if you could really just zero it in on like the decomposition and scaffolding of kind of mundane, even if like somewhat complicated tasks, feels like that would be enough to me to get it to work pretty well. I don't know. What do you think about that? I think it could work well. We just don't really know. We don't really have benchmarks for planning and prioritization yet. And I've had actually debates with open AI people who are researching this actively. Part of my GPT 4.5 thesis is that long inference is kind of the next frontier of OpenAI.

48:01Like what can you do if you gave it a year to inference instead of a few nanoseconds or milliseconds? Anyway, so I would say that you're exactly right in your mental model of what Lindy does and how that interacts or overlaps with Zapier. And I do think that that is very useful. It's still not like that impressive, I guess maybe, but to everyone who's ever had an executive assistant or, you know, wish for an executive assistant, like this is going to approach what you would want as a, as a sort of always on personal assistant in your life. And I do think that that will unlock a level of productivity and quality of life, to be honest, that we just never seen before.

48:42And the fact that we can pay, you know, dollars for this, I think is, is useful for what it's worth. I think Lindy's working out, It's not just reliability, but also economics and security as well. And I do think that one of the pitches of why AI engineering is a thing is because we, the programmers, are going to be the people who put Shogov in a box, is what I say. Right. Like the like. Us fine tuning things down to smaller models, to domain specific models is the safe approach, being able to compose them into usable software systems to put humans first instead of creating sort of potentially ruinous AGI.

49:23I do think that that is a movement that I can stand strongly behind and is like nobody's against it effectively. Like we all we all want this to happen. It cannot happen quickly enough. And when it happens, there's not there's no safety concerns, except to the point where you let agents loose on the Internet with with with no permissioning. And so far, I think everyone, including Flo, who recently had a nice safety discussion turn. I think everyone's being very responsible about it. If you interviewed Connor Lee, Connor's one of the most safety-minded people. He split from me with her because they weren't safe enough.

50:01And he has this very strong criticism of the other GBT team. And I think he's just never met them. All of them are very concerned about safety. And I think it is the engineers who will put Shagov in the box. That's my bottom line. And understanding how to wield these tools for human benefit and not human ruin, I think the responsibility lies very heavily on the engineers. I think I agree with everything you're saying there. I often describe myself as an adoption accelerationist and a hyperscaling pauser. In other words, let's get the benefits that we can from our current systems. Or, you know, I think we're going to go a little bit farther than current.

50:44And that's probably OK. You know, but I'm very nervous about what happens if you put a million times more compute into a model than was put into GPT-4, because I just feel like we have zero ability to know. Reminds me a little bit, too, of Eric Drexler's comprehensive AI services vision. You know, just the idea that if you have narrow, perhaps superhuman, but still just fundamentally narrow systems doing all kinds of jobs in a way that can kind of become a buffer for the world against more generalist systems. And, you know, as such, like it seems really good to try to create these narrow systems that hopefully work really well.

51:27Is there something that going back to the like, what do we expect to work and not work? What do you expect to not work? Is there something that you would say, like, yeah, two, three years from now, I still think, yeah, we're definitely not going to have an AI that you can get to do X. Oh, God. I have a tough time answering this one. Yeah, yeah. Putting upper bounds on these things. One of the reasons I pivoted is very, very hard into this thing is I have to sort of update the algorithm. Whenever I've been wrong, I ask what else I could be wrong about and then I progressively update. It's kind of my sort of Adam optimizer version of momentum-based learning.

52:01but so effectively I think I'm very skeptical about code agents that do PRs so there's like sweep.dev and then there's a bunch of other agent type companies that will promise you know file an issue will turn it into a PR right and it always always works on like very very simple sort of copy change demos and it never works on more complicated things and I do think basically like the problem of software architecture and translating user requirements into code, I would be happy saying that's a five years out thing. It's not an immediate solve thing. I mean, I still think it's a very worthy problem because AI engineer is, instead of AI engineer, like AI engineer is job title, but like AI, like primarily AI engineer.

52:48I do think that's a very useful goal and worthwhile goal because engineers are expensive. And if you can pay an AI bot that's a hundred times cheaper than a regular human engineer, but is maybe 10 % of the capabilities, then you'll probably use it for quite a few situations. I just think like, yeah, the whole promise of like issue to PR is probably going to be overhyped for a long time just because the complexity of what software is. You know, I'm very happy to eat my words because there are quite a few friends pursuing this goal and more power to them, right? Like all human progress is based on unreasonable people.

53:25I think MetaGPT, I don't know if you've covered that paper recently. I think Microsoft also had a similar paper come out where effectively you have AIs play different roles in a sort of agile process in working towards a common outcome in software. Has it shown better human eval results than GPT-4? But again, you're fundamentally not describing the full complexity of what software systems want to optimize for. And I don't think we have nearly enough training data to do that. Like it's very much in the world of intelligent design rather than evolution. I think we can maybe refine this a little bit more offline, but I think I would take the other side of that bet.

54:06I mean, there's, you know, how big of a pull request and, you know, how big of a code base. And I think obviously for it to be interesting, you'd have to have enough tokens that it would be like beyond what a Claude 2 can currently handle. But even so, you know, it does feel to me like I would say in two to three years, it does feel like I would expect us to get there. And still on the basis of like, maybe there's no like eureka moments. I had an episode not long ago with the MedPalm M authors from Google, where it's like they're really closing in on expert level medical question answering, the ability to take a x-ray or another scan and provide like a radiology report at very close to a human level.

54:5340 % of the time their MedPalMem x-ray report was preferred to the actual radiologist. So it's still under 50, but it's like, damn, that's getting really close, you know? So if we can get that far already on problems that hard, I do feel like in a, you know, in a two-year timeframe, some pretty complicated kind of prompt to PR type things should also be possible. So we can And maybe firm that up and put ourselves on record for a little friendly wager. Look, there's a bit of undefined thing here, which is like, how much do you have to specify things in order for a decent PR to come out? And will that prompting effectively approach a new programming language?

55:36Right. Like, you know, English is the new hottest new programming language. But like to spell things out in some level of detail to which the LMs will get it. like are you are you basically just just programming but in a different higher level instruction that is kind of like a domain specific language for llms so yeah that that is open for debate and i do think like some of the approaches i think i've seen gpt engineer uh kind of approach that with their sort of pipeline yaml thing does that count i don't know right like because that's kind of pseudo-prompting that's pseudo-programming already in my mind i want a non-technical person to just manage a software project by themselves, just filing issues.

56:15I would say that's my bar for issue to PR is solved. Non-technical person. But if you need a technical person to come up with some interesting YAML file and prompts in specific ways, then it's nuts. Your hand is too much on the weighing scale at that point. Yeah, okay. I think that makes sense. I have one more topic to discuss, which is probably going to be one of the most talked about talks that is coming up, which is OpenAI just confirmed their talk with me. And we're going to have the first public Dolly 3 and GPT-4 Vision talk. I'm pretty excited about that because that is the signal. Finally, we're moving into multi-modality for everything.

56:55It's going to be a wild world. And I don't really know how to handle it because most of the people I have on, they're all in the text domain. And we're going to be very off balance when new modalities come out. Yeah, I got a little glimpse of that last week in trying to add a feature to a React app. And it was a fascinating experience. I have never prior to this, you know, session last week, programmed in React at all. I've done JavaScript before, but it's been a while since I've really been into it. And so, you know, my first question was, can you explain to me the general structure of a React app?

57:33And then it did. And then it kind of, you know, there were like a couple of file types in my repo that I asked it about. And it explained to me, you know, what those are and said, oh, well, that means you've got these libraries in place and that's what their purpose is. and then I asked it if it could give me a command that I could use to print out the tree structure of all my files so I could paste that back in so it could see the structure of my project so it helped me do that and then you know basically just kind of worked my way through to a working module and along the way because the new image I didn't really have to do this but along the way took a couple screenshots fed those in as well I was like man this is really amazing you know it's kind of understanding like what the defect is in the screenshot and, you know, modifying the code based on, in part on the screenshot.

58:22It definitely is still at the point where like, if you look through this transcript, you'd be like, Nathan, you qualify as a technical person here. This is not, you know, at the level that you're describing, but you can see, you kind of squint at it and see your way there with these screenshots where like, I want it to be doing something else and here's what it is doing, you know, and if it can start to deduce or infer, you know, the right fixes based on just like the visuals of the app not working, then it does seem like, you know, you're not too far away from a non-technical person, maybe not as efficiently, but still kind of being able to like get, you know, something along the lines of what they want, you know, which today is just like unthinkable.

59:02So for those interested, I think McKay Wrigley has been making a lot of us jealous with the GPZ for vision access with some examples of what you can do from like, you know, sort of the vision to code modality. It's not just vision, right? It's also voice. And part of my update that I did for September, and this is something that Ben Thompson of the Strict Techery podcast, I think he's now talking on the Sharp Tech podcast. So he had a private demo of both modalities and I've seen both in person as well. And I kind of agree with this point, which is vision is more impressive individually, but voice is the thing you're going to use every day, all the time.

59:39It's kind of an interesting inversion of expectations because I would count vision as my most important sense out of the six senses or five senses. But voice, the ability to hear that's hands-free and then just always on and understanding what you want. I think that's kind of a thing that I've seen a lot of people here in San Francisco building with. And so my perception, I used to dismiss voice. I was like, oh, like auto GPT has like an 11 labs. That's just for fun. Like who wants an agent talking back when I can just kind of read the text by myself way faster. But the fact that it's hands-free, the fact that it's just always on and the fact that you can imitate sort of human expressions, I think it's going to be remarkable.

1:00:19And I don't know if you've seen, have you seen the Russian language teaching demo of GPT-4 voice mode or chat GPT voice mode? No, I haven't. So Greg Brockman tweeted this maybe like two, three days ago. Duolingo should not exist anymore. This thing is going to teach you languages on a personalized basis much better than any other app in existence. Any human teacher just is way more patient and way more forgiving. You probably can understand multiple languages way better than you can. So I think it's going to be a very, very interesting time for understanding all these new ones out because I'm sure they're working on more that they haven't told us.

1:00:57Yeah. So your comment that Duolingo shouldn't exist anymore. That is kind of what I wanted to take your temperature on next, because I feel like in some sense, I've seen a version of this movie before where going back to early Facebook days, for example, I just happened to be in the same dorm as Zuckerberg in college. And I actually thought it seemed pretty dumb at first. that was like, I'm going to go online and put a picture of myself. So I was a late adopter, actually, of Facebook on campus. Fun fact, Matt Walsh from 6E, he used to be Zuckerberg's CS50 professor. I don't know if he's ever told you that.

1:01:34That did not come up, actually. That's fascinating. So in the social network, like the class that Jesse Eisenberg walks in on and it solves the question and he walks out, that was Matt Walsh. Okay. Small world. Yeah, that's really funny. You know, Facebook initially had this like, come one, come all build on our platform, you know, you get access to everybody's friends, everything's better with friends, it's all going to be social. And then it kind of became clear over time that like, actually, we're going to build the most high value use cases natively on the platform. And you know, the app developers can kind of like get a little bit of stuff around the edges.

1:02:16But today, basically, you can like log in with Facebook. And that's really it. right? There's not like a lot of utility in the Facebook platform. I am wondering if we're starting to hit a moment where OpenAI is going to go a similar direction and start to kind of eat these adjacent use cases that are currently other companies. And so like companies we've talked about, right? Like HumanLoop. I love the product. We're both fans. But it sure seems like from an OpenAI standpoint, you know, they just kind of released the like fine tuning jobs. Yeah. It was like, well, what are you doing there? That's starting to look a little like human loop.

1:02:55Right. And then, so there, I mean, there's a ton of stuff that they could do there and it makes total sense. And it's like, not, they wouldn't obviously be doing it because they're hostile to anybody else in the ecosystem, but just because they're asking their users, you know, what do we, what do you need? And they're just executing on that. And these are things that people need. And then I'm kind of like the same thing might start to happen with, for example, text to speech, you know, we've got some really good companies that have created amazing stuff. I've cloned my voice with both PlayHT and 11 Labs.

1:03:24They both sound awesome. Do you have a favorite? I don't know if you're willing to go on a record. I think it does depend on use case. I would say 11 Labs is probably the favorite for ease of use right now. But if I was doing something more artistic or kind of creative, I wanted more emotion. Yeah, PlaySinks have a lot of diversity. Yeah, the creative direction that you can give it is a little bit more open. So I think, you know, if you're analyzing this from kind of a who's OpenAI posing the biggest problem for, in that analysis, it would be 11 labs probably because OpenAI presumably is going to be more reluctant to like create the voice that's going to, you know, feed into your horror movie or your video, your shooter video game or whatever, where you're going to want kind of a wide range of emotions, at least in the short term.

1:04:13But it seems like quite clear that they're headed for a, I would imagine, tell me if you think this is off, but it seems like, you know, when I think about this dev day coming up in a month, I'm like, well, geez, they just launched voice in the app. They just launched vision in the app. You know, they just launched this preview of fine tuning. They've got vector database built in natively to ChatGPT Enterprise. It seems like all these things are coming to the API. Sorry, I haven't heard about the Vector database. What details are there? I don't have a lot of details, but they have said publicly that, you know, you're going to be able to connect ChatGPT to your information, right?

1:04:54As an enterprise customer. Basically, what I understand that that translates into is you can connect your Google Drive or your Notion or your Dropbox or whatever to their system. I understand they're going to support a number of them out of the box, kind of with their own connectors. And then they presumably will have some way for people to connect their own random data sources into the system as well. And as a ChatGPT Enterprise user, then you can query all your company's Google Drive assets as part of the ChatGPT Enterprise experience. I don't know exactly where that is in terms of whether it's launched or not, but they've sort of said it's coming.

1:05:40And that much, I think, you know, is not like super secret. Beyond that, I don't have a lot of details. But I guess I just feel like all this stuff is coming to the API, right? And it seems like they're going to do this to serve the AI engineer, because, you know, the AI engineer doesn't really want to have to go piece five different services together. Like, you know, today, if I want to create a voice assistant, I've got to like transcribe, send something into the language model, get text back, convert that to audio. And, you know, what would be a lot better is if I could just send audio into OpenAI and get audio streaming back at me.

1:06:18Right. And similarly, like today, I've done this and you've done it, too. Right. We've got to go select a vector database provider and figure out how we want to chunk our stuff and how it gets loaded in there. And then what's our query strategy? And, you know, that is its own art, you know, to figure out, like, how do we generate the query for the vector database to get the right stuff? Because often, like, the question the user asks and the actual contents that you're trying to, you know, find similarity with, those maybe not don't line up so well. So there's, like, whatever, all these kind of different tricks and techniques.

1:06:49And again, I what I really just kind of want is, like, one place to be like, here's my docs. Here's my question. I'm putting it in as an image and an audio file. And by the way, I want my response back as streaming audio and just have kind of OpenAI handle everything, right? Like, does that seem right? It seems like everything's, they're going to kind of just eat all these sort of things around them and the AI engineer benefits. But a lot of these companies, I wonder, for a lot of the speakers that you have, as much as they have led, how do they end up not having been the community R &D department for OpenAI?

1:07:32Look, every large enough platform eventually comes into this tension with their developers. And OpenAI is just coming into that right now. Apple has this long history of Sherlocking, which is an official term, I think, where some very successful community app gets killed effectively by an official Apple version. And I do think OpenAI has promised to not compete with its own developers. They have said that, funny enough, in the deleted transcript of the interview that Sam did with Raza, which is still publicly available, They promised to not compete with developers, but I don't think that promise is ironclad.

1:08:15I don't think they want to stand behind that. That reminds me of the dumb and dumber scene, not to call anybody dumb, but these are just as good as money. These are IOUs. And it's like, yeah, that was easy to say at that round table, but because they're at a thousand people now and they've got product managers. 600. 600? Oh, okay. Yeah. I mean, I happen to have very live stats. That's so I know. I would say that they always want to focus their people on the bigger things, right? Where, you know, there was this recent discussion about the open AI phone. This is what I've been calling it with Jodi Ives.

1:08:55Like working on a phone, that's big and ambitious. You know, working on, you know, A-B testing, fine tuning infrastructure for the enterprise, not that ambitious. Leave that to the lesser engineers in the community to pursue. and I think go for the big things and obviously try to solve problems that everybody has. So I would say, I think, yeah, OpenAI does exist in the tension with the developers that build on top of it. Probably some of the companies will be Sherlock's. I don't think 11 Labs would be a threat. The best thing that startups can do building around the sort of foundation model lab ecosystem, this is not just OpenAI, it's also Anthropik and all the others.

1:09:31It's just proof that you're a really good engineering team because if you're good enough, they'll buy you and pull you into their orbit. Money is free to these people. It doesn't actually matter. Are you choosing to work on interesting problems and do you ship fast enough that you're making an impact? And I think 11 Labs has definitely shown that they can do that. And hopefully enough of the teams that I'm seeing can do that, but not everyone will make it. There's definitely going to be collateral damage. OpenAI has such a huge footprint that it cannot help but realign people. Like when a lot of what Langchain and Guardrails used to do was JSON output validation, right?

1:10:09And like now that it's completely useless because of API and they just have to roll with it. And I think everyone's completely fine. We all understand we're building on top of shaky foundations and it's kind of moving, but that's what you get for being first. You know, in terms of the multimodality stuff that you talked about, I do think like that will be a focus for DevDay. I don't know if they'll be, they'll have the sort of APIs ready by then, but we'll be able to see what you can build with it. in the talk that Simon and Logan will be giving. But this comes back to Rune's post on text being the ultimate universal interface.

1:10:41I do think like text is sort of the king modality. Yes, like whatever species synthesis is being done, whatever sort of text to image or image to text stuff that they're doing, like everything just kind of just passes through text and working on text as like the central pipeline, I do think it has some links to it for the time being. Very happy to eat my words. Ultimately, everything is just tokens within a token space. And text has no particular domain in there, except that we have much more of an understanding of what high quality data looks like in text than in everything else. Yeah, it does seem like if I had to guess, I think things are going to get a lot easier for the AI engineer over the next couple of months with just some really integrated kind of high quality things, you know, such that kind of end of this year, you could probably feed in multiple modalities.

1:11:33You could kind of have them manage your custom data set powering retrieval, but, you know, powered by them. And then they can kind of stream out to you text. Yes, of course, but audio on top of that, perhaps even, you know, images integrated into the responses. And that is going to be quite a difference and quite a simplification. We've got GPT-4 fine tuning as well, which is like another huge one, because to the degree that you have a use case that you can't get to work well yet with any current model, you know, this could be really the thing that allows it to become possible. Things that OpenAI might launch that wouldn't compete with anybody in the ecosystem, but which might actually just be a win for everyone.

1:12:23One concept that I had there is log in with OpenAI. And I think this is like potentially really good for a lot of developers. If you are struggling with this tension between I want to show off how sweet my product is. But to do that, I'm like taking a non-trivial hit in aggregate in terms of token cost. At Waymark, for example, my company, we estimate it's about 15 cents per user that tries the free product in hard cost to us. Some of that is open AI language models. Some of it is image understanding. We have some self-hosted stuff on Amazon. We try to process all their images that we possibly can to create the best profile we can of the user, et cetera, et cetera.

1:13:14I'd rather that the user could kind of pay that free cost instead of us having to pay it. but obviously they would need some mechanism for that. Do you like that idea? And how would you riff on that idea? Do you see that as a plausible thing that they might introduce? Absolutely. I think it makes sense. I just don't know if it meets their bar for interesting work. I think logging with Google is effectively this equivalent thing for logging with OpenAI. But also like implementing SSO provider, I think it's relatively commoditized to begin with and then gets very complex over time as you're starting to manage non-trivial amounts of user data, which maybe OpenAI doesn't want to get into the business of, right?

1:13:53It's more just about choice or alignment with the mission. And I don't know if that helps them build AGI. I think it's a nice quality of life thing for them. Maybe they can get more information as people sort of, you know, use sort of universal login with OpenAI. I do think that if they do build this, then they have effectively decided to become the AI cloud. I think they maybe own AI.com. Don't quote me on that. It might be them or Elon. This would make them the fourth major public cloud, which is very exciting, but also maybe not quite a research organization. And again, I still don't know if they want to do that.

1:14:33I do think it's kind of a choice of paths and they're fortunate enough that they have that luxury and capability. There's many things that they could point their engineers at. And I don't know if building SSO for AI is one of them. could you just kind of offload it to octa or whatever um and someone else could do that as well so i will say one thing which um i think i i like as an idea which is sort of context that kind of lives with you right um and so i think you maybe mentioned this a little bit that does seem to help but probably actually a really character he added does that which which has by far the the longest history with working with this maybe not so much an open eimaxe in this regard.

1:15:12I tend to describe myself as quite an OpenAI Maxie with regards to LN progress. But with regards to like products, I feel like Character has just kind of earned its stripes in trying to be a chatbot much more than OpenAI has done, even though they haven't had, you know, sort of state-of-the-art new models to ship. I don't necessarily want to be an OpenAI Maxie, but it's hard not to be. I write a monthly recap, right? And then my OpenAI News section takes two two screens and then my other frontier model, you know, the, you know, in the frontier model, model form, I have, I have other frontier model updates and it's like five bullet points.

1:15:49Like it's, it's not a competition. Like opening eyes just running away with it so hard. Yeah. I mean, it's just being, being objective makes you an opening eye maxi. Competition right now is just not that deep. Right. I mean, you know, we said just before we started recording, we got to give love to Claude too. I would say Anthropic has definitely shown that they can create a, really high quality models that are not like qualitatively outclassed by open AIs, even though they do fall a little bit short on like the MMLU benchmark and whatnot. But it's like definitely a very good experience to use and something that I do, in fact, use, especially when I need long context or whatever.

1:16:30But they kind of have a strategy of like not leading the market in product releases, as far as I know. And so there really aren't that many other competitors out there right now that could could bring a similarly robust and just useful model forward. I mean, Google is not there. They are maybe going to get there in the not too distant future. But as of now, the models are just not as good. And who else is there, right? So I haven't spent enough time on Palm 2 to actually definitively make that statement. I want to believe in Google. I've met people there. They're very smart. They have all the resources in the world.

1:17:10They just don't have the institutional capacity to shift something interesting for people, unfortunately, so far. I think the research is not the limiting factor there, if I had to guess right now. It's certainly not the compute either. And the fact that they didn't participate in the last Anthropic fundraise and allowed Amazon to kind of set that. Amazon just bid way higher than anyone else was willing to bid. Anthropic is probably worth like$10 billion now or something. And Google, I don't think. they probably want to invest in-house. I mean, I've talked with people involved in the previous rounds and they said it was much more hands-off.

1:17:46So Google in no way owned that relationship as strongly as Microsoft and OpenAI were. My biggest update from that was just that Gemini must be on path to work pretty well because if it weren't, then they would be bidding as much or more than Amazon, right? It seemed like they must have had some inside track, based on the earlier relationship, even if it was fairly, you know, just kind of friendly arm's length, whatever. If they didn't have confidence in their upcoming releases, then hard to imagine how they'd let that one go. Yeah, who knows? Battle of the Titans at stake. Hopefully the tell-all books that someone will write in 10 years.

1:18:26Maybe Walter Isaacs and Phil Rounds. People are wearing their rewind pendants so they're capturing everything. One of the themes that I featured for my monthly recap yesterday was the AI horcrux is what I was saying, or like AI pendants. which I saw a demo from Avi Schiffman of Tab yesterday. I've also talked with a number of people also working on similar projects. I would be willing to wear one of those. So I have an aura on me right now, and it's just kind of logging my heartbeat and whatever, but I'd love to log all my conversations and be able to talk with it. These things are coming. And yes, they will be multimodal, but also I think all of us having an effective digital twin that we can talk to and use as at least just a note-taking thing, if not just exposed to the wider world, I'm very excited by, and it will be built by the AI engineer.

1:19:15So yeah, I will feature probably an entire stage next quarter on, I mean, the next time I do my summit on these things. Because I think in terms of the B2C frontier, if we have a form factor that is beyond the phone, it's going to be the humane clip or the AI pendants or the watch or whatever it is that's just always on you. Yeah, that definitely seems like it's coming before too long as well. So last thing I wanted to ask you about, and I've been kind of asking everybody about this. Listeners to the show will have heard me pitch it before, so I'll try to do it super briefly. Concept is the AI bundle.

1:19:52And I'll give you just kind of brief inspiration from the consumer side and from the app developer side. Consumer side, you know, pretty obvious, right? I've got ChatGPT Pro, I've got Claude Pro, I've got CoPilot, I've got RepliGhostwriter, I've got Perplexity Pro, I've got, you know, other stuff that rewind and, you know, all these things. Right. And my AI, my monthly AI bill is like getting up to the point where it's at or north of what my cable bill used to be. It notably doesn't cover hundreds of other apps that I like might like to use, but, you know, they're all like another 20 bucks a month or whatever.

1:20:28And I'm like, man, do I, how many of these like$20 a month things do I want to buy? Right. So that's the consumer side. on the application developer side, just, you know, representing myself as Waymark, we have this, you know, easy to use video creator. It used to be DIY. Now it's done for you with AI. You tell us what kind of video you want, tell us who you are. We make the video. You just get to sit back and watch it. It's awesome. But we do have this kind of weird tension where, as I said, it costs us 15 cents for every user. We don't want to put it behind a paywall because we want to show it off.

1:21:00You know, people are most likely to buy if they can try it. But then the flip side of that is like, you know, we have the kind of classic, whatever,$30 a month plan. And we need one person to buy for every 200 free trials just to break even on our tokens, right? Before we cover any other cost or make any money whatsoever. So that's a tough situation because it puts us in this weird thing where we're like, when do we gate it? How do we gate it? And then we also see, of course, the behavior that SaaS app developers hate to see, which is where people do the thing they want to do and then immediately cancel.

1:21:34And there's been increasingly a lot of talk about this, like churn is pretty high. Everybody's getting a lot of new customers, interest is really high, but churn is also pretty high. On both ends of this, it seems to me like there could be real value in some sort of bundle. And what I want, I think as a consumer and potentially also as an application developer is something where it's like, hey, pay a hundred bucks a month and get a thousand AI apps. And those would be like your flagship apps, like maybe your JetGPT Pro, and then maybe a bunch of long tail apps, like a Waymark, right? What I figure is if I'm Waymark, which I am, and I'm one of a thousand apps and I'm like, okay, look, I know that there's a lot lot of things going on.

1:22:22Most people don't need us that often. We've got some power users, but we also got a lot of people that just, hey, just one off here and there, right? And they buy and they immediately cancel. And it's not because they didn't like the service. It's just like, I don't want to pay for this every month. So if it was a hundred dollar bundle and we were one of a thousand apps, let's just say we got the average, which is 10 cents a month per app subscriber, kind of in the same way that like ESPN classic, you know, gets a very small fraction of my cable bundle. I think we would do that. And I think we could still kind of upsell the power users into higher tiers or whatever.

1:22:56But for every million, if that was the economics, every million subscribers that the bundle had would represent a million dollars in annual revenue to an app like Waymark, even just getting 10 cents per month per user. Can you see that happening? On the supply and the demand side, it seems like there would be a lot of utility there. Obviously, it takes some doing to pull that trick off. But do you see that being something that could make sense? Or if not, where would it break down? The closest equivalent would be Setapp. Are you familiar with Setapp? No, I don't think so. So there are existing software bundle type things out there.

1:23:39And they do bundle subscriptions. And I do think that they provide some value for a portion of the population. But I don't know how big this is people are pursuing the sort of bundle economics play in traditional SaaS. And maybe you can have an AI bundle SaaS play. I would say some of the questions would be, do you get paid regardless of whether or not people use you? And I think that would very much affect whether or not the extra million comes to you or not. I don't know if you have a quick answer to that one. Yeah. I mean, I think it could be managed in any number of ways. Obviously, for whoever is managing the bundle, you'd want to have some rules or framework in place to keep it somewhat sane.

1:24:20You know, just using the cable bundle as kind of the jumping off point. There's certain channels I never watch, you know, and there's, and other people don't watch the channels that I watch. And my understanding is they all get a per subscriber share, regardless of, you know, whether or not I tune in. But then when it comes up for renegotiation, then there's like different levels, of course, depending on how popular you are. I imagine something like that perhaps managed even purely algorithmically could work. I think you'd have to have some kind of baseline. If you're in, you're going to get something from this to make it appealing to the app developers.

1:24:55I definitely would see it kind of where the actual money flows, I would see shifting around over time with shifting usage patterns. Yeah, exactly. So I would say there's difficulties there. There's difficulties with regards to the differences between what a bundle of a newsletter or a bundle of content would look like. I would say that the people who have probably thought about this the most is Nathan Bashaw, who has now started LexPage as an independent startup, and Ben Thompson of Stratechery, who's written a lot about the media and cable bundle. So they would be much more clear thinkers on this than I am.

1:25:33And I do like the sort of magic of bundle math, right? They say you get more choice, you pay one fee, and then the people also have, the suppliers also have more predictable income and the share and access of income and sort of pooling that distribution in terms of getting people to pay and then getting people to understand what they are and try them out. I will say like for SaaS, the incremental difference between SaaS and content businesses is that SaaS people typically really, really want to own their customer base. When I want to change pricing, do I have to get your approval as Mr. Bundle Guy?

1:26:09And I'm not going to be happy if you push back on something. And so I think there's a lot of these internal politics that happen when it comes to B2B. And then pile that on top of, I'm very immersed in the B2B or enterprise infrastructure environments in San Francisco. Like I, like that's where my background is coming in terms of developer tools. And all of them are not going after the 10,$20 a month subscription, right? They want to go after the 10K, 100K, a million dollar a year contracts that they want to get from other businesses. So effectively, this is around the year. Yeah. I would not expect like your human loop kind of infrastructure players probably to get into something like this.

1:26:56Yeah, it would have to be kind of the way marks. Prosumer B2C, yeah. Yeah, all these, you know, there's a million of these. Like I recently had one of the founders of Gamma on the show. And, you know, just it's like, I don't make that many slides, but when I do, it's cool, right? But I feel bad if I'm just using their free version and, you know, not paying them anything. But I'm also like, I don't necessarily have, you know, hundreds of dollars to drop on every app that I'm interested in trying on an annual basis. So, you know, I either kind of sign up and cancel or I just don't sign up. It's probably not going to happen, but I kind of something about this feels like it should happen to me.

1:27:31So I kind of can't let go of it just yet. Well, the real question is, what is the minimal viable version of this that you can test to find out? Right. Like that's something that one of my mentors has always given me as just life and career advice, which is whenever you feel like you have this idea or whenever you're stuck with a fork in the road, what is the minimal viable step to gain more information for you to be sure? Because you could sit there and have this... It could work. It could not work. I don't know. The only way is to take steps. And so I don't know how to size this down because there is a certain appeal to scale of these bundles.

1:28:09But maybe get five things together and see if that works. Yeah, I think one of the big challenges with this is it does feel to me like you would need anchor products in the bundle to make it compelling. If somebody said, hey, for$10 even, you can have 300 apps you've never heard of, then I'd be like, eh, I'm not sure I need that. But for$100, if I could have ChatGPT Pro and Cloud Pro and Perplexity Pro and a bunch of other things that I've never heard of, then I'm like, yeah, I might go for that because those three, I already know I want. And then whatever else I'll kind of discover along the way.

1:28:50Anything else you want to talk about today before we break? That's a real pleasure. I have just been enjoying the podcast. I would say just keep up what you're doing. There's always room for these sort of long-form deep dives into AI. And I love the passion that you're taking to approach the space. I'd love to have you on at my next conference. Well, thank you very much. Obviously, the feeling is mutual. Why don't you give us the details of the summit and what people can do if they still want to get a ticket and get there. Yeah. I mean, tickets are completely sold out, but you can get an online ticket where it's basically the live stream plus the Slack community discussion at ai.engineer.

1:29:31I really love these short domains, by the way. So just like, yes, engineers at TLD, we bought ai.engineer. So we have the engineer summit, newsletter, got board. The podcast is latent.space. Cool. Well, this has been a lot of fun. Swix, thank you for being part of the cognitive revolution. Thanks for having me.

From the publisher

This is a special double weekend crosspost of AI podcasts, helping attendees prepare for the AI Engineer Summit next week. After our first friendly feedswap with the Cognitive Revolution pod, swyx was invited for a full episode to go over the state of AI Engineering and to preview the AI Engineer Summit Schedule, where we share many former CogRev guests as speakers.

For those seeking to understand how two top AI podcasts think about major top of mind AI Engineering topics, this should be the perfect place to get up to speed, which will be a preview of many of the conversations taking place during the topic tables sessions on the night of Monday October 9 at the AI Engineer Summit.

While you are listening, there are two things you can do to be part of the AI Engineer experience. One, join the AI Engineer Summit Slack. Two, take the State of AI Engineering survey and help us get to 1000 respondents!

Links

* AI Engineer Summit (Join livestream and Slack community)

* State of AI Engineering Survey (please help us fill this out to represent you!)

* Cognitive Revolution full episode with Nathan

* swyx’s ai-notes (featuring Communities in README.md)

* We referenced The Eleuther AI Mafia

* This podcast intro voice was AI Anna again, from our Wondercraft pod!

Timestamps

* (00:00:49) AI Nathan’s intro

* (00:03:14) What is an AI engineer?

* (00:05:56) What backgrounds do AI engineers typically have?

* (00:17:13) Swyx’s Discord AI project

* (00:20:41) Key tools for AI engineers

* (00:23:42) HumanLoop, Guardrails, Langchain

* (00:27:01) Criteria for identifying capable AI engineers when hiring

* (00:30:59) Skepticism around AI being a fad and doubts about contributing to AI

* (00:34:03) AI Engineer Conference speaker lineup

* (00:41:14) AI agents and two years to AGI

* (00:46:04) Expectations and disagreement around what AI agent capabilities will work soon

* (00:50:12) Swyx’s OpenAI thesis

* (00:53:03) AI safety considerations and the role of AI engineers

* (00:56:24) Disagreement on whether AI will soon be able to generate code pull requests

* (01:01:07) AI helping non-technical people to code

* (01:01:49) Multi-modal Chat-GPT and the future implications

* (01:03:33) Nathan living in the same dorm as Mark Zuckerberg

* (01:04:44) Competitive dynamics between OpenAI and other AI model developers

* (01:05:39) Play.ht vs ElevenLabs

* (01:09:20) The tension between platforms and developers building on top of them

* (01:11:40) The best thing startups can do to compete with foundation model providers

* (01:16:26) User identity/authentication services like Login with OpenAI

* (01:19:20) Google vs the other live players

* (01:20:46) AI Horcruxes / Pendants

* (01:22:05) The concept of an AI app bundle for consumers and developers



Get full access to Latent.Space at www.latent.space/subscribe

More from Latent Space: The AI Engineer Podcast

All 247 episodes
[AIE Summit Preview #2] The AI Horcrux — Swyx on Cognitive RevolutionLatent Space: The AI Engineer Podcast · 1 h 30 min
Listen in VO