[High Agency] AI Engineer World's Fair Preview

25 Jun 2024 · 50 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

```markdown

Podcast Episode Summary

High Agency - AI Engineer World's Fair Preview

Overview

In this episode of the **Latent Space

The AI Engineer Podcast**, Raza Habib interviews Shawn Wang (Swyx) to discuss the upcoming AI Engineer World's Fair and reflect on the evolving role of the AI engineer. The conversation touches on themes from Swyx's viral essay "Rise of the AI Engineer," team dynamics, and the future of AI product development.

Key Themes and Concepts

The Role of AI Engineer

  • Definition and Emergence: Swyx argues that the AI engineer role is distinct from traditional ML engineer roles, focusing more on the early phases of product ideation and market validation (zero to one phase).
  • Skill Set:
  • AI Engineers: Need knowledge of APIs, quick iteration, and product-driven thinking. They may not require as deep a mathematical foundation as ML engineers initially, but advanced knowledge becomes important for scaling.
  • ML Engineers: Typically manage end-to-end solutions and require strong mathematical skills to optimize algorithms for specific problems.

Team Composition

  • Optimal team structure for AI product development suggested to be a ratio of 4 AI Engineers to 1 ML Engineer. This reflects the growing trend of needing more product-oriented roles as AI technologies become more accessible.

Trends in AI Startups

  • Vertical vs. Horizontal Startups: Vertical AI startups (focused on specific industries like law or construction) tend to perform better than horizontal ones (general-purpose tools).
  • Demand and Supply: The high demand for AI engineers is contrasted with the lower status associated with the role due to a lower barrier to entry compared to ML engineers.

Discussion Points

AI Engineer World's Fair

  • Event Significance: The fair is framed as a crucial gathering for the AI engineering community, aiming to foster connections and discussions about the future of the field.
  • Multi-Track Format: This year features various tracks focusing on topics such as:
  • Multimodality
  • Evaluation and operations
  • AI in large corporations

Criticisms of New Roles

  • Some experts argue that all engineers should be viewed as AI engineers due to the pervasive nature of AI technologies. Swyx counters this by emphasizing the need for specialization in a rapidly evolving field.

Advice for AI Product Leaders

  • Move Fast: Echoing the "fire, ready, aim" philosophy, Swyx encourages rapid prototyping and shipping of AI products to gather user feedback and iteratively improve.
  • Focus on Vertical Solutions: Building AI tools that serve specific industries or problems can yield better results than attempting to create broad, general solutions.

Trends to Watch

  • Foundation Models: The rise of foundation models is reshaping the landscape, making AI capabilities more accessible.
  • Commodification of Intelligence: Cost of deploying AI systems continues to decrease, which will enable more startups to emerge and innovate.

Conclusion The episode provides valuable insights into the evolving landscape of AI engineering, the significance of specialized roles, and the dynamics of AI product development. As the AI Engineer World's Fair approaches, the discourse around these themes will likely shape the future of AI engineering.

Additional Resources

  • [High Agency Podcast](https://share.transistor.fm/s/890b0ab6)
  • [Latent Space](https://latent.space)
  • [AI Engineer Essay](https://www.latent.space/p/ai-engineer)
  • Event tickets and information can be found at [ai.engineer](https://ai.engineer/).

Listen to the Full Episode For more in-depth insights and discussions, listen to the full episode on your favorite podcast platform. ```

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:29Welcome back, listeners. Raza Habib of Humanloop has now launched his own podcast, High Agency, and Swix recorded a conversation with him that both serves as an examination of the one-year anniversary of the rise of the AI engineer and a preview of the AI engineer World's Fair. We do mention a discount code in the pod, so for any last-minute stragglers, we've opened up a few more tickets for loyal listeners to join us at the fair. Watch out and take care. I think there's some advantage to be early, but you do not get the right to win just by being early. AI engineer is also like sort of a plug or filler gap for all the things that are not traditionally part of the ML engineer skill set.

1:10So what do I mean by that? AI engineer is more of the zero to one phase, whereas the ML engineer is more of the one to end phase. We should adopt like a fire ready aim approach instead of a ready aim fire approach. and where the old ML engineering process was much more deliberative, here you actually win by moving really quickly and getting information from the market from shipping products so that you can iterate much faster. I think we have three foundation models launching, by the way. That's sort of, yeah, I never said that publicly, but. All right, you heard it here first. This is High Agency, the podcast for AI builders.

1:42I'm Reza Habib. Today's a bit of a special episode of the High Agency podcast because I'm joined by Sean Wang, aka SWIX, of Small AI, the host of the Latent Space podcast, and critically the author of the essay, Rise of the AI Engineer, as well as the founder of the AI Engineer World Fair. And we're recording today two weeks before that AI Engineer World Fair, as well as roughly a year on from when Swix first released the essay. And so I wanted to get Swix on the show to talk about the AI Engineer essay, what's happened since then, as well as the upcoming conference and trends. So Sean, it's great to have you on.

2:16Yeah, thanks so much for having me. Excited to come on and talk a little bit about the meta behind the podcast and the essay. Yeah, fantastic. To start with, do you mind? I actually wanted to just read a little extract from the start of the essay to give people who maybe haven't read it the context, right? And it was polemic almost, pitching the claim that there's going to be a new type of role emerging, a different type of engineering role as a result of LLMs and Gen AI. And I think you put it really beautifully. So I'm going to just read a little extract from the start, and then I'd love to get your reflections on it.

2:47now that we're a year on. So what you said was, we're observing a once-in-a-generation shift right of applied AI, fueled by the emergent capabilities and the open-source API availability of foundation models. A wide range of AI tasks that used to take five years and a research team to accomplish in 2013 now just require API docs and a spare afternoon in 2023. However, the devil is in the details. There are no end of challenges in successfully evaluating, applying, and productizing AI. I take this seriously and literally. I think it's a full-time job. I think software engineering will spawn a new sub-discipline, specializing in applications of AI and wielding the emerging stack effectively.

3:27Just as site reliability engineer, DevOps engineer, and data engineer, and analytics engineer emerged, the emerging version of this role seems to be AI engineer. So I think that was a good summary of the thesis and beautifully written. So Sean, my first question to you is, you know, the essay got a lot of attention at the time. inevitably on Hacker News, there's a mixed reception of some people who completely buy it. Some people are critical. Reflecting now, 11 months later, what do you think you got right? What do you think has changed? And I'd be curious to get your take. So first of all, it's very flattering to have it read out by you, because I definitely feel like you were earlier than me, actually, to this persona.

4:04And I think that is a lot of the success of Human Loop as well. Yeah. So I think mostly the general thesis, the diagram, right? Like the one thing I probably did better than in the essay was just that diagram. I knew that I needed to have like a very, very simple diagram that would kind of stick in people's minds. And this is the picture you have in the essay of kind of a spectrum where on one side you have kind of hardcore ML engineers and on the other side, I think you've got sort of front-end developers and there's a new kind of emerging gap. We can link to it from the show notes, but yeah, maybe describe it maybe for people who are listening.

4:37Yeah, I typically describe it as a spectrum from sort of data constraint to product constraint. And then there's a bunch of roles from most sort of ML technical towards least ML oriented, like zero ML oriented. And there's all sorts of different roles beyond that. And I think that the primary feature that I think people should understand is there's a sort of API line between the ML side and the product side, right? The API line used to be internal within the company and now increasingly it's between companies, especially as the cost to create a model increases and model labs become increasingly closed.

5:11Actually, it used to be the case that most companies would hire their own ML engineers and ML research scientists and all that. And now it's increasingly outsourced. And you can sort of decouple that because of the rise of foundation models. And concurrently, there is a rise in specialist engineers on the other side of the API line who are up to date and know the stack because the stack is deepening every single day. It is a full-time job to keep up with the news. And the people who are putting effort into it are going to do better than the people who are generalists and dabbling a little bit in AI.

5:42And I think that once I had that fundamental insight, I think it was very hard to shake the idea that there would be full-timers who want to distinguish themselves. And that can also come a little bit into the origin story of the essay, which was really, I was just trying to create a term that everyone would sort of congregate on, on the buy side and sell side of talent. Because companies were talking to me all the time about, I want to hire for this profile, but the average engineer hasn't spent nearly the amount of time that we will want for an AI specialist role. And then on the software side, on the engineering side, people wanted to know which direction to specialize in and honestly have a little bit of legitimacy because they would never have the background of a research scientist.

6:19They don't have a PhD. They don't have even have the qualifications of a typical data scientist. And I realized that all of those things were irrelevant in the new age of foundation models anyway. So here's an opportunity for a generational shift because of a platform shift. And I've seen this before. So that was the last part of the essay, which was that similar things that I've observed in my career for SRE and DevOps and data is happening for AI. So that's the process I went through, really. One of the lines in that intro that really stuck out for me was you said, a wide range of AI tasks that used to take five years in a research team to accomplish now just require an API doc and an afternoon in 2023.

6:57And I think there's also this very famous XKCD cartoon from maybe like 2010, 2012, that kind of timeframe where someone's asking an engineer, like, oh, can you sort of build an app that uses GPS satellites to tell me my exact location in a national park? And he's like, great, yes. Then the follow-up question is like, and can it also allow me to take pictures of birds and know what they are? And the response is, give me 10 years in a research team, I'll come back to you. And the intuition that people often have for what's hard and what's easy in machine learning has historically been off. Given that there is now possible to do things in hours or days that previously required a research team and years of work, et cetera.

7:36It's clearly not the case that the skills you need to be good at this are like a PhD and maths research and all of those kinds of things. So what are the skills? Like what is the skillset that an AI engineer does need that's different to an ML engineer? I don't have a pat answer to this yet. And that's honestly one thing that I refrain from doing when writing the essay because every successful specification is necessarily underspecified. And I wanted to let people define it themselves and have their own contribution to what it is. I think it's a spectrum. The more skills of the ML engineer you have, the more successful you're going to be as an AI engineer.

8:15It's just not strictly required to get started. But obviously, the more advanced you become as an AI engineer, you're going to want to learn more and more of the ML engineer skills. What does that involve? So for example, to get some kind of AI products up and running really quickly, you just need to prompt, you just need to know how to call a few APIs. But eventually, for example, if you want to put something into production and if you want to create a moat around your products, you probably want to invest into your ops stack, your fine tuning stack, and all the inference and sort of eval capabilities that you may need in order to support that.

8:51And that's increasingly more and more ML engineer. So I definitely think the more deep you go into the moat that you might want to build for an AI product, as well as control over open source you want, I think that those are the primary criteria. If you're primarily wrapping GPT-4, which is totally fine and actually can make decent money doing that, you don't need as much of the ML engineer skills. But the more of the model work that you take on, the more you're going to need to bring that ML engineering in-house. Then I think finally, AI engineer is also like sort of a plug or filler gap for all the things that are not traditionally part of the ML engineer skill set.

9:29So what do I mean by that? Like ML engineers typically did not used to work on agents stuff. And of course, they have been involved in agent research for a while. Of course, they have something to contribute there, but they don't have a unique advantage there. Actually, AI engineers with a completely different base of assumptions and experience and focus on agents might be able to do a better job than the typical ML engineer. I typically call this also a qualitative or anthropological difference between the ML engineer and the AI engineer that I observe. I've been observing the MLOps community for a while and the other sort of ML engineering type data science-y communities.

10:02I've been attending Databricks Data AI Summits and observed the differences between their summits and mine, my conference. And it's just a different type of person. Basically, boil it down to if you're an ML engineer, you work on one-to-end problems, right? You have a baseline, a large amount of data to work with, and you are trying to optimize or build a very specific model for that problem because you have the scale to do so. You have the data and you have the kind of problem where reducing fraud rate from 8 % to 5 % is a big deal. And that makes a lot of money and there should be someone focused on that.

10:35But that's not what an AI engineer would do. AI engineer is more of the zero to one phase, whereas the ML engineer is more of the one to end phase. And it takes probably a different skill set, probably a different mindset to be successful there. So what would be an example of that zero to one phase? What would be some things you have in mind there? The classic example is, can the ML engineer make a cursor or make a co-pilot? Can the ML engineer make a photo AI or all the other generative image companies that have emerged? But is another way to think about this that AI engineers and the kinds of people who are building products with LLMs and foundation models today tend to be closer to thinking about product?

11:11They're closer to full stack. These people have front end skills. Whereas ML engineers, by virtue of the fact that they need to be much more mathematically sophisticated, are much more specialized. For them, the world almost starts and ends with the model. Whereas like for an AI engineer, the world only starts with the model. Right. Actually, the model is kind of given to them. And what they're trying to do is figure out, how do I make this a useful product? Yeah, totally. I draw that in the diagram as the API line, but it's equally just the model, right? Whether or not you're creating the model or you have the model serve to you and you pick up the baton and carry on from there.

11:45So yeah, it's a different mentality. Of course, they can cross-pollinate. Of course, you can do the skills of the others and be successful with the others. But there's a home base where you're most comfortable and finding that these require different personas. And the kind of person that's successful in one doesn't necessarily translate to the other. I think I buy the thesis overall. Like I'm convinced that there is a new type of persona and skillset emerging that's kind of got a little bit of the ML engineer in it, but also kind of drawing from full stack and is more product focused. You just see this in the tools and the languages people are using.

12:16I think JavaScript and front-end tools are way more popular amongst people building compared to like the ML stack or MLOps stack traditionally, which is like much more Python heavy. And I think that shows in the community. The AI engineer is like only one part of, I think, of like a product AI team. I have my own opinions on this, but I'd be curious, like if you were staffing a team to be building an AI product, what composition of skill sets would you be looking for? Oh, boy. That's a fascinating question. I think the AI engineer is definitely in there at the MVP stage. And then it's an open question as to how you develop it as you grow.

12:52I do consider there to, you know, this was the original essay and then six months later at the first summit. I do talk about like the other, the three kinds of AI engineer, which we can talk about. But for team composition, I do think that we've had internal discussions in the Discord. I think it's some kind of ratio of like, let's say four to one in terms of AI engineer to ML engineer for more mature teams. Because AI engineers do the stuff that ML engineers would usually chafe at because it's not really working on the model. But, you know, some kind of nice ratio where you have a decent amount of products, people filling all the long tail gaps and working on product stuff, whereas ML engineers, you know, would get the leeway and capacity to work on these modes of the model.

13:33I think that makes sense. You know, if you want to expand the scope a little bit in terms of your team proposition to designers and product managers and all that, that's your prerogative. but in terms of just your engineering allocation, the general thesis is that the number of AI engineers, because the bar is lower, the number of AI engineers is going to be higher and it's going to be a multiple of ML engineer. And that's an opportunity, obviously, for vendors serving that. But also, I just think it's a factor of just the pipeline that we're not training enough ML engineers, the number of graduates in these programs that would typically qualify people for like the data scientist role.

14:06It's just not going to be enough for the sheer amount of demands that companies have for this role. So therefore, you need to scale the ML engineer by supplementing them with a whole bunch of engineers who actually may be self-sufficient in themselves. It depends how deep this role gets. We don't really know the boundaries yet because it's so new. Yeah, I guess we're less than two years on from ChatGPT and 11 months from the essay. It's certainly early days. What about the non-engineers? We touched on it briefly, but how do you see product managers, domain experts and other people fitting into this?

14:37Yeah, I think actually the product manager and domain expert is going to be very, very key and probably more key than the AI engineer, depending on how good code assistants and code agents get. And one thing you cannot replace is insight into the customer and into the product. And no amount of AI engineering can solve a bad product decision. It's not really good engineering. It's more about just picking the directions for a good product. And so I think that that is probably like the last automatable task on that team. And it's very important. Like either the AI engineer owns that. I always think like that's the ideal sort of one person AI engineer is to embody that sort of product thinking and a little bit of the ML knowledge and enough engineering to be dangerous.

15:23But, you know, on the larger team, yes, absolutely. The PM and domain expert is going to be key because they will provide the inputs for the engineers just like they always have. That part absolutely does not change. They need a counterpart who can tell them what's possible and translate their requirements to current capabilities today. They always needed that. That doesn't change. But I think the AI engineer is going to be much more equipped to tell them what state of the art is with foundation models than the ML engineer, just because they are fundamentally wired to do that because they are product thinkers first and foremost.

15:53Yeah, that really resonates. I mean, I have this kind of thesis with HumanLoop based on what we've been seeing as well, which is that we used to live in a world in which product managers were the experts in customer problems. They wrote the spec, they figured out what was needed. And then it was like entirely the role of engineers to translate that to code. And what I feel like we're seeing more and more of, and I think this trend will continue as the models get better, is product managers and subject matter experts can be much more directly involved in writing prompts and actually creating artifacts that are part of the product.

16:25And so instead of just playing this like definitional transitional role, they're collaborating very directly with the AI engineers actually to bring the product to life, which I don't think was true for like ML before this. Yeah. And I think a collaborative tool that bridges the gaps between the different roles in the team makes sense. I think that's why Human Loop is going to do particularly well in this area. Yeah. Or at least it's why we're making the bet that we are. Yeah, it's a genuine role. The other side of it is that you are not going to be the only person doing this, right? There's going to be a bunch of you because this is an insight that's big enough for multiple teams pursuing this.

17:05And then it's all about execution from there. But that's not really my business. It's yours. It makes sense, though, right? This is the shape of how people are going to collaborate. And this is how domain experts should be involved in the work. They always should have done that. But now it's possible. Now it's easier. And that's great for everyone concerning, including customers. So we've spoken a little bit about the AI engineer role, how it's changed maybe or its emergence over the last year. Before we talk about the upcoming conference and kind of what you're excited about there and what others might be excited about, just briefly, when you introduce a new term like this, inevitably there's going to be some detractors, some critics, some people who try to say, oh, we don't really need a new role here.

17:44What have you found to be the most compelling criticisms? And what do you think people have got wrong? Yeah, the most obvious one actually comes from a friend of mine, Jared, he's VP of AI, VP of products AI, I think, at Purcell. And immediately on the call, you know, I hosted like a Twitter space to discuss the essay and, you know, he was like, I don't think there will be an AI engineer. I think every software engineer is an AI engineer. And I'm like, okay, that is the typical take that is the diametric opposition. Instead of role specialization to feature another product, everyone will have this, you're not special, this is unnecessary labeling.

18:19And I do think that there's some part of every engineer that if you're responsible and up to date with the world, you should be up to speed on AI. The reality is that I don't think it's worked out that way. There are a lot of AI skeptics, something on the order of like 50 % of Hacker News still hasn't really adopted like co-pilot you know like the future is always here but unevenly distributed and there's just going to be always people that take things more seriously than others and that's their prerogative like i do not think everyone should should work on ai i think people should work on distributed systems i think people should work on front end should work on databases there's so many other valuable problems in the world that yes that you should go work on those things and yes you deserve a special job title for that ai engineer is going to be low status for a long time definitely it's going to be low status for compared to the ml engineer and research scientists Just because the barrier of entry is so low, nobody's going to really respect it.

19:08People are always going to question its very existence. And I think that's fine. Push and pulling the boundaries of where the definition is and how much it's going to be, I think it's fine. The primary response I have to that is just wait and see. And I think that for me, creating the shelling point where people can find each other in the job marketplace, that's really the main goal. The secondary goal is that people define skill tree or skill ladder. We're not there yet. We don't have senior staff AI engineer. We do have VPs of AI engineering, though. One of those speakers at the conference is like a VP of AI engineering at MasterCard.

19:43And I was very interested to see that title. It's particularly interesting to see that title at MasterCard, right? At MasterCard. I don't think it was a startup. Like that's a very clear incumbent company. Cool to see they've adopted that title. OpenAI has adopted that title. They're hiring for AI engineers. I have quibbles with their job description, but obviously they have their own take on it and that's fine. What's the OpenAI take on an AI engineer? They want something like five to seven years of ML engineering experience on the job description. And I would never require that. For a role that doesn't involve pre-training, that role that they have listed, it's on the website.

20:16I talked to Shamal, who put that role up. And that role is primarily for fine tuning for their custom models team. And do you need the full MLE experience? Do you need five to seven years? Maybe, maybe not. But it's a requirement that they put up there. And so it's a point of contention, right? Like how much do you need? I would say that they put themselves more on the ML engineer spectrum of an AI engineer. And that's fine. I've always said it's a spectrum, right? So it's a fuzzy nature that I think people get uncomfortable by. They don't like the hype that comes with attaching the word to AI.

20:44You know, often the criticism is also that ML is when it works and AI is when it's sort of magic, right? I forget the exact terminology you're saying for this. But it's true that, like, I am leaning into hype. So a bunch of people propose alternative titles like LLM engineer, cognitive engineer, and so on and so forth. But you just have to go to the lowest common denominator and the natural thing that people want to say that rolls off tongue is AI engineer. They don't say AI developer. They want to say engineer because it is engineering. And I think, you know, having explored all the past, this is the one that, you know, I predicted that people would go to.

21:15And it seems like it's catching to the point now where people complain about it on Hypernews because it's a reality and they find that annoying to them. But yeah, that would be it. Like the legitimacy of do we need a title like this? or should we just reuse existing concepts? And my only response is look at demand and supply, look at the amount of work that actually goes into this thing. Yes, it's going to feel illegitimate now. It's going to feel increasingly less illegitimate every single year. And the whole point of my conference, my podcast, my newsletter, everything I do is to serve this role and to spec it out.

21:47You mentioned that it's maybe a little bit low status today because the barrier to entry is lower than for an ML engineer. Is there any advantage to being early? Like if it's going to be a legitimate role in the future, do you get on the ground floor now and get some advantage from that? I think there's some advantage to be early, but you do not get the right to win just by being early. So I'll elaborate on that. Advantage to being early because there is a lot of papers to keep up on and techniques and history of prompting in the APIs that are out there to keep up on that people just assume is assumed knowledge for anyone in this role.

22:23And so the later you join, the more you're going to have to learn. It's not impossible to learn. People have done it before. It's just going to be harder. It's better to learn the baseline and live with everyone else as you proceed through the timeline at the same pace as the rest of us. But just because you're early doesn't mean you win, right? There's a lot of flame outs in AI. One of the title sponsors of last year's conference was AutoGPT. They're not necessarily as active anymore. They flamed really, really hard. Like April last year, I have this sort of one-year trajectory essay that I'm writing about issue of agents.

22:55You know, they got more stars on GitHub than PyTorch, Bitcoin, Kubernetes, Django combined. And, you know, like selling the promise of AI. And they were early, but were they able to convert that into something lasting? Not really. And I think that's a challenge for a lot of the AI products, right? Like they're going to get a lot of interest because we're in this mode where we're optimistic. stick, we really have the budget or even the time or attention to try things out and to be anticipatory. But then if it doesn't work out, people move on as quickly as they came as well. So we have to be aware that we have to work on deep, sustainable, hard problems and be mindful that hype will be here and be gone the next day.

23:36So assuming that someone's bought into the idea of being an AI engineer, they want to be part of the community, the World Fair is coming up in two weeks. Can you tell me a little bit about the event, how it's changed from the last time you guys organized it and what someone coming should be trying to get out of it? So yeah, the event is my highest stakes expression of my opinion of what AI engineering should be, because obviously everyone should have their own opinion. But this is my chance to sort of gather the community that has gathered around this term. It's also for me an expression of creation of opportunities.

Read the full transcript

24:09I'm always about trying to get people to meet each other in sort of many to many races rather than just through me or meeting to be rate limited by me. So the most obvious transition is the transition from a single track conference to a multi-track conference. And we definitely overdid that one. We went from one track to nine tracks. Didn't really expect that at the start, but just this sheer amount of interest and clear swim lanes of things that people were working on that I thought were worthy of their own conference. I basically picked those things up. Last year, one track. And then this year, we have the sort of standbys, the rag and the code generation stuff, which we also had last year.

24:48But then this year, we added multimodality as its own track. This year, we added evals and ops as their own track, which is super, super popular. Agents, we had a bit of last year, but this year, we had much more organized tracks. And then, for example, I wanted to have tracks to address specific criticisms of AI. For example, a lot of people would say that AI is adopted by startups. And that's absolutely true. But, you know, there's interesting stories to tell in the Fortune 500 and larger scale deployments. So I have straight up AI in the Fortune 500. And just putting that in the title attracted a different kind of speaker and different kind of audience, which is great.

25:21I actually thought last year was too startups focused. That gets you in this sort of naval gazey Silicon Valley type bubble where everyone actually like talks a big game and raises a lot of money, but doesn't actually make a ton of revenue. and uh you know i think the revenue is basically in a capitalist society the only thing that keeps you honest and so i'm really proud to initiate that track uh you know we not everyone in there is actually formerly fortune 500 but they're they're at least household names that everyone would care about coinbase and salesforce and then finally uh we for the first time we're also doing the sort of vps of ai track right like trying to and that's the one that you're speaking on which is addressing some team and leadership level conversations around AI strategy and growing a team, organizing a team and setting principles for what that should look like.

26:08For me, the ultimate win condition for this conference is actually like you don't even go for the talks. You just go because you know everyone else is going and you just show up and you're like, I know I should be going for the talks, but my conversations here are so great. I don't need anything else. That's actually a win for me. But these conferences are expensive affairs and people have to expense it. So I'm definitely creating something for people to expense and to be work relevant and to level up, to find jobs, to hire, to launch their products. I think we have three foundation models launching, by the way.

26:37That's sort of, yeah, I never said that publicly. You heard it here first. Are we going to get the big one? Not the big, big one. No comment. We have decent, so I don't think we've reached that tier where people want to give us their big launch. Every year, we're getting more legit. This is the second iteration. The first year, we had OpenAI and Anthropic, which is a good way to ask. How many people have you got coming this year? We're targeting 2 ,000. We're at 1 ,500 now. I don't know what the exact number is going to be. Typically, the online audience. Last year, our in-person was 500, and the online audience was 20 ,000.

27:11Then the asynchronous audience was 150K-ish for a single one of our top talks. Then this year, we're just basically trying to make everything four times larger. Right. I think relative to almost anyone in the space, you are incredibly well positioned to speak to a lot of different types of people through the Latent Space podcast, through your work as a developer relations engineer. I think you have a pretty unique perspective. And so one thing I wanted to ask you about is, you know, what advice do you have for product and AI and leaders? And I have specific questions there, but I want to give you the generic version of the question first to see where you take it.

27:44So if I'm an AI engineer or I'm a product engineer building today, maybe I'm starting a new project, what are the gotchas? What is the best advice that you would give? Oh, boy. I totally don't feel qualified to do it. I can repeat the wisdom of people much smarter and more experienced and more successful than me because that is my role as, I guess, a content person and community person. Okay, I'll give you a common take and I'll give you a harder take. Okay. Common take would be that you should move fast and much more so than a typical YC startup mentality. If you're taking like three months to ship a thing, what could you do to make it three weeks?

28:23You know, that kind of mentality of try to move fast. Because if your typical development timeline assumptions are based on traditional ways of making AI products, then maybe you should think about what is unlocked by foundation models that you can do differently, that is faster, that is mockable by a really shitty prompt. But whatever, it kind of works and get it out there. A core part of the thesis of the original essay was that you should adopt a fire ready aim approach instead of a ready aim fire approach. and where the old ML engineering process was much more deliberative, here you actually win by moving really quickly and getting information from the market, from shipping products, so that you can iterate much faster.

29:05Maybe, you know, in a way, like the gradient descent is performed in the marketplace of the user rather than in a model weight because you don't need to start with data. You need to gradient descent your product. And that's something that resonates at least from what I've seen as well, in that I've seen a lot of companies ship a V1 that's good enough and use the data they gather in production, the eval data, the feedback data to rapidly improve that. And without that feedback data, it's really hard to do so. So there is a flywheel that you can get going if you're quick to deploy something. How do you trade that off with the caution that some larger companies might have about hallucinations or risks?

29:44Those are very important, but also I think probably overstated for a lot of companies because they can just set the expectation that, hey, this is a beta thing, or this is generative AI products. Wink, wink. You all know the risk that comes with that, but we're just trying our best. And that suffices for most companies that are not named google.com with the responsibility that google.com has. But most companies, you are not Google. Go ahead and experiment. Get out of your own head. Go try things. So you said your non-spicy take was people should definitely focus on moving fast and faster than they might in another circumstance.

30:18What was the spicy take? Yeah, it's not spicy because everyone should be moving faster and regardless of AI or not, right? Like actually that's a universal truth. The spicier take is something I've been sitting on myself for a while and, you know, maybe concerns you a little bit. What kinds of startups are performing better and seeing more success in the market? It's vertical rather than horizontal startups. And a lot of people try to build picks and shovels. And obviously some picks and shovels creators will win. But actually, the people that have the most proprietary data stand out the most, pursue the markets with the highest margin because they pursue very price insensitive, non-technical audiences that just have a problem that's the burning pain points that you can solve with AI.

30:58Like they want an AI solution, no one's solving solutions for them. Those vertical startups actually have the unique insight into their customers that they're making the most money and growing the fastest compared to the horizontal one. We have a company that's launching with us that is working in AI and construction, right? And most software engineers don't really like, you know, the real world. We're much happier building developer tools where, you know, we just talk APIs and can log everything. But like, you know, if you're willing to do the nasty work in the real world, then you get appropriately rewarded.

31:30I think that's entirely appropriate. But at the same time, you also have the most domains of specificity and unique knowledge that you can train your models or whatever. You can pursue high margin markets and you're most likely to not be steamrolled by OpenAI when that day comes when OpenAI does a launch and they're like, whoops, my company's gone. So for many, many reasons, I do think that vertical companies have it right. And I wonder if the horizontal players, there's going to be a shakeout where a lot of them fail. And that's just the brutality of the market. I mean, that's neither here nor there.

32:00But I think that people might underappreciate how easy it is to get started on the vertical stuff compared to on the horizontal stuff where you're one of many. And which vertical AI startups have really caught your attention? Like, what do you have in mind when you say that? Is it something like Harvey AI in the legal space or you mentioned construction? Like, what are other examples? Because the counter argument that people will often give is, you know, some of these vertical AI startups, the worry is that they'll get eaten by improvements in the models over time. Why is that concern wrong? And what are some examples that you are really excited about?

32:33Okay, so maybe we start with the examples, which is always challenging to recite at the drop of a hat. So I think Harvey is definitely well-known as an example. I do think like MidJourney is vertical, like it's serving that sort of creative market. You know, and classically, for those who don't know, MidJourney is making somewhere between$200 to$300 million a year, completely bootstrapped with a 50-person team, something of that order. The numbers might be like plus minus 10 % off, but those are the rough magnitude. And that is one of the most successful startups of all time already. So you have to get creative about what that vertical is.

33:09But I think increasingly you will see other verticals that just come up and take on incumbents directly. I am thinking of perplexity there as a very successful anti-Google, even though they sort of wink-wink. They're like, no, we're friends of everyone. No, they're trying to take on Google. And I think they're doing a decent job. It's an open question as to how profitable they are. But they definitely made a dent in terms of the public persona perception. And then like other kinds of verticals, I could name Peter Levils' photo AI and the sort of room, I forget what it is, I think it's called interior AI, right?

33:39And that focuses on real estate agents doing virtual staging for houses, right? Like whenever they're trying to put houses on sale, like they need that. There's a bunch of people serving those verticals. I mean, yes, one of the verticals is developer tooling. I think the sort of cursors and co-pilots of the world, like focusing on that is like, it's a very clear use case where like you win when everyone basically says like you need to have this or you're behind in that industry. And I think the Harvey customers would say that. I think the cursor and co-pilot customers would say that. And I think like there is a tool for like building the AI, you know, companion or friend or like sort of default tool of every single vertical out there.

34:17Like there's got to be someone for the medical side. I can't really name them just because I'm not that close to that side. But like the same exact thing as Harvey is doing, you would apply on the medical side. But Brightwave, our most recent podcast guest is doing that for hedge funds. It's obvious that people want to do research quickly. A lot of research analysis is very commoditizable using language models. You can do that a lot quicker. And yes, you're going to hallucinate some, but so do your analysts, by the way. And it's no different. So yeah, those are all verticals. And for me, on my very, very small, you know, perch in the world, new summarization is a vertical that no one's actually seriously pursued.

34:53Everyone, you know, builds it as a feature. I'm building it as a product. Like, it's so, it's too easy to win here. Like, it's, that's one thing, something that really, like, puzzles me that, like, you have to work really, really hard in horizontal, like, dev tools type things to, like, win. You know, everything has to be open source. Everything has to be, like, distributed and scalable. and you need all these checkboxes and you need to sock two and whatever in order to get your customer. And then you just build one vertical product that solves the pain point and everyone is raving about it even though they don't really know anything else about it.

35:24And that's fantastic. That's what PMF should look like. You've worked now, I guess, at three different dev tool companies, unicorn companies as in DevRel. I think you've developed a lot of taste for developer tools. You're working on small AI yourself. We just discussed that the horizontal developer tool space for AI is increasingly crowded. There's a ton of different options. If you're a buyer trying to navigate this, there's a lot of noise. So what advice would you have for someone, not on the vendor side, but on the buyer side, about what is smoke and mirrors versus what do people really need?

35:59Are there tools you're excited about? What's the stack, I guess, that people should be thinking about building? I think that there is actually a lot that you should buy from the beginning and then selectively build later on because the community has probably run into more problems than you know of. And actually, it serves you well to buy first, understand your own problems with the protection of the communities to figure this out. And then once you understand where you differ from the community, then you can eject if you want. So obviously, you want to think about lock-in. And it's more expensive to build because you're going to slow yourself down if you try to build everything in-house.

36:38The sort of not-invented-here syndrome is very big. And I think some things like you need some kind of evals platform makes absolute sense. You need some observability and monitoring on your APIs, just like other regular APIs. But now with an AI flavor, some try to charge you a lot for that. And obviously, there's a right price for this. But it does absolutely make sense that you should buy all these things just to move faster. going back to the organizing principle of what makes sense in this world. And then you can build later on if you're like, none of these solutions make sense for our case, or we're seriously overpaying for this.

37:10Fine, you only paid for two months of this. Go back to build, backfill, whatever you need. Only once you've understood the problem and you've explored the features that other people have built to serve other customers, because you're not going to be the only one running into the fact that, oh, you need portable keys or time-nuted keys for your API, key rotation for your OpenAI keys, or you need to track your inputs and outputs. so that you can do evaluations or fine tuning or whatever. Everyone has the same issues. You're not special. Just buy it because you're sharing the development cost with everyone.

37:39And then there's the AI product tooling, which is for you building products for your customers. And then there's the internal productivity tooling, which is a separate thing. And actually, that gets adopted a lot quicker. And that typically comes in the form of developer tool, right? Like in terms of your people using Copilot or Cursor or SourceGraph Kodi or anything of that sort. That's the baseline. That's the most proven one. At least most people in the AI community are using that. But then also there's the other productivity stack of like your meeting summarizers and what have you. You know, I'm not really in the business of picking productivity tooling, but I do think that there is a lot of advantages to that.

38:12Ultimately, the point that we're going to get to here is we should seriously think about virtual employees or like AI employees that would reliably perform parts of the jobs that you would normally assign to a human. That role is going to start small now and then going to grow over time. The people who are most able to take advantage of AI to do that tasks are going to be the most capital leverage, right? Like that they don't have to manage humans to do that. These things can work online for you. And that's fantastic. So like, yeah, we haven't really seen that yet. I think the closest I've seen is Lindy AI from my friend, Booker Velo.

38:43But again, like, you know, it's not really being adopted at massive scale yet. Like every year I have one essay that sends the test of time. This essay is The Sour Lesson, as opposed to The Bitter Lesson by Rich Sutton, which typically talks about the scaling of models and how you should basically never bet against scale. The sour lesson is always actually more about trying to compare humans and AI on that sort of equal footing. When actually they're, due to more of X paradox, they're actually good at different things. And they're just always going to learn differently, develop differently. And you should not try to directly compare humans and AIs.

39:11So like, I think the way that we interact with these AI employees, let's call it, is going to be very different than humans. And actually the analogy isn't going to be that light. Like there are going to be some things that we find where these are killer apps. and that we would never give a human and vice versa for the other side. I think I somewhat agree with that. I mean, for that reason, I think actually the idea of like human level in AI is probably a mistake because they're already superhuman in some dimensions and they're like much weaker in others, right? Like if we kind of look at the graph of skills, there are things that humans find trivial that are really hard for AI.

39:42And so the moment that AIs are able to match humans on the task where they're currently weaker than us, they'll already be superhuman, right? There's not going to be a moment where they're human level because they're already much better than us at many things, retrieval, memory, et cetera. Yeah, exactly. I'm actually going to move on because there's one final thing that I want to chat to you about before we wrap, which is just trends. There is so much noise. There's so much hype. There's so many new papers. There's something released every day. And I think you do a really good job of curating this for the community.

40:13You have your automatic summaries, but you also have the podcast and tweets, et cetera. If someone was going to focus on just a small number of trends or like a couple of things that, you know, they're trying to find the signal from the noise, what are the trends that you think people should really be paying attention to that are going on right now that maybe are slightly overlooked? Or maybe they're not overlooked, but just in the amount of noise there is, it's just hard to follow them. Okay, there are a few categories of this, which I could point out. We do occasional essays on Lanespace. And what I would highlight is the Four Wars essay and then the Research Directions essay, which I'll pick out here.

40:50And then there's other trends that didn't fit those buckets, but I'll just name them and people can go read up on that. The trends are that there are several key battlegrounds that you as a company, if you're going to be an AI company of any consequence, you're going to have a beachfront on this war or you're not really in the game at all because you're fighting a war where there's no opponent. because there's no limited resource and everyone is aligned in the same way. So for example, open source is not really a war, at least in the technical domain. It is in the political domain, but not in the technical domain because most people are pro-open source in tech.

41:22That's not a war. What is a war is fighting over data, fighting over GPUs, fighting over the God model versus the sort of domain-specific model, and then fighting over rag and ops. And so those are the four battlegrounds that I picked out that I think are sort of key battlegrounds where there will be winners, but all the participants cannot all win, and that's what makes it somewhat zero-sum and interesting as a war. That also means that there are several domains where I don't think there are wars yet, which is because they're not interesting or contentious. So like co-generation, very, very important problem, but no war because everyone's trying to make their own headway and no one's really figured out what's the next thing after co-pilot, right?

41:59Maybe it's dev and maybe it's not. Okay, those are the four wars. You can see our reading on that. And then research directions, So moving from the sort of commercial space into the research space. And I think having a filter for what kind of research matters actually really helps to survive on Twitter. Because you're just constantly DDoS by all these influencers going like, this changes everything. And, you know, this week there was like this like paper about how like you don't need matrix multiplications anymore. Oh, my God, NVIDIA is going to zero. And I'm like, what kind of, what are you smoking?

42:32I want that. But so like having a filter for like what is worthwhile research direction is actually important. I put myself on record as like ranking a list of directions, right? So for us, it's long inference as number one. Synthetic data, number two. Alternative architecture is number three. Mixture of experts and merging models, number four. And online algorithms, number five. And yeah, those are the main trends. You know, I can always talk about the other trends that we're seeing. So for example, having like what is a Moore's law of AI and what are the long-term trends you can bet on, right?

43:02And I think we're trying to define this for one of our podcasts that we're doing. But basically, the cost of a 70 MMLU every single year goes down by something like 5 to 10x. And so this is the observed trend, and it's going to keep going down. So it makes sense to build a product that loses money today because you're going to be ahead. You just have to wait it out because of Moore's Law. It's going to bail you out of whatever sort of bad inference ideas you may have. For people who might not know, I assume most of our audience would, but just briefly, what's MMLU? Just let people have the context.

43:32I think it's multimodal. No, it's not multimodal. Massively something. Language understanding. Language understanding. It's not multimodal. It's something about domain specifics. Basically, it's a conglomeration of all the professional exams that you could possibly ever want as a human. It's massive multitask are the two M's. Massive multitask, right? Fun fact, it's created by Dan Hendricks, who is a very, very noted AI safetyist. And that's a whole different discussion because it's very ironic that the thing that he created is like the primary tracker for AGI. Anyway, it's the primary number that every LLM basically compares each other by.

44:10And this is not something that you should fixate on because it's very likely within two to three years we'll move on to something else. Yeah, benchmarks are very much flyby. Once they saturate, we create the next one. Is one way to interpret what you're saying is this kind of Moore's law of AI. MMLU is a specific one you've chosen for now, but it's really a statement more on like the intelligence or the capabilities of the model at professional tasks we care about every year is going down by 5 or 10x or something like that, the cost to achieve the same performance. Yeah, same IQ. You can also call that.

44:40Like our GPT-4 level model was not possible in 2022. And suddenly in 2022, you could pay something like$20 per million tokens for it. Now the cost is$2 per million tokens. It is on its way to 0.5 or 0.25, depending on if you look at DeepSeq MOE, being a legitimate GPT-4 level model, and it will trend towards zero. So, okay, that's great. But then our levels are sort of bar for what acceptable AI intelligence will increase when GPT-5 drops. And so then when GPT-5 starts on a higher curve and goes down again. And that's a very classic sort of cost curve model that's typically what semiconductors, you know, typically operate on, like the different semiconductor process nodes.

45:23And they operate on that curve. And actually you want to understand understand that and build that into your product planning, right? So it's not just cost. I have a few trends here that I haven't really published, but I'll just kind of go through it. It's commodification of intelligence, which is the MMORPG cost going down over time. It's also inference speed as well. So like going from like maybe, you know, 170 to 100 tokens per second to 500. I've had very credible sources tell me that Brock is aiming for 5 ,000 tokens per second. So another 10X from where they're here. And so what do you do differently, right?

45:56Every 10x unlocks a different kind of product. So that's roughly 10 pages of text or a full essay every second in terms of inference speed. Yeah, I used to emphasize every AI API must incorporate streaming because you want to stream out tokens when you're auto-aggressively generating things. But when you're generating 5 ,000 tokens per second, you don't need it. It's crazy. Okay, so I'll keep going. Context is treading to infinity, right? We used to have 4 ,000 token context models, and I thought that was enough. And now we have a million token context models, and I don't know what to do with it.

46:31But people will find the use cases for that. It's really interesting. Multimodal everything is another trend that I call out. I mean, this is obvious now because it's GPT-4-0, but all modalities in, all modalities out is a very nice shorthand for that. There's nuances around that. And then finally, variance is a very interesting trend, which is definitely one of those underrated things that maybe could become more prominent over time. What this is, is basically a lot of the use cases that we talked about for work is temperature zero use cases. You receive this God model from the heavens. And the first thing you want to do is lock it down and make it do your retrieval augmented generation, right?

47:04Like, that's the most boring possible use case. Like, what if you just let it loose and try to make it think of things that you never thought of? you know, that's this whole comparative benefit and you're chaining it down, trying to force it to do this other very unnatural thing for it. What if hallucination was a feature and not a bug? Right. So there's this whole emerging category of temperature tool use cases is what I call it, where they're diametrically opposed to the temperature zero people, where like, hallucination is a feature, like, come on in, like, this is actually really great, because I never thought of that.

47:32And creativity is expensive on my team. And if I can, you know, spend a few dollars to add more creativity, absolutely, I should do that. Yeah, that last one, I think is a particularly interesting one that I resonate with. I think it's not just for creativity, but actually, if we want to be able to create new knowledge from these systems, I don't think we can do it with LLMs alone, but the combination of a model that can act as a conjecture machine, can chuck out possible explanations for things, coupled with some way to actually test that and measure it, I think is a way that you start to build AI or computer systems that can generate new knowledge.

48:05And obviously, you want those operating, as you say, sort of high temperature, not deterministic mode. All right, Sean, I think that's all we're going to have time for today. I feel like I could talk to you for hours and hours. Just before we go, if people want to attend the conference, how do they get tickets? Where do they find it? And where else can they find you on the internet? Yeah, probably the smartest thing I did when writing the essay is I bought the domain. The domain is ai.engineer. You go to ai.engineer and you always see the currently active conference. And I've created the code IAgency for people, for listeners who've come in this far and want to get last minute tickets.

48:38We'll hopefully get this out as soon as possible. But yeah, it's definitely the industry conference to be at. You can find me on Twitter at Swix and our podcast, as well as Latent.Space. We just love the short domains, right? I hope you get high.agency. That's already taken. But we got Latent.Space donated by a listener, actually, which is very fun. You can find more conversations that we have on our podcast. Fantastic. Well, Sean, it's been an absolute pleasure. Thanks for coming on. All right. That's it for today's conversation on high agency. I'm Reza Habib, and I hope you enjoyed our conversation.

49:11If you did enjoy the episode, please take a moment to rate and review us on your favorite podcast platform like Spotify, Apple Podcasts, or wherever you listen, and subscribe. It really helps us reach more AI builders like you. For extras, show notes, and more episodes of high agency, check out humanloop.com slash podcast. If today's conversation sparked any new ideas or insights, I'd really love to hear from you. Your feedback means a lot and helps us create the content that matters most to you. Email me at razahimulub.com or find me at razrasgal on X. Until next time, Raza.

From the publisher

The World’s Fair is officially sold out! Thanks for all the support and stay tuned for recaps of all the great goings on in this very special celebration of the AI Engineer!

Longtime listeners will remember the fan favorite Raza Habib, CEO of HumanLoop, on the pod:

Well, he’s caught the podcasting bug and is now flipping the tables on swyx!

Subscribe to High Agency wherever the finest Artificial Intelligence podcast are sold.

High Agency Pod Description

In this episode, I chatted with Shawn Wang about his upcoming AI engineering conference and what an AI engineer really is. It's been a year since he penned the viral essay "Rise of the AI Engineer' and we discuss if this new role will be enduring, the make up of the optimal AI team and trends in machine learning.

Timestamps

00:00 - Introduction and background on Shawn Wang (Swyx)03:45 - Reflecting on the "Rise of the AI Engineer" essay07:30 - Skills and characteristics of AI Engineers12:15 - Team composition for AI products16:30 - Vertical vs. horizontal AI startups23:00 - Advice for AI product creators and leaders28:15 - Tools and buying vs. building for AI products33:30 - Key trends in AI research and development41:00 - Closing thoughts and information on the AI Engineer World Fair Summit

Video



Get full access to Latent.Space at www.latent.space/subscribe

More from Latent Space: The AI Engineer Podcast

All 247 episodes
[High Agency] AI Engineer World's Fair PreviewLatent Space: The AI Engineer Podcast · 50 min
Listen in VO