[AIE Summit Preview #1] Swyx on Software 3.0 and the Rise of the AI Engineer

7 Oct 2023 · 39 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Latent Space Podcast Episode Notes

Episode Title

[AIE Summit Preview #1] Swyx on Software 3.0 and the Rise of the AI Engineer Podcast Description: A podcast by and for AI Engineers discussing the latest in AI, news, papers, and interviews with leaders in the field.

Episode Summary

In this episode, Swyx delves into the evolution from Software 1.0 through Software 2.0 to Software 3.0, emphasizing the significance of foundation models and the emergence of the AI engineer role. This discussion serves as a precursor to the upcoming AI Engineer Summit, encapsulating the current landscape of AI engineering.

---

Key Topics Discussed

  1. Evolution of Software
  2. Software 1.0:
  3. Traditional, hand-coded software using conditional logic and loops.
  • Software 2.0:
  • Involves machine learning models (e.g., neural networks) trained on datasets.
  • Software 3.0:
  • Characterized by the use of large, pre-trained foundation models, allowing faster deployment without the need for extensive data collection or labeling.
  1. Foundation Models
  2. Definition: Large pre-trained models like GPT-3/4, Claude, and Whisper that can be used via APIs.
  3. Benefits:
  4. Dramatically reduce the time to market for AI products.
  5. Allow developers to skip data collection and training phases.
  1. Putting Foundation Models into Production
  2. Difficulty Levels:
  3. API Calls: Simplest form of implementation.
  4. Local Deployment: Running models locally for personal use.
  5. High-Volume Predictions: Serving models to external users, requiring infrastructure expertise.
  1. Emerging Role of AI Engineers
  2. AI engineers bridge the gap between traditional software engineers and machine learning engineers.
  3. They utilize foundation model APIs to build applications without needing in-depth ML knowledge.
  4. Distinction:
  5. AI Engineers: Use pre-built foundation models.
  6. ML Engineers: Create and train models from scratch.
  1. Economic Implications
  2. Demand for AI engineers is rising while the supply of skilled ML experts is low.
  3. Transition from software engineers to AI engineers likely due to the accessibility of foundation models.
  1. AI Engineering Stack
  2. System of Reasoning: Source of foundation model APIs.
  3. Retrieval Augmented Generation (RAG): Framework that connects models to external data, improving personalization.
  4. AI UX: Innovations in user interaction with AI beyond traditional interfaces.
  1. AI Skepticism and Blame
  2. AI Blame: Companies may attribute failures to AI rather than underlying issues.
  3. Importance of maintaining a balanced view: recognize both the potential and limitations of AI technologies.
  1. Conference Announcement
  2. AI Engineer Summit: Scheduled for October 8-10, featuring speakers from leading companies like OpenAI and Microsoft. Aimed at fostering community and collaboration among AI engineers.

---

Takeaways

  • The shift towards Software 3.0 represents a fundamental transformation in the software engineering landscape, allowing broader participation in AI development.
  • Foundation models drastically lower the barrier to entry for AI projects, enabling more software engineers to explore AI.
  • The AI engineer role is expected to grow significantly, creating a new class of professionals in the tech industry.
  • Continuous innovation in AI UX and the integration of foundation models into practical applications will shape the future of software engineering.

---

Additional Resources

  • [AI Engineer Summit](https://www.ai.engineer/summit/)
  • [State of AI Engineering Survey](https://www.surveymonkey.com/r/aiengineering2023)
  • [Latent Space University](https://www.latent.space/p/lsu-beta)

For a deeper dive into the topics discussed, listeners can access the full transcript [here](https://assets.fireside.fm/file/fireside-images/podcasts/transcripts/3/3911462c-bca2-48c2-9103-610ba304c673/episodes/f/fe1aae1e-4137-4c07-a2c0-34b1769debcb/transcript.txt).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Ladies and gentlemen, it's the Latent Space Weekend Edition. Woohoo! This weekend is a special one, as we are gathering many of our former guests and over 10 ,000 of you for our very first AI Engineer Summit, both in San Francisco and on YouTube. We're all very excited about the summit, and we were even interviewed on a few of our fellow AI podcasts about it. We figured we would cross-post those episodes over the weekend to help you prepare for the exciting lineup of speakers, even if you can't join us in person. First, we'll have an introductory episode recorded with Tejas Kumar of PodRocket, where we introduce the concept of an AI engineer to a generalist engineer audience and go over SWIX's recent conference keynote on Software 3.0.

0:50While you are listening, there are two things you can do to be part of the AI engineer experience. One, join the AI Engineer Summit Slack. Two, take the State of AI Engineering Survey and help us get to 1 ,000 respondents. Both are linked in the show notes, and we would really love to have you. Now here's Swix's conversation about software, 3.0, and the AI Engineer Summit. I'm really excited to get into your talk, especially because we've already talked a little bit about AI and address and things. But before we get into it, let's zero in on scope a little bit and talk about your talk, Software 3.0 and the Emerging AI Developer Landscape.

1:27I've seen the term Web 1.0, Web 2.0, Web 3. Also, like in outside of web dev, there's the industry 3.0 is a trending term. I'm curious, software 3.0, where does that come from? It doesn't have any lineage whatsoever with Web3. So that's a very unfortunate comparison there, I think. But the origin actually comes from a very influential article created by Andre Karpathy about six years ago called Software 2.0. And he was basically trying to articulate the difference between hand-coded software, where we write every single line of code ourselves with like if statements, loops, and whatever, traditional coding paradigms.

2:07And then machine learned code, where you write the layers of what the machine learning model should do, like the architecture of the model. And then you just run it through a lot of data in order to achieve weights. And the weights themselves are the encoding of the knowledge. So he was trying to articulate that difference that the possible space of problems you can tackle with software 1.0 is the problems that you can kind of code for deterministically. And the possible space of problems that you can address with software 2.0 is the stuff that you can address it by machine learning. For example, computer vision and voice recognition.

2:41It's that stuff that you'll never be able to hand code by yourself. And I think there's the fundamental realization that a lot of people should have with regards to how they write software 1.0 code, which is a lot of the times, like what do you do as a programmer, as a software engineer, right? You write some functioning app and then you send it out there. You look at your analytics and your metrics and all that. And then you adjust by adding in some features and adding in some if statements and all that from learning. And essentially what software 2.0 is accelerated learning from data. Whereas in software 1.0, we learn from data through humans in the loop and designers in the loop.

3:15So I think that's a really fundamental realization there that once you realize that sometimes you are just a very slow machine learning model and you're writing all these algorithms, but yourself, sometimes you can just kind of machine learn the algorithms rather than writing them yourself. So how do you proceed from software 2.0 to 3.0 is the arrival of foundation models. And that's the change that has happened more or less in the last three years, enabled by the transformer architecture becoming a thing, which enables deep learning that's parallelizable and massive scale and obviously a lot more money and GPUs and data thrown at this problem.

3:50So now foundation models mean that you do not have to collect a whole bunch of data to create models before you start delivering ML products into production. You can just grab one off the shelf, whether it's open source or closed source, it doesn't really matter. You take a foundation model and then you put that into production and then you can start collecting data to fine tune them if you want to. But otherwise, the time to MVP of an AI product has significantly reduced by orders of magnitude in the software 3.0 paradigm. So hopefully that transition is clear. Software 1.0 is hand-coded code.

4:21Software 2.0 is machine-learned code on your data that you collect. And software 3.0 is just off-the-shelf models when you don't even have to collect the data. Wow, that was an amazing answer. You alluded quite a few times to model architecture, the architecture of a model, etc. Maybe for our listeners, you could have maybe a sentence or two about what model architecture even is and why grabbing one off the shelf is beneficial to this software 3.0 paradigm. There's a lot of ways in which we can take that question. I would say that a very typical model would be, these days, transformer space models, a decoder-only generative P-train model.

4:58A GPT itself is actually a type of architecture that you can reference the GPT-1 and 2 papers from OpenAI. And they actually published open source code to do that. I would say the most definitive open source reference implementation of that kind of architecture is currently from Meta, where they released the Llama 2 code, which is only a few hundred lines of code. It's actually very little code. And all the value of the code has now shifted from the code base to the data weights. And that's a very common paradigm coming from software 1.0 to 3.0, right? Like Meta can open source the code base. It doesn't matter.

5:32It's because there's only a few hundred lines of code. but they are not open sourcing their data set, right? Because that's actually now the much more valuable thing that they're not giving you. They do somewhat open source the weights. It's not fully properly open source, but for most people, they can actually use it in their commercial pursuits and that's what most people care about. So this contrasts, by the way, with some of the other architectures that we might have pursued in the past. So for example, LSTM networks, if you're in traditional NLP or convolutional neural networks, if you're in image recognition, And there's a bunch of other architectures as well.

6:06The RNN is one of the oldest ones and is actually making a little bit of a comeback as a potential challenger to the transformer. But all these are possible architectures where you spec out the model in something like PyTorch. You define the number of layers that you're doing, that you need, and then you run it through a training process, and then you start deploying it. And I'm cutting out a lot, but I do think that AI engineers don't really need to know the internals of these things. They need to know how that impacts the products that they can make. And that's about it. And so I think this is classic engineering where you're not quite a researcher, you're not quite a scientist.

6:39AI NML is at a point where it's crossing over into the engineering sphere where a lot of the rest of us without that training background can actually get pretty far just by knowing how to use end products rather than to make the products ourselves. Yeah, that's a conversation I'm really excited to have. Before we do, I have just a couple more questions based on what you said. And the questions come from wanting to really answer the questions that I know we're going to get from the listeners. You mentioned foundational models. I've heard similar terminology in the space pre-trained models. Is that somewhere close to the same thing?

7:10Are they similar somehow? Yeah, I would say they mostly have 100 % overlap. So some pediatric people might want me to point out that the term is foundation models, not foundational. And then there's also another term that's emerging called the frontier models. And so the frontier models would be foundation models that are extremely cutting edge and the largest of them that most probably come under some kind of regulational scrutiny from Congress or some other government bodies because they are so big that they are potentially civilization threatening. So example foundation models would be GPT-3 and 4, Claude, but also Whisper, also Segment Anything.

7:42All these are foundation models where you can get them off the shelf without actually trading anything, and they zero-shot transfer to tasks that you actually want to put them into use on your apps. And that's what a foundation model is. Stable Diffusion, for example, is also another foundation model. Basically, there are big binary blobs of data, sometimes 4 gigabytes, sometimes 180 gigabytes, depending on how you quantize the models. but these models are just a result of millions and millions of dollars of training these models through GPUs, running them, running data through them based on some predefined recipes such that you get the end result of these blobs of binary data that you can actually run in inference mode through these models in order to make inferences and predictions.

8:24And when we say inferences and predictions, that's a very machine learning term. In terms of products, when we make AI products, it's really much more like generating text or generating images, anything like that. And that's the fun linguistic challenge when you cross fields from the research field into the product field. Because at the end of the day, in the product domain, in the AI engineering domain, we just care about what we can do for our users, right? And once we can produce interesting things for our users that people want to pay for, then we get a lot interested. But we have to respect that there's a lot of prehistory of research terminology that is completely different because they think about things in a different way.

8:57Yeah, I once heard, I think it's Ashi Krishnan say, this term stochastic gradient descent in the ML academic world just means try stuff randomly and have less errors over time. So I would say my favorite spin on stochastic gradient descent is the graduate student descent, which is instead of trying stuff by machines, you just throw graduate students until you find something that works. Yeah. Final question about what you just said. You mentioned nowadays with software 3.0, it's very easy to grab these foundation models and put them into production. What does putting a foundation model of production mean?

9:30How do you practically do that? That's a whole topic of a conference. So I think there's a bit of a U-curve in the difficulty. If you're just wrapping the OpenAI API, then you just call it an API just like you would call any other API in an app. There's not that much difference there apart from maybe you have to be mindful of things like your context limits, your privacy considerations, especially when people are putting in sensitive information into your stuff. And then maybe like your rate limits. Like if someone, there are a lot of bots out there that will scrape any exposed OpenAI endpoints because these are valuable things and expensive token calls.

10:07So if you leave your API endpoints guarded, or if you, God forbid, leave your tokens out there, they will be scraped and used and they'll run up your bill. And then there's the path towards the local llamas, where you run your models locally. And that is quote unquote production just for yourself, right? You're not serving an outside audience. So you're just serving for personal use and you probably want to run it on your local machine. And so that's a one level of difficulty up from just calling an API, because typically you would want to run things like llama.cpp or whisper.cpp locally. And there's a whole stack that needs to be done there.

10:40And then the hardest of all is actually serving your own custom models to a lot of users externally as though you were a model infrastructure platform. And there are many of these out there. You can actually buy them off the shelf or you can set them up yourself. I would say you probably have to be infrastructure expert to be able to be running these by yourself. From what I can tell, it's not that hard. The mechanics of these things, you just have to understand basic principles like saturation, basic principles like what the bandwidth is of individual parts of the model architecture that you've chosen to be able to serve things well but i think ultimately the secrets of high model flop utilization right basically when you buy a gpu you have a certain amount of theoretical flops most people only operate at like 40 to 50 percent model flop utilization so if you want to go down that path you are basically going to have to become a gpu infrastructure expert which i am not in my mind And that is how the landscape flows, right?

11:38Like either you use something off the shelf or you build your own. Can you quickly just define flop for our listeners? Floating point operations. And a lot of this math really is just multiplying matrices again and again. And that's how we do everything from embedding tokens to predicting the next token that eventually ends up towards either building a diffusion model or predicting the next token in a language model. It's all math at the end of the day, which is actually pretty interesting and fun. but very intimidating. So ultimately, every operation reduces to a certain amount of flops. Larger models require more flops to operate.

12:13So how many flops can you generate in order to serve those models quickly? That is the fundamental problem of infrastructure serving. But there's one concern which I will point out for people, which is the ability to batch. And that is primarily the key trick to reducing costs per tenants if you're in a multi-tenant situation. So far, we've talked about half your talk title. We've talked about Software 3.0. I want to spend the rest of our time I'm talking about a topic that I personally am really interested in. We started this discussion earlier, and I'm really excited to continue it. The emerging AI developer landscape.

12:43This is a topic that I absolutely love. And Swix, you'll be proud to know I've been officially wearing the badge of AI engineer since our discussion. But I want to clarify this for everyone. So in your talk, you mentioned AI is shifting right. And you mentioned a new role. I'm curious if you could expand on that just a little bit. Yeah. Basically, I think that the arrival of foundation models make it such that you actually don't need an ML team in-house to ship an AI product. And that is a very different situation than 10 years ago when you would very much need to do that in-house. So what does that mean?

13:18It means that there's a few orders of magnitude more people that will be trying to ship AI products that don't have the traditional ML and research backgrounds to fully understand all of it or to make all of it themselves. But when you give them a shapeable tool like a foundation model, they can actually go make lots of money and make a lot of people happy by building AI products. And which is exactly what we're seeing happen in the indie hacker sphere and increasingly in the B2B sphere. So I basically would call this self-socialization of software engineering. If you imagine a spectrum from left to right, the left would be the research scientists, the people who are innovating on the algorithms and the architectures, like the ones that we talked about.

13:56And then on the further right of those are the machine learning engineers. Those would be the people that are not research scientists, aren't as good at calculus as the other folks, but they are good at model infrastructure and serving in data pipelines and all that fun data science and all stuff. At that point, there's like a permeable boundary, which I like draw a line between. And basically I draw a line at the API layer. Like when stuff gets thrown over the API, whether internally within a company or between companies as an API foundation model lab, then you can consume it on the software engineer side as an API and put that into products.

14:30And so on the right of the spectrum is the software engineers. I do think like the traditional full stack software engineer, the one that is typically front-end only or serverless or front-ending serverless or whatever you call it or full stack web dev it doesn't really matter because there's millions of those out there and don't have any experience with AI and the real argument is that this sort of skills gap between the ML engineer and the research scientist is crossing over into the software fields and there will be a new class of software engineer called the AI engineer that will specialize in the stack because keeping up to date knowing all the latest techniques knowing how to put stuff into production and knowing how to advise companies on what is a bad idea to do because it's not ready yet or whatever.

15:11All of these are the domain of something that will probably be a specialist field, calling it the AI engineering. Just so I understand, in the past, I think even today, frankly, there would be people who hear the term AI engineer and without further clarification would maybe consider it something interchangeable with a machine learning ML engineer. But what I'm hearing is a distinction where the AI engineer is one who consumes foundation models that expose APIs and use those foundation models as primitives for building apps that encompass them and more logic to serve users. So it's different from the ML engineer in that it is not academic.

15:45You don't need a calculus background, but all you need to be able to do is actually be a software engineer as the crude version. And then the specialization is a software engineer who knows how to work with these foundation models exposed over APIs. Is that accurate? Yeah, I think so. And what's fun about this is that I think that AI engineers will be low status for a while because machine learning engineer is a well-established role with a lot of hierarchy and a lot of syllabus and curriculum, most of which is completely unnecessary when it comes to foundation models. So it'll be a very fun, disruptive few years when we figure out what's the right pay or career path of an AI engineer is as it starts to separate from the ML engineer because they're enabled by foundation models.

16:29I'm pretty strongly convicted that this will happen just because this is not a prediction based on tech. This is a prediction based on economics, on pure demand and supply of relative numbers of people and skill sets available. Wow. That's really an interesting perspective. I'm curious if you could speak more to the economics of this. I'm an engineer and the fact that you tell me I can query things over the network or query things in general and build apps and call myself an AI engineer, I'm really happy with. But could you speak a little bit more about the economics of it? Is it just that there's a ton of demand and yeah it's purely demand and supply right like there's a ton of demands not enough machine learning engineers and machine learning research scientists around to supply that demand so an intermediate class will be created and it will probably come from the software engineer side going down the stack rather than the ml engineer side going up the stack just because there's a lot more of the software engineers so that is guaranteed as one thing i think socioeconomically, software engineers also wants a way to jump in on the hype, which is why a lot of people have issues with my usage of the term AI engineer.

17:33A lot of people propose alternatives like LLM engineer or cognitive engineer. The thing is that these things don't roll off the tongue as easily as AI engineer. People want to associate themselves with AI. And so I think the people who do a good job of it will be able to put AI into practice for any company that approaches them, and there'll be very high demands. You know, I think that was a very core inspiration for why I did this, which is that I noticed that a lot of companies were trying to hire this profile of software engineer. And a lot of software engineers wanted to pick up more skills, but they didn't have any other ways to find each other.

18:07And so I think once you have the industry collect and coalesce around a single term that identifies a skillset and interest and maybe a career path eventually, and according to the kind of tools they're confident with, the papers that they might be familiar with. That becomes its own community and its own career in some industries. So I'm pretty interested in growing that. Obviously, I've already made my bets. I definitely will be honest and admit that this is super early, right? Like there are people walking around with the title of AI engineer, but it's definitely still in a smaller minority.

18:41But I do think that it will grow over time and it will probably exceed ML engineer by a lot. So this was backed up by Andre Karkati, who is one of the figureheads of AI. I think he was a co-founder of OpenAI, actually, when he read my piece on the rise of the AI engineer. And he said, yeah, it's probably true that there'll be more AI engineers than ML engineers. And I think this is emerging as a category and we'll have to basically fill out the tech tree. Right now, it's very undefined. So it's just a bunch of people hacking on Twitter and Reddit and hacker news. But over time, the courses will come in, the degrees will come in, the bootcamps will come in.

19:12And I'm very excited to see how that develops. just to be clear since it's so young there isn't currently a senior ai engineer title people can call themselves senior whatever they want uh yeah ultimately your definition of senior is has been around for five years you know the transformers architecture has only been around for six years so uh you'd be that senior on that front but i will say like it's still a lot of software engineering maybe 90 of it is still software engineering and if you qualify for senior software engineer on that side then you just need a bit more training on the ai side to match up.

19:45Yeah, no, the reason I ask is because I'm pretty sure someone's receiving a recruiter email asking for a senior AI engineer with 20 years of experience or something. Your diagram of this on the left hand side is the academics, the machine learning engineers, and on the right hand side is like the front end serverless type of people, I think is so awesome. And I think it can be generalized enough to where really you can use it to reason about a lot of the way the tech industry at large has developed. What do I mean by that? For example, if we think about personal computing, even, right on the left hand side, you had these mainframe folks that was super inaccessible to everyone else.

20:21And then on the far right, you have us today who use personal computers, but at one point in time, computers weren't personal. And they were largely in labs. And over time, that line has expanded. And today we see it commoditized for everybody, mass market commoditization, right? AI is experiencing something similar, where AI to this point is personal. Like I literally use ChatGPT every day. And it wasn't before. I could never run that type of model locally. Or maybe I could and I just didn't know how. Is that a fair way of reasoning about things? And if so, would it be appropriate to consider other things today for maybe people wanting to come up with startup ideas of things that are currently gatekept behind, I don't know, academia or money or access to certain machinery that over time might make it into the mass market, ergo broader consumer style people.

21:08I think that's never really been a hurdle. If you were smart and motivated enough, you would have figured it out at some point. But now it's just, the bar is just constantly getting easier and easier. At the end of the day, I would hesitate to offer startup advice just because I'm not a VC or anything like that. I would say like probably what is still successful and important for startups is to build things that people want and to know your customer better than anyone else and serve them and make them happier, better than anyone else and hopefully pick a growing market that is in demand for that.

21:41I do think that there are very good opportunities in effectively creating a poorer version of what a professional would do. And that's effectively what some of these things offer, right? Like, hey, you can hire an SEO copywriting expert to work on your website copy for like$3 ,000 an hour. Or you can pay chat GPT$20 a month and do an 80 % good job. And that's what automation is going to do. You're really going to take away like the lower tier of all of these specialist roles because now we got with General UKI to do it. And I think, and this is going to make the next things because obviously that's going to take away some people's jobs, but it's going to free them up towards doing higher value add jobs that machines cannot do yet.

22:26Yeah. If I was in such a situation, I'd probably try to use that time to learn AI engineering and get ahead of it. And that may be a prompt for some people listening. Okay, so the AI engineer is a new position that is effectively encompassed by folks consuming foundation models and using them to solve problems. You mentioned the stack, the modern data stack of the AI engineer. What is the stack? Is it outdated yet? Does it change as fast as JavaScript? Is Tailwind good? It's actually changing slower than JavaScript. So for a while, I was saying that JavaScript framework wars are over. and then now we have like Svelte and Solid and Inferno.

23:02I don't know what the new thing is. HTMX is the new thing. Anyway, so I would say it hasn't been that journey. I can go through the stack. So the thing that I announced at the talk and that we're doing a survey on actually is the software 3.0 stack, which I obviously use an idea I brought over from the modern data stack, which you referenced. Modern data stack is this sort of nice composition of how a data engineer should view their tools of the trade. And I think it's a nice map of like what you should learn and be familiar with, or at least consider if you need it within your company for your own needs.

23:33So the software 3.0 stack, instead of data warehouse, I basically have the system of reasoning, is what I call it, instead of like a system of knowledge or system of record. The system of reasoning is like the source of your foundation model, right? Whether it's a foundation model lab, like OpenAI or Anthropic, where it's closed source, but it's best in class and they provide it to you through an API, or it's open source and you have to take care of a lot of the model hosting issues yourself. And that's typically provided through Hugging Face, Replicate, Base10, Modal.com, and Lambda Labs. And so that would be the most valuable ones now, right?

24:05OpenAI has a valuation of$40 billion. Anthropic now probably has a valuation of$10 to$15 billion. Hugging Face has a valuation of$4 billion, and all the others are much smaller. Those are by far the biggest chunk of the value captured so far. And then we can go on to the RAG stack, the Retrieval Augmented Generation stack, because that is the stack that basically personalizes and orchestrates the AI models. So what does that really mean? Receive augmented generation really means that in every natural language model, let's just say like a GPT-4, there's a certain amount of context that lets you put in some extra information or examples such that you can actually personalize your answer towards something that you specifically want.

Read the full transcript

24:48So GPC4 is trained on a lot of web general knowledge facts out there, but it's not going to know specific things about your company. It's not going to know specific things about your products or person. So how are you going to pull in information and paste it in there and generate all that stuff? That is the subject of the retrieval augmented generation stack, or the RAG stack, as they call it. So in that bucket are the VectorDB companies. So most notably Pinecone, which is valued at$750 million. And then there's a long tail of others, Milvus, WeVe, Chroma. All in all, people really like to invest in database companies.

25:22These companies have raised$235 million this year, which is a lot of money. That's more money than MongoDB ever raised its entire lifetime leading up to IPO. So people are just investing very far ahead on the database side of things. And then the piping, the orchestration and application frameworks, the two leaders here are Lanchain and Llama Index. Lanchain has raised$35 million and Llama Index has raised$9 million. And both of them are effectively, we'll connect your LLM adapter, whatever LLM you have, whether it's the closed source ones and open source ones, we'll connect it to your data source, whether it's your Notion, your Slap, your Gmail, your Google Drive, doesn't matter.

25:58And embedded incentive vector database like the Python code, Chroma, VVS, and then we'll insert them into your context whenever you need to generate them. And that will serve as your personalization stack. And that's your bag stack. There are other companies that are focused on eliminating that process, because it's like a janky process that nobody really loves, but it is by far the best in class right now. So he built RAG right into the model itself instead of stitching together all these tools. There's a company called contextual AI that pursues that. The founder is the author of the RAG paper that founded this whole field.

26:29And then there's other open questions as to, can you fine tune new knowledge into existing models? And that is currently completely unknown. So finally, those are the two most established parts of the stack, the system of reasoning and then the RAG stack. The part that is completely open water right now is how people interact with the models and what I've been calling AI UX. I held the very first AI UX meetup in San Francisco. And AI UX is a big portion of the conference that I'm holding in October. And I think this is where front-end engineers should really get excited. Basically, when ChatTBT was announced in November last year, it was mostly a UX innovation.

27:06Instead of the sort of OpenAI playground that was not very inspiring, just being able to thread together a chat and going back and forth seemed to unlock a lot of value for a lot of people. And that caught OpenAI by surprise. The reason that they actually dropped it quietly with a very small blog post is because they didn't think it was going to be a big deal, but it was. So how do we break beyond the chat box? How do we unlock the capabilities of this reasoning and large language models towards more intuitive interfaces apart from just the UX? So one form of that is, for example, GitHub Copilot, where instead of a separate chat box, GitHub Copilot watches as you type and then tries to autocomplete as you type.

27:46That's something that they consciously engineer for over six months to get that experience. Because people find that if you have to context switch back and forth between your code and your chat box, they're not really going to use it as much. But if it just kind of auto appears and you can look at it in the context of your code, then they're much more likely to use it. So I think there's a lot of innovation that's open there. And currently there's no company that's like really owning that, except maybe from Vercel, which just recently released, announced v0.dev, which is their version of ChatGPT for UI generation, right?

28:19That you can type in whatever you want to create and it will create React and Tailwind code for you that you can just copy and paste. Yeah. And this is where I guess the AI engineers also rise up and see what innovations we can drive. Right. I'm inspired to think about the announcement, I think, from OpenAI today, if I'm not mistaken, where they announced audio and video and image inputs for ChatGPT as well. That's rolling out soon. I didn't know, you mentioned ChatGPT was just a UX innovation on top of what already existed. Is that true? Could I query GPT 3.5 and could I have built ChatGPT before OpenAI did?

28:55I guess is my question. So yeah, there's a little bit of debatable facts here. As far as OpenAI is concerned, all the public statements from everybody at OpenAI says that is what they've considered ChatGPT to be. It's a pure AI UX innovation. If you read between the fine lines, it is not actually exactly that because they released GPT 3.5 one day before ChatGPT. And so it was a slightly better model plus a new form of delivery. What I often say is you want to bundle new model with modality. and that's what ChatGPT became for OpenAI, which is the way that they deliver their models to the broader consumers, right?

29:32So today they announced GPC for vision as well as baking in this conversational experience, which means they also, by the way, have entered into the speech synthesis business. So OpenAI Whisper is going from speech to text. Now they have it going the other way as well, text to speech. And all of these are models that they just ship within ChatGPT without releasing an API for it, without open sourcing any of the code, without even publishing any papers. And that really signals them shifting into much more of a product company rather than an infrastructure company. Yeah, I'm really excited about this audio thing, as you mentioned, because this literally is Iron Man's Jarvis in real life, more or less.

30:10You have this thing that I was actually looking at the UX of the audio input feature for ChatGPT, and of course it does this. To me, I saw this feature and I was surprised, but then shortly after I realized, wait a second, this is actually, it makes sense. And the feature I talk about is that when you're done speaking to it, it just automatically submits the text because it's an AI product. Of course, it recognizes when you're done speaking. And I thought that was just really novel. And so you could literally just talk back and forth forever and say, hey, I'm struggling with this. It's amazing to see how good it is.

30:40We've had this before. We have Google Assistant. We have Alexa. We have Siri from all the other companies. Obviously, OpenAI is probably going to do a very good job. But it's not like these things don't exist before. And I think that's ultimately something that I do think a lot about in AI, which is that if you just look at things on a feature by feature basis, you can clone a lot of these features before OpenAI does. But do you have the brand that people trust that says, this is probably one of the best products out there and I'll just trust it default without even evaluating all the other options?

31:08And that's ultimately where you want to get to as a startup or a company. You want to build a brand that people trust where you've done the work and that's mostly good. Even though sometimes it's not that great, I would say Apple has tested their faith with their customers quite a bit. But most Apple customers will trust that whenever something ships, it might be two or three years late in Apple. But it will at least be good. Yeah. All right. Let's wrap up. I wanted to address one, I think, important point before we wrap up. And that is, so there's a number of skeptics around. Of course, I think it's actually good to have a little bit of skepticism around things just to keep us balanced.

31:41and the term AI blame has been going around and it's had some type of impact on a number of industries. So I'm curious if you could speak to one AI blame, but also in contrast to that, and in contrast to some of the skepticism around, what are some examples of AI that have been going really well? And I don't mean well in the sense that OkoChatGP and Copilot helped me code faster, but I mean well in the sense that it's made somewhat of a more, I'd say, meaningful difference in the lives of people. So AI blame is a term I made up. There's a lot of ethical and legal issues. And so specifically what I was thinking about with AI Blame was Chegg, which is a publicly listed company that sells college textbooks in the US.

32:18College textbooks are a huge racket. Everybody knows that you seriously overpay for minorly updated editions of college texts. And you don't really have a choice because you're taking the course, right? So obviously Chegg has not been doing very well. And in the most recent results, they posted really poor performance and they blamed ChatGPT. They said everyone's just using ChatGPT instead of buying textbooks. and the stock went down 50 % on that day. But if you zoom out, the stock has been going down for a while. It's been going down since before JGBT. And so people find AI as a convenient scapegoat for all of society's issues that were going to happen anyway.

32:53And so I do think that's a very funny phenomenon that I think is just common, right? It's very natural to blame things that you don't fully understand or even if you fully understand it, it's just very convenient to have a scapegoat for stuff. We did this kind of thing with Web3 and NFT and things like that. And I'm not saying AI and Web3 is in the same area. What I'm saying is that was new and there were skeptics and it was a scapegoat. So is it fair to say that? In a way, this is the reverse of skepticism, right? This is saying that AI is so successful that it is killing us. There is a reverse of skepticism.

33:25Whereas there's another form of skepticism, which is the Stochastic Pirates group. And this would be the ethical research group that came out of Google Brain that got fired very famously by Jeff Dean because of some disagreements in how the process that the stochastic Paris paper was published. And what these people are saying is that these are just language models. These are just giant matrices that simulate thinking they don't actually think. They generate plausible sounding text, but they don't actually have any real knowledge of the world because they've never lived in the world. They're just trained on text corpuses that we collected off of the internet.

34:01And that's obviously true to some extent. But this is all ultimately the discussion of does the math reflect the territory? It's a very age-old fundamental philosophy. Because if I can talk to you for 45 minutes on a podcast and sound like an AI expert without being an AI expert, how long do I have to talk to ultimately approach being an AI expert? And so there's some amount of like, this isn't real. This isn't actually thinking. It can't be reasonable. We can't treat this as intelligence until it's like human intelligence. And there's some amount of like, just look at the results. It's like language models are now superhuman on most reasoning capabilities, AP, bio, AP, math, whatever, the SAT, the GMAT, the medical exams, the law exams.

34:45It is now superhuman on all those things. At some point, it is genuinely superhuman instead of faking being superhuman. And when that crossover is from being like not really intelligence towards actually it's actually intelligence, it's really up to you to decide. but I can say that most of the time, the Turing test has been so conclusively passed that we are spending more time blocking humans from trying to do things than blocking machines from trying to do things. Yeah, I wish there was some type of committee that would just make the decision for all of us on what or when the intelligence is considered intelligence.

35:21Well, would you trust that committee? So I think cynicism is warranted though. I start off most of my talks by warning that There are a lot of signs of overheatedness, of extremely high expectations of people being very hyperbolic with their AI predictions. And that is just the result of the creator industrial complex is what I always call it. The news cycle needs everything to be at extremes in order for you to pay them attention and to pay them money. So nothing is ever as good as it seems. Nothing is ever as bad as the skeptics might make it out to be. I do think you get the most benefit just trying to build to solve for use cases that your customers have.

36:02And I think projects that you can actually make yourself satisfy and make yourself more productive, I think it's really good. I built a project for myself called Small Developer that generates Chrome extensions. Whenever I want, I just write a prompt for a Chrome extension and it mostly builds it up by itself. And that's super useful for me. And it's a lot of products that I would sell. But I do think that you as an engineer have the ability to experiment like that. and I think it's a wonderful time to start building and exploring. If people are interested in like the full list of like project ideas I have, on Latent Space, we have a little email course.

36:33It's completely free where you can just have like seven days of guided projects through all the major modalities of AI and that's what I've been building with a friend of mine, Noah. Oh my gosh. And that's just a free resource anyone can go sign up to today? Yeah, I just think like a lot of people ask me for like where to start and so this is my answer to where to start. I think that people should just build with projects in mind and obviously as a newsletter author, it's good to collect emails. So I just structured it as a newsletter. Awesome. Yeah, and you're one of the people who pioneered building in public.

37:02So I guess they can also build that in public for learning. Yeah, one of my biggest inspirations was WestBoss's JavaScript 30, where a lot of people learn JavaScript just by building projects on a day-to-day basis over 30 days. And I quite like that. So this is my attempt at building JavaScript 30 for AI. We'll definitely link to your newsletter in the show note captions. So I understand, Swix, you're organizing a conference. It's called the AI Engineer Summit. And I'm excited about it. I know that there's a way for people to participate online for free. I'm curious if you could say a little bit about that and help us understand it.

37:34Also, there will be a link in the show note captions for those who want to attend. Yeah, well, you don't need a link because the domain is something that we sprung for. It's ai.engineer, which is very fun to flex a.engineer TLD. And yeah, it's a two-day conference. It's going to be live stream because we're completely full in person. But we have people from OpenAI, Amazon, Microsoft, Fercel, Notion, MagChain, Llama Index, Fixee, AutoGBT, the biggest open source project of the year. This will be their first ever conference talk. Superbase, Guardrails, everyone notable I could collect in that space, getting all of them in one room and presenting the state of AI engineering.

38:11So if you want to check it out, head to AI engineer, put in your email and see you on the 8th to the 10th. All right, let's wrap it here, Suix. Listen, you are, I think I've said this to you last week, you are legitimately one of the smartest people, if not the smartest person I know. And it's an absolute honor and privilege to be able to talk to you and poke your brain a little bit, ask my silly questions and get your highly insightful answers. Small AI, building in public, the newsletter, all of the work. Thank you for coming on this podcast and enlightening me and so many others listening. Nice to have you.

38:40Yeah, it's been a pleasure.

From the publisher

This is a special double weekend crosspost of AI podcasts, helping attendees prepare for the AI Engineer Summit next week. Swyx gave a keynote on the Software 3.0 Landscape recently (referenced in our recent Humanloop episode) and was invited to go deeper in podcast format, and to preview the AI Engineer Summit Schedule.

For those seeking to ramp up on the current state of thinking on AI Engineering, this should be the perfect place to start, alongside our upcoming Latent Space University course (which is being tested live for the first time at the Summit workshops).

While you are listening, there are two things you can do to be part of the AI Engineer experience. One, join the AI Engineer Summit Slack. Two, take the State of AI Engineering survey and help us get to 1000 respondents!

Full transcript available here!

Links

* AI Engineer Summit (Join livestream and Slack community)

* State of AI Engineering Survey (please help us fill this out to represent you!)

* Podrocket full episode by Tejas Kumar

Show notes

* Explaining Software 1.0, 2.0, and 3.0

* Software 1.0: Hand-coded software with conditional logic, loops, etc.

* Software 2.0: Machine learning models like neural nets trained on data

* Software 3.0: Using large pre-trained foundation models without needing to collect/label training data

* Foundation Models and Model Architecture

* Foundation models like GPT-3/4, Claude, Whisper - can be used off the shelf via API

* Model architecture refers to the layers and structure of a ML model

* Grabbing a pre-trained model lets you skip data collection and training

* Putting Foundation Models into Production

* Levels of difficulty: calling an API, running locally, fully serving high-volume predictions

* Key factors: GPU utilization, batching, infrastructure expertise

* The Emerging AI Developer Landscape

* AI is becoming more accessible to "traditional" software engineers

* Distinction between ML engineers and new role of AI engineers

* AI engineers consume foundation model APIs vs. developing models from scratch

* The Economics of AI Engineers

* Demand for AI exceeds supply of ML experts to build it

* AI engineers will emerge out of software engineers learning these skills

* Defining the AI Engineering Stack

* System of reasoning: Foundation model APIs

* Retrieval augmented generation (RAG) stack: Connects models to data

* AI UX: New modalities and interfaces beyond chatbots

* Building Products with Foundation Models

* Replicating existing features isn't enough - need unique value

* Focus on solving customer problems and building trust

* AI Skepticism and Hype

* Some skepticism is healthy, but "AI blame" also emerges

* High expectations from media/industry creators

* Important to stay grounded in real customer needs

* Meaningful AI Applications

* Many examples of AI positively impacting lives already

* Engineers have power to build and explore - lots of opportunity

* Closing and AI Engineer Summit Details

* October 8-10 virtual conference for AI engineers

* Speakers from OpenAI, Microsoft, Amazon, etc

* Free to attend online



Get full access to Latent.Space at www.latent.space/subscribe

More from Latent Space: The AI Engineer Podcast

All 247 episodes
[AIE Summit Preview #1] Swyx on Software 3.0 and the Rise of the AI EngineerLatent Space: The AI Engineer Podcast · 39 min
Listen in VO