OpenAI researcher on why soft skills are the future of work | Karina Nguyen (Research at OpenAI, ex-Anthropic)

9 Feb 2025 · 1 h 15 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

```markdown

Lenny's Podcast

Product | Growth | Career

Episode Summary

Title

OpenAI Researcher on Why Soft Skills are the Future of Work | Karina Nguyen

Guest: Karina Nguyen, Researcher at OpenAI, ex-Anthropic Host: Lenny Rachitsky

In this episode, Karina Nguyen shares her insights from leading research at OpenAI and her previous role at Anthropic. She discusses the evolving landscape of AI, the importance of soft skills in the future, and shares her thoughts on product development and AI model training.

---

Key Topics Discussed

Introduction to Karina Nguyen

  • Background: Karina's experience in AI research and engineering roles at OpenAI, Anthropic, New York Times, Dropbox, and Square.
  • Current Role: Leading research at OpenAI, contributing to projects like Canvas, Tasks, and the o1 language model.

Challenges in AI Model Training

  • Common Misunderstandings: Model training is more of an art than a science, focusing on data quality and debugging similar to software development.
  • Synthetic Data: The importance of synthetic data in model training, allowing models to learn diverse tasks without hitting data walls.

Product Development at OpenAI

  • Canvas and Tasks: The development process for these products, emphasizing rapid model iteration and user feedback.
  • Collaboration: The integration of researchers, engineers, and designers in product development to ensure robust feature creation.

Future of Work and AI

  • Impact on Jobs: How AI is reshaping roles, with a focus on the declining value of hard skills and the growing importance of soft skills.
  • Role of Creativity: AI's limitations in creativity and aesthetic judgment, suggesting a continued need for human input in these areas.

Building Trust with AI

  • User Interaction: Trust between users and AI through consistent, personalized experiences.
  • Synthetic Data and Evaluations: Using synthetic data to train models for specific behaviors and the role of evaluations in ensuring model accuracy.

Differences Between OpenAI and Anthropic

  • Cultural and Operational Differences: How both companies approach AI development, with Anthropic emphasizing model behavior and OpenAI focusing on creative freedom and risk-taking.

---

Practical Takeaways

  • AI's Role in Future Work: Soft skills like creativity, empathy, and management are becoming more valuable as AI takes over more technical tasks.
  • Importance of Evaluations: Writing evaluations will become a critical part of product development in AI-driven projects.
  • Rising Role of Small Models: Intelligence is becoming cheaper, with small models getting smarter, indicating broader access to AI capabilities.

---

Future Implications

  • AI Agents: The potential of AI agents operating in virtual environments to perform tasks autonomously.
  • Synthetic Data: Continued development in synthetic data will facilitate more robust and diverse model training.
  • New Interactions: Exploring new paradigms beyond chat interfaces for AI interactions, such as collaborative and predictive models.

---

Connect with Karina Nguyen

  • Twitter: [@karinanguyen_](https://x.com/karinanguyen_)
  • LinkedIn: [linkedin.com/in/karinanguyen28](https://www.linkedin.com/in/karinanguyen28)
  • Website: [karinanguyen.com](https://karinanguyen.com/)

Connect with Lenny

  • Newsletter: [Lenny's Newsletter](https://www.lennysnewsletter.com)
  • Twitter: [@lennysan](https://twitter.com/lennysan)
  • LinkedIn: [linkedin.com/in/lennyrachitsky](https://www.linkedin.com/in/lennyrachitsky/)

---

Sponsors

  • Enterpret: Transform customer feedback into product growth.
  • Vanta: Automate compliance, simplify security.
  • Loom: Screen recorder for video messages.

---

Additional Resources

  • Transcript: [Podcast Episode Transcript](https://www.lennysnewsletter.com/p/why-soft-skills-are-the-future-of-work-karina-nguyen)
  • Referenced Articles and Tools:
  • [Synthetic Data Guide](https://mitsloan.mit.edu/ideas-made-to-matter/what-synthetic-data-and-how-can-it-help-you-competitively)
  • [Canvas Introduction](https://openai.com/index/introducing-canvas/)
  • [Anthropic Context Windows](https://www.anthropic.com/news/100k-context-windows)

---

Produced and marketed by Penname. For sponsorship inquiries, email podcast@lennyrachitsky.com. ```

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Not only are you working at the cutting edge of AI and LLM, you're actually building the cutting edge. When I first came to Andar right, I was like, oh, that I really love from the engineering. And then the reason why I switched to research is because I realized, oh my god, cloud is getting better at the end. Cloud is getting better at like coding. I think cloud can like develop new apps. What skills do you think will be most valuable going forward for product teams in particular? Theater, thinking. And you kind of want to like generate a bunch of ideas and like filter through them. And not just load the best product experience.

0:33I think it's actually really, really hard to teach the model how to be aesthetic or really good visual design. All I thought to be extremely creative in the way they write. What do you think people most misunderstand about how models are created? When you taught the model some of the self -knowledge, you actually don't have a physical body to operate in the physical world. The model would get like extremely confused.

0:57Today, my guest is Karina Nguyen. Karina is an AI researcher at OpenAI, where she helped build canvas tasks, the O1 chain of thought model and more. Prior to OpenAI, she was at Anthropic, where she let work on post -training and evaluation for the cloud three models, built a document upload feature with 100K context windows and so much more. She was also an engineer at New York Times, was a designer at Dropbox and at Square. It's very rare to get a glimpse into how someone working on the bleeding edge of AI and LLMs operates and how they think about where things are heading. In our conversation, we talk about how teams at OpenAI operate and build products.

1:35What skills she thinks you should be building as AI gets smarter, how models are created, why synthetic data will allow models to keep getting smarter and why she moved from engineering to research after realizing how good LLMs are gonna be at coding. If you enjoy this podcast, don't forget to subscribe and follow to your favorite podcasting app or YouTube. It's the best way to avoid missing feature episodes and it helps the podcast tremendously. With that, I bring you Karina Nguyen. This episode is brought to you by Interpret. Interpret unifies all your customer interactions from gong calls to zen des tickets, to Twitter threads, to App Store reviews and makes it available for analysis.

2:13It's trusted by leading product orgs like Canva, Notion, Lume, Linear, Monday .com and Strava to bring the voice of the customer into the product development process, helping you build best in class products faster. What makes Interpret special is its ability to build and update customer specific AI models that provide the most granular and accurate insights into your business. Connect customer insights to revenue and operational data in your CRM or data warehouse to map the business impact of each customer need and prioritize confidently and empower your entire team to easily take action on use cases like wind loss analysis, critical bug detection and identifying drivers of churn with Interpret's AI assistant wisdom.

2:53Looking to automate your feedback loops and prioritize your roadmap with confidence like Notion, Canva and Linear, visit ENTER, P -R -E -T, dot com slash Lenny to connect with the team and to get two free months when you sign up for an annual plan. This is a limited time offer that's Interpret .com slash Lenny. This episode is brought to you by Vanta and I am very excited to have Christina Casiopo, CEO and co -founder Vanta joining me for this very short conversation. Great to be here, big fan of the podcast and the newsletter. Vanta is a long time sponsor of the show but for some of our newer listeners, what is Vanta do and who is it for?

3:32Sure, so we started Vanta in 2018, focused on founders, helping them start to build out their security programs and get credit for all of that hard security work with compliance certifications like SAC -2 or ISO -2701. Today, we currently help over 9 ,000 companies, including some startup household names like Atlassian, Ramp and Lengchain, start and scale their security programs and ultimately build trust by automating compliance, centralizing GRC and accelerating security reviews. That is awesome. I never have experienced that these things take a lot of time and a lot of resources and nobody wants to spend time doing this.

4:10That is very much our experience but before the company, in some extent, during it, but the idea is with automation with AI with software, we are helping customers build trust with prospects and customers in an efficient way. And you know, our joke, we started this compliance company so you don't have to. We appreciate you for doing that and you have a special discount for listeners. They can get $1 ,000 off Vanta at Vanta .com slash Lenny. That's V -A -N -T -A .com slash Lenny for $1 ,000 off Vanta. Thanks for that, Christina. Thank you.

4:45Karina, thank you so much for being here. Welcome to the podcast. Thank you so much, Lenny, for enlightening me. I'm very excited to have you here because not only are you working at the cutting edge of AI and LLM's, you're actually building the cutting edge of AI and LLM's. You recently launched this feature which basically, the first agent feature of OpenAI. I also just did this survey. I don't know if you know about this. I did a survey of my readers and asked them what tools to use everyday in your work and most use. And Chatchy BT was number one above Gmail, above Slack, above anything else.

5:1990 % of people said that use Chatchy BT regularly. It's absurd. It wasn't around two years ago. Yeah. Also, we're recording this the week that OpenAI announced Stargate, which is this half trillion dollar investment in AI infrastructure. So it's just like a lot happening, constantly in AI and you have a really unique glimpse into how things are working, where things are going, how work gets done. So I have a lot of questions for you. I want to talk about how you operate and how you work at OpenAI, where you think things are going, what skills are going to matter more and less in the future, and also just where things are going broadly.

5:53So how does that sound? Sounds great. Thank you so much. Yeah. I was extremely lucky to join early days on topic and kind of learned a lot of things there and I joined OpenAI around like eight months ago. So yeah, I'm excited to do that more. Okay. I'm going to definitely ask you about the differences between those, but I want to start more technical and just dive right in. I want to talk about model training. People always hear about models being trained. These big models, how much data takes, how long it takes, how much money it takes, how we're running out of data, which I want to talk about.

6:30Let me just ask you this question. What do you think people most misunderstand about how models are created? Model training is more an art than a science and in a lot of ways like we as model trainers, think a lot about data quality. It's one of the most important things in model training is how do you ensure the highest quality data for certain interaction model behavior that you want to create. But the way you debug models is actually very similar, the way you debug software. So one of the things that I've learned early days at Antoport was like we've discovered, especially with like cloud -to -e training when you taught the model some of the soft knowledge of like, hey, like you actually don't have a physical body to operate like in the physical world.

7:21But then at the same time we had data that kind of taught the model some of the function calls, which is like, this is how you start the alarm. As a model would get extremely confused about whether it can set an alarm, but it doesn't have a body in the physical world. So it's like the model gets confused and sometimes it's like over a few years. So sometimes it says, I don't know, like, sorry, I cannot help you. As there is always a balanced trader off between how do you make the model to be more helpful for users, but also not being harmful in other scenarios. So it's always about like, how do you make the model more robust and like operate across like variety of diverse scenarios?

8:09That is so funny and I never thought about that. Most of the data that it's trained on is kind of like assuming it's like a human describing the world and how they operate and there's, it assumes there's a body and you could do things and the model told you don't have a body. Yeah. Okay. I wanna talk a little bit about data while we're on this topic, I know you have strong opinions here. There's kind of this mean that models are gonna stop getting smarter because they're running out of data. They're trained in a large part on the internet and there's only one internet and they've already been trained on it.

8:37What more can you show them about the world? And there's this trend of synthetic data, this term synthetic data. What is synthetic data? Why do you think this important? Do you think it's gonna work? I think there are two questions here. We can unpack on one at a time, but people say if you're hitting the data wall, I think people think more in the terms of like pre -trained large models that are trained on the entire internet to predict the next talking. But what actually the model is learning during that process is actually how do you compress the compression algorithm here? The model learns to compress a lot of knowledge and it learns how to model the world.

9:21So the next prediction of the world like teach me how to drive basically and you only have like, if you words that will match that a car. So the model actually learns about the world in itself. So it's like it's modeling human behavior, sometimes it's modeling. And when you talk to like pre -trained models, which are very, very large, they're actually extremely diverse and extremely creative because you can talk to almost any Reddit user through a pitching model. But I think what's happening right now was like new paradigm of like, oh, one series is of like the scaling in post -training itself is not hitting the wall.

10:08And that's because basically you went from like raw data sets from pre -trained models to infinite amount of tasks that you can teach the model in the post -training world via reinforcement learning. So any task, for example, like how to search the web, how to use the computer, how to write, wow, like all sorts of tasks that you like trying to teach the model, all the different skills. And that's why we say like there's no data wall or whatever because there will be infinite amount of tasks. And that's how the model becomes extremely super diligent. And we are actually getting saturated in all benchmarks.

10:52So I think the bottleneck is actually in evaluations that we don't have all the frontier like EVA's like, I don't know, GPTA, which is like a Google proof fashion answering web PhD level. And the better work is like getting to like, I don't know, more than like 60, 70%, which is what HD gets. So it's like literally hitting the wall and like eval's. I'm gonna follow both those threads. So the first is on this idea of synthetic data is a simple way to understand it that the models are generating the data that future models are trained on. And you ask it to generate all these ways of doing stuff, all these tasks as you described and then the newer models trained on this data that the previous model generated.

11:39Sometimes synthetic like you read it. So this is like an active research area is like, how do you can, synthetic can start like new tasks model to like learn? Sometimes, you know, like when you develop products, you get a lot of like data from the product and like use if you thought and you can use that data to like this like questioning world. Sometimes you still want to like use like human, and human data because actually some of the tasks can be like really, really hard to achieve. Like experts like only know like set of knowledge about like some chemicals like biological knowledge. So like you actually need to tap into the expert knowledge a lot.

12:26So yeah, I think to me like synthetic data training is more for like product is like a rapid model iteration for similar product outcomes. And we can dive more into it, but the way we made canvas and tasks and like new like part of features, feature, feature, feature, was mostly done by synthetic training. It was actually get into that. That's really interesting. I want to talk about e -vals, but let's follow that thread. So talking about how this helped you create canvas. So when that first came to OpenAI, I really had this idea of like, okay, like it would be really cool for chat to be to actually like change the visual interface, but also change like the way it is with people.

13:13So going from like being a chatbot to more collaborative agent and the collaborator, it is like a step towards like more gentle existence that become like innovators ultimately. And so the entire team of like applied engineers, designers, products like research kind of like got like formed in the air, almost out of like nothing, it's just like a collection of people who just like got together and the rapidly centered each way to do each other. Actually like Kevin's was like one of the, I would say like the first project of OpenAI where researchers and applied in here started working together from the very beginning of the product development cycle.

14:00And I think like there's a lot of things that we have learned on the way, but I definitely came to with the mindset of like, we need to do like really rapid model situation, such that like it would be much easier for engineers to, you know, work with the latest model possible, but also learn from like use a few back or like early like internal dog food, how to be improved model very rapidly. And you know, it's really hard to like, kind of like figure out like how people, when you deploy a product, how people would be able to like use it. And so like the way you use them directly, trained the model is basically figuring out like, what are the most core behaviors that you wanted this product feature to do?

14:50And for Canvas, for example, it was, it came down to like three main behaviors. It was, how did you trigger Canvas for prompts like write me along assay when the user intention is mostly like, iterating over long documents or write me a piece of code or when to not trigger Canvas for prompts like, can you tell me more about president like, I don't know, some of the general questions, so you don't want to let trigger Canvas because the user intention is mostly getting answer and not necessarily like iterates over the long document. The second behavior is, how do you teach the model to update the document when the user asks?

15:37So one of the behaviors that the, taught the model is actually have like the, some agency on autonomy to literally go to the document and like select specific sections and either deleted or edit. So highlight it and rewrite certain sections. So sometimes the model, sometimes the user would just like say, change the second paragraph to be something friendlier and you would have to like each the model to literally find the second paragraph in the document and change it to a friendly tool. So basically you teach both like how to trigger like, edit itself, but also how do you teach the model to get higher quality edit for that document.

16:21In case of like coding, for example, there's also like the question of like, how good the model is of like completely rewriting the document versus like having like very specific targeted edits. So that's like another like layer of decision boundary within like edit itself. Is like select the entire document and like rewrite completely or you want to like have like very targeted custom behavior. And you know, like when you first launched the model, we would bias the model towards like more rewrite because we thought the quality of the rewrite were like much higher, but over time you like kind of shift in based on like, usually that and was the learning from each other's deployment.

17:02Lastly, the third behavior that we taught and that actually the model is hard to make comments on any document. So the way we use that is like, we would use a one model to produce, to like simulate like user conversation, let's say like write me a document about x ,y, but then we used a one to like produce the document and then we kind of injected like user prompt to be like, oh, make some comments critique my piece of writing or critique the space of writing that you just made. And then we taught the model to like make comments on the document on like very specific document. So it's like also like what kind of comments you want the model to make?

17:49Like do they make sense or not? Like how do you teach the quality of that? And it all came down to like measuring progress via very robust evolves. But yeah, this is how you would use like a one like kind of synthetic solution for like the screening. Okay, this is so interesting. So you talk about this idea of teaching the model and you mention how it's using synthetic data to teach the model different behaviors. Is a simple way to think about it. Basically that's where you do that by showing it what success looks like using basically evalzes. Is that the simple way to think about it? Like here's what you doing the successfully would look like and that teaches it, okay, I see this is what I should do.

18:30Yeah, great. Yeah, amazing. Yeah, you got it. Got it. I want to start unpacking what your day -to -day looks like as you're building these sort of things. Is it like you sitting there talking to some version of chat GPT crafting these evils? Sometimes I do that, but sometimes I do sit for a time to fit. I think I learned this so much from Andabit. It's like people spend so much time just like prompting models and like quality the way we buy a bunch all the time. And you actually get at a lot of new ideas. How do you make the model better? It's like, oh, like this response is kind of weird. Like why is it doing this?

19:09And you start like debugging or something or like you start like thinking out like new methods or like how do you teach the model to respond in a different way and like have better personality, let's say. So it's the same thing of like how personality is made like in the models with those like very similar methods. But yes, I think my time out of there I have changed. I think when the first came I was like mostly like research I see work. So I was like building a lot of like, I was like writing code, like you know, changing models, writing evals, working with PMs and like designers to like learn teach them how to like even think about like invitations.

19:50I think it was like really cool experience. And I think it was like an adoption of like how do you like do this like prior management of like AI features or like AI models. Yeah, but now it's like mostly like, you know, like management. I'm like mentorship. I'm still like doing SC like research code after like 4 PM, although, but yeah, I just kind of like changed. All right, don't talk too much about being a manager because everyone's firing their managers. Who needs managers anymore? That's the what I hear now just kidding. It's interesting that so much of your time was spent on teaching product teams how evals integrate and how important it is.

20:36And I've heard this a few times and I haven't personally experienced it yet. So I think it's an important threat to follow is just how writing these evaluations is going to become increasingly an important part of the job of product teams, especially when they're building AI features and working with them. So can you just talk a bit more about what that looks like? Is it like sitting there with an Excel spreadsheet? Basically showing like here's the input, here's the output, here's how good the result was. Let's talk about what that actually looks like very practically. It suddenly depends on what you're developing, but there are various types of like evaluations.

21:08So sometimes I do ask product managers or there's also like new role that we have like model designers to kind of like go through some of the users feedback maybe or like think of like various like user conversations that should have triggered like under this time of senses it should trigger canvas. And then you have this like ground truthly ball of like, okay, with this conversation it should look to grab canvas under this conversation it should not trigger canvas. And you have this like very binary, the tremendous that kind of like you all that for like this is about it, behaviors is like this.

21:46When we were launching tasks for example, like how do you make correct schedules? It's like actually really hard for the model, but if you both out like some of the deterministic evaluations, then it's like, okay, like if the user says like 7 p .m. it should like demoral should say 7 p .m. So if you can like have a deterministic evolves with it's like paths or fail. So yeah, and like the way it works is like, sometimes I ask product managers just like go create like the goals you'd like, have different tabs and like what's the current behavior, what's like the ideal behavior and like why or like some nodes.

22:27And sometimes we usually use it for evolve. Sometimes we use it for training because like if you give this structure to like a one model, it can probably figure out like how to teach itself a good behavior. And I think there are certain type of like evals that is kind of more prevalent is like human evaluations. And you can have specific trainers or you can have like internal people to when you have like a conversation of the prompt and then you have like various completion of models, you can't choose the win rate, which model is the best, which model produced the highest quality comment or edit.

23:10And then you can have like continuous win rates. And as you develop new models, it should always like win over the previous models. So it depends on what you want to measure. So interesting like basically what I'm hearing and there's something I'm learning about as I talk to people is product development might move from this like here's a spec PRD, let's build it together and then cool, let's review it are we happy with this too. From that to hey AI build this thing for me. And here's what correct looks like. And I'm spending all my time on what is correct look like any valves essentially. You definitely want to like measure progress to your model and this is where evals is because like you can have prompted model as a baseline already.

23:57And if the most robust evolved is the one where prompted baselines get the lowest score or something. And then because then you know like, oh, if you trained a good model, then it should like just like hill climb and that evil of the time, well not like also like regressing on like other intelligence evals. So it's like I think it's more what that's that's what I'm saying like it's more than art than science is like, okay, like if you optimize the model for this behavior, like you kind of don't want to like brain damage in like other areas of intelligence or this is happening like all the time in every lab in every like research team.

24:34I would say like pampering is like also a way to like prototype like new part of good years. Like early days at Andabra Gleene was working like file uploads feature. I remember just like you know prompting the model to just like, I mean when we were like launching like a hundred key contacts, I was just like prototyping this in a local browser which I did the demo like people really, really loved it. And they just like wanted like API for like file uploads or something. And then that's when it clicked to me like, I also like run the blog post on tabular like it clicked to me like pampering is like a new way of like product development or prototyping for designers and for like product measures.

25:20For example, one of the features that I want to do is like have like personalized recommend personalized starter prompts. So whenever you come to like cloud, like it should like recommend you like starter prompts based on what your interests are. And so like you can literally do it like prompting for that. And another feature was like generating titles for the conversations. It's a very small like micro -experience, but I'm really proud of the way you did that was because we took like five latest conversation from the user like us the model like was the style of the user. And then like for the next kind of new conversation the generated title will be of the same like style which is like really little like micro -experiences like those.

26:11That's so cool. Did you do that at the topic or at OpenAI? I don't talk like. Okay, cool. I love the file upload feature that cloud has by the way. Oh, Chatchy PT doesn't have that yet, is that right? I think it has. I think like the way it's implemented is like very different. Okay, maybe it's the PDF feature because I use it all the time with the cloud. Okay, so it needs to get on that. Man, it's while how many features you built that I use every day and then many people use every day. This prototyping point you made is really important. It's something that comes up a ton on this podcast also of how that is maybe the way that AI has most impacted the job of product builders.

26:44Recently it's just prototyping. Instead of going from showing just like here's a PRD, here's a design, PM's more and more, just here's the part of the idea that I have and it's working and you can play with it. Yeah. Yeah. Okay, I want to spend a little more time on how you operate. So you talked about you built this in launch of this tasks. Features, is that the way you just grab your tasks? Yeah. So talk about how that emerged and let's better understand just how you collaborate with product teams and how open AI works in that way. Whatever you can share there. I think Canvas and tasks are going into the bucket of objects where it's more like short or medium terms.

Read the full transcript

27:23And actually the way Canvas and tasks came about to be was like, it started with like one person prototyping. And creating like a spec, it's kind of like PRD, is like creating a spec of like the behavior of the model. I don't think like tasks is like extremely like ground -breaking ground -breaking feature in necessarily. What makes it like really cool is because the models are so general, model can now search, they can like write sci -fi stories, they can like search for stocks, they can like summarize the news every day because the models are so general like giving something familiar to people that like, you know, notifications is like very familiar.

28:11Like having reminders is like very familiar. So like feeling like a form factor for the people who like very familiar, Canvas like Canvas, very just like a bulldoffs is very familiar. But then you add like mychical AI moment and it becomes like very powerful. By the way, it comes usually like operationally like, yeah, size is like a prototype, like literally prompted prototype or like how you would want to like the model to behave. For like tasks, for example, like you kind of like need to design, a little bit like design, design systems design that you need like, okay, like well, if this, if the user says like, remind me to go to lunch like at 8 a .m.

28:51tomorrow. Okay, what kind of information does the model needs to extract from that prompt in order to create a reminder? And so this is how you like design like a stack for a new feature, like a tool, Canvas and tasks out all tools. So it's like, how do you like create the tool stack? And then it's like mostly like like, developing JSON schema. I was like, okay, like from this prompt, maybe the model should extract like the time that the user requested. And then you're thinking about like, what, which one right do you want the time to be? And then like, how do you want the model to like notify you?

29:30Is like basically, if the user should give instruction to the model, and then this instruction would like fire off like every day or something at the particular time. So for example, if you say like search like every day, I want to like learn, no, but the latest AI news, the models should rewrite into like, okay, like search for the latest AI news. And this will, this task will get fired at that particular time that the user requested. And then you know, your design is like tool stack. And then actually, I don't know, like if you like sometimes like, it's like through conversations, I like, I don't like people ask me to like join the, you know, like team and they're like, oh my god, we need to be searchers.

30:20We need like some support. Like we need to train the models. Also, there's like, what kind of this is like mostly like, I just pitched the idea of like, it got staffed quite immediately during the break. So, you know, like a, like dependent project. And usually it was staffing is like mostly like a product manager, model designer, actual product designer, a couple of researchers don't like by children like applied engineers, the puzzle, the complexity of projects. And then like, you know, it took for tasks, it took like, like two months or so to go from a year to one basically. A lot. For canvas, this was like four or five months, I guess, to go from zero to one.

31:05But yeah, I'm like, you know, you teach product managers how to like both evolves. And like maybe, you know, how do we not only like ship the better feature, but how do we think like, well, good term like, what kind of like cool features did you want tasks to have? Like I think it would be nice for tasks to be like, a little bit more personalized. It would be nice to have like, to create tasks via voice and on a mobile, right? Like, so you kind of need to like, this is how you get like research route now right here is like, thinking like how the feature will be developed in the future. And then from there, let's like, you like start creating data sets like with eWAS, you want to make sure that goes well.

31:49And then like, you need to have like a trade off between like what methods you want to use. And the reason why I really love like synthetic, like relying on synthetic data instead of like collecting data from humans is because it's like much more scalable. It's cheap. Let's have like, you literally sample from the model. And you teach the core behaviors of the model that will generalize on to all sorts of diverse coverage. And when you launch the better feature, you learn some my family users that you can like all your synthetic sets can be can be shifted into distribution of how the users behave and on the part behavior.

32:27And this is how you improve. And this is what happens to when you are from better to J. This episode is brought to you by LUM. LUM lets you record your screen, your camera, and your voice to share video messages easily. Record a LUM and send it out with just a link to gather feedback at context or share an update. So now you can delete that novel length email that you are writing. Instead, you can record your screen and share your message faster. LUM can help you have fewer meetings and make the meetings that you do have much more productive. Meeting start with everyone on the same page and and early.

33:05Problem solved time safe. We know that everyone isn't a one take wonder when it comes to recording videos. So LUM comes with easy editing and AI features to help you record once and get back to the work that counts. Save time, align your team, stay connected, and get more done with LUM. Now part of it last year, the makers of JIRA. Try a LUM for free today at LUM .com slash LENI. That's L -O -O -M dot com slash LENI. Something that I want to help people understand, and I don't even 100 % understand this, is what's the simplest way to understand the job of a researcher versus say a model designer and other folks involved.

33:44Like what's the simplest way to understand what researchers do at OpenA? The project that I discovered mostly like product oriented, research is mostly product research. Another part component of my team is actually more like longer term exploratory projects. And it's more about developing new methods, understanding those methods, and a variety of circumstances. So basically developing methods, you kind of like need to follow very similar kind of like recipe of like building e -balls. But it's like most of us thinking it evolves. You kind of want to have like other distribution, or like if you want to like measure generalization, if I don't need to like after thoughts.

34:26But it's basically more science in a way where you know, it's the talk about synthetic data like one of the hardest things about synthetic data is like how do you make it like more diverse diversity and synthetic data is like one of the most important questions. Right now, and it's like exploring like ways to in check like diversity as a general method that will work for all is like a one of the black research operations. Other ones is like more like the bottom making new capabilities. I feel like it's always about like, you know, like you you work on this like new method. And you have like signs of life that it's working.

35:03I think you think of like how do you make it more general or you think of like how do you make it very useful or like. And this is how like longer term projects become more like media and like time project. That makes sense. Essentially working on developing ways to make the models smarter. Or oh, four or five or six ways to like a one was a big breakthrough right the way it operates where it's not just here's your answer. It actually thinks and has right takes time to think through the process of coming up with an answer. Okay. Yeah. Very helpful. Speaking of that, I was thinking about the future where things are going.

35:38I want to spend some time on just this insight that basically you are building the cutting edge of AI like at the very bleeding edge of where AI is going and where it is. And so I'm very curious to hear just your take on how you think things are going to change in the world and how people work based on where you see things are going. And I know it's a broad question, but let's say like in the next three years, how do you see the world changing how do you see people's way of working changing. It's a very humbling experience to be in both labs. I guess like to me when they first came to and I was like, I will love from the engineering.

36:16And then like the reason why I switched to like research is because I realized at that time is like, oh my god, like cloud is getting better like for and ends. Like cloud is getting better like coding. I think cloud can like develop new apps or something. And so like it can like develop new features for the thing that I'm working. So it's like, it was kind of like this meta realization where it's like, oh my god, like they were all does actually changing and they're like, when we first like launched a hundred key context at that time. Obviously, you know, I'm thinking about like from factors that's like, yeah, like file uploads were like very natural way familiar to people.

36:55But you can imagine because just like make like infinite chats in the cloud that AI up, right. Like as if like it's like a honeykey context. But because like file uploads is like foreign follows function is like the form factor of the file uploads kind of enable people to just like literally upload anything the books of like any reports financial and like ask any task to the moral. And I remember it was like, you know, enterprise customers like like financial customers are like really interested in that is like, oh wow, like actually be. It's actually one of the very common tasks that people do in that setting is like kind of crazy to like see how some of the redundant tasks are getting like automated basically by the select smart models.

37:47And they're entering them the error where I actually don't know for example, sometimes like is a one gives me the correct answer or not because I'm not an expert in that field. And it's like, I don't even know how to verify the outputs of the models is because like on my experts, not like they can like verify this. So yes, so basically there are trends that are going on the first trend is the cost of reasoning and intelligence is drastically going down. I had a blog post about this. Maybe I should update them like latest benchmarks because at that time like MMO everybody was like doing like. So like one benchmark interview like who could saturated the benchmarks like now we need to like do the same flop but was with another like frontier evolve.

38:40But the cost of intelligence is like going down because it's becomes not much cheaper smart small models are becoming even smarter than like large models and that's because of like the distillation research. This happened with a clotty high school. I was like working on like pursuing a lot of high school and I realized it was much smarter than like clot tool which was like way you know bigger let's not add. But like the power of like small models become very intelligent and fast L cheap. We are moving towards that world that has multiple implications but the news that like people who have more access to AI and that's really good.

39:25But like builders and developers will have much better access to AI. But also it means like all diverse that has been like bottleneck by intelligence will be kind of unblocked. So anyone like I'm thinking about a healthcare right like if I have instead of going to doctor I can like ask Chai GPT or give chat GPT a list of symptoms and ask me like which like would I have like a cold flu or like something else like I can literally get the access to like doctor almost and there's been some like research studies around that. Yeah, there's a New York time story about that where they compared doctors to doctors using Chai GPT to just chat GPT and just just chat GPT was the best.

40:15Yeah, like like doctors made it worse. Yeah, yeah, that's crazy. I like like education. I think I will have friends if like I had the tool like Chai GPT and when I was like young and like would learn something like that. But it's like people can now learn almost anything from this model so they can learn new language. They can learn how to build new look ups like I write anything that you want and like I'm so like it's humbling to like have like launch canvas and like bring that thing to the people enable them to do something else that they couldn't have ever before. I think this is something like magical around this experience.

40:57So education has will have massive implications like I guess like scientific research right like I think it's like the dream of like any I research is like I've made a research. It's kind of scary I'd say which makes me think that like people management will fade you know it's like one of the hardest things to it's like emotional intelligence for the model also like create a creativity in itself is like one of the hardest things. So writers I don't think like people should be worried as much I think it's like I think it's a lot of like redundant tasks for people. This is awesome. Okay. I want to fall this start for sure and it's funny that the way you described is like you're an engineer and then drop it.

41:40And you're like okay. Claude is going to be very good at engineering. This isn't going to be a potentially career long term. So I'm going to move into research and may I is going to need me for a long time to build it to make it smarter. I would say we still have like I think canvas you have still have like a really cool like front engineers that I really like you know people who like really have all like interaction design like interactions. There's like I don't think like models are there yet like I think if but we can get the models to like this top 1 % of like front end something for sure.

42:15So I want to move on to next along these lines is just and this is just speculation but what skills do you think will be most valuable going forward for product teams in particular. So folks are listening and they're like okay this is scary. What should I be building now to help me stay ahead and not be in trouble down the road. What skills do you think are going to be most more and more important to build. Yeah I think like creative thinking like you kind of want to like come up like generate a bunch of ideas and like filter through them and not just like build the best product experience listening.

42:56You know you want to like build something that like the most general model will not replace you and oftentimes you build something and you make it really really good for like specific set of users and actually the mode is now in like your user feedback. The mode is like more and like you whether you listen to them like whether you you can like rapidly it's like the mode is like in here. I don't think like we are yet to like there are so many ideas. I think that abundance of like ideas that you look like about is like I wouldn't be worried. I feel like in fact like I just think like people in AI fields are like I wish they were like a little more creative and like connecting thoughts across like different like feel to something like that to like develop really cool new like generation.

43:50A new part times of interactions with the AI like I don't think we've cracked this problem at all. A couple of years ago I was like telling some people I was like you know you kind of want to like build for the future so it's like it doesn't necessarily matter whether the model is good or not good right now. But you can build product ideas such that like by the time the models will be really good it will work really well. I think it just like happened not really like for example like I don't know like right like the cloth artifacts and I feel like early days of canvas was like back in like 2022 like before Chichibit.

44:33Like writing ID was like I know it's just but I feel like cloth 1 .3 model itself was like not there to like made like really extreme good like high quality edits for example like coding. And I feel like I see like start up as like car sir and it's like doing super well like I must because you like iterate so fast they like events like new ways of like training models they move really fast they listen to like user is like massive distribution is like yeah it's kind of. That's really helpful actually so what I'm hearing is that soft skills essentially are going to be more and more important powerful to talk about management leading people being creative and coming up with innovative insights listening.

45:20There's a post I wrote that I'll link to right look I try to analyze what AI how AI will impact product management and we're actually very lined. And my sense was the same thing that soft skills are going to become more and more important and the things that are going to be replaced is the hard skills which is interesting because usually people value the hard skills like coding design writing really well and it's interesting that AI is actually really good at that because it's taking a bunch of data synthesizing it and writing creating a thing versus all these fuzzy things around of what influences convinces people to do things in a lining and listening like you said creativity.

45:58Anything along those lines come up as I say that. I think it's actually really really hard to just the model how to be aesthetic or like do like visual really good like visual design or like how to be extremely creative in the way they're right. I think like I still think like chagype kind of sucks of like writing and that's because it's like it's like bottlenecks by this like creative reasoning. I think like prioritization is like one of the most important like I think like for a manager I feel like I actually like AI research progress is bottlenecks by like management like research management is because you have like.

46:35The question set of compute and you need to like allocate the compute to the research path that you feel the most. Commits about was like you need to like really you need to have like a really high conviction and the research that to put the compute and like it's more like return on investment kind of situation is like okay yeah like I'm thinking a lot about like like like okay like how do I across all my projects which projects a higher priority is like prioritization and also like on the lower level is like which experiments are really important to run right now and which are not in like cut through the line so I see like prioritization.

47:12Priorization communication like management people's kids like empathy like understanding people like I know like collaboration like I think like canvas wouldn't be like an amazing launch if these wasn't like about like people and I think it's the wonderful group of people and like. I got a chance to like rub with like people like Lee Byron who's like a cook for you know like GraphQL and like some of the best like Apple designers and it's like so cool to like see and like how do you create this like collaboration between people is just like something that's still human I think. Let me just follow us are a little bit because I imagine people listening are like okay but once we have a GI or SGI it'll do all this it's you know it's like there's a world where like why isn't all this done.

48:01I think it's easy to just assume all that I'm curious this idea of creativity and listening why you think AI isn't good at it other than it's just very hard to train it to do this well is there anything there just like why this is especially difficult for AI now lambs to get good at. I think currently it's difficult. For many reason I think it's still like an act of like research area and there's something that like I think my team is like working on is like okay I like how do you teach the model to be like more creative and like the writing and actually like. I think he like doesn't your paradigm of life the models.

48:40I think more should actually lead to like better writing and itself but like when it comes down to like idea generation or like. Discommunicating of like what is the good like visual design and not I feel like it hasn't had learned like examples from like people to discommunicated very well I do think it's because like. You know there are not many people who are like actually like really like it's not like accessible to like model so learn from these people like us. So I guess like that's why it's often yeah make sense basically there's not enough of you yet researchers. Teaching it to do these things slash people that have incredible taste and creativity that can teach these things you could argue this will come but I'm not we don't need to keep going down that thread.

49:35Let me ask you a specific question in this poster road I made this argument that a lot of people disagreed with that strategy is something that AI tooling will become increasingly great at and take over there's the sense that that's a thing that people will continue to be much better at and you. Can't offer to AI basically developing your strategy telling you what to do to win my cases isn't strategy just take all the inputs all the data available. Understand the world around you and come up with a plan to win feels like I would be like an L I would be incredibly smart at this what's your take I think so to I think like again like we teach the model all sorts of like tools and like capabilities and like reasoning right and it's like when it comes down to like.

50:19I was just a for canvas right now you're very cool to like for the models just like aggregate all the feedback from users like some arrived me like the top five like most painful flow flows like music experiences and then like the model itself is like very capable of like like thinking of like knowing how it's being made. Figure out like how to like treat a data sets for itself to like train on it and I think like me a far away from that kind of like soft improvements models becoming like self improved by a like then like the part of development is basically kind of like self improving like it's kind of like it's all like organisms and something yeah like I gave like strategies like it's more like data analysis and like.

51:09Coming up was like. Like I think what matters are really good is like. Like connecting the dots I think it's like okay if you have users you back from this source but you also have an internal like dashboard with metrics and then you have you know like other kind of like you back. Or like inputs and then like it can co create like a plan for you like look at the nation's event and it is like one of the most common use cases for actually people is like coming up was like this sort of things that makes sense like essentially. A human can only comprehend so much information at once and look at so much data once to synthesize takeaways and as you said these context windows are huge now here's all the information what's the most important thing I should do.

51:59Yeah same was like scientific research is because like you like ideally the model would be both like suggest like ideas with new ideas like it already on the experimental like given the empirical results of the previous experiments like how do you. Like come up with like new ideas or like the methods yeah man. Okay so just to close the loop on this conversation this part of the threat is the skills you're suggesting people focus on building and leaning into soft skills like creativity managing influence collaboration looking for patterns. Is that generally where your mind is that yeah I'm meeting a lot of all like how do you make a relationship more effectively and I think this is more most like management I guess.

52:45It's like how do you organize like research teams like generally teams like combine. Compose teams such that they will be at the maximally succeed like at the maximum like performance of what can possibly. Like you can like literally create like the next generation of computers is just like the matter of conviction and like the way you manage to that is like scaling organizations or like scaling product research as. Yeah I think what it like you're basically building this thing and not efficiently doing it is like limiting the potential of the human species right now is. Right mismanagement within the research team and open and I and then I can some of these other models.

53:32Yeah I've kind of crazy to think about holy moly okay so speaking of a drop making opening out you've worked at both very few people have worked at both companies and seen how they operate. Care is just what you've noticed about the differences between these two how they operate how they think how they approach stuff what can you share along those lines. It's more similar than different obviously there is a lot of like. There are some differences also comes to like nuances or to culture. I really love on top of it and they have a lot of friends there and also love opening out and I still love a lot of friends.

54:05So it's like it's not about like enemy I feel like there's like in the animals all like yeah the competitors that's like enemies that's actually like one big community and like a people like doing the same thing. I would say would have learned from on top of it is this like real care and craft towards like model behavior model cost model training. And I've been thinking a lot about like what makes cloud cloud and what makes strategy and strategy and it's like I just sounds like operational processes that kind of leads to the outputs to the model is the output and model and it's like the reason why collage has so much more personality and like.

54:52It's more like a librarian I don't know what I don't I am like visualizing a cloud being like a librarian like a very like naughty or something. It's because I feel like it's the reflection of the creators who like making this model and like a lot of like details around like the character and the personality and like whether the model should follow up on this question or like not like. Was to correct like ethical behavior for the model to like in the scenarios like a lot of like craft and like truly did like the results and this is where I learned that part of like art I guess I don't know that I was in the darkness like much smaller like when I joined it was like what like 70 people when I last did those times people like obviously the culture changed so much.

55:44I really enjoyed being like early days sort of like lives and like people. Neuriches as a family but like the culture shifted I would say like under I learned from a topic that like they much better at like focusing and like prejudices and like very very hard like very high school practice. I just like to do it like that I think like opening eyes like much more in a way to and much more like risk take hardest in terms of like products or like research actually you know like. I don't really can like your full time job can be just like teaching the model how to be like creator writers and it's like this some luxury in this like research freedom that that comes to scale maybe I don't know.

56:30But it gives you it's like you'll help I feel like I have much more creative like product freedom to do almost anyone I guess within like opening I like of all especially to be into like the region that is like more like yeah probably what is off it is yeah that's how I was thinking about is feels like opening eyes more bottoms up distributed people bubble up ideas try stuff there's more law and that emerged leads to more products launching imagine more things just kind of being tried versus more of a let's just make sure everything we do is awesome great and craft and thinking deeply about every right investment that's really interesting I've never heard it described this way.

57:11Karina we've covered so much ground this is going to help a lot of people with so many ways of thinking about where the future is going before we get to a very exciting lighting around I'm curious if there's anything else they think might be helpful to share get into one of my regrets I guess when I was early days at on top of was that like I think there was like some luxury of the time this pretty China to be key to actually like. Comment was like a bunch of ideas on like prototype was almost every day and I think like we did a lot of cool ideas like cloud and slack was actually one of the first like tool use like products like a cloud could operate in like your workplace now it's like kind of cool but you like at cloud to summarize the threat so maybe you have a good entire conversation with someone and then you want to like a summary.

58:04What happened like you can see like at the class and crisis also was really fun to like even like to read on the model itself is like when you just like talk to the model in like slug forever and if you're just on the social element it's kind of what was kind of a journey and like. This discord like people learned so much about the content and like how do you work with like clouds actually one of the features that was like early tasks part of time is like you know every Monday clock which is like summarized the entire channel or like every Friday we just like summarized like bunch of channels and give like the news about the organization or something so it's and it's going to be a lot of fun to do.

58:48It's kind of like really cool like form factor I think I think about like phone factors like a really important question like in AI especially we haven't even figured out like how to be create like an awesome like products experience was like O series models is like the paradigm between like synchronous real time given answer paradigm into like more synchronous paradigm of like agents working on the background but then now the questions like the agents should be. So you know you're both trust that you're right and trust both over time which is like this humans and you know you start this collaboration which is why like a collaborative like this collaboration model was like you and the models like some important because you both trust and the model are certain preferences so that it can become more personalized and it will start predicting the next like action that you want to take on the computer or something and it's like kind of like more predictive much more.

59:48And from like personal computers like personal model basically here you know that's why is it not a thing that seems like such an obvious feature that every element should have is a slack bought version of them is that is that a thing I can have you installers that not a thing right now. I know that cloud and thought was unsaturated in like 20 or 20 years on thing but I think I think I think it was like after Chitra Pt it was mostly like the focus on like consumer just cases like enterprise cases. I think the one like I think the form fucked up like cloud and slack is like was kind of constrained.

1:00:24A little bit when you want to develop new features. I want that. I know that GIGP have like slack part to I or not like maybe it will come back. All right I would I would pay for that. Any other memories from that time of early days because that's a really special place to have been as early days and the topic any other memories or stories from that time that might be interesting to share. I think the very first launch when they fall like when clips and use again was like a hundred key context launch is like when the models could input the entire book and give you like summer you know the book or something.

1:01:03Of the entire financial or like have like multi files financial reports and then like give you an answer to the question to very specific question. I think there was something in there that kind of like oh my god this is like a really cool new capability not like model capability but more like the capability that came from the product form factor itself rather than like the model capability as much. I think like other project types that you were thinking about like you know like there's like one part of the clockwork spaces and it's like kind of the same like idea like. Claude and I would have to share workspace and that share work space like a document and he like to read in the book and I feel like sometimes the ideas like primary is log and they'll lock for like two years.

1:01:58Just like in this case it's interesting there's a milestones that kind of open up our view of what is happening and where things are going. Chatchy PT I think was the first of just like wow this is much better than I would have thought you talked about one hundred K context windows or get up a little book and ask questions have it summer is actually use that all time when I have interview guests and they wrote a book. I sometimes don't have time to read the whole book so I use it to help me understand what the most interesting parts are and then I actually dive into the book just to be clear. And then I don't know maybe like voice was another one where you could talk to say chatchy PT there are any other moments there that you're like wow this is much better than I thought it was going to be.

1:02:39Yeah I think like computer use agents like the model operating the desktop and you can essentially think of like you know new kind of like experience where the model can learn the way you browse and from that preference it can just like browse as just like you. It's kind of like simulation simulated that person up and it's actually very similar to the idea of like okay like maybe some album doesn't have a lot of like time maybe I want to like talk to like his team leader's like his simulation and ask like or like for example like yeah like I really appreciate some of the technical mentorship like yeah cool like but he doesn't have a lot of time so it's like I really want to like ask him this question is like kind of respond like simulated environments like those will be really cool.

1:03:37It's a great place to plug Lenny but I have one of those it's trained on all of my podcast the newsletters and I it's on many models that I know which one exactly they use but it's exactly that it's and it's not even me it's all the guests that have been on the podcast. I'm just like I wrote and you could just ask it how do I grow my product how do I develop a strategy and it's actually shocking we could do feel like it reflects. Yeah, it's the best part of it is you can talk to it it's built there's an 11 labs voice version that's trained on my voice on the phone spot gas and it's actually very good and people like I've told me they sit there for hours talking to it.

1:04:15Wow and somebody told it interview me like I am on Lenny's podcast ask me questions about my career and he did a half hour podcast episode with Lenny but that's so fun. It's incredible future is wild. Yeah I think content transformation is like you know like I would imagine sometime like you know but you generally decipher a story in canvas like you can like transform this into like audio book. I would go like where I have like very natural like content transformation like one media to another medium. I think like one of my early inspiration is like one of the last episodes of like Westworld where I want to explore but where Dolores comes to her work at the time and she comes to like this like new workspace and she starts like writing a story and then she like virtual reality starts like creating on the fly so I don't want to hate that.

1:05:21Yeah, critical. Wow speaking of medium I guess I was wondering if I should go in the structure but real quick Kevin Wile slash Kevin wheel I don't know exactly that her pronounce was last name the CPU of I think real wheel okay okay let's just say that okay we know he was he did a panel at the Lenny in front summit last year and he made this really fascinating point that Chad is a really interesting interface for these tools because they're just getting smarter and smarter smart and smart smarter and Chad continues to work as a paradigm to just interact with them similar to a human you could talk to Albert Einstein you could talk to someone not very smart and it's all conversation still and so it's a really flexible way to interact with increasingly good intelligence at some point it will not be so great and you're talking about all these ways that you're adding additional ways to interact but it's interesting chat proof to be a really powerful layer on top of all stuff.

1:06:21Yeah, that's really cool. I feel like Chad also has the social element which is like very humane it's like you sometimes want to like get into a group chat and like having conversations with the eyes kind of like a group chat and it's out for some massive thing. Actually this idea of like how do you both like features like this like I see tasks as like this like general kind of like feature that will scale very nicely as the models would develop like new capabilities as well as like like the model will be able to like do better like searches and like you know create new like come up with like more creative like writing on like render you know react apps and like HTML like apps and like you can have like every day and you have a new section that we ended up going down is this idea of your agents using a computer.

1:07:22I know this is actually something you are going to launch today the day recording it which will be out by the time this comes out call operator you talk about this very cool feature that people will have access to. Yeah, so I unfortunately did not work on that but I'm really really excited about like this launch. It's basically an agent that can complete the task in its own like virtual computer like in its own virtual environment. You can do any literally task like ordering me a book on Amazon and then ideally the model will either like follow up with you like which book do you want or like know you so well that it will like start recommending like oh here's the five books that I might recommend to you to buy and then like you hit like yeah help me help me buy and then the model goes off into its own like virtual little browser and like complete the task and by the book and the Amazon and then if you give the model like credentials kind of cards obviously it comes with like a lot of trust and like safety and then it will just complete the thing for you.

1:08:33It's a virtual assistant. It's interesting and this just sounds like obviously this should happen like why is this not a yet a thing which is also mind blowing that we're just assuming this should exist like just some AI doing things for you on a computer. It just has to do like it's absurd. It's actually really hard and I think like you're just still cracking this but I feel like I don't know if he was like tuple it's like a band programming.

1:09:04But I don't know if you love their programming. So if you shop or fight is this a member came up on a podcast episode. Yeah so it's a very cool product where you can just like call anyone at any time and then like share screen and the other person can like have access to this stream and like start like literally operating your computer and it's very like real time like the lead is just like very high quality and it's just like I kind of want the same is like I want to like pair program with like my model and like the model should you want to talk to me like draw like very specific like section in my code and yes go tell me like obviously to me and you can have like different modes.

1:09:48It's like right here it was like a product right here for you. I don't know people sent you most of the old dad. It sounds like a startup just got birthed. Yes from someone listening to this. You mentioned that it's very hard to do this agent controlling a computer as you and helping out what makes it so hard for whatever however much you can explain briefly. My two bit is like because right now the models operating on like pixels instead of like language or whatnot like pixels is actually really really hard for the models because like perception or visual perception. I think there's still a lot of like multiple like research that's going on.

1:10:31But I think like language scaled so much like easier compared to like multi models because of that. Another was like thing that I just like my team is working on is like how do you derive human intent very correctly. It's like sometimes like doesn't model know enough information to ask about question or like to complete the task. You kind of don't want like an agent to like go off for like 10 minutes and then come up with like an answer that you didn't even want that actually creates like much more versus experience. And this is comes with like teaching the model like like people skills is like you know like what do people like like kind of like creating like the my style model of the user and like care about the user in order to ask certain questions like actually that part is like hard to the models.

1:11:27And I'm related to what we talked about earlier this kind of the soft skill people skills pieces not were these models are strong yet. Okay. I'm going to skip the lighting ground. I want to ask just one question from the lighting round something fun. Okay, so when AI replaces your job, Karina, I'm curious what you're going to give you a stipend gives you a monthly stipend. Here's your salary for the month. What would you want to do? What do you want to spend your time on? What will you be doing? In a future world. I think thinking about this. Oh, sometimes I have I really have a lot of jobs options.

1:12:04I would love to be a writer I think I think that would be super cool. You should like write like short stories like sci fi stories. No, I really like art history. So you know, it's like conservationist to like in the museums, which is like try to preserve like art paintings, but just like painting through a lot of things. I think that would be really cool to do. Yeah, that sounds beautiful. I don't know. What I'm hearing is you need to nerf these models to not get very good at writing so that you can continue. Although at that point, you already need to do it from like you don't need people to buy you're just doing it for fun.

1:12:48So it doesn't even matter if they're incredibly good at writing or art or conservation. Oh, man, what an episode or conversation. What a wild time we're living in. Karina, thank you so much for being here. Two final questions. We're confused by new online if they want to reach out and follow up on anything. And how can listeners be useful to you? You can sign me. I'm a Twitter. I'm in. You can also shoot me a email on my website. And my team is hiring. And so like I'm looking for research engineers, research scientists, as well as like machine learning engineers, like people who come from a party.

1:13:25And engineers want to learn about model training, I'm actually hiring for like my team, my team is called the frontier part of research. And the train models be developed new methods, but for part of the oriented outcomes. What a place to work, holy moly. What's the best way for people to apply for these very lucrative roles? I think you can shoot me a DM on Twitter or. I'm yet to create a job description. Okay, this is the job description. Are you going to fly into like post training? Okay, this you're going to get a flood of DMs. I hope you're prepared. Karina, thank you so much for being here.

1:14:02This was incredible. Thank you so much, Lenny. Bye everyone. It was fun. Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite podcast app. Also, please consider giving us a rating or a leaving review as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show at Lenny's Podcast dot com. See you in the next episode.

From the publisher

Karina Nguyen leads research at OpenAI, where she’s been pivotal in developing groundbreaking products like Canvas, Tasks, and the o1 language model. Before OpenAI, Karina was at Anthropic, where she led post-training and evaluation work for Claude 3 models, created a document upload feature with 100,000 context windows, and contributed to numerous other innovations. With experience as an engineer at the New York Times and as a designer at Dropbox and Square, Karina has a rare firsthand perspective on the cutting edge of AI and large language models. In our conversation, we discuss:

• How OpenAI builds product

• What people misunderstand about AI model training

• Differences between how OpenAI and Anthropic operate

• The role of synthetic data in model development

• How to build trust between users and AI models

• Why she moved from engineering to research

• Much more

—

Brought to you by:

• Enterpret—Transform customer feedback into product growth

• Vanta—Automate compliance. Simplify security

• Loom—The easiest screen recorder you’ll ever use

—

Find the transcript at: https://www.lennysnewsletter.com/p/why-soft-skills-are-the-future-of-work-karina-nguyen

—

Where to find Karina Nguyen:

• X: https://x.com/karinanguyen_

• LinkedIn: https://www.linkedin.com/in/karinanguyen28

• Website: https://karinanguyen.com/

—

Where to find Lenny:

• Newsletter: https://www.lennysnewsletter.com

• X: https://twitter.com/lennysan

• LinkedIn: https://www.linkedin.com/in/lennyrachitsky/

—

In this episode, we cover:

(00:00) Introduction to Karina Nguyen

(04:42) Challenges in model training

(08:21) Synthetic data and its importance

(12:38) Creating Canvas

(18:33) Day-to-day operations at OpenAI

(20:28) Writing evaluations

(23:22) Prototyping and product development

(26:57) Building Canvas and Tasks

(33:34) Understanding the job of a researcher

(35:36) The future of AI and its impact on work and education

(42:15) Soft skills in the age of AI

(47:50) AI’s role in creativity and strategy development

(53:34) Comparing Anthropic and OpenAI

(57:11) Innovations and future visions

(01:07:13) The potential of AI agents

(01:11:36) Final thoughts and career advice

—

Referenced:

• What’s in your stack: The state of tech tools in 2025: https://www.lennysnewsletter.com/p/whats-in-your-stack-the-state-of

• Anthropic: https://www.anthropic.com/

• OpenAI: https://openai.com/

• What is synthetic data—and how can it help you competitively?: https://mitsloan.mit.edu/ideas-made-to-matter/what-synthetic-data-and-how-can-it-help-you-competitively

• GPQA: https://datatunnel.io/glossary/gpqa/

• Canvas: https://openai.com/index/introducing-canvas/

• Barret Zoph on LinkedIn: https://www.linkedin.com/in/barret-zoph-65990543/

• Mira Murati on LinkedIn: https://www.linkedin.com/in/mira-murati-4b39a066/

• JSON Schema: https://json-schema.org/

• Anthropic—100K Context Windows: https://www.anthropic.com/news/100k-context-windows

• Claude 3 Haiku: https://www.anthropic.com/news/claude-3-haiku

• A.I. Chatbots Defeated Doctors at Diagnosing Illness: https://www.nytimes.com/2024/11/17/health/chatgpt-ai-doctors-diagnosis.html

• Cursor: https://www.cursor.com/

• How AI will impact product management: https://www.lennysnewsletter.com/p/how-ai-will-impact-product-management

• Lee Byron on LinkedIn: https://www.linkedin.com/in/lee-byron/

• GraphQL: https://graphql.org/

• Claude in Slack: https://www.anthropic.com/claude-in-slack

• Sam Altman on X: https://x.com/sama

• Jakub Pachocki on LinkedIn: https://www.linkedin.com/in/jakub-pachocki/

• Lennybot: https://www.lennybot.com/

• ElevenLabs: https://elevenlabs.io/

• Westworld on Prime Video: https://www.amazon.com/Westworld-Season-1/dp/B01N05UD06

• A conversation with OpenAI’s CPO Kevin Weil, Anthropic’s CPO Mike Krieger, and Sarah Guo: https://www.youtube.com/watch?v=IxkvVZua28k

• Tuple: https://tuple.app/

• How Shopify builds a high-intensity culture | Farhan Thawar (VP and Head of Eng): https://www.lennysnewsletter.com/p/how-shopify-builds-a-high-intensity-culture-farhan-thawar

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email podcast@lennyrachitsky.com.

—

Lenny may be an investor in the companies discussed.



Get full access to Lenny's Newsletter at www.lennysnewsletter.com/subscribe

More from Lenny's Podcast: Product | Career | Growth

All 287 episodes
OpenAI researcher on why soft skills are the future of workLenny's Podcast: Product | Career | Growth · 1 h 15 min
Listen in VO