How LLMs Are Reshaping Recommendation Systems

18 Aug 2026 · 48 min · 24 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How LinkedIn rebuilt its newsfeed using LLM-style sequence modeling (predicting what a user wants next like next-token prediction), plus LLM retrieval and natural-language “policies” to steer content quality. It also covers engineering challenges: inference cost/scale, evaluation, context construction, and combining LLMs with traditional signals.

Guests

Tim Yerka, VP of Engineering at LinkedIn (consumer products like feed/search/profile); joined 13 years ago as an early AI engineer building the feed; works on massive-scale ranking for ~1.3B members. Matt Merrill, software engineering leader with 20+ years scaling backend/cloud/distributed systems; architects and leads at Dept Agency.

Key claims

Sequence models capture content paths better than independent item scoring; LLMs don’t replace quantitative signals (e.g., popularity/math); “policies” defined by product managers/tastemakers guide quality and are evaluated via A/B tests and human-in-the-loop/auto eval loops; engagement weights are personalized per user intent.

Notable examples

“Zero-shot” onboarding; unconnected suggested content (topical, beyond your network); filtering language on CPU; popularity counts treated as numeric signals; detecting “AI slop”/inauthentic content and limiting low-quality/unoriginal posts.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction to LLMs in Recommendation Systems

0:00 to 0:44

Learn how LLMs are changing the landscape of recommendation systems.

“Newsfeeds and recommendation systems have long relied on deep learning architectures that score each candidate item independently.”

LinkedIn's Feed Revamp with AI

1:40 to 2:54

Explore the challenges and innovations involved in revamping LinkedIn's feed.

“I'm Matt Merrill, and I am here with Tim Jerka, VP of Engineering from LinkedIn.”

Challenges of Scaling User Engagement

2:54 to 5:04

Understand the hurdles LinkedIn faces with its growing user base.

“Like what was the overall problem and why was it challenging?”

Enhancing Content Relevance with AI

5:04 to 6:36

Discover how LinkedIn is improving content relevance for users using AI techniques.

“So was the primary driver new user stickiness and a nice side benefit was existing user engagement or was it all of it?”

Sequence-Based Interaction for Recommendations

6:36 to 7:59

Learn about the use of sequential interactions to boost content relevancy.

“Does it, does it also include that or is that outside the scope of what you're doing?”

Building Unique Models for LinkedIn

7:59 to 9:38

Insight into how LinkedIn builds tailored models for its recommendation system.

“That's then your like deep learning model.”

Experimentation Process for AI Models

9:38 to 11:23

Examine LinkedIn's continuous experimentation to optimize its models.

“Like it pulls from concepts from Lodlager, including some stuff that we've built internally.”

Defining Engagement Metrics

11:23 to 13:00

Understand how engagement is measured and personalized on LinkedIn.

“it's what we call a multi-head model with a multiple objective optimization.”

Evaluating Content Quality with Policies

13:00 to 14:00

Explore how LinkedIn evaluates content quality using defined policies.

“That's something that I have personally been, I am not working at nearly the scale you are, but I have been doing a lot in my job is just how you define what is good in this space.”

Leveraging LLMs for Content Retrieval

14:00 to 17:56

Learn how LLMs can enhance content retrieval by using defined policies.

“But for something like retrieval, where you're using LLMs to actually retrieve content that is relevant to a person, you actually want to be looking at some sort of, we call them policies, right?”
Show all 24 chapters

Leveraging LLMs for Content Retrieval

18:01 to 19:02

Learn how LLMs can enhance content retrieval by using defined policies.

“Your GitHub Actions bill is now a function of how much AI code you generate.”

The Role of Engineers in LLM Integration

19:37 to 23:28

Explore the diverse roles engineers play in integrating LLMs into systems.

“Let me start with the people who really are moving 1 ,000 miles a minute on this stuff are people who are generalists who can traverse more parts of the stack.”

Scaling LLM Architectures for Performance

23:28 to 28:00

Learn about the strategies used to scale LLM architectures efficiently.

“And so it's not like each of these is a permanent large investment.”

Batch Processing in Recommendation Systems

28:00 to 29:09

Learn how batching updates can optimize data flow in recommendation systems.

“I'm having trouble wrapping my mind around how you would batch some of this stuff.”

Infrastructure for AI Models

29:10 to 30:28

Explore the infrastructure needs for running AI models effectively at scale.

“Like what does this look like in terms of hardware and infrastructure and cloud?”

Evaluation vs. Inference in AI

30:29 to 33:18

Understand the differences between evaluation and inference in AI models.

“And so that kind of eval, you're not serving that to like a billion people.”

Challenges of Context in LLMs

33:19 to 35:57

Learn about the importance of context and token management in LLMs.

“With time, they fine tune the models to be pretty good at like, you know, math.”

Balancing Old and New Technologies

35:58 to 37:57

Discover the balance between new LLM technologies and traditional methods.

“So that new set of problems, what does your team lose sleep over now that maybe they didn't even just one or two years ago?”

Human-in-the-Loop for ML Systems

37:58 to 40:49

Learn how human input is integrated into modern ML systems for better outcomes.

“Now it's much more about when you have exposed so many controls.”

Ethical Considerations in AI

40:50 to 42:01

Understand LinkedIn's commitment to responsible AI and ethical engineering practices.

“that the kind of feed distribution, feed quality is equitable across different audiences and creators using like a series of statistical like checks.”

Trust and Credibility on LinkedIn

42:01 to 43:30

Learn about the importance of verified identities in building trust on LinkedIn.

“For example, we have like 100 million members verified on LinkedIn.”

Preventing Low-Quality Content

43:31 to 44:32

Discover how LinkedIn combats low-quality and unoriginal content through various tools.

“to make sure we avoid broadly promoting them on our platform.”

Skills for the Future of Software Development

44:33 to 47:18

Explore the essential skills needed for effective collaboration and decision-making in software development.

“good system design, knowing when to use the right tool for the right job, when to use the right size tool for the right job, input sanitization.”

Curiosity and Critical Thinking

47:19 to 47:31

Understand why curiosity and critical thinking are vital in the evolving tech landscape.

“What I just heard is curiosity and critical thinking are absolutely key.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Newsfeeds and recommendation systems have long relied on deep learning architectures that score each candidate item independently. As LLMs have matured, they have opened up a fundamentally different approach, where a system can reason about content the way it reasons about language. However, that power comes with a fresh set of engineering challenges around cost, scale, and evaluation. LinkedIn recently rebuilt its newsfeed to treat content recommendation as a sequence modeling problem. The general approach is to predict what a user will want next, much like an LLM predicts the next token in a sentence.

0:38Tim Yerka has worked at LinkedIn for 13 years and is currently a VP of engineering. In this episode, Tim joins Matt Merrill to discuss how LinkedIn re-engineered its feed, how the team combines LLMs with traditional signals, managing inference costs at massive scale, steering content quality using natural language policies, and more. Matt Merrill is a software engineering leader with over 20 years of experience building and scaling software teams across enterprise and product-focused organizations. His background is in backend development, cloud architecture, and distributed systems design. He currently architects and delivers software products and leads a team of engineers at Dept Agency.

1:22You can learn more about his work at code.theothermattm.com.

1:39Hey, everybody. I'm Matt Merrill, and I am here with Tim Jerka, VP of Engineering from LinkedIn. So today we're going to talk a little bit about what LinkedIn's doing with revamping their social feed with AI. But before we start, Tim, can you tell us a little bit about yourself and your current role at LinkedIn? Thanks for having me, Matt. As you mentioned, my name is Tim Jerka. I am VP of Engineering at LinkedIn, and I work on our consumer products. So what that would be is things like the LinkedIn feed, search experience, profile, pretty much anything you will engage with when you open up the LinkedIn app.

2:10And I actually started here 13 years ago as an engineer on the feed. I was an AI engineer, one of the first AI engineers building the LinkedIn feed. So it's been pretty amazing to go from actually building the thing to now see the thing operating at just massive scale with 1.3 billion members. It's been quite a journey. So recently, you guys have rolled out some pretty major updates to how you're making the LinkedIn feed relevant to users. And I know that's welcome news to me. And it's probably welcome news to a lot of people who may not have been otherwise engaged with LinkedIn because of content that was irrelevant or they thought was annoying.

2:44So I'm really excited to hear about what you did and also how it intersects with AI tooling, which I think this audience will really like. So let's start with just what was the problem, right? Like what was the overall problem and why was it challenging? I mean, I can maybe start by rooting it in a problem of scale. Like we've, since just 2020, we've gone from about like 690 million members to 1.3 billion members. We've scaled considerably. And we have about seven new members joining LinkedIn every single second. And so with that scale comes a lot of challenges. Number one, you have tons of new industries coming on the platform.

3:20People in new roles, new opportunities. They're looking for something different from our platform. And there's a lot of people coming with evolving expectations of what they need from their professional social network. Like starting from 2020, I mean, the world of work has changed a lot. We went through a global pandemic. We then went through the entire like AI LM revolution, which we're probably ostensibly still in the middle of. And people are just turning to LinkedIn more than ever to get this information about like, how do I navigate this? How do I reskill? How do I learn about what's happening, keeping pace with things that are happening in my industry?

3:54And so the problems we had to solve were twofold. One is make sure that we could kind of scale to this very kind of granular topical matching and representation of the entire LinkedIn feed ecosystem and make sure we're giving people the right content at the right time with timely insights. The other is with seven new members joining every second, you have to do a pretty good job in the first few seconds of a person's session, narrowing down what they actually care about. And so we almost call this like the zero shot experience. And both of these things together are kind of challenges we were solving.

4:25Number one is like, you need to get that signal very quickly to help the people get the insights that they need within a few seconds of onboarding. And they're not necessarily like endlessly scrolling on LinkedIn. They're trying to get a job done. They're joining to find a job, to learn something, and you have to really quickly get them that information to make sure you're maximizing the utility of LinkedIn. For people who are diehard users, I mean, there's a lot of things that even you alluded to, like we started seeing the emergence of AI slop, engagement bait, like we needed fine-grained control to make sure that we are focusing on the highest value conversations, people who are authentic, verified, and making sure that we have the capabilities to keep the ecosystem steered in that direction.

5:05And so this new foundation not only kind of up-leveled the fundamental infrastructure modeling capabilities on retrieval ranking, it also gave us these fine-grained controls to solve a lot of these problems from like faster time to signal world awareness, being able to do the zero shot session kind of stuff, but also being able to steer the system to make sure that we're giving the highest value content to our members. So was the primary driver new user stickiness and a nice side benefit was existing user engagement or was it all of it? It was across the board because one thing that like you'll see is there's kind of two classes of content that we surface.

5:44You know, LinkedIn has historically been this very network-oriented product where like I connected with a bunch of people. But the people you connected with aren't always going to be the people that are giving you the cutting-edge insights for where you're trying to learn. And so you'll actually see in your feed a lot more what we call unconnected content, suggested content, that is supposed to be really topically oriented to like, hey, I'm trying to learn about what supply chain risk means in the context of like, you know, like anthropic and cloud models. Like, how do I go deeper on that? What does that mean for my company?

6:09Or an electrical engineer trying to learn about modular reactors, like try and go deeper there. So like you need to get deeper into where the subject matter experts are, even if they're beyond your network. And so it's not just solving the short term zero shot signal. It's also be able to map the interesting content based on where you're trying to learn much more quickly. Yeah. And does that also extend to your outreach, right? Like I know I get emails with content that I might be interested in from LinkedIn. Does it, does it also include that or is that outside the scope of what you're doing?

6:39It's in scope, but right now we're starting with rolling this out in the feed and we're iterating a lot there. This stuff kind of, I mean, it cascades across our products. Like this is stuff that we build out and make sure that we have this feature rich capability everywhere. You might encounter content on LinkedIn. Cool. So let's dive right into the interesting stuff. So like one of the things I was reading about, there's a really interesting blog article that you guys posted about this goes into a lot of detail, which is awesome. you're using sequential time-based interactions to drive relevancy.

7:10And this is not something I personally have seen much about in the AI space. So can you talk a little bit more about that? Yeah. I mean, the underlying premise is actually quite similar to how you would think about LLMs, which is LLMs are really just sequence models for language. They're trying to predict, given the token, what are the next tokens that come after? And then stringing together like a comprehensive sentence, paragraph, whatever the context window is. You can do the same thing with content in the feed or any kind of item in a recommender system where you're trying to predict sequences of items versus sequences of words.

7:45And so there's a lot of benefit of this compared to like the old style MLP architectures where like you might have heard of these like feature crossing, DCNV2, deep and wide feature crossing, where you're like, you dump tons of features into the model. You try to find all these feature interactions. That's then your like deep learning model. And every time you give it a new item, it weights all the features in that model and it scores the item. And it's like, here's the score of this one and this one and this one. And then you sort. Sequence-based models are actually saying, contextually, we're going to learn that once somebody has viewed a piece of content about modular reactors, then they're probably going to go deeper in direction X or Y.

8:22It starts like modeling out a sequence of where you would go next, based on all the data points that we have of every single expert on LinkedIn reading about any number of topics. And so it has a lot of benefits then of starting to go extrapolate, you might be starting at like an introduction to PyTorch, but like you go like 15 levels down the sequence, you might actually be trying to build like a re-implementation of like an LLM locally on your machine. And so like, can you reconstruct that sequence of how people would go down a content path and help kind of understand the engagement patterns from that vantage point versus de novo, like rescoring every single item, you know, just using all the features in the model.

9:02That is super interesting. And also just, we'll get into this a little bit, just mind blowing at that scale. If you can't answer this, totally understand, but did you build this on any sort of open source or more widely available model? Or is this like something you guys have completely come up with R &D from your own perspective? As with anything, we pull from, you know, the research that we see around us. There's a lot that has to be uniquely built for LinkedIn. And so it's not like anything off the shelf works for us. It requires us to piece together something unique. To my knowledge, this is like fairly unique at this scale across like this kind of the system of the sequence based models and ranking and then the LLM based retrieval.

9:38Like it pulls from concepts from Lodlager, including some stuff that we've built internally. In terms of open source, like we do leverage like open source components in what we build. And it's honestly through experimentation that we figure out where we have to build things in-house versus open source. You try lots of different frontier models. You try lots of different approaches. And you kind of see what converges to solve the specific problems that we have in our scale, our professional context, the kind of fine-grained controls that we need. It's tough for me to visualize what this architecture might look like.

10:10But I'm assuming you're using multiple models in different places. I mean, we're continuously. like if you think about the kind of life cycle of we have offline environments where we will train models using some interesting like an engineer has an idea of like adding a certain feature to the model or trying a different model architecture offline we can run like thousands of these tests offline and replay them on historical data and say like did we do a better job ranking items in terms of like engagement then we have to actually like bring these online to see like do the members actually like it like not in simulation but like in in the real world and so there we then like we will split up our traffic and we'll run hundreds of experiments while there so at any given time we're running tons of tests to try all sorts of stuff and you may have seen some of this already you know publicly like making sure that we're trying to focus on like authentic voice conversations making sure that we're not you know rewarding ai slop making sure that our suggested content is really well targeted to interests and being able to better model member affinity to different interests.

11:11Like all these things are variations of tests that we're running at any given point in time. And what's the engagement signal? Is it clicks? Is it time? Is it all of it? Yeah, so as a standard, you know, in industry, it's what we call a multi-head model with a multiple objective optimization. And so it's kind of a little bit of everything. You want to make sure that you have predictors for different ways that people will engage with things. That means positive predictors of like, are you going to be commenting on this are you going to be resharing it are you going to be spending time on it also negative predictors are you going to skip it are you going to be like is it going to be in a viewport for a second just pass over it that's like a negative prediction that we try to do as well and you combine all these together into one multi-objective optimization to actually then ultimately optimize for i mean for us it's really legitimately trying to get the most interesting conversations in front of people that we feel like it's gonna be time well spent but it's a combination of all these factors that then ladders up to that.

12:08That makes a lot of sense that what defines engagement is a very complicated thing in and of itself. So yeah, using these models to help track that makes sense. And maybe to like illustrate it even more, the weights of all these different actions are personalized on a per user basis. If you're somebody on LinkedIn, who's like in marketing, like you're trying to get like follower reach and comments, like you're trying to achieve something different than somebody who is prepping for their job interview is like, I just want the content that's going to show me how to like nail the job interview at LinkedIn, right?

12:39You're not going to be looking to get like a post and to get comments and like reach. And so you have to adjust the model and how it works depending on the actual intent of the person using the model. And it's literally by user. This is we're not talking about roles or personas or anything like that. It's very unique to each user. It's fully personalized. Yeah. Wow. That is awesome. So let's double click a little bit into that evaluation. That's something that I have personally been, I am not working at nearly the scale you are, but I have been doing a lot in my job is just how you define what is good in this space.

13:14Right. And we just touched on that a little bit. So I think the thing is like over time, how does this happen and who is doing this, right? Like I can't imagine, well, I don't know, maybe I'm wrong, but I can't imagine they've got armies of you, you know, PhD trained in AI doing this. So are these engineers, are they product people? And like, how often are you retraining at full scale versus just incremental improvement over time? I'd love to hear more about that. So in terms of training, we incrementally retrain constantly. So like we're pulling in the latest data as soon as we get it, retraining, updating the model.

13:51The more interesting part is the evaluation of how you're actually sure that this stuff is good. And I mean, there's a number of ways you do it. There's the more traditional ways you eval, which is like A-B tests, and you look at your kind of like metrics and you're like, okay, what happened to A-B? be. But for something like retrieval, where you're using LLMs to actually retrieve content that is relevant to a person, you actually want to be looking at some sort of, we call them policies, right? You define a policy for what consists of high quality content. This is something that's defined by our product managers.

14:22And these can be evaluated using people like, you know, we've called tastemakers, people who are like really well attuned to the product strategy and where we want to take things and make sure that content is timely and insightful, et cetera. So you'll have like a timeliness policy and like type of policy, but also you can start encoding this into like auto eval. You can have GitHub copilot CLI essentially doing revs of like, Hey, I looked at a bunch of content that was surfacing retrieval. It missed the policy in these places. Like, do you agree with me or do I need like some feedback to like iterate on, make sure the policy is reflecting what you intend.

14:57And so you get into that kind of like active learning loop. And so we do both. And it's been pretty incredible to just be able to have that capability. It gives you just much more sense and fine brain control in terms of, are we actually doing what we aspire to do? Like, is this actually timely content? Is this actually interesting for this particular audience? You know, in the old days of ML, you kind of hoped that that was the emergent property of your system. You know, like you put out a model and it's like, we optimize for this thing. And we hope that like it surfaces interesting stuff now you can actually kind of validate it you can say like you know for people in nuclear engineering are you actually surfacing stuff that is relevant to their industry or is this something that is either like general purpose humor or something that's not as fine-tuned and so you can get into that level of granularity and it's sounding like it's kind of happening in in parallel almost yeah which is wild it's almost like an organism Awesome.

15:50That's very interesting. And about the who, you said it's product managers, tastemakers. How are they inputting that signal in to correct that? What does that look like from a user experience? I mean, there's a number of ways they will directly edit. I mean, at the end of the day, it's a text file that's checked in to GitHub, you know, and like they've written a policy almost you can think of as a prompt. We've built out or agenda tooling around that. So like they can iterate on that prompt by getting some examples from GitHub Copilot CLI and it's like, hey, here's a few examples. Here's what I think I would score them based on what you told me.

16:28And they can be like, no, you missed the point on this one. And the Copilot might then update some of the words of the policy to make sure that then it's reflecting kind of the intent of what we're trying to do. So this is a very engineering centric, like CLI based process. So your product owners are very technical. Everything is version checked and like, you know, yeah. Wow. Oh man, that sounds great. So we can see live over the course of a week, places where we've evolved our opinion on like, you know, what constitutes inauthentic content, like people giving kind of AI slop style replies, and we can fine tune to make sure that we address that.

17:07You're building agents that can write code, summarize documents and automate workflows, but they're missing one thing, awareness of the world around them. X-Weather combines enterprise-grade weather intelligence with agent-ready APIs, natural language capabilities, and an MCP server built for tools like CLAWD, Codex, Copilot, and modern IDEs, so your agents can adapt workflows, automate responses, and make better decisions based on real-world conditions. Backed by Vaisala, whose instruments fly on NASA missions to Mars, X-Weather delivers trusted data and unique insights that go beyond conditions to actual impact, from real-time lightning strikes to road surface forecasts.

17:43Start with 15 ,000 free API calls every month and pay only for what you use as you grow. Your full weather stack for developers by developers. Start building for free today at xweather.com. This episode of Software Engineering Daily is brought to you by Warp Build. AI is writing more code than ever, which means GitHub Actions is running more than ever. Your GitHub Actions bill is now a function of how much AI code you generate. And every engineer knows the feeling. You push a commit, and then you wait. Warp Build makes GitHub Actions twice as fast at half the cost, with a one-line change to your workflow.

18:18Linux, macOS, and Windows Runners, in Warp Builds Cloud or your own, enterprise-ready, SOC 2 Type 2 attested, and trusted by teams like Sky from Comcast, Bitcoin, and Braintrust AI. Get started with$50 in free credits at warpbuild.com. You're shipping faster than ever with AI coding agents, but those agents don't vet the packages they pull in, and they don't have security context built in. Ori by Endor Labs fixes that. It plugs directly into your editor via MCP, catching vulnerabilities, blocking malicious packages, and flagging exposed secrets in real time. No separate tool to switch to, no dashboard to babysit, security that fits how you actually build.

18:56Teams using Ori see 10 times fewer security tickets and six times faster fixes. Free for developers. Get started at www.endorlabs.com slash A-U-R-I. And I'm very curious about like the day-to-day for an engineer in here, right? So I mean, even your product people are working in like a code-based workflow. So what is like an average engineer doing in this environment? because I think that's where a lot of people are wondering, like, okay, how do I fit in with my existing knowledge? I might not be a PhD in AI. How do I fit in and start to prepare myself for what this world might look like? There's no average engineer.

19:38That's fair. Let me start with the people who really are moving 1 ,000 miles a minute on this stuff are people who are generalists who can traverse more parts of the stack. So they not only understand how the AI models work and how to fine tune like an LLM to distillation. They also understand how to integrate that into like the back ends of the infrastructure. Maybe they even know how like prompt engineering works and like all that kind of is a super powerful skill set for people to kind of cross over the boundary of decision-making into like, what does this product actually look like? However, that hasn't negated the role of having some specialized skills.

20:15So to get an ML model to production in front of this many people, you need to make sure that you have nearline pipelines that are taking all the tracking events from the client and like propagating them in nearline to all these feature stores to make sure i know this particular post by you know jaylen bronson was like you know like the 25 ,000 time in the new york metro area and these metro areas which is all in real time being updated to then know where we target this stuff so like all these like nearline signals need to get adjusted to the You then need to run this massive model in inference on a GPU fleet.

20:49And so you need inference engineers that are optimizing the QPS throughput of the stuff because we don't want to overspend and buy every NVIDIA GPU out there. You can do a lot of things to increase the throughput of these systems. Sometimes you can offload some of the stuff to CPU workloads, keep it in the GPU workload. You have infrastructure engineers looking at that. You have AI engineers that are training and deploying. and a lot of engineers that, like I said, are generalists that are then making sure that these things all kind of work seamlessly together as we ramp to our members. And so the average engineer is somewhat of a misnomer because you still need a multifaceted team.

21:22And we almost think of it as kind of like these pods of teams where you might have like a product manager that's helping shape some of the taste and you have a designer that's designing the feed card and infrastructure engineer and a backend engineer. And so one of these pods might be like working on a particular model end to end and bringing it to like the members. That's fascinating. At a system of this scale, how many pods and how many people roughly on a pod? I mean, it was a couple hundred engineers to bring this to life, but like not everything's conducive to like these pod style. Like you have some deep infrastructure optimization work, but like where you're moving fast on a new product idea and like some way to like improve the feed, that's where like these pods can move really fast between like an idea iterating on it, deploying it and having that closed kind of loop ecosystem where the team is making their decisions and running with it and using agentic tools to develop.

22:11That makes sense to me. You need the generalists to be able to react quickly, put those puzzle pieces together fast and get it presentable to the user, so to speak, or even out to the user. And then you need those really deep specialties to be able to tweak the fine details to make sure it's not too slow, that it's not completely irrelevant, that it's not costing too much. Yeah, that's really interesting. It sounds like that's a lot of people to be able to get this done as opposed to, No, we actually did it with less people, which maybe it's just me wanting to see this through rose-colored glasses and save our industry.

22:45But I'm curious what your take is there. I mean, these are like, you know, systemic shifts that we've been working on for like 12 to 18 months and like iterating our way into this. And you're shifting from one model architecture and infrastructure to like this GP fleet. Like that does take build out. Like there is the upfront build out. Our infrastructure team is phenomenal. They work on lots of things beyond this. Like they've been working on feed in like heads down mode for the last couple of years. But that's not the only thing they do. Like once we got the build out going, you know, we are running and operating this with a much smaller team and focused on like using the platform to build interesting new features and getting interesting content in front of members.

23:21Our infrastructure teams are then moving on to like other interesting building out of like agentic platforms and they're kind of reading the next wave. And so it's not like each of these is a permanent large investment. You got to do the build out to then leverage this platform to do interesting things and build compelling experiences. Yeah, the investment's got to come in there. Yeah. So let's go back to scale because the scale is just absolutely incredible. And I think that there's kind of two things that I'm thinking about. One is just how is this architecture inherently set up to scale just from a conceptual perspective?

23:55And then there's also the hardware infrastructure part of it, too, which you've kind of hinted at. And I'm curious if you can describe a little bit of like how the interplay and how you thought about that to be able to create something of this scale with this cutting edge technology. When we started this, it's kind of like try to do everything on like GPUs and like see the best experience we can create for our members at the end of the day. Like you start there and you quickly realize, well, this isn't going to scale. And so that's where you find something that's compelling at small scale, test to 1%, 2%, 5 % of members and say, okay, there's something here.

Read the full transcript

24:30And then you do a lot of iteration. You start with a GPU native architecture, but then you're like, hey, there's lots of these things that we can probably offload to CPU. Not everything needs to be a GPU workload. You figure out how to make your embedding generation pipelines GPU accelerated, but then making sure that you maybe do some sort of batching. So like, you're not like just streaming everything directly into the system, but maybe you're aggregating on like 20 minute windows to like try to batch things together and reduce the kind of continuous compute and updating the model. There's tons of things that we've done, not just like the easiest thing to do is like change your model size.

25:05I think like, you know, like everybody talks, like use a smaller model. But like, I think we've focused a lot more on like the system design. So you can do things like you can build shared context batching. there's techniques like custom attention kernels that we were using to reduce the per request compute costs. That's not a concept I'm that familiar with. You said it was a custom attention kernel? Yeah. I mean, these are, so you are now probably pushing me like to start tapping in some of my infrastructure partners that like, honestly, like work more directly at the GPU level, like literally going down to the, to the CUDA kernel level and optimizing the CUDA kernel for the specific kind of computations we're doing at inference time, like they speed up to that level to make sure that we get the throughput that we expect.

25:49And so you're going really deep into the stack to make this happen. And I'm humble enough to say, I don't know what I don't know. Like, you should talk to some of our amazing infrastructure engineers that can probably give you a whole podcast just on the GPU acceleration that they did. Because folks like Ali Nakvi on the team have just done a phenomenal job there. I would love to. Let's scratching the surface of there, but that is fascinating. And you just go back real quick to you said that there were certain things that you offloaded to the CPU. What are some examples of that? Like, because that's something that makes sense to me, but like having something to grab onto, I think might be useful.

26:27I mean, there's a lot of places where like you can do pre-processing before you load everything into GPU memory. Once things are in GPU memory, like you want that to be the thing that you're using to then score a massive amount of items. But there are things like you can filter on different facets. You can say like, hey, this member is coming from California. We know the vast majority of content is going to be in English for this member. You can do language filtering. You don't need a GPU to do language filtering. You can do this in a simple CPU index and just do a faceted filter on that. And so that's where you're starting to cut down on some of the unnecessary compute by using cheaper hardware and CPUs before you load everything else into GPU memory and then do like your large scale inference.

27:10Am I oversimplifying by saying like that type of stuff can be broken down into more procedural code that's run on CPUs? Or is that, is it still somehow using models and, you know, embeddings and things like that to be able to do that? I think a lot of the stuff that's easily handed off to CPU is filtering and that kind of like non-model based stuff. There are also like models that just don't require GPU. Like you can have simpler modeling techniques that can give you like a approximation of like this item is going to be completely out of scope for us to score for this particular member and you don't need to do the full like 100 billion parameter model like like for every single item and so there are ways to break up even like the modeling capability of it yeah so it's kind of like simple logical passes to get to the things that need exactly yeah okay cool and the batching piece is interesting too I'm having trouble wrapping my mind around how you would batch some of this stuff.

28:07Like, how does that look in like a data flow? You said something about batching embeddings. Yeah, I mean, well, there's lots of places we use batching. I mean, you can like do batch scoring of items. You can also do like just how you update the features. You can do that in batch as well. So you don't do as many writes. Like if you think of the simplest thing in our Nealight pipelines, you have how many people liking on LinkedIn every second. I actually don't know off the top of my head, but it's going to be some exorbitant number in terms of likes per second. When you're updating the feature counter, you don't really care if it's like updated from like 100 to 105 in the next millisecond.

28:41You can stream in the next 10 minutes of like data just so you know the trajectory. And that way you're doing a write to database only once versus, you know, like doing it with every. This is like more traditional software engineering and pipeline engineering. Okay. But that makes a lot more sense. It's basically just using those models and those GPUs very intelligently because they're costly. Exactly. How about infrastructure, right? Like I can't even like fathom or, you know, you're owned by Microsoft. So you've got hopefully that at your disposal, but like I would get your own GPU farms. Like what does this look like in terms of hardware and infrastructure and cloud?

29:19I mean, we really have anything at our disposal, but like using frontier models, like without fine tuning them at all. is pretty inefficient because they're not task-aware for what you're trying to do. And so we run this in our data centers, right? We have our GPU fleets in our data centers. We run models that we have fine-tuned on LinkedIn data to make sure that they're more context-aware. And so they can be more performant on the tasks that we care about. They don't have to be fully general-purpose. That would be wasteful for ranking a feed. And so that makes them more suitable to then run at the scale that we need them to and on our own GPU fleet.

29:57There's like the reason I'm kind of like pausing here a little bit is it's super context dependent in terms of what part of the life cycle we're in for this kind of stuff. Like sometimes you might use a frontier model to do the kind of stuff we're talking about in terms of evaluation. Like you want to eval the quality of like you actually probably want the best model that's as close to like, you know, human intelligence as possible to be doing an eval against like a policy. because that's like a super nuanced thing of like, I'm telling you these things about what makes it interesting to a member and what makes it timely and relevant.

30:28What is it and isn't AI slop, right? And so that kind of eval, you're not serving that to like a billion people. You're using it to score kind of a sample of your data and say like, how are we doing here? Whereas for actually like running the inference, you want to fine tune the model and tweak it such that it can run on the GPU fleet in a more effective way. So it's very elastic. Yeah. is what i'm hearing yeah elastic in terms of both the techniques that we use depending on the life cycle of where we are iterating like eval versus inference but also in terms of the kind of compute footprint that you have to employ whether it be kind of you know running frontier models in the cloud or you're running the actual feed ranking model on our like gpus in the data center and you know in terms of just like you know i'm trying to imagine a scenario where like linkedin gets hit particularly hard i don't know there's like a zero day or something and everybody goes on and starts looking for LinkedIn articles.

31:18Like, is it just any other infrastructure scale problem? You just need more when that happens? Is there something else that might be more nuanced? I mean, what you're describing is like, it goes into all your end-to-end software engineering and data-centered engineering. It's like, you have Edge, you have some systems at Edge. Like, if it truly is what you described, an attack, right? Like, we're going to be blocking it at the Edge, not by the time it gets to, like, our inference containers. If we have a surge in traffic, yeah, we just provision more capacity elastically. We plan for that kind of stuff.

31:45we regularly run load tests to make sure we understand what kind of like our capacity is and what even like unexpected peak capacity would be to make sure that we can we can serve it so that gets back into just good systems and software engineering good procurement good redlining of your system caching yeah yeah exactly prediction okay cool it's good some of us mere mortals can do that stuff just kidding it's almost like some of the software engineering skills still apply in this new new agentic era unbelievable unbelievable yeah i i thought claude could do it all but maybe i'm wrong so this is really kind of getting into the details but i'm going to use this as kind of a way to kind of dive into some of these gotchas so one of the things that i read about in this article was that you were finding that numeric data like the count of an article would act like plain text and completely missed that it was a popularity signal and i think a lot of engineers, and you mentioned this in the article, are just dumping structured data into prompts and assuming the model just gets it.

32:46And I think many times, like what I have seen is that it kind of fakes it pretty well. But then when you start scratching beneath the service, it's like, what's going on? Like, I don't think it really gets that. So do you have like a good, like a good story or a good example like that? And like what people can learn from that and how you, the mistakes that you made so you can help some other folks about those types of patterns. I mean, even that is like a pretty great example of like maybe two things jump to mind. One is even like the item popularity features like even the early days of some of these like frontier models, they were not doing a great job with math, you started with like kind of rag to like format the task to a calculator, you know, and just you would literally have Python eval like the value and get it back to you.

33:28With time, they fine tune the models to be pretty good at like, you know, math. But like, when we say LLM, you can't like run rag at serve time for feet, right? You can't call out a bunch of tools. That'd be way too slow. You're scoring a feat. And so this is where you build out parallel capabilities. You have your LLM-based capabilities. And you can still use traditional counting systems that you would use in, I'll say, the olden days to track popularity and project out, is this on the upswing in terms of virality? What is the pace in which people are engaging with it? And so we had to combine these systems together.

34:04Like when we say like LLM, it's not like magical LLM solved everything. It's like LLM in combination with a lot of other features that are like more quantitative and maybe aren't easily understandable by the LLM. So we use these as feature interactions with sequence based model, etc. All this together forms kind of the model that we then end up using. The thing that we struggled with the most when it came with these LM systems, like the new problems that we created for ourselves, is the scale of the system is really dependent on what you put in the context for the LM. If you think about it, the LM is looking at the member context.

34:40Who is this person? What does their profile look like? Where do they work? What have they stated as their interest on their profile? What have they been reading in the feed in the last few sessions? And then the context of the item. Who wrote that particular post? What is the post about? If there's an image in it, describe the image, caption it. And if you just represent everything as just pure tokens and dump it in, you're just going to blow up the system. And so you have to be very judicious of almost cutting down and summarizing and figuring out what are the most important tokens that we factor in to make the context as compact as possible, but not too compact, that it then starts losing precision and the whole purpose that we would be using the LLM.

35:19And that's actually turned into like an entire, both for LLMs themselves, but also we talked about the sequence-based models where we, how do you actually like train the model against sequences? What sequences do you look at? Do you sample certain sequences more than others? Like not everything in the sequence is equally important. And so really careful construction of the context window, the tokens you put into it, items that you put in the sequence. These are all things that are like an entire pod is working on this to figure out how to maximize the information gain here. What I'm hearing, right, and this lines up with the gut that I've been feeling is that it's a big art of knowing when to employ LLM-based technologies and when to employ old-fashioned technologies and how to tie them all together.

36:05and i think one of the big misconceptions outside of the engineering community and you know i hope most engineers understand this now is that you can't just dump everything into an llm and just give it a bunch of markdown and expect it to do what it needs to do at scale it might work for you at your desk using claude you can probably hack it but when you want to actually make a system no way another way to put this is it takes a ton of refinement the amount of iterations like i mentioned on like the policy that we have etc it requires a lot of guiding the system to get to the outcome that you want especially at scale like you can make it work for a specific instance and probably get something pretty useful my favorite quote is like if you have an agentic system where like each step of the agentic system works 99.5 percent of the time and it's like a 20 stage agentic workflow then like the n10 workflow is useless because across 99.5 success across 20 steps, you're basically like, you know, the system is completely useless.

37:01And so we try to like augment the system with ways to ground it in being a bit more performant by giving it like popularity signals or making sure we have the right context in it, making sure we're finding the policy such that it was representative of the types of content that we want to select. All right. So moving on to something a little bit more open-ended, I saw in one of your write-ups about this project, you said that moving to LLMs didn't really solve your old problems so much as to swap them for a new set, which makes total sense based on this conversation. So I'm finding the exact same thing.

37:31So that new set of problems, what does your team lose sleep over now that maybe they didn't even just one or two years ago? It's a, I mean, I don't lose a lot of sleep. Good for you. I do because I have a two-year-old and a six-year-old, so I lose sleep from that. Makes sense. You're operating a lot more in this realm of trying to provide a lot of human-in-the-loop input to make sure that the system works as you expect. It's something that, like, I think with previous ML systems, you would really just train on large amounts of data, and it would be like the model is what it is, and you see the emergent behaviors of the model.

38:09Now it's much more about when you have exposed so many controls. You can use an LLM. You can use Evaldi's policies. like that's a lot of degrees of freedom that you have in terms of developing your product and it means that there's a lot higher pace of change in a lot of parts of the stack because now product managers are stepping in and they're stating an opinion about like iterating on the quality of this subsystem and then like an engineer is changing the model at the same time while somebody's iterating on the policy and then all these things come together in like an integration environment it's a lot more surface area that you have to make sure works together cohesively and you have the right mechanisms like the right operational rigor cohesion to bring this together in a cohesive way to like an end experience and so that's probably the thing that is both maybe a problem but also super exciting is like you expose all these controls in terms of what you can do to like improve the product experience but it also means a lot more coordination and so we've built up agendic systems to help us coordinate across some of these different things in terms of how we deploy them and they interact yeah and getting the balance right of being able to react quickly and not making that process too cumbersome so that you can't get changes out fast yeah that's really interesting there's other things where it's like we had to build up a lot of bespoke tooling for like understanding when somebody was like hey this recommendation didn't resonate with me what happened in the system you know you'd have a lot of you'd look at all the coefficients of the model and you'd be like i think you know these features were the reason that this was recommended this way some of the the plus side of this is like you can now actually ask the model to reason and say like why was this recommended to explain it, right?

39:40And that's somewhat magical in terms of the pace with which you can improve the product. You can understand where there's misunderstandings and like maybe we have not provided enough guidance for the model to be able to operate in the way that we expect it to. Right. And this problem set too, right? Like it's not like it's a medical device. You're serving content to people. So it's such a great use case. I do get kind of scared with the medical stuff, but yeah. So it's all about the right tool for the right problem. Speaking of that, one thing that also just not necessarily technical, but also very important to talk about is just ethical use of this.

40:13And I did see that you guys at LinkedIn have what's called a commitment to responsible AI, and that's really nice to hear. And can you talk about what that is and how you folded it into the engineering of this architecture? So yeah, responsible, it's a top priority for LinkedIn. Like fundamentally, it comes into like the kinds of signals that you're using in the model, right? So we make sure to use professional signals and like engagement behavior only, not like sensitive demographic attributes in predicting things. And then with every single feed model that we ship, we run kind of regular audits to make sure that the kind of feed distribution, feed quality is equitable across different audiences and creators using like a series of statistical like checks.

40:59So that's kind of how we incorporate it into internal and workflow to make sure that at the end of the day, like the ranking that we're doing, et cetera, is equitable for every single member. And in terms of guardrails, like, you know, just making sure that malicious content doesn't spread or anything like that. Yeah. Okay. That too. I mean, trust is a top, like, I think when I think about responsibly, I there's kind of like how, how systems take into account demographic variables that when you talk about trust, it's making sure that you're, You have the right defenses to protect members against malicious content, scams, spam, all that kind of stuff.

41:34Both of these are separate and super critical technical things to get right in an ecosystem. Yeah. I'm sensing like the pattern here is almost like in many cases in different ways, like defense in depth, whether it's your architecture doing a couple of things a different way or having a bunch of different guardrails. That's really interesting. On that front, I think there's a lot of things that kind of like are organically come from LinkedIn's professional context in terms of some of these safeguards. For example, we have like 100 million members verified on LinkedIn. Like, when you come to LinkedIn, and you see a post from a verified member, like, you know, this is a real person.

42:09And so, you know, you're getting credible, verified professionals, you know, that's ultimately the goal is that the system can find the most useful, credible information and then service to you. you can trust that like, yeah, okay, this person is who they say they are, and they're like credible in their field. And that's just kind of like a implicit thing in our ecosystem, just by virtue of the fact that we are this professional platform where people verify their identity and say, here's what I've worked on. Here's who I am. Here's where I went to school. And I feel like that gives a certain level of depth to the experience in addition to the content itself, in terms of your confidence as a member viewing that and saying, yeah, okay, I don't have to necessarily double fact check this or worry about if I'm getting something from somebody who's a malicious actor.

42:58We lean really hard into that from a trust standpoint. And I may be wrong here, but you all are pretty, this might not be the right word, but protective of your APIs, right? As far as I know, you do a lot to prevent bots from just like putting out AI slop content. Is that correct? Like you are looking for humans entering content into your site, your platform, right? That's right. So we use a variety of different tools to make sure that we detect what we call like low quality and unoriginal content. And then we take action on those to make sure we avoid broadly promoting them on our platform. And so a lot of the kind of bots, AI slop, like they tend to post exactly this kind of unoriginal stuff.

43:41and you can kind of sense it. And that's not what we're about. We want to make sure that, obviously, members should use these tools, but we want them to use the tools to bring the kind of authentic, original perspectives to LinkedIn, not like super repetitive stuff through automation. That's a no-go for us. Yeah, my analogy to the old world is like you're using LLM technology as input sanitization to not have to deal with it later in the pipe. Exactly. Like people using browser extensions or like scripts or third-party tools to like leave automated comments is just like, and first of all, it's officially not allowed in our terms of service, even though people do it anyway.

44:20And we make sure to remove those from the like most relevant view of comments. So they won't show up in most relevant. We make sure we limit the reach of those. And for repeat offenders, we might limit like access to LinkedIn altogether. It's refreshing to hear, you know, some of the basics here. good system design, knowing when to use the right tool for the right job, when to use the right size tool for the right job, input sanitization. Don't just throw everything in the same pile. It's nice to know that the basics still stand. All right. So the last question, I like to end some of these interviews with something for the audience to reflect on and how they may be able to relate this to their work.

44:59So when you look at this whole project and as it still goes on, And what's the skill you found your team or yourself wishing that you had more of? And I think that what I'm trying to get at is for the folks listening, where should they try to point themselves next to make themselves more valuable or potentially try to fill that type of hole? Because this is probably a bellwether for what's going on. It's a great question. I always start from a premise that whenever you're using these agentic coding tools, they don't know team boundaries. They don't know code repo boundaries. They know none of that.

45:33And that's part of what makes them so effective is they can start taking a problem that you formulate in your head and it's like, I can solve this by navigating an end-to-end code base or universe and make it happen. And I think a lot of software development teams have been built up with charters and understanding, here's my role and here's my infrastructure partner's role and here's my data science partner's role. And those walls are really, they're melting. the people who are who are most effective are those that can cut across these contexts and make sure that even if you're not a domain expert in like an adjacent kind of capacity that you know enough about it that you can kind of figure out how the broader picture stitches together like that's super super critical and so you know you still have infrastructure specialists but like more and more they understand like the ai models that you're using and architecture is because like you know you need to you can't just wait for like some ai engineer to to mosey along and tell you kind of the performance characteristics of the next idea.

46:31And I think that extends even further beyond the technical realm, which is when decision making happens this quickly, you're building so quickly and you're starting to build in adjacent places and you're able to get these ideas out in front of your members. You need to have internalized the context of the problem you're working in, you know, like a bit more of the product, like vision and like what you're actually trying to build, making sure you understand the business context that you operate within. because the bottlenecks then become like, are you able to just make those calls yourself? Is your batting average high enough that you're gonna just make these calls?

47:04And more often than not, you're gonna make the right call because you're gonna immediately convert that decision into code. Like that is the most pure way to develop a product is if you can like have good intuition, be a good tastemaker, convert that decision into code and then get that in front of members. That is kind of the life cycle. Yeah, that is really well said. What I just heard is curiosity and critical thinking are absolutely key. Being an order taker and expecting what to do to be given to you will not work in this new world.

From the publisher

News feeds and recommendation systems have long relied on deep learning architectures that score each candidate item independently. As LLMs have matured, they have opened up a fundamentally different approach, where a system can reason about content the way it reasons about language. However, that power comes with a fresh set of engineering challenges around cost, scale, and evaluation.

LinkedIn recently rebuilt its news feed to treat content recommendation as a sequence modeling problem. The general approach is to predict what a user will want next, much like an LLM predicts the next token in a sentence.

Tim Jurka has worked at LinkedIn for 13 years and is currently a VP of Engineering. In this episode, Tim joins Matt Merrill to discuss how LinkedIn re-engineered its feed, how the team combines LLMs with traditional signals, managing inference costs at massive scale, steering content quality using natural language policies, and more.

Sponsorship inquiries:
sponsor@softwareengineeringdaily.com

The post How LLMs Are Reshaping Recommendation Systems appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
How LLMs Are Reshaping Recommendation SystemsSoftware Engineering Daily · 48 min
Listen in VO