Self-supervised User Profile Generation for Personalization

9 Jun 2026 · 22 min · 11 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Self-supervised generation of personalized AI user profiles for assistants, avoiding prompt “blank slate” behavior and expensive human labeling.

Guests/backgrounds

No specific guest bios are provided; the episode is a “Deep Dive” discussion featuring researchers named in the paper: Clark Bingswan-Ju (UAQ), Tong Zhao, and Neil Shaw.

Key claims

Current personalization relies on hidden prompt preambles built from interaction-history summaries, but training accurate profile generators needs costly supervised rewards from labeled downstream tasks. The proposed BMP (Bidirectional User Modeling via Profiles) trains a profile generator self-supervised from raw interaction logs only, using GRPO (Group Relative Policy Optimization) without explicit human reward models.

Notable examples

LAMP benchmark subtasks like personalized citation generation, movie tagging, personalized emails, and news summaries; bidirectional “guess who” tests using a judge LLM with multipositive NDCG and in-batch “free negatives.”

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Current Limitations of AI Models

1:49 to 2:26

Discuss the limitations of current AI models in personalizing user interactions.

“That's Clark Bingswan-Ju, UAQ, Tong Zhao, and Neil Shaw.”

Understanding User Profiles in AI

2:26 to 3:56

Learn how AI generates user profiles through interaction history.

“Under the hood, developers are already trying to personalize these models.”

Challenges of Supervised Learning

3:56 to 5:44

Examine the difficulties associated with supervised learning in AI.

“If I understand the current paradigm correctly, developers are having to define what good looks like for every single specific thing a user might ask.”

Introducing BMP Framework

5:44 to 6:15

Discover the BMP framework, a self-supervised approach for user profiling.

“It is incredibly expensive, the resulting data is sparse, and it simply cannot scale to the complexity of human life.”

Mechanics of GRPO Method

6:15 to 8:01

Delve into the mechanics of the GRPO method for profile generation.

“It completely sidesteps the human annotation bottleneck.”

Bidirectional Ranking Test Explained

8:01 to 11:41

Explore the bidirectional ranking tests that validate user profiles.

“It's self-correcting based on the group.”

Scoring Mechanism and NDCG

11:41 to 13:52

Understand the scoring mechanism using NDCG for profile accuracy.

“There are zero human graders checking if it wrote a good email or suggested a good movie.”

In-Batch Ranking and Free Negatives

14:00 to 16:16

Learn how in-batch ranking optimizes data processing in machine learning.

“In machine learning, gathering millions of wrong answers, what developers call negative examples, usually requires heavy compute to generate synthetic data or mine hard negatives.”

Evaluating the BMP Framework on LAMP

16:16 to 17:48

Discover how the BMP framework was tested and its impressive results.

“And we know this specific bidirectional approach works because the researchers didn't just propose the math.”

Democratization of Personalized AI

17:48 to 19:12

Understand the implications of self-supervised learning for AI development.

“it drastically lowers compute costs using free negatives, and it demonstrably beats the old expensive methods on rigorous benchmarks.”
Show all 11 chapters

The Psychological Implications of AI Profiling

19:12 to 20:38

Explore the unsettling questions raised by AI's ability to predict behavior.

“It rests the technology away from the center and puts it back in the hands of the user.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00You know that feeling when you sit down at your computer and you open up your favorite artificial intelligence assistant and you ask it for something simple like maybe a draft of an email or some research for a project right and what it's a great you know that's a great you know that's a great you know that's great. So you end up spending the next 20 minutes practically begging the system to understand your preferences. Oh, constantly. They're constantly reminding it like, no, I need this to be highly technical or no, I already know the basics. Skip the introduction. It is a I mean, it's a uniquely modern frustration.

0:48It really is. You were talking to one of the most advanced technologies in human history. Yeah. And yet it feels like it has a kind of digital amnesia every single time you open a new chat window. It is a profound friction point in human-computer interaction right now. We call these systems assistants, but think about a real human assistant. They learn about you over time, you know? Right, they pick up on your quirks. Exactly, they pick up on your habits, your tone, the unspoken context of your workflow. Right now, well, most of the large language models you interact with are essentially starting from a blank slate.

1:24They're fundamentally isolated from who you are as an individual user. Because they're just trained on everything, right? Yeah, they are trained on the statistical average of the entire internet, which means, by default, they sound like everybody and therefore nobody. And solving that exact problem, teaching these systems how to actually understand you without making you do all the heavy lifting, is our mission for today's Deep Dive. It's an exciting one. We are looking at a brand new research paper submitted in June 2026 by a team of researchers. That's Clark Bingswan-Ju, UAQ, Tong Zhao, and Neil Shaw.

1:58And they have proposed a completely groundbreaking framework that aims to fix this generic AI problem once and for all. And do it efficiently. Right. The kicker is they are doing it in a way that doesn't require expensive manual human handholding to teach the system how to personalize its answers. Okay, let's unpack this, because before we get into this new solution, we really need to establish why things are so clunky right now. Good idea. Under the hood, developers are already trying to personalize these models. So how does that current standard actually operate? Well, to understand the baseline we are working from, you have to look at the prompt engineering layer.

2:38A large language model only knows what is directly in its context window. Right. It's immediate memory for that specific chat. Exactly. So to make an LLM understand you, developers typically try to aggregate your interaction history. They take the things you've searched for, the links you've clicked, past prompts you've written, and they run that raw data through a summarization process. Okay, to build a profile. Right, to generate a natural language memory or a user profile. Ah, so behind the scenes, before I even see the system respond, they prepend a hidden paragraph to my prompt. If I ask, what are some good dinner ideas?

3:14The system is actually seeing a massive invisible preamble that says, like, this user is a vegetarian who loves spicy food and usually cooks in under 30 minutes. Now answer their question. That is the standard architecture, yeah. But the critical failure point isn't the prepending itself. It is how we train the AI to write that hidden user profile in the first place. Oh, okay. How so? Well, currently, to get a model to generate a genuinely accurate, useful profile of you, developers rely heavily on explicit rewards, and those are derived from labeled downstream tasks. Right there is the massive bottleneck.

3:49Yeah. Because relying on supervised learning for downstream tasks is just, it's an absolute nightmare for scalability. It really is. If I understand the current paradigm correctly, developers are having to define what good looks like for every single specific thing a user might ask. You've hit the nail on the head. Think about the sheer volume of human labor required for that supervised learning. It must be staggering. If a developer wants the AI to be good at recommending movies based on a user profile, human graders have to manually look at thousands of user profiles. And then they grade whether the resulting movie recommendations were actually accurate.

4:23Oh, wow. So if they want the AI to be good at drafting emails. Then humans have to manually grade the email writing task against the profile. You are demanding explicit supervised reward signals for the system to learn if the profile it generated was actually useful for that specific task. That is just mathematically impossible to scale. You cannot hire enough human annotators to cover the infinite number of tasks a user might throw at an AI. No, you really can't. It reminds me of hiring a human personal assistant. But instead of letting them naturally observe your preferences over time, you have to sit down and write a highly specific 100-page instruction manual for every single store they might ever visit on your behalf.

5:06Yeah, exactly. Like, here is the grading rubric for the grocery store. Here is the grading rubric for the dry cleaners. Here is the rubric for the post office. You are doing all the cognitive work. And honestly, I have to push back on the industry here. Isn't that the complete opposite of artificial intelligence? If we have to manually grade and supervise every single downstream task just so the machine can write a decent profile about us, where is the actual intelligence? It feels like massive human labor dressed up as autonomous code. That is a highly accurate critique of the last few years of AI development, honestly.

5:40We have been brute forcing intelligence with human annotation. Yeah, brute forcing is the perfect word for it. It is incredibly expensive, the resulting data is sparse, and it simply cannot scale to the complexity of human life. And that brings us to the core innovation of this new paper. It's a framework called BMP, which stands for Bidirectional User Modeling via Profiles. Okay, BMP. What's fascinating here is BMP is completely self-supervised. Self-supervised, meaning it completely severs the reliance on those human graders? Like, no explicit rewards for downstream tasks at all? None. Zero.

6:14BMP trains a profile generator without any downstream labels whatsoever. It completely sidesteps the human annotation bottleneck. Instead of relying on graders to say, yes, this profile helped write a better email, BMMP simply uses your raw interaction logs, just the pure unfiltered history of what you naturally do online. Okay, but this is where the mechanics get a bit tricky for me. How does a model learn from just raw logs without an external signal telling it what's good or bad? Yeah. You still need to train the neural network to actually emit the text of the profile. Right. The researchers tackle this by using a specific reinforcement learning method called GRPO.

6:52That stands for Group Relative Policy Optimization. Okay. They use this to train a large language model to emit a free-form textual profile based purely on your history. Let's dig into GRPO for a second because reinforcement learning usually still requires some kind of reward model, right? Okay. To tell the system if it's doing a good job. How is GRPO functioning here without a human-labeled reward model? So GRPO is incredibly efficient because it eliminates the need for a separate massive value model that usually doubles the compute cost in traditional reinforcement learning. Oh, I didn't realize that.

7:26Yeah. Instead of comparing a generated profile against some absolute external gold standard, GRPO samples a group of different generated profiles for the same user. Let's say it generates four different attempts at describing your personality based on your clicks. Okay, four drafts. Right. It evaluates all four, scores them, and then normalizes those scores relative to each other. The model essentially learns by looking at its own attempts and saying, oh, attempt C was much better than attempts A, B, and D. I will adjust my internal weights to generate more profiles like C. Okay, I follow the group relative part.

8:01It's self-correcting based on the group. But that brings up the million-dollar question. What is the scoring mechanism? When it compares attempts A, B, C, and D, how does the system actually test if attempt C is an accurate representation of me? Ah, to figure out if the generated profile is objectively accurate, B-U-M-P introduces a rigorously mathematical, brilliantly simple test. It evaluates the profile using a bidirectional in-batch ranking objective. Here's where it gets really interesting. Bidirectional in-batch ranking. Walk me through the first direction of this test. So, Direction 1 asks a very pointed question.

8:36Can the generated profile successfully pick out your specific interactions from a massive crowd? Okay. The framework employs a secondary, smaller LLM to act as a judge. This judge takes the textual profile BUMMP just generated for you, uses it as a search query, and looks at a massive data set of held-out interactions. And held-out just means data it hasn't seen yet. Exactly. Held out simply means real things people have typed or clicked on that the system hasn't trained on yet. The judge has to rank your specific interactions higher than the interactions of all the other users in that data set based solely on reading the text of your profile.

9:14Wait, so let me make sure I'm visualizing this right. It's almost like playing a massive high stakes game of guess who. Yeah, that's a good analogy. Let's say the AI generated a profile that says, this user is a software engineer who is obsessed with mid-century modern furniture and repairing vintage espresso machines. Okay, very specific. Right. In direction one, I hand the judge that specific personality profile and I point to a giant anonymous list of thousands of internet search histories. If the profile is accurate, the judge should be able to scan that massive list and say, ah, here is a search for compiling a Python script, followed by a search for a 1960s T credenza, followed by a search for replacing a gasket on a lever espresso machine.

9:57This anonymous search history must belong to this profile. It has to literally pick my anonymous history out of a lineup. That is a perfect visualization. To use a slightly more technical metaphor, it is like a cryptographic key in a lock. Oh, I like that. The profile is the key. The key has been cut correctly by the generator. It should flawlessly slide into and unlock your specific sequence of interactions out of thousands of wrong locks. That is direction one. Okay, so if it can find my needle in the haystack of anonymous search histories, I assume the true test of robustness is whether it works in reverse.

10:29Can a single search history point back to my profile? Is that direction two? You've deduced the exact logical flip. Yes. Direction two reverses the query and the target. Now, the judge takes just one single isolated interaction from your held-out history. Just one. Just one. Let's use that single search for replacing a gasket on a lever espresso machine. It uses that isolated interaction as the query. Then it looks at a massive lineup of different generated user profiles. The test is, can that single interaction successfully rank your true profile above the profiles of all the other users? Oh, that is a much harder test.

11:04Yeah. So if I hand the AI judge just one isolated crumb of my digital footprint, it needs to look at a massive lineup of distinct personalities and determine that the person who searched for this obscure espresso part is definitely the mid-century modern software engineer. Exactly. Not the teenage competitive gamer and not the retired accountant from Florida. Both directions must work perfectly. The profile must predict the isolated interactions, and the isolated interactions must perfectly match the ridges of the profile. If the generated profile passes this bidirectional test, the system mathematically knows it has created an accurate, robust representation of you.

11:41And again, notice what is completely absent from this process. There are zero human graders checking if it wrote a good email or suggested a good movie. None. The system is just validating the profile against the raw objective reality of what you actually clicked and typed. But I want to go back to the scoring mechanism for this matching game. You mentioned it scores the group of profiles and normalizes them. What is the actual mathematical metric the AI judge uses to keep score during this test? The judge evaluates both directions of this test using a specific metric called multipositive NDCG.

12:15NDCG. I'm going to need you to unpack that acronym because that sounds incredibly dense. It does. It stands for normalized discounted cumulative gain. It is a standard metric in information retrieval like search engines, but we need to break down the words to understand why it is so powerful here. Let's start with cumulative. Okay, let's start with cumulative. It is cumulative because the judge isn't just looking for one right answer. It adds up the scores for finding all of your held out interactions in the pile. Okay, cumulative makes sense. You get credit for every correct match. What about the discounted part?

12:48The discount is the penalty for being imprecise. NDCG cares deeply about ranking order. If the AI judge uses your profile and correctly identifies your interaction but ranks it at position number 10, buried beneath nine incorrect guesses, the system applies a heavy mathematical discount to the score. Ouch. Yeah. The lower down the list, the correct answer is the closer the score drops to zero. NDCG gives exponentially higher rewards if your interaction is ranked at absolute number one. That makes perfect sense. I mean, if a search engine puts the right answer on page two, it's essentially useless.

13:22Right. So it gets heavily discounted. And normalized. Normalized simply means the final score is scaled to fit between 0 and 1. That way, it can be easily compared across different users who might have wildly different amounts of data. Ah, got it. So the system calculates this NDCG score for direction 1, then calculates it for direction 2, and combines them into a single dense reward signal per training rollout. That dense reward is what feeds back into the GRPO reinforcement learning to make the profile generator smarter. The mechanics of the scoring are brilliant, but wait to calculate that NDCG score.

13:57You don't just need the right person. You need a massive board full of wrong people to compare them against to see how well it ranks. You do. In machine learning, gathering millions of wrong answers, what developers call negative examples, usually requires heavy compute to generate synthetic data or mine hard negatives. Where is BOMP getting all the wrong answers for this massive lineup? That brings us to what might be the most elegant compute optimization in the entire framework. They utilize something called in-batch ranking. Okay, in-batch ranking. When a large language model is training, it isn't processing one user at a time.

14:32It is processing a large batch, say, 256 or 512 users simultaneously in parallel matrices. Ah, I see where this is going. Yes. So where does the algorithm get the wrong answers to test your profile against? It simply leverages the other users currently sitting in that very same training batch. They provide what the researchers call free negatives. Free negatives. So if the system is processing my data, and it is concurrently processing the data of a college student in London and a chef in Tokyo in the exact same batch, it just points to the student's data and the chef's data as the wrong answers for my test.

15:10And conversely, my data automatically becomes the wrong answer for their tests. Exactly. The researchers completely eliminate the computational overhead of generating synthetic data sets or mining for negative examples. The data is already loaded into the GPU memory for the current training cycle. That is so smart. It is. So by using a batch size of, say, 256, every single user gets one positive match in 255 mathematically free negative matches. Every single training example yields rich supervision from raw interaction logs alone without adding any extra compute cost for negative sampling. That is just breathtaking efficiency.

15:48The system gets smarter simply by being in a crowded room of data, by bouncing everyone's information off of everyone else's. Every single user serves as the perfect foil for the person standing next to them. Perfectly said. You don't have to build a synthetic environment. The reality of multiple distinct users is enough to train the system. So the end result is an AI that learns to perfectly profile users completely on its own with no expensive human labeling just by comparing us to each other. That is the ultimate promise of self-supervised learning. And we know this specific bidirectional approach works because the researchers didn't just propose the math.

16:23They proved it. Oh, they were in tests. Yes. They rigorously evaluated the BMP framework on the LAMP benchmark. Let's talk about the proof. The LAMP benchmark that stands for language model personalization, right? Right. What exactly does that test? The LAMAPI benchmark is a comprehensive suite designed to evaluate how well an LLM can adapt to a user. It includes subtasks like personalized citation generation, where the AI has to predict which academic papers a specific user would cite based on their past publications. Oh, that's incredibly specific. It is. It also includes personalized movie tagging, predicting how a specific user would categorize a film.

16:59It includes generating personalized emails and news summaries. It is a very rigorous gauntlet. And how did BMP fare against the traditional models that use all those expensive human-labeled downstream rewards? The results were definitive. BMP didn't just work as a conceptual theory. It matched or outright outperformed state-of-the-art methods and closed-source APIs that heavily rely on those incredibly expensive labeled data sets. It achieved top-tier personalization performance across the LAMP benchmark while requiring absolutely zero task labels during training. The profile generator, trained purely on self-supervised raw logs, created a better context window than systems trained by armies of human annotators.

17:41So what does this all mean? We have a new mathematical framework to train AI to deeply understand us. It is vastly faster, it completely circumvents the need for human graders, it drastically lowers compute costs using free negatives, and it demonstrably beats the old expensive methods on rigorous benchmarks. It checks all the boxes. What is the actual tangible impact for the listener who's just tired of their AI sounding like a generic robot? If we connect this to the bigger picture, the most profound impact is a radical democratization of personalized AI. Because Bon & P removes the massive financial and temporal cost of human task labeling, deeply personalized systems are about to become incredibly cheap to develop.

18:22Because right now that's totally gated. Yes. Over the last few years, only the giant tech monopolies have had the capital resources to hire armies of human annotators to build highly personalized systems. But with a self-supervised framework like BMP, you don't need an army. And you don't need billions of dollars in reinforcement learning infrastructure. You just need the raw interaction logs. Which means smaller developers, independent researchers, and open source communities can finally build AI models that actually understand us on an intimate level. Precisely. It paves the way for AI that instantly adapts to your unique workflow, your specific tone, and your individual idiosyncrasies.

19:04And because the training pipeline is so much cheaper and lighter, this level of personalization can increasingly happen locally. Locally. Like right on my laptop. Exactly. Imagine a model running entirely on your local device, learning your profile through BMP's bi-directional testing, rather than requiring your private interaction data to be sent off to a tightly controlled tech monopoly server just to figure out how you like your emails drafted. It rests the technology away from the center and puts it back in the hands of the user. That is a fundamental paradigm shift. We are moving from AI being this rigid top-down utility where we have to constantly bend to its generic nature to an AI that is fluid, bottom-up, and naturally molds itself around the specific contours of our digital lives.

19:48They're beautiful, really. It is. It is the difference between buying a mass produced suit off the rack that kind of fits everybody but looks terrible on you and having a master tailor who silently observes exactly how you move and stitches something perfectly to your unique dimensions. That's a great analogy. And all of it happens behind the scenes, mathematically, securely, just by analyzing the raw history of how you interact with the world. And, well, stepping back from the computer science of it all, the success of the BMMP framework raises a lingering, almost psychological question for us to consider.

20:19Oh. This framework proves that an algorithm without any human guidance or explicit context can look at a scattered, seemingly chaotic history of isolated clicks, searches, and interactions and perfectly map them to a cohesive core profile of who you are. The bidirectional test works. It's a bit unsettling when you phrase it like that. The isolated interactions flawlessly predict the profile and the profile flawlessly predicts the interactions. We like to think of ourselves as so spontaneous, you know, complex, and totally unpredictable in our daily lives. But the math in this research suggests otherwise.

20:55If our digital footprints are so internally cohesive that a self-supervised algorithm can map them flawlessly using just a cryptographic key and lock mechanism, we have to ask ourselves, are human personalities just highly predictable statistical patterns? Wow. We view our daily choices as independent, spontaneous events, like what article we decide to read, what specific product we search for, how we phrase a subtle question. But to the system, they are just inevitable expressions of a baseline equation. That is a heavy thought. And it makes you wonder what happens when the self-supervised profile, generated entirely by a machine, knows what you are going to click, search, or type better than your conscious mind does.

21:35If the algorithm can define your parameters so clearly, who is actually steering the ship when you sit down at your keyboard? Well, on that note, we are going to wrap up our deep dive for today. Thank you for joining us on this intellectual journey, unpacking the intricate mechanics of how these advanced systems are learning to map the contours of our behavior without us even realizing it. We will leave you to explore that final question on your own and perhaps take a closer look at your own digital footprint to see just how predictable you might be. Until next time.

From the publisher

This paper describes a self-supervised framework called BUMP, which is designed to improve how large language models deliver personalized content. Traditionally, creating user profiles for search and recommendation tasks requires expensive, human-labeled data to train the system. To solve this, researchers developed a method that uses a bidirectional ranking objective to learn directly from raw interaction logs without manual supervision. By comparing a user's generated profile against their actual history, the system creates a dense reward to refine the model's accuracy. This approach allows the AI to summarize interaction histories into natural language descriptions that are as effective as those produced by more costly, supervised methods. Ultimately, the source demonstrates that personalization can be achieved efficiently by training models to recognize the unique patterns in a user's own digital footprint.

More from Best AI papers explained

All 475 episodes
Self-supervised User Profile Generation for PersonalizationBest AI papers explained · 22 min
Listen in VO