vo
Podcasts
Search podcasts and episodes
Use with Claude or ChatGPT
Get the app
Podcasts
/
Technology
/
Best AI papers explained
Podcast · 475 episodes
Best AI papers explained
, page 4
by
Enoch H. Kang
· English
Cut through the noise. We curate and break down the most important AI papers so you don’t have to.
Technology
Follow in VO
All episodes, page 4
Search Best AI papers explained episodes
Search
Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings
How pairwise comparisons can rank AI model accuracy even when the “judge” model lacks subject-matter knowledge, using hidden structural signals rather than formatting.
13 Jun 2026 · 20 min · 9 chapters
Critical Batch Size for LLM Policy Optimization
The episode explains “critical batch size” for reinforcement learning from verifiable rewards in LLM post-training, focusing on GRPO (Group Relative Policy Optimization), and how it determines when…
11 Jun 2026 · 19 min · 11 chapters
Self-supervised User Profile Generation for Personalization
Self-supervised generation of personalized AI user profiles for assistants, avoiding prompt “blank slate” behavior and expensive human labeling.
9 Jun 2026 · 22 min · 11 chapters
From Augmentation to Reconstruction: Guiding the AI Disruption to the Good Place
The episode argues that AI’s “disruption” hasn’t hit daily life yet because society is still using human-era workflows.
7 Jun 2026 · 22 min · 11 chapters
Self-Distilled Agentic Reinforcement Learning
Training LLM agents for long multi-step tasks without compounding errors.
7 Jun 2026 · 22 min · 10 chapters
Subliminal Learning Is Steering Vector Distillation
Subliminal learning in AI—how a student model can inherit a behavioral trait (e.g., “love owls”) from a teacher even when the teacher outputs only sanitized random digits, explained as steering…
5 Jun 2026 · 23 min · 8 chapters
Subsidizing Sequential Search
How “search subsidies” work in directed search, and how AI assistants in an “agentic economy” tokenize search costs, letting platforms profit by inducing excess (computationally inefficient)…
5 Jun 2026 · 20 min · 11 chapters
Meta-Harness: End-to-End Optimization of Model Harnesses
Meta-Harness argues AI “harnesses” (context managers, memory/tool orchestration code) matter as much as the frozen LLM weights, and can be end-to-end optimized by letting an agent debug itself using…
2 Jun 2026 · 18 min · 10 chapters
Self-Improving Language Models with Bidirectional Evolutionary Search
Bidirectional Evolutionary Search (BES) to help language models escape “entropy shell” failure modes in hard, creative, multi-step reasoning by combining evolutionary operators with dense goal…
1 Jun 2026 · 21 min · 11 chapters
Generative Modeling via Drifting
The episode explains “drifting models” for generative modeling that remove iterative inference-time denoising (a “1NFE” single network pass) by learning an antisymmetric attraction/repulsion…
31 May 2026 · 22 min · 10 chapters
Instance-Optimal Estimation with Multiple LLM Judges on a Budget
Budget-constrained evaluation of an LLM using multiple “LLM-as-a-judge” models.
31 May 2026 · 21 min · 12 chapters
Robust AI Personalization Will Require a Human Context Protocol
The episode argues that today’s AI personalization is surveillance-based and often misreads user intent due to inferred-preference “inversion,” siloed data “fragmentation,” and resulting AI…
29 May 2026 · 23 min · 13 chapters
Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning
CMU “equilibrium reasoners” (EQR) for scalable reasoning by treating the model’s latent space like a terrain where iterative “marbles” roll toward an attractor; correct answers correspond to the…
27 May 2026 · 18 min · 8 chapters
Position: The Pre/Post-Training Boundary Should Govern IP in Industry–Academia ML Collaborations
How to resolve IP deadlock in industry–academia ML collaborations using a contract template (PBOS) built around a “pre/post-training boundary.”
25 May 2026 · 13 min · 7 chapters
MEMO: Memory as a Model
MEMO (“Memory as a Model”) proposes separating an AI’s frozen reasoning (“executive model”) from a swappable, dedicated memory (“memory model”) so systems can learn new facts without retraining or…
24 May 2026 · 18 min · 10 chapters
Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces
Multi-agent AI marketplaces can crash or rot without “economic alignment.” The episode uses Agent Bazaar simulations to show (1) B2C price-war bankruptcies and (2) C2C “lemon market” fraud via Sybil…
23 May 2026 · 23 min · 14 chapters
General Preference Reinforcement Learning
General Preference Reinforcement Learning (GPRL) argues that AI alignment fails when training uses a single scalar reward (like RLHF reward models), enabling reward hacking via verbosity and other…
23 May 2026 · 22 min · 11 chapters
Explaining and Preventing Alignment Collapse in Iterative RLHF
How iterative RLHF can trigger “alignment collapse,” where an AI exploits reward-model blind spots and retraining amplifies the deception; proposes Foresighted Policy Optimization (FPO) to prevent it…
21 May 2026 · 21 min · 10 chapters
Curriculum Learning-Guided Progressive Distillation in Large Language Models
Curriculum Learning-Guided Progressive Distillation (CLPD) for training small LLMs, arguing that “genius teachers” can harm students unless teacher strength is coupled to data difficulty.
19 May 2026 · 16 min · 7 chapters
Think Twice, Act Once: Verifier-Guided Action Selection For Embodied Agents
Why embodied robots fail despite strong multimodal “digital genius,” and how VEGAS (Verifier-Guided Action Selection) improves reliability by replacing greedy decoding with verifier-guided…
19 May 2026 · 26 min · 12 chapters
How Much Should a Conversational Recommender System Converse?
How conversational AI shopping assistants decide how many clarifying questions to ask, balancing “preference matching” against user “abandonment hazard,” and how this differs by product type and…
17 May 2026 · 22 min · 11 chapters
FUSE: Ensembling Verifiers with Zero Labeled Data
The episode explains FU-SE (Fully Unsupervised Score Ensembling), a Stanford/Google method for ranking AI-generated answers using multiple AI verifiers without any labeled data or answer keys.
14 May 2026 · 20 min · 8 chapters
EVOLM: Self-Evolving Language Models through Co-Evolved Discriminative Rubrics
EVOLM (evolutionary language model) trains a language model to generate its own grading rubrics via co-evolution, avoiding human-written or opaque “scalar reward” grading.
14 May 2026 · 23 min · 14 chapters
Personalized Alignment Revisited: The Necessity and Sufficiency of User Diversity
Why AI personalization can fail even when users provide their personal data; the episode argues that success requires “decision-relevant user diversity” (full-rank variation in user-specific…
12 May 2026 · 22 min · 14 chapters
Previous
1
…
3
4
5
…
20
Next
Newest first · 24 per page