Best AI papers explained cover art
Podcast · 475 episodes

Best AI papers explained, page 9

by Enoch H. Kang · English

Cut through the noise. We curate and break down the most important AI papers so you don’t have to.

All episodes, page 9

PPI-SVRG: Unifying Prediction-Powered Inference and Variance Reduction for Semi-Supervised OptimizationSemi-supervised optimization that unifies Prediction-Powered Inference (PPI) with Stochastic Variance Reduced Gradient (SVRG) to reduce variance using abundant unlabeled data plus a small labeled set.5 Feb 2026 · 17 min · 8 chapters
When Models Don’t Collapse: On the Consistency of Iterative MLEWhether “model collapse” (AI trained on its own outputs degrading in a feedback loop) is inevitable for iterative Maximum Likelihood Estimation (MLE), and what mathematical conditions prevent or…3 Feb 2026 · 16 min · 6 chapters
An orthogonal learner for individualized outcomes In markov decision processesOff-policy reinforcement learning for personalized, long-term treatment strategies modeled as Markov decision processes (MDPs), addressing “The Curse of the Horizon” (compounding error over long time…3 Feb 2026 · 18 min · 11 chapters
Shaping capabilities with token-level data filteringThe episode argues that AI safety should shift from post-training “muzzles” (e.g., RLHF) to proactive token-level data filtering, especially for dual-use biology/bioweapon risk.1 Feb 2026 · 12 min · 7 chapters
Self-Improving Pretraining: using post-trained models to pretrain better modelsThe episode explains “self-improving pretraining,” a method to train better AI models without relying on ever more high-quality human data, aiming to overcome the “data wall” around 2026.1 Feb 2026 · 15 min · 9 chapters
Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success“Success conditioning” (selective imitation of past successes) as a mathematically exact, conservative optimization method for AI alignment and learning, explained via trust-region geometry and…31 Jan 2026 · 20 min · 11 chapters
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM ReasoningThe episode explains Trajectory Bellman Residual Minimization (TBRM), a value-based training method for LLM reasoning that replaces step-by-step “micromanagement” with an objective that minimizes…31 Jan 2026 · 18 min · 8 chapters
GameTalk: Training LLMs for Strategic Multi-Turn ConversationUniversity of Cambridge “GameTalk” research on training LLMs as strategic multi-turn agents that can bluff and negotiate, including deception without explicit programming.30 Jan 2026 · 16 min · 8 chapters
Reinforcement Learning via Self-DistillationReinforcement Learning via Self-Distillation (SDPO) replaces sparse pass/fail rewards with “rich feedback” from verifiers (e.g., compiler errors) and uses self-distillation to turn error text into…30 Jan 2026 · 14 min · 7 chapters
Self-Supervised Contrastive Learning is Approximately Supervised Contrastive LearningSelf-supervised contrastive learning (no labels) can approximate supervised contrastive learning, because “false negatives” vanish as the number of classes grows; the learned representations exhibit…28 Jan 2026 · 15 min · 6 chapters
On the alignment between supervised and self-supervised contrastive learningExplains why self-supervised contrastive learning (CL) without labels can still produce semantic categories, and shows a theoretical/empirical link to a supervised objective called NSCL…28 Jan 2026 · 16 min · 11 chapters
Rethinking the value of multi-agent work-flow: a strong single agent baselineChallenges “agentic AI” hype by arguing that many multi-agent workflows are inefficient clones.24 Jan 2026 · 17 min · 10 chapters
Greedy Sampling Is Provably Efficient for RLHFRLHF training for LLMs argues that “greedy sampling” (always pick the currently best option from observed data) is provably efficient, removing the need for expensive exploration bonuses, because KL…24 Jan 2026 · 13 min · 6 chapters
A Generalization Theory for Zero-Shot PredictionExplains zero-shot prediction (ZSP) for vision-language models (e.g., CLIP) using “indirect prediction path” theory: image-to-text-to-label, and why it works or fails.24 Jan 2026 · 15 min · 8 chapters
Learning to Discover at Test TimeTest-time training for “discovery” (TTT Discover) where an AI updates its weights during problem solving instead of only searching with a frozen model.23 Jan 2026 · 16 min · 10 chapters
How Does the Pretraining Distribution Shape In-Context Learning? Task Selection, Generalization, and RobustnessExplains how a model’s pretraining distribution shapes in-context learning, focusing on task selection vs generalization and robustness under distribution shifts.23 Jan 2026 · 19 min · 8 chapters
Highlighting What Matters: Promptable Embeddings for Attribute-Focused RetrievalWhy CLIP-style image-text retrieval fails on specific attributes (e.g., “zero people,” “brick vs stone,” “morning vs afternoon”), and how promptable (question-conditioned) image embeddings fix it by…20 Jan 2026 · 14 min · 9 chapters
Activation Reward Models for Few-Shot Model AlignmentActivation Reward Models (ActivationRMs) for aligning few-shot models by steering internal activations toward “safety/truth,” avoiding heavy RLHF or gullible LLM-judge scoring.20 Jan 2026 · 16 min · 8 chapters
Reward is enough: LLMs are in-context reinforcement learnersIn-context reinforcement learning (ICRL) for LLMs: “reward is enough” to enable test-time learning inside the conversation history, without changing model weights.19 Jan 2026 · 11 min · 6 chapters
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPOThe episode analyzes why RLHF (reinforcement learning from human feedback) and DPO (direct preference optimization) can have different performance, focusing on “performance gaps” under model…19 Jan 2026 · 14 min · 9 chapters
The End of Reward Engineering: How LLMs Are Redefining Multi-Agent CoordinationThe “reward engineering” failure mode in reinforcement learning—robots exploit poorly specified reward functions (e.g., dropping screws repeatedly) and multi-agent training breaks due to credit…18 Jan 2026 · 18 min · 12 chapters
PRL: Process Reward Learning Improves LLMs’ Reasoning Ability and Broadens the Reasoning BoundaryProcess Reward Learning (PRL) trains LLMs to improve reasoning by giving dense, step-by-step rewards instead of only rewarding final answers.18 Jan 2026 · 16 min · 10 chapters
Coverage Improvement and Fast Convergence of On-policy Preference LearningThe episode explains how “on-policy preference learning” can improve large language models faster than offline DPO by using self-generated, scored data.17 Jan 2026 · 15 min · 11 chapters
Stagewise Reinforcement Learning and the Geometry of the Regret LandscapeStagewise reinforcement learning; why agents improve in sudden “jumps” due to a tradeoff between regret (performance) and internal complexity (geometry of the regret landscape), extending singular…16 Jan 2026 · 13 min · 3 chapters