vo
Podcasts
Search podcasts and episodes
Use with Claude or ChatGPT
Get the app
Podcasts
/
Technology
/
Best AI papers explained
Podcast · 475 episodes
Best AI papers explained
, page 9
by
Enoch H. Kang
· English
Cut through the noise. We curate and break down the most important AI papers so you don’t have to.
Technology
Follow in VO
All episodes, page 9
Search Best AI papers explained episodes
Search
PPI-SVRG: Unifying Prediction-Powered Inference and Variance Reduction for Semi-Supervised Optimization
Semi-supervised optimization that unifies Prediction-Powered Inference (PPI) with Stochastic Variance Reduced Gradient (SVRG) to reduce variance using abundant unlabeled data plus a small labeled set.
5 Feb 2026 · 17 min · 8 chapters
When Models Don’t Collapse: On the Consistency of Iterative MLE
Whether “model collapse” (AI trained on its own outputs degrading in a feedback loop) is inevitable for iterative Maximum Likelihood Estimation (MLE), and what mathematical conditions prevent or…
3 Feb 2026 · 16 min · 6 chapters
An orthogonal learner for individualized outcomes In markov decision processes
Off-policy reinforcement learning for personalized, long-term treatment strategies modeled as Markov decision processes (MDPs), addressing “The Curse of the Horizon” (compounding error over long time…
3 Feb 2026 · 18 min · 11 chapters
Shaping capabilities with token-level data filtering
The episode argues that AI safety should shift from post-training “muzzles” (e.g., RLHF) to proactive token-level data filtering, especially for dual-use biology/bioweapon risk.
1 Feb 2026 · 12 min · 7 chapters
Self-Improving Pretraining: using post-trained models to pretrain better models
The episode explains “self-improving pretraining,” a method to train better AI models without relying on ever more high-quality human data, aiming to overcome the “data wall” around 2026.
1 Feb 2026 · 15 min · 9 chapters
Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success
“Success conditioning” (selective imitation of past successes) as a mathematically exact, conservative optimization method for AI alignment and learning, explained via trust-region geometry and…
31 Jan 2026 · 20 min · 11 chapters
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
The episode explains Trajectory Bellman Residual Minimization (TBRM), a value-based training method for LLM reasoning that replaces step-by-step “micromanagement” with an objective that minimizes…
31 Jan 2026 · 18 min · 8 chapters
GameTalk: Training LLMs for Strategic Multi-Turn Conversation
University of Cambridge “GameTalk” research on training LLMs as strategic multi-turn agents that can bluff and negotiate, including deception without explicit programming.
30 Jan 2026 · 16 min · 8 chapters
Reinforcement Learning via Self-Distillation
Reinforcement Learning via Self-Distillation (SDPO) replaces sparse pass/fail rewards with “rich feedback” from verifiers (e.g., compiler errors) and uses self-distillation to turn error text into…
30 Jan 2026 · 14 min · 7 chapters
Self-Supervised Contrastive Learning is Approximately Supervised Contrastive Learning
Self-supervised contrastive learning (no labels) can approximate supervised contrastive learning, because “false negatives” vanish as the number of classes grows; the learned representations exhibit…
28 Jan 2026 · 15 min · 6 chapters
On the alignment between supervised and self-supervised contrastive learning
Explains why self-supervised contrastive learning (CL) without labels can still produce semantic categories, and shows a theoretical/empirical link to a supervised objective called NSCL…
28 Jan 2026 · 16 min · 11 chapters
Rethinking the value of multi-agent work-flow: a strong single agent baseline
Challenges “agentic AI” hype by arguing that many multi-agent workflows are inefficient clones.
24 Jan 2026 · 17 min · 10 chapters
Greedy Sampling Is Provably Efficient for RLHF
RLHF training for LLMs argues that “greedy sampling” (always pick the currently best option from observed data) is provably efficient, removing the need for expensive exploration bonuses, because KL…
24 Jan 2026 · 13 min · 6 chapters
A Generalization Theory for Zero-Shot Prediction
Explains zero-shot prediction (ZSP) for vision-language models (e.g., CLIP) using “indirect prediction path” theory: image-to-text-to-label, and why it works or fails.
24 Jan 2026 · 15 min · 8 chapters
Learning to Discover at Test Time
Test-time training for “discovery” (TTT Discover) where an AI updates its weights during problem solving instead of only searching with a frozen model.
23 Jan 2026 · 16 min · 10 chapters
How Does the Pretraining Distribution Shape In-Context Learning? Task Selection, Generalization, and Robustness
Explains how a model’s pretraining distribution shapes in-context learning, focusing on task selection vs generalization and robustness under distribution shifts.
23 Jan 2026 · 19 min · 8 chapters
Highlighting What Matters: Promptable Embeddings for Attribute-Focused Retrieval
Why CLIP-style image-text retrieval fails on specific attributes (e.g., “zero people,” “brick vs stone,” “morning vs afternoon”), and how promptable (question-conditioned) image embeddings fix it by…
20 Jan 2026 · 14 min · 9 chapters
Activation Reward Models for Few-Shot Model Alignment
Activation Reward Models (ActivationRMs) for aligning few-shot models by steering internal activations toward “safety/truth,” avoiding heavy RLHF or gullible LLM-judge scoring.
20 Jan 2026 · 16 min · 8 chapters
Reward is enough: LLMs are in-context reinforcement learners
In-context reinforcement learning (ICRL) for LLMs: “reward is enough” to enable test-time learning inside the conversation history, without changing model weights.
19 Jan 2026 · 11 min · 6 chapters
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
The episode analyzes why RLHF (reinforcement learning from human feedback) and DPO (direct preference optimization) can have different performance, focusing on “performance gaps” under model…
19 Jan 2026 · 14 min · 9 chapters
The End of Reward Engineering: How LLMs Are Redefining Multi-Agent Coordination
The “reward engineering” failure mode in reinforcement learning—robots exploit poorly specified reward functions (e.g., dropping screws repeatedly) and multi-agent training breaks due to credit…
18 Jan 2026 · 18 min · 12 chapters
PRL: Process Reward Learning Improves LLMs’ Reasoning Ability and Broadens the Reasoning Boundary
Process Reward Learning (PRL) trains LLMs to improve reasoning by giving dense, step-by-step rewards instead of only rewarding final answers.
18 Jan 2026 · 16 min · 10 chapters
Coverage Improvement and Fast Convergence of On-policy Preference Learning
The episode explains how “on-policy preference learning” can improve large language models faster than offline DPO by using self-generated, scored data.
17 Jan 2026 · 15 min · 11 chapters
Stagewise Reinforcement Learning and the Geometry of the Regret Landscape
Stagewise reinforcement learning; why agents improve in sudden “jumps” due to a tradeoff between regret (performance) and internal complexity (geometry of the regret landscape), extending singular…
16 Jan 2026 · 13 min · 3 chapters
Previous
1
…
8
9
10
…
20
Next
Newest first · 24 per page