vo
Podcasts
Search podcasts and episodes
Use with Claude or ChatGPT
Get the app
Podcasts
/
Technology
/
Best AI papers explained
Podcast · 475 episodes
Best AI papers explained
, page 2
by
Enoch H. Kang
· English
Cut through the noise. We curate and break down the most important AI papers so you don’t have to.
Technology
Follow in VO
All episodes, page 2
Search Best AI papers explained episodes
Search
Q-Learning with World Models
QWM (Q-Learning with World Models) for faster, safer robotic learning by combining model-free Q-learning with short-horizon world-model lookahead to avoid compounding simulation bias.
23 Aug 2026 · 25 min · 15 chapters
Conformal Language Modeling via Posterior Sampling
Conformal language modeling via posterior sampling, a MIT statistical method to reduce LLM hallucinations during text generation by conditioning token sampling on a calibrated “high-trust” region,…
20 Aug 2026 · 22 min · 10 chapters
BoNVoyage: Learning Better Rewards without Ranking
Explains why RLHF reward models trained with pairwise rankings (Bradley-Terry) fail under adversarial distribution shift, enabling reward hacking, and presents Bon Voyage, a method to train reward…
20 Aug 2026 · 22 min · 10 chapters
Demystifying Agent Skills: Why They Work—Until They Don’t
Why AI “agent skills” (distilled procedural checklists) improve performance, and when they fail.
18 Aug 2026 · 20 min · 12 chapters
Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence
Epistemic stability of AI “judges” (LLMs used for grading, moderation, and reward modeling) under silence, pressure, and persistence; argues that standard accuracy tests are misleading because models…
15 Aug 2026 · 20 min · 11 chapters
Predicting Neural Scaling Laws without Training: A Data Manifold Oracle
The episode explains the “Data Manifold Oracle” (DMO), a training-free method to predict neural scaling laws (the “floor” error limit and the “slope” learning rate) from raw text using standard…
15 Aug 2026 · 22 min · 12 chapters
Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing
A July 2026 research framework for making interpretable, bias-resistant discoveries from unstructured text using sparse autoencoders, high-dimensional multiple-hypothesis testing (KFWER), Gaussian…
11 Aug 2026 · 24 min · 14 chapters
Overcoming the Incentive Collapse Paradox
The “incentive collapse paradox” in AI-human workflows: if pay depends only on final accuracy, near-perfect AI makes workers rationally free-ride (exert less effort), potentially requiring infinite…
11 Aug 2026 · 21 min · 11 chapters
Position: Modular Memory is the Key to Continual Learning Agents
Continual learning agents that learn over a lifetime without catastrophic forgetting, using modular memory (core model + working memory + long-term memory) and two regimes: external interaction and…
10 Aug 2026 · 27 min · 14 chapters
Harness RL is Meta-Learning: Training to Self-Improve at Test Time
Harness RL for meta-learning/self-improvement at test time by moving adaptation from updating model weights to revising the “harness” (instructions, memory, tool rules, verification).
8 Aug 2026 · 22 min · 9 chapters
Escaping the Nash Trap: Structural Estimation and Alignment of Strategic Reasoning in Large Language Models
The episode argues that large language models often fail in strategic settings because they assume opponents are perfectly rational “Nash-type” optimizers, creating “Nash traps” where the…
7 Aug 2026 · 22 min · 10 chapters
When Does LeJEPA Learn a World Model?
When LeJEPA (Joint Embedding Predictive Architecture) learns a reliable “world model,” i.e., linearly identifiable latent variables from pixel observations, and when that guarantee fails in real…
7 Aug 2026 · 23 min · 17 chapters
Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems
The episode explains “role drift” in compound LLM systems trained with outcome-only reinforcement learning, where modules abandon their intended roles while still achieving high final accuracy.
3 Aug 2026 · 21 min · 13 chapters
Do you really need to pretrain Q-functions for online RL fine-tuning?
Whether online reinforcement learning fine-tuning needs a pre-trained Q-function (critic) and why naive offline Q pretraining can hurt.
1 Aug 2026 · 24 min · 10 chapters
The Evolution of Digital Search: From Blue Links to Delegated Decision-Making
The episode argues that traditional keyword “blue link” search is dying and being replaced by AI-native delegated decision-making, where consumer agents interpret intent, evaluate options invisibly,…
29 Jul 2026 · 19 min · 8 chapters
Ask, Don’t Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement
How to evaluate LLM outputs reliably without opaque “holistic” scores, using binary yes/no checklists (Binavol) to catch specific errors and enable automated self-correction.
28 Jul 2026 · 5 min · 4 chapters
Understanding Reasoning from Pretraining to Post-Training
How reasoning emerges from the standard AI pipeline (pretraining + supervised fine-tuning + reinforcement learning), using chess and math as controlled testbeds; includes scaling laws, compute…
24 Jul 2026 · 22 min · 12 chapters
A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
Whether LLM “self-explanations” (the reasons they give alongside answers) are faithful to their actual internal decision logic, and whether explanations improve prediction of future behavior.
23 Jul 2026 · 15 min · 9 chapters
Reject, Resample, Repeat: Understanding Parallel Reasoning in Language Model Inference
How inference-time “parallel reasoning” works for language models using sequential Monte Carlo (SMC) with process reward models (PRMs), including math guarantees and why theory can fail on real…
19 Jul 2026 · 23 min · 11 chapters
Rethinking the Evaluation of Harness Evolution for Agents
Evaluates “automatic harness evolution” for AI agents—whether letting a model rewrite its prompts/tools/control logic actually makes it smarter, or just exploits benchmark evaluation.
19 Jul 2026 · 23 min · 13 chapters
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning
Compositional generalization in language-model reasoning—why supervised fine-tuning (SFT) can make models rigid, and how reinforcement learning (RL) “untangles” reasoning into reusable skill and…
18 Jul 2026 · 19 min · 11 chapters
Position: Interpretability can be actionable
The episode argues that AI interpretability must become actionable—moving from passive “black box” observation to concrete interventions that improve real-world behavior, safety, and usability.
17 Jul 2026 · 25 min · 12 chapters
High-accuracy sampling for diffusion models and log-concave distributions
A 2026 theoretical breakthrough on high-accuracy sampling for diffusion models and log-concave distributions using only first-order information (gradients), avoiding the “polynomial wall” of standard…
17 Jul 2026 · 22 min · 9 chapters
Causal Inference with Video Features as Treatments
Causal inference for video persuasion—isolating the frame-by-frame causal effect of a fleeting visual feature (treated as a “treatment”) on changing human emotion, while correcting for confounding…
15 Jul 2026 · 22 min · 11 chapters
Previous
1
2
3
…
20
Next
Newest first · 24 per page