vo
Podcasts
Search podcasts and episodes
Use with Claude or ChatGPT
Get the app
Podcasts
/
Technology
/
Best AI papers explained
Podcast · 475 episodes
Best AI papers explained
, page 8
by
Enoch H. Kang
· English
Cut through the noise. We curate and break down the most important AI papers so you don’t have to.
Technology
Follow in VO
All episodes, page 8
Search Best AI papers explained episodes
Search
Prescriptive Scaling Reveals the Evolution of Language Model Capabilities
Prescriptive scaling in AI—using a “capability boundary” (an S-curve) to predict model performance from pre-training compute, and to diagnose whether gains come from real reasoning or from issues…
24 Feb 2026 · 18 min · 8 chapters
Experiential Reinforcement Learning
Experiential Reinforcement Learning (ERL), a training paradigm that replaces sparse-reward “button-mashing” with a human-like loop: try, reflect, retry, and store lessons across episodes; then remove…
23 Feb 2026 · 23 min · 8 chapters
Learning Personalized Agents from Human Feedback
Personalized agents from human feedback (PAHF) that adapt to changing user preferences using a read-write memory loop, avoiding “static personalization” failures like preference drift.
21 Feb 2026 · 15 min · 7 chapters
Learning to summarize user information for personalized RLHF
Explains why RLHF personalization can produce “vanilla/beige” responses, then argues PLUS (preference learning via natural-language summarization) fixes it by replacing vector user profiles with…
20 Feb 2026 · 18 min · 11 chapters
Intrinsic Credit Assignment for Long Horizon Interaction
The episode explains the paper “Intrinsic Credit Assignment for Long Horizon Interaction,” focusing on teaching AI agents “curiosity” via Delta Belief Reinforcement Learning (Delta Belief RL).
20 Feb 2026 · 18 min · 10 chapters
Learning to Continually Learn via Meta-learning Agentic Memory Designs
The “Groundhog Day” problem in agentic AI—stateless foundation models lose context between sessions—and how ALMA (Automated Meta Learning of Memory Designs for Agentic Systems) uses a meta-agent to…
20 Feb 2026 · 20 min · 13 chapters
Why Self-Rewarding Works: Theoretical Guarantees for Iterative Alignment of Language Models
The episode explains a mathematical proof for “self-rewarding language models” (SRLMs): models that act as both policy (generator) and reward model (judge), aiming to iteratively align/improve…
19 Feb 2026 · 18 min · 9 chapters
PAD: Personalized Alignment of LLMs at Decoding-Time
PAD (Personalized Alignment of LLMs at Decoding Time) tackles the “alignment tax” where RLHF-trained LLMs drift to a safe, average “committee” voice.
19 Feb 2026 · 14 min · 9 chapters
The Reward Model Selection Crisis in Personalized Alignment
Personalized alignment via reward models (RLHF) is flawed because reward-model ranking accuracy often doesn’t translate into better generation; self-evaluation enables reward hacking and “circular…
19 Feb 2026 · 16 min · 7 chapters
Causal-JEPA: Learning World Models through Object-Level Latent Interventions
Causal-JEPA (C-JEPA) for learning world models by forcing object-level causal reasoning instead of “cheating” via pixel interpolation.
18 Feb 2026 · 15 min · 9 chapters
How Sampling Shapes LLM Alignment: From One-Shot Optima to Iterative Dynamics
How the “sampling problem” in preference-based LLM alignment (choosing which answer pairs to compare) shapes model behavior, and how iterative alignment feedback loops can cause collapse or…
17 Feb 2026 · 16 min · 8 chapters
Deriving neural scaling laws from the statistics of natural language
Explains a research breakthrough claiming neural scaling laws can be derived from statistical properties of natural language, not from model architecture.
15 Feb 2026 · 19 min · 7 chapters
Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL
Reasoning Cache, a short-horizon RL-trained method that lets small autoregressive models “pause, summarize, delete, and continue,” extending effective reasoning from ~16k tokens to ~512k while…
15 Feb 2026 · 15 min · 6 chapters
Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL
Orbit, a framework for online reinforcement-based in-context learning that lets LLM agents learn from failures across episodes (solving the “static aftershipping problem”/amnesia between chats).
14 Feb 2026 · 15 min · 13 chapters
Divide-and-Conquer CoT: RL for Reducing Latency via Parallel Reasoning
The episode explains “divide-and-conquer chain-of-thought” (DC-CoT), an RL-trained method that reduces LLM latency by parallelizing reasoning.
12 Feb 2026 · 16 min · 7 chapters
Owning the AI Pareto Frontier — Jeff Dean
How to “own the AI Pareto frontier” by balancing capability vs cost/latency, using distillation, sparsity, low energy/data movement, retrieval over huge context windows, and fast agentic workflows.
12 Feb 2026 · 16 min · 8 chapters
Learning to Reason in 13 Parameters
The episode explains “Learning to Reason in 13 Parameters,” arguing that a 7B-scale language model can dramatically improve math reasoning by updating only 13 parameters (TinyLoRA), using…
11 Feb 2026 · 19 min · 9 chapters
Nearly Optimal Active Preference Learning and Its Application to LLM Alignment
How to reduce the cost of RLHF for LLM alignment by choosing which human preference labels to collect, using “nearly optimal active preference learning” that targets the decision boundary (“zero…
8 Feb 2026 · 17 min · 9 chapters
Language Model Circuits Are Sparse in the Neuron Basis
The episode argues that LLM “black box” interpretability may be easier than current practice suggests.
8 Feb 2026 · 16 min · 6 chapters
Rethinking the Trust Region in LLM Reinforcement Learning
The episode argues that PPO (Proximal Policy Optimization) is poorly suited for LLM reinforcement learning because its trust-region mechanism uses token-level ratio clipping, which over-penalizes…
8 Feb 2026 · 16 min · 8 chapters
Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward
Learning LLM personalization from user edits, reframing “frustrated rewriting” as training signal via edit cost and a late ensemble of SFT, DPO, and cost-based RL.
8 Feb 2026 · 14 min · 7 chapters
Self-distillation enables continual learning
Self-distillation fine-tuning (SDFT) for continual AI learning that avoids catastrophic forgetting by using the model itself as both “teacher” and “student,” turning in-context learning into…
7 Feb 2026 · 20 min · 11 chapters
Maximum Likelihood Reinforcement Learning
Maximum Likelihood Reinforcement Learning (MaxRL) argues standard reinforcement learning (RL) is only the first term in a Maclaurin-series expansion of the true maximum-likelihood objective.
6 Feb 2026 · 16 min · 10 chapters
In-Context Algorithm Emulation in Fixed-Weight Transformers
The episode argues that fixed-weight transformer LLMs can “learn” new tasks via in-context algorithm emulation: the prompt effectively selects and runs algorithmic subroutines inside the frozen…
5 Feb 2026 · 17 min · 12 chapters
Previous
1
…
7
8
9
…
20
Next
Newest first · 24 per page