Best AI papers explained cover art
Podcast · 475 episodes

Best AI papers explained, page 8

by Enoch H. Kang · English

Cut through the noise. We curate and break down the most important AI papers so you don’t have to.

All episodes, page 8

Prescriptive Scaling Reveals the Evolution of Language Model CapabilitiesPrescriptive scaling in AI—using a “capability boundary” (an S-curve) to predict model performance from pre-training compute, and to diagnose whether gains come from real reasoning or from issues…24 Feb 2026 · 18 min · 8 chapters
Experiential Reinforcement LearningExperiential Reinforcement Learning (ERL), a training paradigm that replaces sparse-reward “button-mashing” with a human-like loop: try, reflect, retry, and store lessons across episodes; then remove…23 Feb 2026 · 23 min · 8 chapters
Learning Personalized Agents from Human FeedbackPersonalized agents from human feedback (PAHF) that adapt to changing user preferences using a read-write memory loop, avoiding “static personalization” failures like preference drift.21 Feb 2026 · 15 min · 7 chapters
Learning to summarize user information for personalized RLHFExplains why RLHF personalization can produce “vanilla/beige” responses, then argues PLUS (preference learning via natural-language summarization) fixes it by replacing vector user profiles with…20 Feb 2026 · 18 min · 11 chapters
Intrinsic Credit Assignment for Long Horizon InteractionThe episode explains the paper “Intrinsic Credit Assignment for Long Horizon Interaction,” focusing on teaching AI agents “curiosity” via Delta Belief Reinforcement Learning (Delta Belief RL).20 Feb 2026 · 18 min · 10 chapters
Learning to Continually Learn via Meta-learning Agentic Memory DesignsThe “Groundhog Day” problem in agentic AI—stateless foundation models lose context between sessions—and how ALMA (Automated Meta Learning of Memory Designs for Agentic Systems) uses a meta-agent to…20 Feb 2026 · 20 min · 13 chapters
Why Self-Rewarding Works: Theoretical Guarantees for Iterative Alignment of Language ModelsThe episode explains a mathematical proof for “self-rewarding language models” (SRLMs): models that act as both policy (generator) and reward model (judge), aiming to iteratively align/improve…19 Feb 2026 · 18 min · 9 chapters
PAD: Personalized Alignment of LLMs at Decoding-TimePAD (Personalized Alignment of LLMs at Decoding Time) tackles the “alignment tax” where RLHF-trained LLMs drift to a safe, average “committee” voice.19 Feb 2026 · 14 min · 9 chapters
The Reward Model Selection Crisis in Personalized AlignmentPersonalized alignment via reward models (RLHF) is flawed because reward-model ranking accuracy often doesn’t translate into better generation; self-evaluation enables reward hacking and “circular…19 Feb 2026 · 16 min · 7 chapters
Causal-JEPA: Learning World Models through Object-Level Latent InterventionsCausal-JEPA (C-JEPA) for learning world models by forcing object-level causal reasoning instead of “cheating” via pixel interpolation.18 Feb 2026 · 15 min · 9 chapters
How Sampling Shapes LLM Alignment: From One-Shot Optima to Iterative DynamicsHow the “sampling problem” in preference-based LLM alignment (choosing which answer pairs to compare) shapes model behavior, and how iterative alignment feedback loops can cause collapse or…17 Feb 2026 · 16 min · 8 chapters
Deriving neural scaling laws from the statistics of natural languageExplains a research breakthrough claiming neural scaling laws can be derived from statistical properties of natural language, not from model architecture.15 Feb 2026 · 19 min · 7 chapters
Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RLReasoning Cache, a short-horizon RL-trained method that lets small autoregressive models “pause, summarize, delete, and continue,” extending effective reasoning from ~16k tokens to ~512k while…15 Feb 2026 · 15 min · 6 chapters
Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RLOrbit, a framework for online reinforcement-based in-context learning that lets LLM agents learn from failures across episodes (solving the “static aftershipping problem”/amnesia between chats).14 Feb 2026 · 15 min · 13 chapters
Divide-and-Conquer CoT: RL for Reducing Latency via Parallel ReasoningThe episode explains “divide-and-conquer chain-of-thought” (DC-CoT), an RL-trained method that reduces LLM latency by parallelizing reasoning.12 Feb 2026 · 16 min · 7 chapters
Owning the AI Pareto Frontier — Jeff DeanHow to “own the AI Pareto frontier” by balancing capability vs cost/latency, using distillation, sparsity, low energy/data movement, retrieval over huge context windows, and fast agentic workflows.12 Feb 2026 · 16 min · 8 chapters
Learning to Reason in 13 ParametersThe episode explains “Learning to Reason in 13 Parameters,” arguing that a 7B-scale language model can dramatically improve math reasoning by updating only 13 parameters (TinyLoRA), using…11 Feb 2026 · 19 min · 9 chapters
Nearly Optimal Active Preference Learning and Its Application to LLM AlignmentHow to reduce the cost of RLHF for LLM alignment by choosing which human preference labels to collect, using “nearly optimal active preference learning” that targets the decision boundary (“zero…8 Feb 2026 · 17 min · 9 chapters
Language Model Circuits Are Sparse in the Neuron BasisThe episode argues that LLM “black box” interpretability may be easier than current practice suggests.8 Feb 2026 · 16 min · 6 chapters
Rethinking the Trust Region in LLM Reinforcement LearningThe episode argues that PPO (Proximal Policy Optimization) is poorly suited for LLM reinforcement learning because its trust-region mechanism uses token-level ratio clipping, which over-penalizes…8 Feb 2026 · 16 min · 8 chapters
Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and RewardLearning LLM personalization from user edits, reframing “frustrated rewriting” as training signal via edit cost and a late ensemble of SFT, DPO, and cost-based RL.8 Feb 2026 · 14 min · 7 chapters
Self-distillation enables continual learningSelf-distillation fine-tuning (SDFT) for continual AI learning that avoids catastrophic forgetting by using the model itself as both “teacher” and “student,” turning in-context learning into…7 Feb 2026 · 20 min · 11 chapters
Maximum Likelihood Reinforcement LearningMaximum Likelihood Reinforcement Learning (MaxRL) argues standard reinforcement learning (RL) is only the first term in a Maclaurin-series expansion of the true maximum-likelihood objective.6 Feb 2026 · 16 min · 10 chapters
In-Context Algorithm Emulation in Fixed-Weight TransformersThe episode argues that fixed-weight transformer LLMs can “learn” new tasks via in-context algorithm emulation: the prompt effectively selects and runs algorithmic subroutines inside the frozen…5 Feb 2026 · 17 min · 12 chapters