Best AI papers explained cover art
Podcast · 475 episodes

Best AI papers explained, page 3

by Enoch H. Kang · English

Cut through the noise. We curate and break down the most important AI papers so you don’t have to.

All episodes, page 3

What Does Thompson Sampling Optimize?Exploit–explore tradeoffs in multi-armed bandits, focusing on what Thompson sampling optimizes and how new research “opens the black box” using Markov decision processes and Bellman equations.15 Jul 2026 · 22 min · 10 chapters
Globally Convergent Offline Reinforcement Learning with Smoothed Bellman Residual MinimizationOffline reinforcement learning safety via Globally Convergent Offline RL using Smoothed Bellman Residual Minimization (OffGladius), addressing the “deadly triad” (offline data + neural nets +…13 Jul 2026 · 12 min · 5 chapters
LLM-as-a-Verifier: A General-Purpose Verification FrameworkLLM-as-a-Verifier proposes a general-purpose framework to verify LLM outputs by extracting internal uncertainty (logits) instead of using discrete AI “judges,” addressing a “verification crisis”…10 Jul 2026 · 20 min · 12 chapters
How Much Do Language Models Memorize?How much do transformer language models memorize vs generalize, and what mathematical threshold triggers “grokking” (learning rules) instead of rote storage.9 Jul 2026 · 24 min · 9 chapters
Position: Uncertainty Quantification in LLMs is Just Unsupervised ClusteringThe episode argues that uncertainty quantification (UQ) in large language models is mis-specified: mainstream UQ methods effectively measure internal consistency via unsupervised clustering of…7 Jul 2026 · 22 min · 12 chapters
Position: Agents Should Invoke External Tools ONLY When Epistemically NecessaryWhen AI agents should use external tools (search, APIs, code) versus rely on internal reasoning, arguing tools shouldn’t be used unless epistemically necessary.6 Jul 2026 · 12 min · 6 chapters
From conversations to mechanisms: aligning advertiser Incentives in ai-powered product recommendationsHow AI preference-alignment methods (used in chatbots via RLHF/DPO) can fail when they assume a fixed logistic “link function” for human choices, and how semi-parametric preference optimization (SPO)…5 Jul 2026 · 22 min · 14 chapters
Is one layer enough? Training a single transformer layer can match full-parameter RL trainingThe episode explains research on reinforcement learning post-training (RLVR) in large transformer LLMs, arguing that training only a subset of layers—especially middle layers—can match or beat…4 Jul 2026 · 23 min · 11 chapters
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM trainingThe episode discusses a Harvard paper, “RL Excursions During Pre-Training,” arguing that reinforcement learning (RL) can be applied early during LLM pre-training—challenging the standard sequential…2 Jul 2026 · 22 min · 10 chapters
Language Generation with Feedback: Queries and MistakesHow language-generation learning fails under “no feedback” math, but becomes robust when feedback is added; the episode centers on theorems using language generation in the limit, set-based…1 Jul 2026 · 20 min · 9 chapters
Quantifying Theoretical AI Alignment Guarantees: Receiver-Utility Bounds in Bayesian PersuasionBayesian persuasion limits for misaligned AI that has an information advantage and tries to steer a rational receiver’s decisions, quantified via receiver-utility bounds.1 Jul 2026 · 22 min · 9 chapters
SPIRAL: Learning to search and aggregateThe episode argues that AI “capability” is jagged: models can solve hard math (e.g., Erdos unit distance problem) yet fail on simple logic.29 Jun 2026 · 22 min · 10 chapters
Qwen-AgentWorld: Language World Models for General AgentsQwen- AgentWorld describes “language world models” that let general AI agents simulate interactive computer environments internally before acting, reducing errors like clicking the wrong…27 Jun 2026 · 21 min · 10 chapters
When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning?When offline reinforcement learning can learn efficiently using only trajectory-level feedback (final outcomes or preferences) instead of step-by-step rewards, and when it becomes impossible.27 Jun 2026 · 19 min · 11 chapters
SuperThoughts: Reasoning Tokens in SuperpositionThe episode explains “SuperThoughts,” a framework to speed up LLM reasoning by generating two tokens per step using “superposition,” while preserving accuracy via confidence-based adaptive fallback.26 Jun 2026 · 19 min · 7 chapters
First-Explore PPO : Learning Meta-Exploration with Proximal Policy OptimizationExplore-vs-exploit in reinforcement learning under deceptive rewards, and a new method called First-Explore PPO (FEPPO) that learns meta-exploration efficiently.25 Jun 2026 · 23 min · 9 chapters
Self-Distillation for Data-Scarce Language Model PretrainingData-scarce pretraining for language models; why multi-epoch reuse overfits, and how self-distillation (teacher-student with identical architectures) improves learning using “dark knowledge” (soft…24 Jun 2026 · 22 min · 11 chapters
Meta-Harness for Agent-State ConstructionMeta-harness for agent-state construction (SAIT): an outer-loop AI that writes and debugs the “harness” (working-memory management layer) for other agents, optimizing what gets…21 Jun 2026 · 23 min · 9 chapters
ExpRL: Using Reference Solutions as Rewards for LLM Mid-TrainingThe episode explains the XPRL (Exploratory RL for LLM Mid-Training) paper: using hidden human reference solutions as reward scaffolds to overcome the “initialization bottleneck” in sparse pass/fail…21 Jun 2026 · 21 min · 12 chapters
Valid Inference with Synthetic Data via Task ExchangeabilityHow to make valid statistical inference from AI-generated synthetic data despite hallucinations/bias, using a Stanford paper’s “task exchangeability” to calibrate confidence intervals with historical…18 Jun 2026 · 13 min · 5 chapters
GRPO is Secretly a Process Reward ModelExplains how GRPO (Group Relative Policy Optimization) trains AI for multi-step reasoning, why it effectively behaves like a hidden process reward model (PRM), what flaw harms learning (“imbalanced…17 Jun 2026 · 21 min · 15 chapters
Agentic InteractionsThe episode reviews research on “agentic interactions,” testing whether delegating negotiations to AI agents removes human quirks or magnifies them, and how prompt-writing drives outcomes.17 Jun 2026 · 19 min · 10 chapters
A Unifying View of Attention Sinks: Two Algorithms, Two SolutionsVision-transformer “attention sinks” show near-vertical stripes in attention maps.16 Jun 2026 · 23 min · 8 chapters
From AGI to ASIPost-AGI roadmap from AGI (median human-level cognition) to ASI (consistently outperforming large coordinated human expert groups over long periods), and the theoretical ceiling UAI via the AIXI…14 Jun 2026 · 24 min · 11 chapters