vo
Podcasts
Search podcasts and episodes
Use with Claude or ChatGPT
Get the app
Podcasts
/
Technology
/
Best AI papers explained
Podcast · 475 episodes
Best AI papers explained
, page 3
by
Enoch H. Kang
· English
Cut through the noise. We curate and break down the most important AI papers so you don’t have to.
Technology
Follow in VO
All episodes, page 3
Search Best AI papers explained episodes
Search
What Does Thompson Sampling Optimize?
Exploit–explore tradeoffs in multi-armed bandits, focusing on what Thompson sampling optimizes and how new research “opens the black box” using Markov decision processes and Bellman equations.
15 Jul 2026 · 22 min · 10 chapters
Globally Convergent Offline Reinforcement Learning with Smoothed Bellman Residual Minimization
Offline reinforcement learning safety via Globally Convergent Offline RL using Smoothed Bellman Residual Minimization (OffGladius), addressing the “deadly triad” (offline data + neural nets +…
13 Jul 2026 · 12 min · 5 chapters
LLM-as-a-Verifier: A General-Purpose Verification Framework
LLM-as-a-Verifier proposes a general-purpose framework to verify LLM outputs by extracting internal uncertainty (logits) instead of using discrete AI “judges,” addressing a “verification crisis”…
10 Jul 2026 · 20 min · 12 chapters
How Much Do Language Models Memorize?
How much do transformer language models memorize vs generalize, and what mathematical threshold triggers “grokking” (learning rules) instead of rote storage.
9 Jul 2026 · 24 min · 9 chapters
Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering
The episode argues that uncertainty quantification (UQ) in large language models is mis-specified: mainstream UQ methods effectively measure internal consistency via unsupervised clustering of…
7 Jul 2026 · 22 min · 12 chapters
Position: Agents Should Invoke External Tools ONLY When Epistemically Necessary
When AI agents should use external tools (search, APIs, code) versus rely on internal reasoning, arguing tools shouldn’t be used unless epistemically necessary.
6 Jul 2026 · 12 min · 6 chapters
From conversations to mechanisms: aligning advertiser Incentives in ai-powered product recommendations
How AI preference-alignment methods (used in chatbots via RLHF/DPO) can fail when they assume a fixed logistic “link function” for human choices, and how semi-parametric preference optimization (SPO)…
5 Jul 2026 · 22 min · 14 chapters
Is one layer enough? Training a single transformer layer can match full-parameter RL training
The episode explains research on reinforcement learning post-training (RLVR) in large transformer LLMs, arguing that training only a subset of layers—especially middle layers—can match or beat…
4 Jul 2026 · 23 min · 11 chapters
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training
The episode discusses a Harvard paper, “RL Excursions During Pre-Training,” arguing that reinforcement learning (RL) can be applied early during LLM pre-training—challenging the standard sequential…
2 Jul 2026 · 22 min · 10 chapters
Language Generation with Feedback: Queries and Mistakes
How language-generation learning fails under “no feedback” math, but becomes robust when feedback is added; the episode centers on theorems using language generation in the limit, set-based…
1 Jul 2026 · 20 min · 9 chapters
Quantifying Theoretical AI Alignment Guarantees: Receiver-Utility Bounds in Bayesian Persuasion
Bayesian persuasion limits for misaligned AI that has an information advantage and tries to steer a rational receiver’s decisions, quantified via receiver-utility bounds.
1 Jul 2026 · 22 min · 9 chapters
SPIRAL: Learning to search and aggregate
The episode argues that AI “capability” is jagged: models can solve hard math (e.g., Erdos unit distance problem) yet fail on simple logic.
29 Jun 2026 · 22 min · 10 chapters
Qwen-AgentWorld: Language World Models for General Agents
Qwen- AgentWorld describes “language world models” that let general AI agents simulate interactive computer environments internally before acting, reducing errors like clicking the wrong…
27 Jun 2026 · 21 min · 10 chapters
When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning?
When offline reinforcement learning can learn efficiently using only trajectory-level feedback (final outcomes or preferences) instead of step-by-step rewards, and when it becomes impossible.
27 Jun 2026 · 19 min · 11 chapters
SuperThoughts: Reasoning Tokens in Superposition
The episode explains “SuperThoughts,” a framework to speed up LLM reasoning by generating two tokens per step using “superposition,” while preserving accuracy via confidence-based adaptive fallback.
26 Jun 2026 · 19 min · 7 chapters
First-Explore PPO : Learning Meta-Exploration with Proximal Policy Optimization
Explore-vs-exploit in reinforcement learning under deceptive rewards, and a new method called First-Explore PPO (FEPPO) that learns meta-exploration efficiently.
25 Jun 2026 · 23 min · 9 chapters
Self-Distillation for Data-Scarce Language Model Pretraining
Data-scarce pretraining for language models; why multi-epoch reuse overfits, and how self-distillation (teacher-student with identical architectures) improves learning using “dark knowledge” (soft…
24 Jun 2026 · 22 min · 11 chapters
Meta-Harness for Agent-State Construction
Meta-harness for agent-state construction (SAIT): an outer-loop AI that writes and debugs the “harness” (working-memory management layer) for other agents, optimizing what gets…
21 Jun 2026 · 23 min · 9 chapters
ExpRL: Using Reference Solutions as Rewards for LLM Mid-Training
The episode explains the XPRL (Exploratory RL for LLM Mid-Training) paper: using hidden human reference solutions as reward scaffolds to overcome the “initialization bottleneck” in sparse pass/fail…
21 Jun 2026 · 21 min · 12 chapters
Valid Inference with Synthetic Data via Task Exchangeability
How to make valid statistical inference from AI-generated synthetic data despite hallucinations/bias, using a Stanford paper’s “task exchangeability” to calibrate confidence intervals with historical…
18 Jun 2026 · 13 min · 5 chapters
GRPO is Secretly a Process Reward Model
Explains how GRPO (Group Relative Policy Optimization) trains AI for multi-step reasoning, why it effectively behaves like a hidden process reward model (PRM), what flaw harms learning (“imbalanced…
17 Jun 2026 · 21 min · 15 chapters
Agentic Interactions
The episode reviews research on “agentic interactions,” testing whether delegating negotiations to AI agents removes human quirks or magnifies them, and how prompt-writing drives outcomes.
17 Jun 2026 · 19 min · 10 chapters
A Unifying View of Attention Sinks: Two Algorithms, Two Solutions
Vision-transformer “attention sinks” show near-vertical stripes in attention maps.
16 Jun 2026 · 23 min · 8 chapters
From AGI to ASI
Post-AGI roadmap from AGI (median human-level cognition) to ASI (consistently outperforming large coordinated human expert groups over long periods), and the theoretical ceiling UAI via the AIXI…
14 Jun 2026 · 24 min · 11 chapters
Previous
1
2
3
4
…
20
Next
Newest first · 24 per page