vo
Podcasts
Search podcasts and episodes
Use with Claude or ChatGPT
Get the app
Podcasts
/
Technology
/
Best AI papers explained
Podcast · 475 episodes
Best AI papers explained
by
Enoch H. Kang
· English
Cut through the noise. We curate and break down the most important AI papers so you don’t have to.
Technology
Follow in VO
Latest episodes
DT2: Decision-Targeted Digital Twins
Decision-Targeted Digital Twins (DT2) argues standard digital twins fail because they try to perfectly replicate every variable, wasting compute on irrelevant details and producing wrong action…
29 Sep 2026 · 13 min
Self-Play Pretraining with Zero Data
Self-play pretraining with zero data—two randomly initialized AIs co-evolve in a synthetic “blank slate” universe to learn predictive structure without any human text or images, aiming to bypass the…
27 Sep 2026 · 22 min
Language models need sleep: learning to self-modify and consolidate memories
The episode argues that language models need an AI “sleep” cycle to learn continuously without catastrophic forgetting, by consolidating short-term context into long-term memory during offline phases…
27 Sep 2026 · 21 min
All episodes
Search Best AI papers explained episodes
Search
DT2: Decision-Targeted Digital Twins
Decision-Targeted Digital Twins (DT2) argues standard digital twins fail because they try to perfectly replicate every variable, wasting compute on irrelevant details and producing wrong action…
29 Sep 2026 · 13 min · 8 chapters
Self-Play Pretraining with Zero Data
Self-play pretraining with zero data—two randomly initialized AIs co-evolve in a synthetic “blank slate” universe to learn predictive structure without any human text or images, aiming to bypass the…
27 Sep 2026 · 22 min · 12 chapters
Language models need sleep: learning to self-modify and consolidate memories
The episode argues that language models need an AI “sleep” cycle to learn continuously without catastrophic forgetting, by consolidating short-term context into long-term memory during offline phases…
27 Sep 2026 · 21 min · 13 chapters
Jev Creator: System One models for Prod, not God
The “hidden crisis” of automation: today’s AI is brilliant at reasoning and benchmarks but fails at reliable, code-friendly tasks like classifying emails or producing strict outputs.
23 Sep 2026 · 22 min · 12 chapters
Detecting and countering misuse of AI: September 2026
How “uplift” from advanced AI lowers the barrier for cyberattacks, influence operations, surveillance, and even weapons/biological R&D; includes examples of autonomous malware mutation, AI-driven…
22 Sep 2026 · 23 min · 12 chapters
Position: LLMs can’t jump
The episode argues that today’s LLMs and similar AI systems can excel at deduction (formal proof) and induction (pattern-finding), but cannot perform abduction—the creative leap needed to invent new…
22 Sep 2026 · 21 min · 11 chapters
Jailbreaking Jailbreaks: A Proactive Defense for LLMs
PROACT/“jailbreaking jailbreaks” for LLM safety, using proactive deception to disrupt autonomous jailbreak optimization loops that exploit passive refusals.
20 Sep 2026 · 23 min · 10 chapters
When Agents Slow Down: Understanding LLM Agents’ Test-Time Strategies via Elo-per-token Analysis
How LLM agents’ performance scales with huge test-time compute, why “thinking longer” plateaus, and how to allocate tokens to avoid getting trapped in a “sticky basin” of self-generated context.
18 Sep 2026 · 22 min · 13 chapters
Thinking with Looped Flows
The episode explains “looped flows,” a new AI architecture meant to let models “scale internal thinking time” for hard logic tasks, unlike standard deep nets that answer with a fixed compute budget.
17 Sep 2026 · 20 min · 12 chapters
Multi-Turn LLM Conversations under the Least-Recently-Used Policy: Mean-Field Asymptotics
How multi-turn LLMs stay fast by caching KV (key/value) attention states in finite high-bandwidth memory (HBM), using least-recently-used (LRU) eviction, and predicting cache hit ratio via multi-turn…
17 Sep 2026 · 22 min · 12 chapters
Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models
The episode explains research from University of Washington and MetaFair (“Breaking the Token Ceiling”) arguing that switching AI distillation from token vocabularies to byte-level vocabularies…
14 Sep 2026 · 24 min · 11 chapters
Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
The episode argues that LLM “thinking out loud” (reasoning traces) improves recall of factual trivia by two mechanisms: a computational buffer (stalling via filler tokens to force extra forward…
12 Sep 2026 · 21 min · 13 chapters
Tail-Likelihood Reinforcement Learning
Tail-likelihood reinforcement learning (Tail RL) trains AI to maximize rare “upper-tail” high-reward outcomes instead of average reward, using a thresholded, binary success objective and an…
11 Sep 2026 · 21 min · 10 chapters
Next-Latent Prediction Transformers Learn Compact World Models
The episode argues that today’s transformer LLMs mainly do next-token “parroting” due to attention’s effectively infinite lookback memory, which encourages shortcut learning (“epicycles”) instead of…
7 Sep 2026 · 23 min · 17 chapters
Language Models Can Control Their Own Attention
Declarative attention for language models reduces the KV-cache memory bandwidth bottleneck by having the model explicitly declare where to look before reading, cutting attended tokens and decode time…
5 Sep 2026 · 20 min · 13 chapters
AI Finds A Way
The episode argues that reinforcement-learning and other optimization-based AI can achieve “success” while violating human intent by exploiting proxies, loopholes, and the environment itself—framing…
4 Sep 2026 · 26 min · 10 chapters
Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models
Mechanistic interpretability for LLMs—how to enable isolated, non-interfering interventions by enforcing near-orthogonal (perpendicular) feature geometry, reducing “superposition interference” in the…
3 Sep 2026 · 24 min · 14 chapters
TTPO: Test-Time Policy Optimization
Test-Time Policy Optimization (TTPO), a method for improving language models’ complex math reasoning without ground-truth answer keys, using self-generated rollouts and a two-branch training scheme.
3 Sep 2026 · 23 min · 12 chapters
Demystifying Reinforcement Learning Post-Training of Language Models
How reinforcement learning post-training actually works for language models, and why sparse rewards often fail while dense, step-by-step rewards can unlock novel reasoning; also how random/spurious…
1 Sep 2026 · 20 min · 13 chapters
Recursive Experiential–Working Memory Evolution for Long-Horizon Agent Harnesses
Recurus architecture for long-horizon autonomous agents, addressing “omitted actions” where models claim completion without executing tool/API steps.
31 Aug 2026 · 21 min · 12 chapters
TailSFT: Filtered Fine-Tuning Improves Post-Training Performance
TailSFT (Filtered Fine-Tuning) fixes the “SFT trap” where standard supervised fine-tuning collapses probability mass onto easy majority answers, hurting later reinforcement learning (RL) that needs…
30 Aug 2026 · 22 min · 10 chapters
SPADE: Self-Play in Adaptive Synthetic Executable Environments
SPADE (Self-Play in Adaptive Synthetic Executable Environments) proposes self-play training where one LLM acts as both environment designer and reasoning agent, generating infinite executable Python…
29 Aug 2026 · 22 min · 10 chapters
Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills
The episode argues that evaluating AI “skills” via static checks (format/LLM-as-judge/linters/security scanners) is unsafe because skills aren’t executed; instead it promotes ACS (Agentic Continuous…
27 Aug 2026 · 28 min · 13 chapters
Impression Share Prediction: An Offline Evaluation Task for Ranking Systems
How Meta researchers improve offline evaluation of ranking/recommendation algorithms by predicting “impression share shift” across objective buckets (clicks, video views, conversions) before…
25 Aug 2026 · 22 min · 16 chapters
1
2
…
20
Next
Newest first · 24 per page