vo
Podcasts
Search podcasts and episodes
Use with Claude or ChatGPT
Get the app
Podcasts
/
Technology
/
Best AI papers explained
Podcast · 475 episodes
Best AI papers explained
, page 7
by
Enoch H. Kang
· English
Cut through the noise. We curate and break down the most important AI papers so you don’t have to.
Technology
Follow in VO
All episodes, page 7
Search Best AI papers explained episodes
Search
Provable and practical in-context policy optimization for self-improvement
Test-time scaling and self-reflection in language models via in-context policy optimization (ICPO), using a multi-armed bandit framing to let the model optimize its own reasoning on the fly without…
17 Mar 2026 · 21 min · 13 chapters
Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models
Explains why long-form generation derails (teacher forcing + distribution shift) and why common fixes like RLVR can “hack” rewards, then introduces energy-based fine-tuning (EBFT) that matches…
16 Mar 2026 · 23 min · 13 chapters
Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights
The episode explains new research claiming that “guess-and-check” can work for large, heavily pretrained AI models because pretraining reshapes the loss landscape into a dense “thicket” of nearby…
14 Mar 2026 · 20 min · 12 chapters
AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization
AdaEvolve (UC Berkeley, Bespoke Labs) argues that algorithm/code search with LLM-guided evolutionary methods fails due to fixed “static schedules,” and proposes hierarchical adaptive optimization:…
14 Mar 2026 · 20 min · 10 chapters
∇−reasoner: LLM reasoning via test-time gradient descent in latent space
Deep dive into “NABLA Reasoner,” a framework for LLM test-time reasoning using test-time gradient descent in latent space (Differentiable Textual Optimization, DTO), aiming to improve multi-step…
14 Mar 2026 · 21 min · 11 chapters
Inference for Regression with Variables Generated by AI or Machine Learning
The episode explains why the common “two-step” workflow—use AI/ML to generate variables from text/audio, then run a standard OLS regression—can produce biased causal estimates and misleading…
12 Mar 2026 · 22 min · 11 chapters
Fast KV Compaction via Attention Matching
The episode explains the “context window” bottleneck in LLMs, focusing on KV cache memory growth, and a new method to compress KV cache ~50x in seconds without losing reasoning.
12 Mar 2026 · 23 min · 10 chapters
Position: stop anthropomorphizing intermediate tokens as reasoning/thinking traces!
Argues that “reasoning traces” (intermediate tokens like “aha”) in large reasoning models are not human-like thought, and that treating them as genuine logic misleads users and researchers.
11 Mar 2026 · 19 min · 9 chapters
Code World Models for General Game Playing
Code World Models (CWMs) for General Game Playing—using an LLM to translate natural-language game rules into executable Python “world” simulators, then letting classical planning (MCTS) play within…
8 Mar 2026 · 22 min · 8 chapters
Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought
How chain-of-thought prompting changes a transformer’s optimization behavior, letting it implement multi-step gradient descent rather than a single-step update; includes related “loop transformer”…
7 Mar 2026 · 17 min · 7 chapters
Task Descriptors Help Transformers Learn Linear Models In-Context
How “task descriptors” in prompts (e.g., “translate English to French”) change the internal math of in-context learning in transformers, including provable behavior with infinite data and measured…
7 Mar 2026 · 19 min · 15 chapters
Equivalence of Context and Parameter Updates in Modern Transformer Blocks
Explains a mathematically proven view of in-context learning: a prompt is equivalent to applying temporary “implicit weight patches” inside transformer blocks, even for modern bias-free architectures.
7 Mar 2026 · 21 min · 9 chapters
Learning without training: The implicit dynamics of in-context learning
Explains Google research on in-context learning (ICL): how a frozen transformer “learns” a new task from the prompt without updating permanent weights, by applying implicit, temporary parameter…
7 Mar 2026 · 24 min · 15 chapters
Causal Identification from Counterfactual Data: Completeness and Bounding Results
Causal AI for counterfactual identification, introducing “counterfactual randomization” (Layer 2.5) and a new algorithm (CTFIDU+) that is complete for realizable queries; when exact identification is…
7 Mar 2026 · 20 min · 15 chapters
Is Cosine-Similarity of Embeddings Really About Similarity?
Netflix researchers argue that cosine similarity between learned embeddings may not reflect true semantic similarity because hidden scaling introduced by regularization can make cosine scores…
6 Mar 2026 · 22 min · 17 chapters
Diffusion LLMs are Natural Adversaries for any LLM
The episode explains a research paper, “Diffusion LLMs are Natural Adversaries for Any LLM” (Technical University of Munich), arguing that jailbreak vulnerabilities are “data-specific” and can be…
5 Mar 2026 · 25 min · 16 chapters
Are you going to finish that? A Practical Study of the Partial Token Problem
The episode explains the “partial token problem,” where language models tokenize text into chunks and can fail when a user prompt ends mid-token (e.g., pausing after an incomplete word or punctuation…
4 Mar 2026 · 19 min · 7 chapters
Language Models Struggle to Use Representations Learned In-Context
“Inert map” problem in language models—models can encode a representation learned from in-context data but fail to use it for rule-based navigation and adaptive world modeling.
2 Mar 2026 · 19 min · 12 chapters
LLMs are Bayesian, In Expectation, Not in Realization
In-context learning in LLMs violates exchangeability: shuffling the order of the same demonstrations can change predictions.
1 Mar 2026 · 19 min · 10 chapters
Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs
Reflective test-time planning for embodied LLM robots—moving from “static oracle” behavior (repeat the same failed plan) to trial-and-error learning via simulated reflection, hindsight regret, and…
27 Feb 2026 · 18 min · 10 chapters
LLMs Can Learn to Reason Via Off-Policy RL
Off-policy RL for large language model (LLM) reasoning, proposing OAPL (Optimal Advantage Based Policy Optimization with Lagged Inference) to fix training instability caused by “on-policy”…
27 Feb 2026 · 20 min · 11 chapters
Test-Time Training with KV Binding Is Secretly Linear Attention
Test-time training (TTT) with KV binding—supposed to “learn/memorize” new context during inference—but the episode argues it’s actually equivalent to linear attention, enabling simpler and faster…
27 Feb 2026 · 18 min · 8 chapters
Unified Latents (UL): How to train your latents
Unified Latents (UL) by Google DeepMind Amsterdam—how to train latent representations for diffusion-based image/video generation more efficiently, with controllable “information capacity”…
26 Feb 2026 · 20 min · 9 chapters
Spectral Bellman Method: Unifying RL Representation and Exploration
Spectral Bellman Method (SBM) for reinforcement learning representation and exploration, addressing the “deadly triad” instability from function approximation + bootstrapping + off-policy learning.
25 Feb 2026 · 21 min · 10 chapters
Previous
1
…
6
7
8
…
20
Next
Newest first · 24 per page