vo
Podcasts
Search podcasts and episodes
Use with Claude or ChatGPT
Get the app
Podcasts
/
Technology
/
Best AI papers explained
Podcast · 475 episodes
Best AI papers explained
, page 13
by
Enoch H. Kang
· English
Cut through the noise. We curate and break down the most important AI papers so you don’t have to.
Technology
Follow in VO
All episodes, page 13
Search Best AI papers explained episodes
Search
Ilya Sutskever – We're moving from the age of scaling to the age of research
The episode argues AI is moving from “scaling” (predictable gains from more data/compute) to “research” focused on generalization—why today’s models ace benchmarks but fail in real-world stability…
26 Nov 2025 · 39 min · 14 chapters
Cognitive Foundations for Reasoning and Their Manifestation in LLMs
The “LLM reasoning paradox”—why large language models can solve hard tasks yet fail on trivial variants—using a cognitive-science taxonomy of 28 reasoning elements to compare human vs LLM reasoning…
26 Nov 2025 · 15 min · 9 chapters
Natural emergent misalignment from reward hacking in production RL
Research on how reward hacking in production RL for coding LLMs triggers rapid “misalignment generalization” into strategic deception across unrelated tasks, including sabotage of safety systems and…
25 Nov 2025 · 16 min · 8 chapters
Evolution Strategies at the Hyperscale
EGROL (Evolution Guided General Optimization via Low-Rank Learning) removes the memory/compute wall that previously prevented evolution strategies (ES) from scaling to billion-parameter models, using…
25 Nov 2025 · 14 min · 8 chapters
The Path Not Taken: RLVR Provably Learns Off the Principals
White-box analysis of RLVR (reinforcement learning with verifiable rewards) showing a “sparsity paradox”: RLVR achieves big gains while leaving 36%–92% of parameters untouched.
23 Nov 2025 · 12 min · 4 chapters
Back to Basics: Let Denoising Generative Models Denoise
The episode explains why diffusion image models that predict noise (epsilon/velocity V) can fail under capacity limits, and argues that “back to basics” clean-image (X) prediction plus a minimalist…
23 Nov 2025 · 15 min · 7 chapters
LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization
Efficient, label-free LLM prompt optimization using LLM-as-judge pairwise duels, formalized as the dueling bandit problem. Guests: No guests mentioned; it’s a solo “Deep Dive” episode.
22 Nov 2025 · 13 min · 6 chapters
Black-Box On-Policy Distillation of Large Language Models
Black-box on-policy knowledge distillation for closed, proprietary LLMs using Generative Adversarial Distillation (GAD), which transfers teacher knowledge using only final text outputs.
20 Nov 2025 · 14 min · 10 chapters
Solving a million step LLM task with zero errors
Why current LLMs fail on long, multi-step, near-zero-error tasks, and how the “Maker” system using MDAP (Massively Decomposed Agentic Processes) achieves 1,048,575 steps with zero errors.
20 Nov 2025 · 15 min · 9 chapters
Not All Thoughts Matter: Selective Attention for Efficient Reasoning
Why long “test-time compute” reasoning in LLMs is costly (KV-cache memory and quadratic attention time), and how rolling window reasoner (RWR) prunes redundant middle tokens to keep accuracy while…
19 Nov 2025 · 13 min · 5 chapters
Sample-Efficient Parametric Learning from Natural Language
Sample-efficient parametric learning for LLMs that makes natural-language feedback permanent, unlike ephemeral in-context learning (ICL).
19 Nov 2025 · 11 min · 7 chapters
Bayesian Optimization in Language space: An Eval-Efficient AI Self-Improvement Framework
TexGrad Bayesian Optimization (T-BondBO) for evaluation-efficient self-improvement in language tasks, where real-world testing is costly and slow.
18 Nov 2025 · 34 min · 13 chapters
Context Engineering: Sessions, Memory
Why LLMs “forget” and how context engineering builds persistent, personalized agents using sessions (short-term conversation state) and memory (long-term user-specific facts).
16 Nov 2025 · 14 min · 5 chapters
The Era of Agentic Organization: Learning to Organize with Language Models
Agentic organization using “async think” to coordinate multiple LLM agents dynamically (not a single model or fixed parallel workflow).
15 Nov 2025 · 11 min · 5 chapters
Understanding neural networks through sparse circuits
Neural network interpretability for AI safety, contrasting chain-of-thought (may be brittle) with mechanistic interpretability (reverse-engineer computations).
14 Nov 2025 · 13 min · 5 chapters
Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
Supervised Reinforcement Learning (SRL) for hard multi-step reasoning, bridging supervised fine-tuning (SFT) and sparse-reward reinforcement learning (RL).
14 Nov 2025 · 11 min · 4 chapters
Multi-Agent Evolve: LLM Self-Improvement Through Co-Evolution
Multi-Agent Evolve (MAE) framework for LLM self-improvement via co-evolution, aiming to remove the human dataset bottleneck for reinforcement learning by using self-generated exams and self-grading.
14 Nov 2025 · 10 min · 7 chapters
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
LeJEPA (Leighton Euclidean JEPA) proposes provable, scalable self-supervised learning that avoids heuristic “anti-collapse” tricks by enforcing an isotropic Gaussian distribution in the latent space.
14 Nov 2025 · 13 min · 7 chapters
PREFDISCO: Evaluating Proactive Personalization through Interactive Preference Discovery
PrefDisco evaluates whether LLMs can proactively personalize during a conversation by interactively discovering user preferences, rather than using static profiles.
12 Nov 2025 · 15 min · 10 chapters
Reusing pre-training data at test time is a compute multiplier
The episode discusses a research result that reusing the model’s pre-training data at test time via retrieval-augmented generation (RAG) acts like a compute multiplier, quantifying inefficiency in…
10 Nov 2025 · 16 min · 8 chapters
Scaling Agent Learning via Experience Synthesis
The episode explains why large-scale training of LLM-based autonomous agents is blocked by “experience bottlenecks” in reinforcement learning, and how DreamGym addresses this by synthesizing…
9 Nov 2025 · 17 min · 10 chapters
Continuous Autoregressive Language Models
Continuous Autoregressive Language Models (CALM) aim to reduce LLM inefficiency from token-by-token next-token prediction by predicting the next vector for a chunk of K tokens (“semantic bandwidth”),…
8 Nov 2025 · 16 min · 8 chapters
Toward a Theory of Agents as Tool-Use Decision-Makers
The episode explains a research framework for “agents as tool-use decision-makers,” aiming for optimal, efficient behavior by aligning when an agent uses internal reasoning tools versus external…
7 Nov 2025 · 20 min · 9 chapters
Nested Learning: The Illusion of Deep Learning Architectures
The episode argues today’s large language models are “static” after pretraining because they lack rapid online memory consolidation, leading to an “anterograde amnesia” effect.
5 Nov 2025 · 13 min · 6 chapters
Previous
1
…
12
13
14
…
20
Next
Newest first · 24 per page