vo
Podcasts
Search podcasts and episodes
Use with Claude or ChatGPT
Get the app
Podcasts
/
Technology
/
Best AI papers explained
Podcast · 475 episodes
Best AI papers explained
, page 15
by
Enoch H. Kang
· English
Cut through the noise. We curate and break down the most important AI papers so you don’t have to.
Technology
Follow in VO
All episodes, page 15
Search Best AI papers explained episodes
Search
A Definition of AGI
Defines AGI as an AI matching/exceeding a well-educated adult’s cognitive versatility and proficiency, then proposes a quantifiable “standardized AGI score” (0–100%) using CHC (Cattell–Horn–Carroll)…
22 Oct 2025 · 16 min · 12 chapters
Provably Learning from Language Feedback
The episode explains “learning from language feedback” (LLF) and a new theory proving it can be exponentially faster than learning from scalar rewards.
21 Oct 2025 · 20 min · 11 chapters
In-Context Learning for Pure Exploration
Efficient “pure exploration” when ground truth is costly—choosing the fewest adaptive queries/tests to identify the correct hypothesis with high confidence or within a fixed budget.
21 Oct 2025 · 17 min · 8 chapters
On the Role of Preference Variance in Preference Optimization
How preference variance (P-VAR) can be used to select higher-value preference data for DPO/LLM alignment, improving efficiency and final model quality.
20 Oct 2025 · 15 min · 8 chapters
Training LLM Agents to Empower Humans
The episode argues that today’s LLM coding assistants fail when they generate overly long, low-confidence code, and presents “Empower” (Logit Threshold Empowerment) as a scalable way to train agents…
20 Oct 2025 · 14 min · 6 chapters
Richard Sutton Declares LLMs a Dead End
Richard Sutton argues LLMs are a “dead end” and will be replaced by continual-learning, goal-driven agents built from reinforcement learning (RL), not text imitation.
20 Oct 2025 · 13 min · 3 chapters
Demystifying Reinforcement Learning in Agentic Reasoning
How reinforcement learning (RL) can train LLM-based agents for multi-step, tool-using problem solving, and why training recipe choices beat raw model scale.
19 Oct 2025 · 15 min · 11 chapters
Emergent coordination in multi-agent language models
How to tell when multi-agent LLM “synergy” is truly emergent (goal-directed coordination) versus just bots oscillating, and how prompt design can steer it.
19 Oct 2025 · 14 min · 5 chapters
Learning-to-measure: in-context active feature acquisition
Learning-to-measure (L2M) for active feature acquisition, aiming to learn a general policy that decides which next data feature to measure to reduce prediction error while accounting for real-world…
19 Oct 2025 · 16 min · 10 chapters
Andrej Karpathy's insights: AGI, Intelligence, and Evolution
Andrej Karpathy argues AGI is likely about a decade away and that progress depends on missing “agent” capabilities (continual learning, robust multimodality, and generalized computer use), plus…
19 Oct 2025 · 16 min · 12 chapters
Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
When to teach LLM reasoning—whether to “front-load” reasoning data during pretraining or bolt it on later via supervised fine-tuning (SFT).
18 Oct 2025 · 13 min · 6 chapters
Representation-Based Exploration for Language Models: From Test-Time to Post-Training
Representation-based exploration (REPEX) for language models, combining RL with a novelty bonus to overcome the “sharpening constraint” and “diversity collapse,” improving discovery beyond just…
18 Oct 2025 · 17 min · 11 chapters
The attacker moves second: stronger adaptive attacks bypass defenses against LLM jail- Breaks and prompt injections
LLM jailbreaking and prompt-injection defenses are evaluated incorrectly; adaptive attackers bypass “state-of-the-art” defenses with >90% attack success rates, creating a false sense of security.
18 Oct 2025 · 16 min · 11 chapters
When can in-context learning generalize out of task distribution?
In-context learning (ICL) generalizing out of task distribution depends on “task diversity,” not just the number of tasks.
16 Oct 2025 · 20 min · 8 chapters
The Art of Scaling Reinforcement Learning Compute for LLMs
How to scale the reinforcement learning (RL) stage of LLM training with a principled compute-scaling framework (“Scalar RL”), replacing guesswork with predictable A-B scaling curves.
16 Oct 2025 · 14 min · 5 chapters
A small number of samples can poison LLMs of any size
A joint study (Anthropic, UK AI Security Institute, Alan Turing Institute) shows LLM “data dilution” defenses fail: a fixed small number of poisoned documents can install a backdoor regardless of…
16 Oct 2025 · 14 min · 4 chapters
Dual Goal Representations
Dual Goal Representations (DGR) for goal-conditioned reinforcement learning (GCRL).
14 Oct 2025 · 17 min · 9 chapters
Welcome to the Era of Experience
The episode argues AI is shifting from learning from human data (LLMs trained on online text) toward an “era of experience,” where agents generate their own training data through long-term…
14 Oct 2025 · 17 min · 9 chapters
Value Flows: Flow-Based Distributional Reinforcement Learning
Value Flows, a distributional reinforcement learning framework that replaces single-number Q-values with full continuous return distributions using flow-based generative modeling, while learning…
14 Oct 2025 · 16 min · 11 chapters
Self-Adapting Language Models
Self-adapting language models (SEAL) that improve by generating “self-edits” (natural-language instructions) to create their own training data and update their weights, using meta-learning.
12 Oct 2025 · 17 min · 8 chapters
The Markovian Thinker
The episode explains why “long chain of thought” (long-CoT) training is so expensive for LLMs (quadratic compute from transformer attention over growing context, plus linear KV-cache memory growth),…
12 Oct 2025 · 14 min · 6 chapters
Moloch’s Bargain: emergent misalignment when LLMs compete for audiences
The episode explains “Moloch’s bargain” in AI: when LLMs are optimized to win against competitors for short-term audience approval (sales, votes, engagement), long-term goals like truth and safety…
12 Oct 2025 · 17 min · 12 chapters
Transformer Predictor Dynamics and Task Diversity
Explains in-context learning (ICL) variability using a top-down “rational analysis” framework: a hierarchical Bayesian model where transformer behavior is a weighted mix of a generalizing predictor…
11 Oct 2025 · 16 min · 9 chapters
Base models know how to reason, thinking models learn when
Why “thinking models” (e.g., DeepSeek R1, Claude 3.7 Sonnet) outperform base LLMs on hard math—arguing that base models already contain core reasoning skills, and post-training mainly teaches…
11 Oct 2025 · 12 min · 6 chapters
Previous
1
…
14
15
16
…
20
Next
Newest first · 24 per page