Best AI papers explained cover art
Podcast · 475 episodes

Best AI papers explained, page 17

by Enoch H. Kang · English

Cut through the noise. We curate and break down the most important AI papers so you don’t have to.

All episodes, page 17

Linear Transformers Implicitly Discover Unified Numerical AlgorithmsA linear transformer trained on masked matrix block completion implicitly discovers a unified numerical solver, EGLE (Emergent Algorithm for Global Low-Rank Estimation), a two-line update rule…29 Sep 2025 · 14 min · 7 chapters
Regularizing Extrapolation in Causal InferenceRegularizing extrapolation in causal inference when source and target groups differ (positivity violations).27 Sep 2025 · 15 min · 10 chapters
DoubleGen - Debiased Generative Modeling of CounterfactualsDoubleGen, a framework for generating unbiased counterfactuals (“what if” outcomes) from biased observational data by debiasing confounding and avoiding misspecification.27 Sep 2025 · 13 min · 6 chapters
What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoTEffective reasoning in large reasoning models (LRMs) is driven by structure and failure management, not by longer chain-of-thought (CoT) or more “review” tokens.27 Sep 2025 · 17 min · 10 chapters
Compute as Teacher: Turning Inference Compute Into Reference-Free SupervisionCompute as Teacher (CAT) turns a model’s inference-time exploration into reference-free supervision for training specialized skills, reducing the “supervision gap” when gold labels are scarce or…27 Sep 2025 · 16 min · 8 chapters
Learning without training: The implicit dynamics of in-context learningExplains in-context learning (ICL) as implicit optimization during inference, not “magic” or explicit weight updates.24 Sep 2025 · 14 min · 7 chapters
Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base ModelTests whether reinforcement learning with verifiable rewards (RLVR) truly increases LLM reasoning capacity beyond the base model, or only improves sampling efficiency within existing capabilities.24 Sep 2025 · 13 min · 7 chapters
Open Problems in Mechanistic InterpretabilityOpen problems in mechanistic interpretability (MI): how to reverse-engineer neural networks’ internal computations to enable trust, control, and scientific discovery despite “black box” behavior.21 Sep 2025 · 19 min · 12 chapters
Maestro: Joint Graph & Config Optimization for Reliable AI AgentsThe episode explains a study on “thought anchors,” sentence-level reasoning steps that disproportionately determine how an LLM reaches its answer, making chain-of-thought more interpretable and…21 Sep 2025 · 12 min · 5 chapters
Thought Anchors: Which LLM Reasoning Steps Matter?“Thought anchors” for LLM chain-of-thought—identifying which reasoning sentences most causally determine final answers in hard math, using intermediate-level abstraction rather than token-level…21 Sep 2025 · 16 min · 9 chapters
RL's Razor: Why Online RL Forgets LessCatastrophic forgetting in continual learning, and a MIT paper’s claim that “RL’s razor” explains why online reinforcement learning (RL) forgets less than supervised fine-tuning (SFT).7 Sep 2025 · 25 min · 12 chapters
Why Language Models HallucinateWhy large language models hallucinate (produce plausible but false statements) and why hallucinations persist due to training objectives and benchmark evaluation incentives; proposes changing…6 Sep 2025 · 18 min · 8 chapters
ALFA: Aligning LLMs to Ask Good Questions A Case Study in Clinical ReasoningALFA (Alignment Via Fine-Grained Attributes) teaches LLMs to ask better follow-up questions for clinical reasoning by decomposing “good questions” into attributes, synthesizing targeted training…6 Sep 2025 · 16 min · 12 chapters
Sample Efficient Preference Alignment in LLMs via Active ExplorationThe episode discusses the paper “Sample Efficient Preference Alignment in LLMs via Active Exploration,” arguing that aligning LLMs to be helpful/harmless is costly because it requires many human…6 Sep 2025 · 15 min · 5 chapters
Adventures in Demand Analysis Using AIUsing AI to estimate consumer demand price elasticity from product details, improving on traditional econometric methods that understate price effects due to confounding factors.4 Sep 2025 · 14 min · 5 chapters
Memento: Fine-tuning LLM Agents without Fine-tuning LLMsMemento proposes “fine-tuning LLM agents without fine-tuning LLMs” by enabling continuous, real-time learning for autonomous LLM agents using an external memory (case bank) rather than retraining the…1 Sep 2025 · 19 min · 9 chapters
On the Theoretical Limitations of Embedding-Based RetrievalTheoretical and empirical limits of embedding-based (single-vector) retrieval models, showing they cannot represent all possible “top-K relevant document sets” for a query unless the embedding…31 Aug 2025 · 17 min · 12 chapters
Performance Prediction for Large Systems via Text-to-Text RegressionA Google Research approach for predicting performance of large, complex industrial systems using text-to-text regression with regression language models (RLMs), avoiding lossy tabular feature…30 Aug 2025 · 16 min · 8 chapters
Demystifying the Visual Quality Paradox in Multimodal Large Language ModelsMultimodal large language models (MLLMs) can perform better on some vision-language tasks when images are degraded (blur/noise/fog) rather than pristine, due to a “visual quality paradox.” Guests:…30 Aug 2025 · 17 min · 10 chapters
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RLChain-of-Agents (COA) proposes end-to-end “agent foundation models” (AFMs) that simulate multi-agent collaboration inside one model, avoiding slow multi-agent back-and-forth while retaining planning,…30 Aug 2025 · 20 min · 13 chapters
Compute-Optimal Scaling for Value-Based Deep RLCompute-optimal scaling for value-based deep reinforcement learning, focusing on how to allocate limited compute between model size, batch size, and update-to-data (UTD) ratio to improve data…25 Aug 2025 · 16 min · 12 chapters
LLM-based Conversational Recommendation Agents with Collaborative Verbalized ExperienceHow LLM-based conversational recommendation agents can learn user preferences from dialogue history, reflect on past interactions, and improve diversity via multi-agent debate.23 Aug 2025 · 17 min · 9 chapters
Signal and Noise: Evaluating Language Model BenchmarksThe episode explains why common LLM benchmarks often fail to predict performance at production scale, and introduces a “signal and noise” framework to quantify benchmark reliability.23 Aug 2025 · 12 min · 4 chapters
Breaking Feedback Loops in Recommender Systems with Causal InferenceRecommender systems can create harmful feedback loops because their recommendations shape the user data used for training, biasing future recommendations and potentially causing homogenization (“rich…21 Aug 2025 · 13 min · 6 chapters