Best AI papers explained cover art
Podcast · 475 episodes

Best AI papers explained, page 18

by Enoch H. Kang · English

Cut through the noise. We curate and break down the most important AI papers so you don’t have to.

All episodes, page 18

RAG is Dead, Context Engineering is King: Building Reliable AI SystemsArgues “RAG is dead” as a misleading label; retrieval still matters, but teams should replace demo-driven “alchemy” with “context engineering” to reliably assemble the right information for LLMs.20 Aug 2025 · 20 min · 9 chapters
A Survey of Personalization: From RAG to AgentThe episode surveys how AI personalization is added to Retrieval-Augmented Generation (RAG) and how that evolves into personalized agent systems, using the academic paper “Survey of Personalization:…20 Aug 2025 · 25 min · 11 chapters
Facilitating the Adoption of Causal Infer-ence Methods Through LLM-Empowered Co-PilotTreatment effect estimation (TE)—estimating causal impact of a “treatment” X on an outcome Y—especially from observational data where correlation ≠ causation.19 Aug 2025 · 22 min · 10 chapters
Performance Prediction for Large Systems via Text-to-Text RegressionA Google/Cornell research paper on “text-to-text regression” using regression language models (RLMs) to predict performance of large, complex systems (e.g., Google’s Borg compute scheduler) by…16 Aug 2025 · 19 min · 12 chapters
Sample More to Think Less: Group Filtered Policy Optimization for Concise ReasoningThe episode explains “length explosion” in LLMs—why models trained with reward signals can become overly verbose—and how the ARCS paper “Sample More to Think Less: Group Filtered Policy Optimization…15 Aug 2025 · 28 min · 8 chapters
DINOv3: Vision Models for Self-Supervised LearningMeta AI’s DINOv3, a self-supervised vision model (“universal visual encoder”) trained without human pixel labels, emphasizing improved dense/patch-level features for tasks like segmentation, depth,…15 Aug 2025 · 20 min · 9 chapters
Agent Lightning: Training Any AI Agents with Reinforcement LearningAgent Lightning (Microsoft Research) trains multi-step AI agents with reinforcement learning so they improve reliably in messy real-world tasks, without rewriting existing agent code.14 Aug 2025 · 20 min · 12 chapters
Computational-Statistical Tradeoffs at the Next-Token Prediction BarrierThe episode explains “error amplification” in next-token prediction (autoregressive language models) and imitation learning, focusing on a 2025 PMLR paper, Computational-Statistical Tradeoffs at the…14 Aug 2025 · 12 min · 5 chapters
From Model Weights to Agent Workflows: Charting the New Frontier of Optimization in Large Language ModelsThe episode explains how AI optimization is shifting from tuning single LLM “model weights” (e.g., RLHF with KL-regularized objectives) to optimizing modular LLM agents (workflow programs with…12 Aug 2025 · 17 min · 11 chapters
Is Chain-of-Thought Reasoning a Mirage?Whether chain-of-thought (CoT) prompting in large language models reflects genuine reasoning or a “brittle mirage” caused by training-data pattern matching; includes a “data distribution lens” and…12 Aug 2025 · 19 min · 14 chapters
Agentic Web: Weaving the Next Web with AI AgentsThe “agentic web” paradigm shift where AI agents autonomously perceive, plan, and execute goal-driven tasks across the internet, replacing manual clicking/searching.11 Aug 2025 · 22 min · 8 chapters
The Assimilation-Accommodation Gap in LLM IntelligenceThe episode argues that today’s LLM “intelligence” is mainly next-token prediction plus assimilation into fixed or prompted schemata, not true accommodation (restructuring schemas when they fail).10 Aug 2025 · 23 min · 10 chapters
The Minimalist AI Kernel: A New Frontier in ReasoningThe episode argues that AI reasoning can be separated from brute-force scaling via a “reasoning core” and that an LLM can function as an operating system (LLMOS) that orchestrates external tools.6 Aug 2025 · 19 min · 12 chapters
Statistical Rigor for Interpretable AIStatistical rigor for mechanistic interpretability (MI) in AI—how to reverse-engineer neural networks by finding features and circuits, while avoiding statistical pitfalls from dependent “cheap” data…6 Aug 2025 · 18 min · 11 chapters
Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of ValueFull-stack alignment argues that aligning AI only to its operator’s local objective (e.g., profit or engagement) can still harm society; it proposes “thick models of value” (TMV) to preserve and…4 Aug 2025 · 22 min · 10 chapters
A foundation model to predict and capture human cognitionThe Nature (July 2, 2025) paper “A Foundation Model to Predict and Capture Human Cognition” introduces Centaur, a foundation model fine-tuned on human behavioral data to predict and simulate human…4 Aug 2025 · 19 min · 7 chapters
Generative Recommendation with Semantic IDs: A Practitioner’s HandbookGenerative recommendation using Semantic IDs (SIDs), focusing on Snap Inc.’s open-source GRD framework.4 Aug 2025 · 17 min · 8 chapters
Hierarchical Reasoning ModelThe episode explains a new “Hierarchical Reasoning Model” (HRM) for multi-step AI reasoning, arguing that flat LLM-style chain-of-thought can be brittle and computationally shallow, while…4 Aug 2025 · 12 min · 5 chapters
Test-time Offline Reinforcement Learning on Goal-related ExperienceGoal Conditioned Test Time Training (GCTTT) for offline reinforcement learning—dynamically fine-tuning a goal-conditioned policy during evaluation using only relevant, high-value past experience,…4 Aug 2025 · 14 min · 5 chapters
Interpreting Chain of Thought: A Walkthrough and DiscussionHow to interpret LLM chain-of-thought using “Thought Anchors” (a circular, color-coded graph of sentence-level reasoning) plus counterfactual importance and resampling sentence-to-sentence importance…4 Aug 2025 · 15 min · 6 chapters
The wall confronting large language modelsThe episode argues that large language models (LLMs) face a “wall” of diminishing returns: scaling up parameters/data yields rapidly worsening reliability and uncertainty, making them unsuitable for…4 Aug 2025 · 18 min · 9 chapters
COLLABLLM: LLMs From Passive to CollaborativeThe episode argues current LLMs feel “passive” because training rewards next-turn answers, not long-term task success, leading to inefficient back-and-forth when user intent is unclear.31 Jul 2025 · 18 min · 12 chapters
A decade's battle on dataset bias: are we there yet?The episode discusses a modern re-test of the “Name That Dataset” experiment, showing that deep neural networks can reliably identify which dataset an image came from, even when humans struggle.29 Jul 2025 · 16 min · 9 chapters
GEPA: Generative Feedback for AI System OptimizationGEPA (Genetic Pareto) is a prompt optimizer for LLM system optimization that targets sample inefficiency in reinforcement-learning-style prompt tuning (e.g., GRPO/GRPO-like RLVR), which can require…29 Jul 2025 · 15 min · 7 chapters