Best AI papers explained cover art
Podcast · 475 episodes

Best AI papers explained, page 11

by Enoch H. Kang · English

Cut through the noise. We curate and break down the most important AI papers so you don’t have to.

All episodes, page 11

Detailed Balance in Large Language Model-Driven AgentsLLM-driven agents (with memory, tool use, code, iteration) are modeled as state-space dynamics that obey “detailed balance,” implying an underlying macroscopic potential function learned during…28 Dec 2025 · 12 min · 6 chapters
Learning to reason in LLMs by expectation maximizationThe episode explains a research paper that frames LLM “chain-of-thought” training as a latent-variable model approximated by expectation maximization (EM).28 Dec 2025 · 14 min · 7 chapters
Exploratory Causal Inference in SAEnceExploratory causal inference (ECI) for modern science using learned measurements from foundation models and sparse autoencoders, addressing the “Matthew effect” and a paradox where higher power…25 Dec 2025 · 15 min · 10 chapters
Detailed balance in large language model-driven agentsLLM-driven agents (with memory, tool use, and code) exhibit “detailed balance,” implying a macroscopic physical law: transitions between agent states follow a global potential function.24 Dec 2025 · 12 min · 5 chapters
The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified AutoencodingThe episode explains the PRISM hypothesis and its model implementation, Unified Autoencoding (UAE), which aim to unify semantic (meaning) and pixel (visual detail) representations in one latent space…24 Dec 2025 · 16 min · 7 chapters
Adaptation of Agentic AIAgentic AI adaptation roadmaps—four paradigms (A1, A2, T1, T2) for making AI autonomous and improving it, plus safety risks and the move toward hybrid co-adaptation.23 Dec 2025 · 13 min · 6 chapters
Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL FinetuningChallenges with standard behavioral cloning (BC) as pretraining for robotics RL fine-tuning; proposes Posterior Behavioral Cloning (POSTBC) to improve “demonstrator action coverage” and sample…22 Dec 2025 · 11 min · 6 chapters
Let’s (not) just put things in Context: Test-Time Training for Long-Context LLMsLong-context LLMs fail at “needle in a haystack” retrieval due to score dilution in self-attention; the episode explains the math and presents Query-only Test-Time Training (QTTT) as a fix.21 Dec 2025 · 14 min · 7 chapters
TabPFN-2.5: Advancing the State of the Art in Tabular Foundation ModelsTabPFN-2.5, a “tabular foundation model” for structured data that aims to eliminate per-dataset hyperparameter tuning while improving accuracy, calibration, and scalability.20 Dec 2025 · 15 min · 10 chapters
What’s In My Human Feedback? Learning Interpretable Descriptions of Preference DataExplains “WIMHF” (What’s In My Human Feedback), a method to make human preference data interpretable by discovering sparse, human-readable features behind which response wins in pairwise comparisons.19 Dec 2025 · 16 min · 8 chapters
Bolmo: Byteifying the Next Generation of Language ModelsThe episode explains tokenization as a foundational barrier in LLMs and introduces BOLMO, a byte-level “tokenization-free” (UTF-8 bytes) open LLM family that removes subword tokenization issues via…19 Dec 2025 · 13 min · 7 chapters
What happened with sparse autoencoders?The “sparse autoencoder (SAE) saga” in AI interpretability—why SAEs initially looked like they could recover monosemantic features, how evaluation/analysis traps broke that optimism, and how the…17 Dec 2025 · 30 min · 17 chapters
What Matters Right Now in Mechanistic InterpretabilityMechanistic interpretability (MI) must pivot to keep up with newer, more capable “agentic” and reasoning models, focusing on diagnosis for safety rather than surgical control.16 Dec 2025 · 33 min · 15 chapters
CLaRa: Bridging Retrieval and Generation with Continuous Latent ReasoningRetrieval-Augmented Generation (RAG) inefficiencies and hallucinations; introduces CLaRa (Continuous Latent Reasoning) to unify retrieval and generation via continuous compressed “memory tokens,”…16 Dec 2025 · 15 min · 7 chapters
Self-Improving AI and Human Co-Improvement for Safer Co-SuperintelligenceWhether AI should pursue autonomous self-improvement (recursive, potentially unbounded) versus “co-improvement” with humans to achieve safer “co-superintelligence,” accelerating discovery while…16 Dec 2025 · 13 min · 10 chapters
Towards a Science of Scaling Agent Systems / Google DeepmindDeepMind research argues there’s no universal “more agents is better” rule; multi-agent LLM performance depends on task structure and coordination topology, with major cost/fragility tradeoffs.15 Dec 2025 · 16 min · 10 chapters
Emergent hierarchical reasoning in LLMs through reinforcement learningHow reinforcement learning reveals emergent hierarchical reasoning in LLMs, explaining “aha moments” and length scaling, and proposing Hierarchy Aware Credit Assignment (HICRA/IICRA) to target…14 Dec 2025 · 13 min · 8 chapters
AI revolution finally comes to Relational foundational models for structured dataRelational foundation models (RFMs) for structured enterprise data—treating databases as graphs so a frozen, pre-trained transformer can make fast, forward-looking predictions without manual feature…13 Dec 2025 · 15 min · 10 chapters
REFRAG: Rethinking RAG based DecodingReFRAG (Rethinking RAG-based decoding) speeds up retrieval-augmented generation by compressing RAG context to avoid quadratic TTFT latency and KV-cache memory blowups.13 Dec 2025 · 14 min · 4 chapters
Provable Long-Range Benefits of Next-Token PredictionA complexity-theoretic argument that autoregressive next-token prediction training can provably yield long-range structural coherence, despite being locally optimized and computationally bounded.12 Dec 2025 · 12 min · 5 chapters
Jeff Dean on TPUs, AI Research, and FundingHow Google’s TPUs and the software stack (JAX, XLA, Pathways) enable planetary-scale AI training, and why public academic research and new funding models are essential for future breakthroughs.12 Dec 2025 · 38 min · 18 chapters
Latent Debate: surrogate framework for Interpreting LLM Thinking“Latent Debate” proposes a surrogate framework that models an LLM’s hidden internal “support vs attack” conflict to predict and explain hallucinations, aiming to bridge capability vs reliability.11 Dec 2025 · 15 min · 8 chapters
Distribution-calibrated inference time compute for thinking llm-as-a-judgeHow to make “thinking LLM as a judge” evaluations reliable by transforming noisy, stochastic judge votes into calibrated ratings using distribution-calibrated inference-time compute (ITC) and a…11 Dec 2025 · 12 min · 4 chapters
Principled RL for diffusion LLMs emerges from sequence level perspectiveWhy standard token-level RL (e.g., GRPO/PTO) fails for diffusion LLMs, and how ESPO (ELBO-based sequence-level policy optimization) fixes it by treating the whole sequence as one action and using a…11 Dec 2025 · 12 min · 6 chapters