Best AI papers explained cover art
Podcast · 475 episodes

Best AI papers explained, page 14

by Enoch H. Kang · English

Cut through the noise. We curate and break down the most important AI papers so you don’t have to.

All episodes, page 14

GST-UNet: A Neural Framework for Spatiotemporal Causal Inference with Time-Varying ConfoundingGST-UNet, a neural framework for spatiotemporal causal inference with time-varying confounding, combining UNet-style spatial modeling with iterative G-computation to estimate counterfactual effects…5 Nov 2025 · 18 min · 8 chapters
Beyond a million tokens: benchmarking and enhancing long-term memory in llmsLong-term memory in LLMs beyond million-token context windows; introduces Beam benchmark and Light memory architecture to reduce contextual forgetfulness and improve long-range coherence.4 Nov 2025 · 15 min · 10 chapters
Agentic Economic ModelingAgentic economic modeling (AEM) uses LLMs as “economic agents” to generate counterfactual decision data, then corrects systematic LLM bias with a small amount of real human data, enabling econometric…3 Nov 2025 · 14 min · 7 chapters
Emergent Introspective Awareness in Large Language ModelsWhether large language models’ self-reports about “thoughts/feelings” are genuine introspection or confabulation, and how Anthropic researchers test “emergent introspective awareness” using concept…3 Nov 2025 · 16 min · 7 chapters
Can Large reasoning models self-train?Self-rewarded training (SRT) for large reasoning models—letting models “grade their own homework” using self-consistency/majority voting to create pseudo-rewards, aiming to reduce reliance on…1 Nov 2025 · 12 min · 4 chapters
ALITA-G: Self-Evolving Generative Agent for Agent GenerationAlita-G, a self-evolving generative agent framework that turns a general-purpose agent (“masquerade”) into a reusable domain expert by generating, abstracting, and curating specialized toolkits (“MCP…1 Nov 2025 · 16 min · 10 chapters
Self-improving LLM agents at test-timeTest-time self-improvement (TTSI) for LLM agents—learning during inference only for uncertain cases, using a three-step loop (H self-awareness, G self-data augmentation, T parameter-efficient…30 Oct 2025 · 19 min · 8 chapters
Offline RL by Reward-Weighted Fine-Tuning for Conversation OptimizationOffline reinforcement learning for LLM conversation optimization via reward-weighted fine-tuning.30 Oct 2025 · 15 min · 9 chapters
Language models are injective and hence invertibleThe episode argues that standard decoder-only transformers are structurally lossless because the mapping from prompt tokens to internal hidden states is (almost surely) injective, so hidden states…30 Oct 2025 · 12 min · 5 chapters
ReasoningBank: Scaling Agent Self-Evolving with Reasoning MemoryReasoningBank, a Google Cloud AI research memory framework for LLM agents that prevents “forgetfulness” by storing structured, transferable reasoning strategies distilled from both successes and…29 Oct 2025 · 15 min · 11 chapters
RLAD: Training LLMs to Discover AbstractionsRLAD (reinforcement learning with abstraction discovery) trains LLMs to solve hard, unseen reasoning problems by first generating concise “reasoning abstractions” (high-level strategies/guardrails)…29 Oct 2025 · 16 min · 15 chapters
How to Train Your Advisor: Steering Black-Box LLMs with ADVISOR MODELSAdvisor models for steering “black-box” LLMs via a trainable small model that generates per-request natural-language steering instructions, optimized with reinforcement learning while the large model…29 Oct 2025 · 13 min · 7 chapters
Self-improving LLM agents at Test-TimeThe episode discusses test-time self-improvement (TTSI) for LLM agents, aiming to reduce information overload and agent training cost by adapting only when the model is uncertain.27 Oct 2025 · 23 min · 11 chapters
KL-Regularized Reinforcement Learning is designed to Mode CollapseKL-regularized reinforcement learning (RL) fine-tuning for LLMs can mathematically force mode collapse, reducing output diversity.27 Oct 2025 · 16 min · 8 chapters
How do LLMs use their depth?How LLMs use their depth during inference, arguing for a “guess then refine” workflow plus “complexity-aware depth use” (more layers for harder parts).27 Oct 2025 · 12 min · 7 chapters
Thought Communication in Multiagent CollaborationThought communication (ThoughtCom/TCOM) for multiagent collaboration, enabling LLM agents to share latent “intent drivers” directly instead of exchanging language tokens, avoiding language…27 Oct 2025 · 17 min · 10 chapters
Reasoning with Sampling: Base Models Outperform RLThe episode argues that “power sampling” can make base LLMs outperform or match RL post-training methods (e.g., GRPO) for reasoning tasks using only inference-time sampling—no new training, data,…26 Oct 2025 · 16 min · 10 chapters
Continual Learning via Sparse Memory FinetuningContinual learning for LLMs using sparse memory fine-tuning (SMF) to prevent catastrophic forgetting.26 Oct 2025 · 14 min · 5 chapters
Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary PreferencesPluralistic AI alignment using Direct Preference Optimization when user preferences vary in unobserved ways; argues standard binary preference data (A vs B) is mathematically insufficient and…24 Oct 2025 · 12 min · 9 chapters
The Coverage Principle: How Pre-Training Enables Post-TrainingThe episode explains why next-token pre-training that minimizes cross-entropy (or sequence-level KL) can still fail after fine-tuning for coding/reasoning, and introduces the “coverage principle.” It…24 Oct 2025 · 16 min · 12 chapters
The Era of Real-World Human Interaction: RL from User ConversationsThe episode argues alignment and personalization for AI assistants should shift from static RLHF (human-annotated A/B comparisons) to RLHI, reinforcement learning from human interaction using real…24 Oct 2025 · 14 min · 6 chapters
Agent Learning via Early ExperienceEarly experience for training autonomous language agents—using the agent’s own explorations and failures to learn without dense external rewards, bridging imitation learning (IL) and reinforcement…24 Oct 2025 · 13 min · 5 chapters
Demystifying the Mechanisms Behind Emergent Exploration in Goal-conditioned RLDemystifies emergent exploration in goal-conditioned RL via Single Goal Contrastive RL (SGCRL), arguing exploration comes from learned representation structure (not bigger networks), using InfoNCE…22 Oct 2025 · 15 min · 11 chapters
Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior“Rewriting history” interventional analyses for causal study of how pretraining data affects LLM factual knowledge.22 Oct 2025 · 19 min · 7 chapters