Best AI papers explained cover art
Podcast · 475 episodes

Best AI papers explained, page 12

by Enoch H. Kang · English

Cut through the noise. We curate and break down the most important AI papers so you don’t have to.

All episodes, page 12

Algorithmic Thinking TheoryAlgorithmic Thinking Theory for LLMs on hard math (e.g., IMO-style problems), explaining why single-shot “pass@1” is low but multi-attempt “pass@K” can be high, and formalizing how to assemble…10 Dec 2025 · 17 min · 11 chapters
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language ModelsHow pre-training, mid-training (continued pre-training), and RL post-training interact to improve reasoning in language models, and when RL truly adds new reasoning vs merely polishing existing…10 Dec 2025 · 14 min · 4 chapters
Natural language actor-critic: Scalable off-policy learning in language spaceNatural Language Actor-Critic (NLAC) for training LLM agents with scalable off-policy reinforcement learning in language/action space, addressing long-horizon credit assignment and sparse rewards.9 Dec 2025 · 14 min · 5 chapters
Beyond the Transformer: Titans, MIRAS, and the Future of Infinite ContextLong-context AI memory beyond transformers, focusing on Google Research’s Titans architecture and the Miras framework for “test-time memorization” (updating long-term knowledge during inference) to…7 Dec 2025 · 39 min · 20 chapters
On the Limits of Test-Time Compute: Sequential Reward Filtering for Better InferenceHow to improve large language model accuracy at inference using limited “test-time compute” (TTC), focusing on sequential reward filtering (RF-SEC-BON) versus parallel Best-of-N (BON).7 Dec 2025 · 14 min · 4 chapters
The Universal Weight Subspace HypothesisThe Universal Weight Subspace Hypothesis argues that many deep neural networks trained on very different tasks converge to a shared, low-dimensional “universal” subspace in their weight matrices,…7 Dec 2025 · 16 min · 13 chapters
Stabilizing Reinforcement Learning with LLMs: Formulation and PracticesHow to stably combine reinforcement learning with LLMs by justifying a token-level surrogate objective for sequence-level rewards, and preventing failure modes when the key approximation breaks.7 Dec 2025 · 15 min · 7 chapters
Benchmarking In-context Experiential Learning Through Repeated Product RecommendationsWhether frontier LLMs can do in-context experiential learning—improving strategy over repeated interactions using ambiguous feedback—tested via a new product-recommendation benchmark (BLA).4 Dec 2025 · 16 min · 7 chapters
Training LLMs for Honesty via ConfessionsTraining LLMs to be honest via a separate “confession” output that self-reports noncompliance, deception, and uncertainty, countering reward misspecification and reward hacking.4 Dec 2025 · 16 min · 11 chapters
STOIC REASONER: Dual-Mode Transformers that Compress to Think and Decompress to SpeakExplains “Stoic Reasoner” (Soft Token Implicit Context Reasoner), a dual-mode transformer that compresses LLM reasoning into continuous “soft tokens” (latent, silent thinking) and only occasionally…4 Dec 2025 · 12 min · 5 chapters
E-GEO: A Testbed for Generative Engine Optimization in E-CommerceThe episode explains “E-GEO” (e-commerce generative engine optimization): how to improve a product’s rank in LLM-curated shopping recommendations, shifting from classic SEO (ranked links) to GEO…4 Dec 2025 · 33 min · 21 chapters
1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching CapabilitiesThe episode explains how reinforcement learning (RL) can be scaled to 224–1024-layer “1000 layer networks” using scaled contrastive RL (CRL), producing emergent goal-reaching behaviors and large…4 Dec 2025 · 15 min · 9 chapters
Treatment Effect Estimation for Optimal Decision-MakingHow to estimate conditional average treatment effects (CATE) for better real-world decisions, arguing that best statistical CATE fit can be policy-worse when the CATE model class is restricted;…4 Dec 2025 · 14 min · 7 chapters
Pass@K Policy Optimization: Solving Harder Reinforcement Learning ProblemsPass@K Policy Optimization (PKPO) for reinforcement learning of LLMs on objective, high-stakes tasks (code/unit tests, theorem correctness).3 Dec 2025 · 14 min · 8 chapters
Debugging misaligned completions with sparse-autoencoder latent attributionDebugging LLM misalignment by using sparse autoencoder (SAE) latent attribution to identify causal internal features, avoiding “model diffing” limits (correlation, need for sibling models, expensive…2 Dec 2025 · 30 min · 12 chapters
Building Effective AI Agents \ AnthropicHow to build reliable “AI agentic systems” from Anthropic, emphasizing least complexity, transparency, and especially the agent-computer interface (ACI) for tool reliability.2 Dec 2025 · 39 min · 15 chapters
How to Correctly Report LLM-as-a-Judge EvaluationsHow to report evaluations when an LLM judges another model’s outputs, arguing that “raw accuracy” is statistically biased and that confidence intervals must account for judge error.2 Dec 2025 · 12 min · 5 chapters
In-Context Learning with Hypothesis-Class GuidanceExplains in-context learning (ICL) using a testbed called ICL-HCG, where an instruction prefix describes the “hypothesis class” (the rule set to search for).2 Dec 2025 · 13 min · 6 chapters
Selecting Belief-State Approximations in Simulators with Latent StatesHow to choose belief-state approximations for POMDP simulators when you can’t perfectly reset hidden latent states, and how that choice affects Q-value estimation rollouts.1 Dec 2025 · 11 min · 6 chapters
Latent Collaboration in Multi-Agent SystemsLatent Collaboration in Multi-Agent Systems (Latent MAS) replaces slow, error-prone text-based communication (“TextMAS”) with instantaneous sharing of continuous latent “thoughts” between agents…29 Nov 2025 · 13 min · 7 chapters
CausalPFN: Amortized Causal Effect Estimation via In-Context LearningThe podcast explains the research paper “CausalPFN: Amortized Causal Effect Estimation via In-Context Learning” (RxC 1.2506.07918), arguing that a single transformer model can automate causal effect…28 Nov 2025 · 28 min · 14 chapters
DELTA: How Does RL Unlock and Transfer New Algorithms in LLMs?Whether RL fine-tuning makes LLMs learn genuinely new algorithms or only reveal latent skills; the episode centers on the “Delta” benchmark (Distributional Evaluation of Learnability and…28 Nov 2025 · 11 min · 9 chapters
Self-Boost via Optimal Retraining: An Analysis via Approximate Message PassingHow to optimally do iterative self-boost/retraining when supervised labels are noisy, using a Bayes Optimal Aggregator Function derived with Approximate Message Passing (AMP), including an on-sager…27 Nov 2025 · 15 min · 12 chapters
Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMsPrompted Policy Search (PropES) uses an LLM as the optimizer in reinforcement learning, merging linguistic “manuals” and numerical reward history to speed policy search and add interpretability…27 Nov 2025 · 31 min · 14 chapters