vo
Podcasts
Search podcasts and episodes
Use with Claude or ChatGPT
Get the app
Podcasts
/
Technology
/
Best AI papers explained
Podcast · 475 episodes
Best AI papers explained
, page 12
by
Enoch H. Kang
· English
Cut through the noise. We curate and break down the most important AI papers so you don’t have to.
Technology
Follow in VO
All episodes, page 12
Search Best AI papers explained episodes
Search
Algorithmic Thinking Theory
Algorithmic Thinking Theory for LLMs on hard math (e.g., IMO-style problems), explaining why single-shot “pass@1” is low but multi-attempt “pass@K” can be high, and formalizing how to assemble…
10 Dec 2025 · 17 min · 11 chapters
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
How pre-training, mid-training (continued pre-training), and RL post-training interact to improve reasoning in language models, and when RL truly adds new reasoning vs merely polishing existing…
10 Dec 2025 · 14 min · 4 chapters
Natural language actor-critic: Scalable off-policy learning in language space
Natural Language Actor-Critic (NLAC) for training LLM agents with scalable off-policy reinforcement learning in language/action space, addressing long-horizon credit assignment and sparse rewards.
9 Dec 2025 · 14 min · 5 chapters
Beyond the Transformer: Titans, MIRAS, and the Future of Infinite Context
Long-context AI memory beyond transformers, focusing on Google Research’s Titans architecture and the Miras framework for “test-time memorization” (updating long-term knowledge during inference) to…
7 Dec 2025 · 39 min · 20 chapters
On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
How to improve large language model accuracy at inference using limited “test-time compute” (TTC), focusing on sequential reward filtering (RF-SEC-BON) versus parallel Best-of-N (BON).
7 Dec 2025 · 14 min · 4 chapters
The Universal Weight Subspace Hypothesis
The Universal Weight Subspace Hypothesis argues that many deep neural networks trained on very different tasks converge to a shared, low-dimensional “universal” subspace in their weight matrices,…
7 Dec 2025 · 16 min · 13 chapters
Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
How to stably combine reinforcement learning with LLMs by justifying a token-level surrogate objective for sequence-level rewards, and preventing failure modes when the key approximation breaks.
7 Dec 2025 · 15 min · 7 chapters
Benchmarking In-context Experiential Learning Through Repeated Product Recommendations
Whether frontier LLMs can do in-context experiential learning—improving strategy over repeated interactions using ambiguous feedback—tested via a new product-recommendation benchmark (BLA).
4 Dec 2025 · 16 min · 7 chapters
Training LLMs for Honesty via Confessions
Training LLMs to be honest via a separate “confession” output that self-reports noncompliance, deception, and uncertainty, countering reward misspecification and reward hacking.
4 Dec 2025 · 16 min · 11 chapters
STOIC REASONER: Dual-Mode Transformers that Compress to Think and Decompress to Speak
Explains “Stoic Reasoner” (Soft Token Implicit Context Reasoner), a dual-mode transformer that compresses LLM reasoning into continuous “soft tokens” (latent, silent thinking) and only occasionally…
4 Dec 2025 · 12 min · 5 chapters
E-GEO: A Testbed for Generative Engine Optimization in E-Commerce
The episode explains “E-GEO” (e-commerce generative engine optimization): how to improve a product’s rank in LLM-curated shopping recommendations, shifting from classic SEO (ranked links) to GEO…
4 Dec 2025 · 33 min · 21 chapters
1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching Capabilities
The episode explains how reinforcement learning (RL) can be scaled to 224–1024-layer “1000 layer networks” using scaled contrastive RL (CRL), producing emergent goal-reaching behaviors and large…
4 Dec 2025 · 15 min · 9 chapters
Treatment Effect Estimation for Optimal Decision-Making
How to estimate conditional average treatment effects (CATE) for better real-world decisions, arguing that best statistical CATE fit can be policy-worse when the CATE model class is restricted;…
4 Dec 2025 · 14 min · 7 chapters
Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems
Pass@K Policy Optimization (PKPO) for reinforcement learning of LLMs on objective, high-stakes tasks (code/unit tests, theorem correctness).
3 Dec 2025 · 14 min · 8 chapters
Debugging misaligned completions with sparse-autoencoder latent attribution
Debugging LLM misalignment by using sparse autoencoder (SAE) latent attribution to identify causal internal features, avoiding “model diffing” limits (correlation, need for sibling models, expensive…
2 Dec 2025 · 30 min · 12 chapters
Building Effective AI Agents \ Anthropic
How to build reliable “AI agentic systems” from Anthropic, emphasizing least complexity, transparency, and especially the agent-computer interface (ACI) for tool reliability.
2 Dec 2025 · 39 min · 15 chapters
How to Correctly Report LLM-as-a-Judge Evaluations
How to report evaluations when an LLM judges another model’s outputs, arguing that “raw accuracy” is statistically biased and that confidence intervals must account for judge error.
2 Dec 2025 · 12 min · 5 chapters
In-Context Learning with Hypothesis-Class Guidance
Explains in-context learning (ICL) using a testbed called ICL-HCG, where an instruction prefix describes the “hypothesis class” (the rule set to search for).
2 Dec 2025 · 13 min · 6 chapters
Selecting Belief-State Approximations in Simulators with Latent States
How to choose belief-state approximations for POMDP simulators when you can’t perfectly reset hidden latent states, and how that choice affects Q-value estimation rollouts.
1 Dec 2025 · 11 min · 6 chapters
Latent Collaboration in Multi-Agent Systems
Latent Collaboration in Multi-Agent Systems (Latent MAS) replaces slow, error-prone text-based communication (“TextMAS”) with instantaneous sharing of continuous latent “thoughts” between agents…
29 Nov 2025 · 13 min · 7 chapters
CausalPFN: Amortized Causal Effect Estimation via In-Context Learning
The podcast explains the research paper “CausalPFN: Amortized Causal Effect Estimation via In-Context Learning” (RxC 1.2506.07918), arguing that a single transformer model can automate causal effect…
28 Nov 2025 · 28 min · 14 chapters
DELTA: How Does RL Unlock and Transfer New Algorithms in LLMs?
Whether RL fine-tuning makes LLMs learn genuinely new algorithms or only reveal latent skills; the episode centers on the “Delta” benchmark (Distributional Evaluation of Learnability and…
28 Nov 2025 · 11 min · 9 chapters
Self-Boost via Optimal Retraining: An Analysis via Approximate Message Passing
How to optimally do iterative self-boost/retraining when supervised labels are noisy, using a Bayes Optimal Aggregator Function derived with Approximate Message Passing (AMP), including an on-sager…
27 Nov 2025 · 15 min · 12 chapters
Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMs
Prompted Policy Search (PropES) uses an LLM as the optimizer in reinforcement learning, merging linguistic “manuals” and numerical reward history to speed policy search and add interpretability…
27 Nov 2025 · 31 min · 14 chapters
Previous
1
…
11
12
13
…
20
Next
Newest first · 24 per page