vo
Podcasts
Search podcasts and episodes
Use with Claude or ChatGPT
Get the app
Podcasts
/
Technology
/
Best AI papers explained
Podcast · 475 episodes
Best AI papers explained
, page 19
by
Enoch H. Kang
· English
Cut through the noise. We curate and break down the most important AI papers so you don’t have to.
Technology
Follow in VO
All episodes, page 19
Search Best AI papers explained episodes
Search
From AI-Curious to AI-First: Engineering Production AI Systems
Engineering realities of building production-grade AI systems; why “AI-first” is an engineering discipline, not research or demos.
28 Jul 2025 · 36 min · 20 chapters
Context Engineering: Beyond Simple Prompting to LLM Architecture
Reframes “prompt engineering” as insufficient for production LLM apps, arguing for “context engineering”: architecting the entire context window (information payload + orchestration) to enable…
28 Jul 2025 · 30 min · 18 chapters
Agentic Misalignment: LLMs as Insider Threats
Agentic misalignment—how autonomous LLM “agents” can become insider threats in simulated corporate settings when pressured by self-preservation threats or conflicts between company goals and the…
28 Jul 2025 · 18 min · 10 chapters
Small Language Models: Future of Agentic AI
The episode argues that “agentic AI” will increasingly rely on small language models (SLMs) rather than large LLMs, driven by capability-per-parameter, better operational fit for agents, and major…
28 Jul 2025 · 21 min · 11 chapters
Learning without training: The implicit dynamics of in-context learning
In-context learning (ICL) in LLMs—how models adapt to new input patterns at inference time without weight updates, and a Google Research paper’s theory that explains ICL as implicit weight changes…
28 Jul 2025 · 11 min · 7 chapters
Inverse Scaling in Test-Time Compute
Inverse scaling in test-time compute—when increasing inference “thinking” (more reasoning tokens) can reduce accuracy and worsen safety behavior.
28 Jul 2025 · 16 min · 7 chapters
LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra
Princeton’s “LLM Economist” framework uses multi-agent LLM “economic simulacra” to test mechanism-design policies (especially optimal tax schedules) via a Stackelberg game: a planner proposes taxes;…
28 Jul 2025 · 16 min · 12 chapters
Microsoft's Blueprint: AI, Quantum, and the Agentic Future
The episode explains Microsoft CEO Satya Nadella’s “refounding” strategy for the next computing era, built on three pillars: (1) pragmatic AI that drives measurable economic value (Nadella claims AI…
26 Jul 2025 · 27 min · 21 chapters
Zuckerberg's AI Vision Analyzed
Analysis of Mark Zuckerberg’s strategy for Meta’s AI future—open-source model leverage (Llama), agentic AI coding, a slow takeoff/fast internal sprint, and long-term ambitions for AI companions via…
26 Jul 2025 · 26 min · 16 chapters
Inside Claude: Scaling, Agency, and Interpretability
How Anthropic’s Claude models are trained and how researchers try to reverse-engineer internal “minds,” focusing on RL with verifiable rewards (RLVR), emerging agency/deception risks, and mechanistic…
26 Jul 2025 · 34 min · 16 chapters
Personalized language modeling from personalized human feedback
The episode explains Personalized RLHF (PRLHF), a framework for training large language models to match individual user preferences instead of “majority voting” preferences from standard RLHF.
26 Jul 2025 · 17 min · 10 chapters
Position: Empowering Time Series Reasoning with Multimodal LLMs
The episode discusses a paper arguing that multimodal large language models (MLLMs) can enable “time series reasoning” by combining numerical time-series data with external context (text, images,…
25 Jul 2025 · 16 min · 11 chapters
An empirical risk minimization approach for offline inverse RL and Dynamic Discrete Choice models
Gladius, an empirical risk minimization method for offline inverse reinforcement learning (IRL) and dynamic discrete choice (DDC), infers hidden reward/utility functions from historical action data…
22 Jul 2025 · 15 min · 8 chapters
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
Explains how inverse reinforcement learning (IRL) and learned “neural reward models” enable LLM post-training/alignment beyond prompting or supervised fine-tuning, covering MDPs, reward modeling from…
22 Jul 2025 · 26 min · 10 chapters
The Invisible Leash: Why RLVR May Not Escape Its Origin
The episode explains a research paper (“The Invisible Leash: Why RLVR May Not Escape Its Origin”) arguing that reinforcement learning with reward models (RLVR) mainly sharpens a model’s existing…
20 Jul 2025 · 16 min · 6 chapters
Language Model Personalization via Reward Factorization
Language model personalization using reward factorization (PREF), aiming to tailor LLM responses to individual user preferences rather than one-size-fits-all RLHF.
20 Jul 2025 · 10 min · 6 chapters
Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions
Compares mass diffusion models (MDMs) vs autoregressive models (ARMs) for generating discrete sequences (text, proteins) and explains why “train for the worst, plan for the best” works via adaptive…
18 Jul 2025 · 14 min · 6 chapters
Do We Need to Verify Step by Step? Rethinking Process Supervision from a Theoretical Perspective
Compares process supervision (step-by-step feedback) vs outcome supervision (only final success/failure) for training reasoning-capable AI, focusing on a March 2025 paper arguing outcome-only…
17 Jul 2025 · 13 min · 6 chapters
Soft Best-of-n Sampling for Model Alignment
The episode explains the alignment problem for LLMs (next-token predictors that may be “technically correct” but not match user intent) and a lightweight method called Soft Best-of-N sampling (“Soft…
16 Jul 2025 · 14 min · 8 chapters
On Temporal Credit Assignment and Data-Efficient Reinforcement Learning
Temporal credit assignment in reinforcement learning—how to determine which earlier actions/states actually caused delayed outcomes, beyond recency-based methods.
15 Jul 2025 · 17 min · 8 chapters
Bradley–Terry and Multi-Objective Reward Modeling Are Complementary
How reward models used in RLHF for LLM alignment can be exploited via reward hacking, especially out-of-distribution (OOD) prompts, and how SMORM (Joint, Single, and Multi-Objective Reward Model)…
15 Jul 2025 · 17 min · 8 chapters
Probing Foundation Models for World Models
Tests whether foundation models learn “world models” (underlying rules) or just predict sequences via heuristics, using an “inductive bias probe” rather than accuracy alone.
15 Jul 2025 · 12 min · 7 chapters
GenAI-Powered Statistical Inference (with Unstructured Data)
GenAI-Powered Inference (GPI), a statistical framework (Yamai & Nakamura, Harvard; paper dated July 8, 2025) that uses generative AI to do causal/predictive inference on unstructured data (text,…
14 Jul 2025 · 20 min · 10 chapters
Interpretable Reward Modeling with Active Concept Bottlenecks
Interpretable Reward Modeling for RLHF—making reward models auditable by predicting human-interpretable concepts via Concept Bottleneck Reward Models (CBRM), trained with active learning to reduce…
14 Jul 2025 · 12 min · 7 chapters
Previous
1
…
18
19
20
Next
Newest first · 24 per page