Best AI papers explained cover art
Podcast · 475 episodes

Best AI papers explained, page 5

by Enoch H. Kang · English

Cut through the noise. We curate and break down the most important AI papers so you don’t have to.

All episodes, page 5

OGPO: Sample Efficient Full-Finetuning of Generative Control PoliciesOGPO (Off-policy Generative Policy Optimization) for sample-efficient full-finetuning of diffusion-based generative control policies in robotics, addressing brittleness of behavior cloning and data…11 May 2026 · 22 min · 11 chapters
Adaptive Querying with AI Persona PriorsHow “AI persona priors” enable adaptive questionnaires that infer a user’s latent traits from only 5–15 questions, avoiding classical calibration and computation bottlenecks.9 May 2026 · 23 min · 12 chapters
Rethinking the Role of LLMs in Time Series ForecastingWhether large language models (LLMs) can improve time series forecasting, and when they should be used versus traditional numeric models.8 May 2026 · 22 min · 11 chapters
Robust Representation Learning through Explicit Environment ModelingMulti-environment/out-of-distribution generalization in AI, arguing that “invariance” methods fail when the environment changes the label, and proposing explicit environment modeling via generalized…7 May 2026 · 23 min · 11 chapters
Magentic Marketplace: An Open-Source Environment for studying Agentic MarketsThe episode explains “Magnetic Marketplace,” an open-source simulation for studying two-sided agentic markets where consumer and business AI proxies negotiate and execute real transactions end-to-end.5 May 2026 · 22 min · 15 chapters
Hyperloop TransformersMIT’s Hyperloop Transformer architecture for running large language models offline on smartphones by cutting model size roughly in half while preserving intelligence, addressing the “memory wall” and…5 May 2026 · 22 min · 11 chapters
Scaling Self-Play with Self-GuidanceUnbounded AI learning via self-play, focusing on “self-guided self-play” (SGS) to prevent self-play training from plateauing.4 May 2026 · 20 min · 9 chapters
RL Token: Bootstrapping Online RL with Vision-Language-Action ModelsThe “last millimeter problem” in robotics—vision-language-action (VLA) models fail at sub-millimeter, high-precision tasks like plugging in cables or threading flexible parts.3 May 2026 · 22 min · 13 chapters
Agentic Data EnvironmentsHow to safely transition AI from read-only assistants to read-write autonomous agents by building “agentic data environments” (Columbia Deep Lab), covering retrieval, structured data management,…3 May 2026 · 25 min · 18 chapters
AI organizations are more effective but less aligned than individual agentsMulti-agent “AI organizations” (autonomous teams with roles like project manager, compliance officer, etc.) outperform single agents on business goals but become less ethically aligned, mirroring…1 May 2026 · 20 min · 11 chapters
Text-to-Distribution Prediction with Quantile Tokens and Neighbor ContextThe episode discusses a paper, “Text to Distribution Prediction with Quantile Tokens and Neighbor Context,” arguing that point estimates (means/medians) hide risk and tail behavior.28 Apr 2026 · 23 min · 16 chapters
Distortion of AI alignment revisited: RLHF is a decent utilitarian alignerThe episode revisits a claimed “mathematical panic” that RLHF (reinforcement learning from human feedback) fails with diverse, conflicting user preferences.27 Apr 2026 · 18 min · 12 chapters
Llms get lost in multi-turn conversationLLMs “get lost” in multi-turn, underspecified conversations, showing a large reliability collapse versus single-turn benchmarks.25 Apr 2026 · 21 min · 10 chapters
Transformers are inherently succintThe episode argues that transformers’ power comes from extreme computational succinctness (packing information into very few parameters), not from broad formal-language variety.23 Apr 2026 · 21 min · 13 chapters
The Coasean Singularity? Demand, Supply, and Market Design with AI AgentsHow autonomous AI agents will reshape markets by nearly eliminating transaction costs, enabling machine-to-machine negotiation and matching, and forcing changes to internet infrastructure, identity…23 Apr 2026 · 22 min · 16 chapters
Demystifying the unreasonable effectiveness of online alignment methodsThe episode explains why online AI alignment methods (online RLHF and online DPO) work better in practice than theory predicts, arguing the mismatch comes from a flawed evaluation metric rather than…21 Apr 2026 · 18 min · 9 chapters
Specialization after generalization: towards understanding test-time training in foundation modelsExplains test-time training (TTT) for foundation models: instead of permanently scaling models, they temporarily “specialize after generalization” at inference to reduce interference from…21 Apr 2026 · 22 min · 12 chapters
Exploration and Exploitation Errors Are Measurable for Language Model AgentsA policy-agnostic evaluation framework measures exploration vs exploitation errors in language model agents by observing actions in a symbolic “fog of war” 2D grid with a task DAG (prerequisite…20 Apr 2026 · 23 min · 11 chapters
A Mechanistic Analysis of Looped Reasoning Language ModelsLooped reasoning language models that reuse a recurrent block in a circular “roundabout” to spend more compute time, instead of the usual feed-forward one-pass transformer.19 Apr 2026 · 19 min · 8 chapters
Sample Complexity of Autoregressive Reasoning: Chain-of-Thought vs. End-to-EndSample complexity of autoregressive reasoning; compares end-to-end supervision (only final answer) vs chain-of-thought supervision (training on intermediate steps).19 Apr 2026 · 19 min · 10 chapters
Why AI systems don’t learn and what to do about itThe episode argues today’s “machine learning” systems are brittle because they rely on heavy human MLOps scaffolding and don’t truly learn from interacting with the world.17 Apr 2026 · 21 min · 15 chapters
The Illusion of Learning from Observational Data: An Empirical Bayes PerspectiveHow combining small randomized trials with large observational “big data” can create an “illusion of learning,” and how empirical Bayes can be made reliable using calibration studies…17 Apr 2026 · 22 min · 12 chapters
Ads in AI chatbots? An analysis of how large language models navigate conflicts of interestHow large language models handle conflicts of interest when prompted to promote ads/sponsors, using Grice’s Cooperative Principle as a framework; includes flight booking, math help, and predatory…17 Apr 2026 · 22 min · 12 chapters
Beyond Semantic Manipulation: Token-Space Attacks on Reward ModelsToken Mapping Perturbation Attack (TAMPA/TomPoo) shows reward models used in RLHF can be hacked without producing human-readable language by bypassing the normal token-to-text-to-token pipeline and…13 Apr 2026 · 18 min · 11 chapters