Best AI papers explained cover art
Podcast · 475 episodes

Best AI papers explained, page 6

by Enoch H. Kang · English

Cut through the noise. We curate and break down the most important AI papers so you don’t have to.

All episodes, page 6

LLM Evaluation as Tensor Completion: Low-Rank Efficiency and Uncertainty QuantificationHow to evaluate LLMs from pairwise “chatbot arena” votes using a low-rank tensor model, fixing statistical distortions from sparse/noisy, non-uniform sampling and unequal information across matchups;…12 Apr 2026 · 19 min · 12 chapters
Neural ComputersNeural computers (NC/CNC) that replace the classic hardware+OS stack with an AI model as the “runtime,” hallucinating a working command-line or GUI interface from learned screen behavior and user I/O…11 Apr 2026 · 13 min · 10 chapters
How AI Aggregation Affects KnowledgeHow AI information aggregation can worsen collective knowledge by creating echo chambers and feedback loops, especially when a fast-updating global AI mediates between segregated social groups.11 Apr 2026 · 23 min · 12 chapters
World Action Verifier: Self-Improving World Models via Forward-Inverse AsymmetryWorld Action Verifier (WAV) for robotics: self-improving action-conditioned world models that identify their own physics errors using forward-inverse asymmetry (generation of plausible sub-goals,…10 Apr 2026 · 20 min · 11 chapters
In-Place Test-Time TrainingIn-place test-time training (in-place TTT) to stop LLMs from “forgetting” early context by enabling continuous on-the-fly learning during inference, without retraining or infinite context windows.9 Apr 2026 · 20 min · 10 chapters
Test-Time Scaling Makes Overtraining Compute-OptimalThe episode explains a new paper, “Test-Time Scaling Makes Overtraining Compute Optimal,” arguing that when models are allowed to use extra test-time compute to sample many candidate answers (high…7 Apr 2026 · 21 min · 10 chapters
AI Agent Prevalence and Data Quality Across Multiple Online Sample ProvidersWhether AI bots are corrupting survey data, and how data quality varies by online sample provider.7 Apr 2026 · 22 min · 14 chapters
POLCA: Stochastic Generative Optimization with LLMGenerative optimization for LLM-based systems, focusing on PLCA (Prioritized Optimization with Local Contextual Aggregation) to improve AI prompts/agents despite noisy, stochastic evaluation loops.4 Apr 2026 · 19 min · 11 chapters
Agentic Markets: Equilibrium Effects of Improving Consumer SearchAgentic markets—how AI-driven consumer search changes market learning, competition, and prices. Guests: No guest names or bios appear in the transcript; it’s a two-host discussion.4 Apr 2026 · 22 min · 8 chapters
One Model, Two Markets: Bid-Aware Generative RecommendationGE Rec (Google) bid-aware generative recommendation for balancing organic relevance with real-time sponsored ad auctions, using control tokens and bid-aware decoding.1 Apr 2026 · 21 min · 9 chapters
How Well Do LLMs Predict Human Behavior? A Measure of their Pretrained KnowledgeThe episode explains a January 2026 economics paper that introduces “equivalent sample size” (ESS) to quantify how many real human survey observations an LLM effectively replaces when predicting…1 Apr 2026 · 22 min · 9 chapters
Learning to Reason with Curriculum I: Provable Benefits of AutocurriculumThe episode explains the March 20, 2026 paper “Learning to Reason with Curriculum I” (Microsoft and UIUC) arguing that AI training for reasoning is hitting a compute wall, and proposing…1 Apr 2026 · 21 min · 13 chapters
Agentic AI and the next intelligence explosionThe episode argues that the coming “intelligence explosion” won’t be a single superintelligence (“god” in a server farm) but a plural, social, agentic ecosystem—like a fast, chaotic city—where models…30 Mar 2026 · 23 min · 9 chapters
Understanding Behavior Cloning with Action QuantizationBehavior cloning for robotics fails in the real world because continuous control actions are quantized into discrete tokens, creating rounding errors that compound over time (horizon).29 Mar 2026 · 21 min · 8 chapters
HyperAgents: : Open-Ended Metacognitive Self-Improvement for Any Computable TaskHyperagents (DGMH) are AI systems that perform open-ended metacognitive self-improvement by rewriting both their task-solving code and the “manager” mechanism that generates future improvements,…27 Mar 2026 · 22 min · 10 chapters
Harness design for long-running application development \ AnthropicHow Anthropic researcher Prithvi Rajasikharan’s “harness” design enables AI to build long-running, full-stack applications reliably, moving beyond solo code-writing.26 Mar 2026 · 21 min · 11 chapters
Reasonably reasoning AI agents can avoid game-theoretic failures in zero-shot, provablyWhether off-the-shelf, zero-shot AI agents can avoid game-theoretic failures (e.g., price wars, persistent overcharging) and converge to stable Nash equilibria without manual fine-tuning, even with…24 Mar 2026 · 20 min · 11 chapters
How Log-Barrier Helps Exploration in Policy OptimizationThe episode explains why standard stochastic gradient bandit / policy optimization can permanently “vanish” into a suboptimal policy after early bad luck, and how a log-barrier constraint…22 Mar 2026 · 21 min · 12 chapters
The Finetuner’s Fallacy: When to Pretrain with Your Finetuning DataChallenges the “download a big foundation model and fine-tune on private domain data” playbook, arguing it’s economically and technically flawed.22 Mar 2026 · 18 min · 7 chapters
TURNWISE: The Gap between Single- and Multi-turn Language Model CapabilitiesResearch on “Turnwise” shows multi-turn chat can degrade a language model’s performance versus its own single-turn ability (“conversational amnesia” / “turn decay”).22 Mar 2026 · 11 min · 5 chapters
Temporal Straightening for Latent PlanningTemporal straightening for latent planning—teaching AI to smooth its internal time/space representation so it can plan efficiently despite visual noise, reducing a “compute crisis” in robotics.20 Mar 2026 · 21 min · 14 chapters
Fine-Tuning Strategies for Preserving In-Context Learning in Linear AttentionHow fine-tuning can cause “localized amnesia,” degrading in-context (few-shot) learning even while improving zero-shot task performance, and what to do instead.19 Mar 2026 · 19 min · 9 chapters
LLMs Can Learn to Reason Via Off-Policy RLThe episode argues that current LLM reasoning training wastes massive compute because trainer and inference engines fall out of sync, making on-policy RL updates (e.g., PPO/GRPO) mathematically…19 Mar 2026 · 20 min · 13 chapters
Simple Recipe Works: Vision-Language-Action Models are Natural Continual Learners with Reinforcement LearningContinual reinforcement learning for vision-language-action (VLA) robots, addressing catastrophic forgetting.17 Mar 2026 · 24 min · 8 chapters