Best AI papers explained cover art
Podcast · 475 episodes

Best AI papers explained, page 10

by Enoch H. Kang · English

Cut through the noise. We curate and break down the most important AI papers so you don’t have to.

All episodes, page 10

Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language ModelsProposes “conditional memory” to complement Mixture-of-Experts (MoE).16 Jan 2026 · 14 min · 6 chapters
Learning Latent Action World Models In The WildLearning latent action world models (LAMs) from massive unlabeled “in-the-wild” video to enable planning without action labels. Guests: No named guests; the episode is a single host-led deep dive.16 Jan 2026 · 14 min · 10 chapters
From Unstructured Data to Demand Counterfactuals: Theory and PracticeDemand estimation for differentiated products using unstructured-data proxies (images/text) in random-coefficients logit/BLP-style models; proposes a post-estimation bias correction plus proxy…14 Jan 2026 · 14 min · 6 chapters
In-context reinforcement learning through bayesian fusion of context and value priorIn-context reinforcement learning with SPICE-E (SPICY) to overcome behavior-policy bias and noisy offline data using Bayesian fusion of a context evidence term with a value prior, plus posterior-UCB…14 Jan 2026 · 12 min · 7 chapters
Digital RedQueen: Adversarial Program Evolution in Core War with LLMsThe episode explains the “Digital Red Queen” (DRQ) algorithm, which uses an LLM (GPT-4.1 mini) to evolve Core War programs in a self-play arms race, aiming for robust generalists rather than brittle…14 Jan 2026 · 14 min · 7 chapters
Extending the Context of Pretrained LLMs by Dropping Their Positional EmbeddingsThe episode explains why pretrained LLMs struggle with long-context “context walls,” attributing it to rotary positional embeddings (RoPE) enabling efficient training but causing an extrapolation…13 Jan 2026 · 12 min · 8 chapters
Representation-Based Exploration for Language Models: from test-time to post-trainingRepresentation-based exploration (REPEX) for language models to enable deliberate, conceptually novel discovery instead of “sharpening” via RL; covers inference-time selection and post-training…12 Jan 2026 · 14 min · 5 chapters
NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and GenerationNextFlow, a unified sequential (decoder-only transformer) model that combines text/image understanding and generation, aiming to match diffusion-model visual quality while retaining LLM-like…10 Jan 2026 · 15 min · 8 chapters
RelayLLM: Efficient Reasoning via Collaborative DecodingRelayLLM proposes “token-level collaborative decoding” so a small language model can request brief, precise help from a larger LLM during generation, avoiding all-or-nothing routing.10 Jan 2026 · 13 min · 8 chapters
A Unified Definition of Hallucination, Or: It’s the World Model, StupidDefines “hallucination” for LLMs via a unified WVP framework: hallucination is an observable mismatch between a model’s claims and an explicitly defined reference world model (W), given the model’s…8 Jan 2026 · 12 min · 4 chapters
Deep sequence models tend to memorize geometrically; it is unclear why.Deep sequence models (transformers, Mamba) store knowledge in two competing ways: fast local associative memory (next-token/neighbor lookups) and slower “geometric memory” (embeddings whose dot…8 Jan 2026 · 13 min · 6 chapters
From Entropy to Epiplexity: Rethinking Information for Computationally Bounded IntelligenceThe episode argues that classical information theory (Shannon/Kolmogorov) assumes an unlimited-compute observer, so it mismeasures what AI learns.8 Jan 2026 · 14 min · 7 chapters
Diffusion Language Models are Provably Optimal Parallel SamplersThe paper “Diffusion Language Models Are Provably Optimal Parallel Samplers” argues diffusion language models (DLMs) can achieve provably optimal inference latency and memory for sampling, and that…7 Jan 2026 · 12 min · 9 chapters
Universal Reasoning ModelUniversal Reasoning Model (URM) for hard algorithmic reasoning benchmarks (ARC-AGI1/2, Sudoku), arguing that smaller Universal Transformer (UT) architectures beat much larger vanilla LLMs by using…6 Jan 2026 · 14 min · 10 chapters
Recursive language modelsRecursive language models (RLMs) to overcome “context rot” and extend reliable reasoning from ~200k tokens to 10M+ tokens by externalizing long prompts into a persistent workspace the model can…6 Jan 2026 · 16 min · 7 chapters
Adapting fast and slow: transportable circuits for few shot learningCross-domain generalization via causal transportability for zero-shot and few-shot learning, contrasting “fast” vs “slow” adaptation.4 Jan 2026 · 15 min · 8 chapters
Position: Probabilistic Modelling is Sufficient for Causal InferenceArgues that causal inference (interventions and counterfactuals) can be done with standard probabilistic modeling by writing a full joint probability over “observed,” “intervened,” and…3 Jan 2026 · 12 min · 4 chapters
End-to-End Test-Time Training for Long ContextEnd-to-End Test-Time Training (TTT-E2E) for long-context LLMs, aiming for full-attention-like performance up to 128K tokens while keeping constant-cost inference.3 Jan 2026 · 14 min · 7 chapters
Parallel Token Generation for Language ModelsParallel Token Prediction (PTP) for language models to reduce slow sequential token-by-token generation latency by predicting multiple interdependent future tokens in a single forward pass.2 Jan 2026 · 16 min · 11 chapters
Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL FinetuningPosterior Behavioral Cloning (POSTBC) improves reinforcement learning (RL) fine-tuning by better pretraining policies, focusing on sample efficiency in expensive continuous robotics.31 Dec 2025 · 16 min · 12 chapters
Activation oracles: training and evaluating llms as general-purpose activation explainersActivation oracles (AOs) as “universal translators” that decode an LLM’s internal activations into natural-language answers for auditing and safety.30 Dec 2025 · 15 min · 10 chapters
Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learningHow “internal reinforcement learning” leverages temporally abstract sub-goals already encoded in frozen autoregressive models to enable efficient hierarchical RL under sparse rewards.29 Dec 2025 · 14 min · 6 chapters
Joint-Embedding vs Reconstruction: Provable Benefits of Latent Space PredictionSelf-supervised learning “label problem” and a theory comparing reconstruction-based SSL (SSLRC) vs joint-embedding SSL (SSLJE), showing why more data can’t fix bad augmentations.29 Dec 2025 · 14 min · 6 chapters
Monitoring Monitorability/ OpenAIThe episode argues that as frontier AI becomes more autonomous, “monitorability” is a load-bearing safety layer: reliable, verifiable detection of harmful or misaligned behavior during deliberation…28 Dec 2025 · 14 min · 7 chapters