PLAY PODCASTS
Daily Paper Cast

Daily Paper Cast

2,319 episodes — Page 2 of 47

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

Sep 15, 202620 min

Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

Sep 15, 202620 min

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

Sep 15, 202621 min

PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models

Sep 15, 202622 min

Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction

Sep 15, 202622 min

NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

Sep 11, 202624 min

SenseNova-U1.5: Towards Native Unified Visual Intelligence

Sep 11, 202622 min

EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents

Sep 11, 202620 min

SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem

Sep 11, 202622 min

Show-Harness: Just a VLM Agent Can Play Robots

Sep 10, 202621 min

Programmable World Model

Sep 10, 202619 min

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

Sep 10, 202621 min

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

Sep 9, 202619 min

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

Sep 9, 202621 min

Omni Interaction Agent Technical Report

Sep 9, 202622 min

DriveZero: End-to-End Driving Beyond Human Demonstrations

Sep 9, 202622 min

GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

Sep 9, 202620 min

Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout

Sep 9, 202620 min

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

Sep 9, 202619 min

Miles v0.1: Production-Level Post-Training

Sep 9, 202620 min

SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution

Sep 9, 202618 min

LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

Sep 4, 202622 min

LatentPress: Context Compression Beyond Text and Vision

Sep 4, 202620 min

Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM

Sep 4, 202620 min

Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning

Sep 4, 202620 min

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

Sep 4, 202622 min

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

Sep 4, 202622 min

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

Sep 4, 202622 min

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

Sep 3, 202621 min

SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

Sep 3, 202620 min

EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction

Sep 3, 202621 min

It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning

Sep 3, 202620 min

Language Models Can Control Their Own Attention

Sep 3, 202620 min

On the Design Fundamentals of Pixel Text Representation Learning

Sep 3, 202619 min

StudentSim: Training LLM-based Student Simulators

Sep 2, 202620 min

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

Sep 2, 202621 min

UI-Venus-2 Technical Report

Sep 2, 202623 min

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

Sep 2, 202621 min

ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

Sep 2, 202620 min

H3-World: Turning Language Understanding into World Control

Sep 2, 202621 min

Hi-Q: Hierarchical Evidence-guided Query Refinement for Multi-Hop Question Answering

Sep 2, 202619 min

DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

Sep 1, 202625 min

GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling

Sep 1, 202620 min

Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling

Sep 1, 202622 min

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

Sep 1, 202622 min

PaperGym: Rubric-Centered Evolution for Research-Plan Generation

Sep 1, 202622 min

Normalized Low-Rank Adaptation

Sep 1, 202619 min

CogEvol: Towards Efficient and Reliable Learning Environment Generation

Sep 1, 202618 min

Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models

Aug 28, 202620 min

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

Aug 28, 202621 min