
Daily Paper Cast
2,319 episodes — Page 6 of 47
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement
Jul 22, 202619 min
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
Jul 22, 202621 min
Apple-$π$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
Jul 22, 202622 min
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
Jul 21, 202619 min
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM
Jul 21, 202619 min
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Jul 21, 202620 min
Loop the Loopies!
Jul 21, 202619 min
xHC: Expanded Hyper-Connections
Jul 21, 202620 min
Cura 1T: Specialized Model for Agentic Healthcare
Jul 21, 202620 min
On-Policy Delta Distillation
Jul 21, 202619 min
RecGPT-V3 Technical Report
Jul 21, 202622 min
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
Jul 18, 202624 min
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
Jul 18, 202619 min
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
Jul 18, 202619 min
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
Jul 18, 202615 min
BadWAM: When World-Action Models Dream Right but Act Wrong
Jul 18, 202621 min
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation
Jul 18, 202622 min
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
Jul 18, 202620 min
Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes
Jul 18, 202621 min
From Pixels to States: Rethinking Interactive World Models as Game Engines
Jul 18, 202617 min
UniVR: Thinking in Visual Space for Unified Visual Reasoning
Jul 18, 202618 min
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
Jul 17, 202619 min
Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation
Jul 17, 202619 min
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
Jul 17, 202620 min
KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill
Jul 17, 202622 min
OvisOCR2 Technical Report
Jul 17, 202620 min
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails
Jul 17, 202619 min
MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors
Jul 17, 202618 min
GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch
Jul 17, 202621 min
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding
Jul 16, 202620 min
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation
Jul 16, 202619 min
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
Jul 16, 202619 min
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models
Jul 16, 202619 min
Weak-to-Strong Generalization via Direct On-Policy Distillation
Jul 15, 202619 min
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
Jul 15, 202622 min
4D Human-Scene Reconstruction from Low-Overlap Captures
Jul 15, 202619 min
LightMem-Ego: Your AI Memory for Everyday Life
Jul 15, 202618 min
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Jul 14, 202620 min
Scalable Visual Pretraining for Language Intelligence
Jul 14, 202618 min
Video Generation Models are General-Purpose Vision Learners
Jul 14, 202620 min
Vidu S1: A Real-Time Interactive Video Generation Model
Jul 11, 202622 min
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
Jul 11, 202624 min
UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
Jul 11, 202626 min
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
Jul 11, 202625 min
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning
Jul 10, 202626 min
Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation
Jul 10, 202625 min
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Jul 10, 202621 min
Infinite Worlds with Versatile Interactions
Jul 10, 202619 min
RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation
Jul 9, 202625 min
RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation
Jul 9, 202621 min