ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training Paper • 2609.00188 • Published 10 days ago • 51
GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling Paper • 2608.29335 • Published 12 days ago • 67
CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing Paper • 2608.17566 • Published 23 days ago • 16
Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers Paper • 2607.28611 • Published Jul 30 • 22
BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms Paper • 2607.26497 • Published Jul 30 • 52
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation Paper • 2607.27816 • Published Jul 30 • 34
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing Paper • 2607.27616 • Published Jul 30 • 39
Vera: A Layered Diffusion Model for Content-Preserving Video Editing Paper • 2606.23610 • Published Jun 22 • 13
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models Paper • 2606.03988 • Published Jun 3 • 127
MilliVid: Hierarchical Latents for Long-Range Consistency in Video Generation Paper • 2606.09056 • Published Jun 8 • 8
minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models Paper • 2605.30263 • Published May 28 • 62
CoVEBench: Can Video Editing Models Handle Complex Instructions? Paper • 2606.08415 • Published Jun 7 • 53
Running on Zero Agents Featured 77 VGGT-Omega Demo 🌀 77 3D reconstruction from images/video with VGGT-Omega
SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer Paper • 2605.15178 • Published May 14 • 91