OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining Paper • 2609.07398 • Published 12 days ago • 77
GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation Paper • 2609.05588 • Published 15 days ago • 59
DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution Paper • 2608.31106 • Published 19 days ago • 102
LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering Paper • 2608.28281 • Published 22 days ago • 106
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation Paper • 2608.13489 • Published Aug 13 • 101
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model Paper • 2509.09372 • Published Sep 11, 2025 • 259
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Paper • 2608.01964 • Published Aug 3 • 185
LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies Paper • 2606.15768 • Published Jun 14 • 6
Peak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic Assessment Paper • 2607.13941 • Published Jul 15 • 2
OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data Paper • 2606.30019 • Published Jun 29 • 19
BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding Paper • 2606.31315 • Published Jun 30 • 77
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Paper • 2606.17030 • Published Jun 15 • 48
DreamX-World 1.0: A General-Purpose Interactive World Model Paper • 2606.16993 • Published Jun 15 • 116
Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution Paper • 2606.10917 • Published Jun 9 • 77
TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation Paper • 2605.22355 • Published May 21 • 179
Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos Paper • 2605.18233 • Published May 18 • 93
MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation Paper • 2512.18181 • Published May 7 • 88
Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation Paper • 2604.18168 • Published Apr 20 • 95
Elucidating the SNR-t Bias of Diffusion Probabilistic Models Paper • 2604.16044 • Published Apr 17 • 72