arXiv AI By Xin Zhou, Cong Miao

Astronex-World 1.0: Real-Time Interactive World Model Foundation

Read the original on arXiv AI →

Astronex-World 1.0 is an open, controllable video world‑model foundation that can generate future visual states from text prompts or initial images, supporting frame‑aligned camera trajectories, continuous actions, and text events during rollout. It offers both a bidirectional model for full‑context generation and a causal model with block‑causal attention, built on the Wan2.2‑TI2V‑5B prior, and achieves real‑time 832×480 video at 24 fps. The model was trained over five stages on two NVIDIA L20 GPUs, scores 73.5 on WBench Navi and 70.0 on WBench Full, and outperforms several larger competitors while providing interfaces for embodied intelligence and autonomous driving.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 10

BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression

arXiv:2606. 10135v1 Announce Type: cross Abstract: Transitioning bidirectional video diffusion models into an autoregressive paradigm improves the interactivity of video world models, but existing causal pipelines need many stages (control fine-tuning, autoregressive training, causal initialization, few-step distillation) and still trail bidirectional models in quality due to error accumulation.

By Shaohao Rui, Xiaofeng Mao, Zhanyu Zhang, Peijia Lin, Yansong Zhu, Yibo Zhang, Haibin Wan, Weijie Ma
arXiv Computer Vision
Sep 4

Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation

The paper introduces a large‑scale synthetic data pipeline built on Unreal Engine to generate action‑conditioned, multi‑view video for training world models. The system operates in two stages: real‑time physics simulation records trajectories, then offline rendering produces high‑quality video. It includes a distributed production framework with task partitioning, automated filtering, and a 25‑server cluster, yielding over 2,600 hours of 1080p and 6,000 hours of 720p video from 429 levels and 40 characters.

By Haoyu Wang, Songchun Zhang, Haoran Li, Haoyang Huang, Zeyue Xue, Nan Duan
arXiv AI
Aug 24

ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation

ForeTime‑VLA is a causal vision‑language‑action policy that distills future‑aware representations from a frozen Fast‑WAM teacher, enabling it to anticipate contact events during conveyor‑belt manipulation. The method compresses current and future video latents into a 64‑dimensional target, uses an eight‑frame history encoder to predict this target along with manipulation phase and time‑to‑transition, and conditions a VLM prefix on future tokens and phase. On a deduplicated conveyor‑belt dataset, ForeTime‑VLA reduces test MAE by 2.63% and L2 by 3.02%, while real‑robot experiments show significantly higher grasp success rates compared to the next‑best reference. whyItMatters":"The approach demonstrates that distilling future‑token knowledge from a world‑action model can improve dynamic manipulation performance without the computational cost of running the teacher at inference time."

By Siyuan Ma, Yutian Zhang, Boshi Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Xiaojin Huang
arXiv Computer Vision
Sep 11

World in World: Explore the World with World Models

World in World introduces a training‑free, inference‑time interface that transforms diverse control signals—such as source‑video observations, target‑view projections, geometry renderings, and retrieved states—into camera‑ and time‑labelled visual states. These states are processed by a frozen causal video model’s self‑attention, enabling tasks like camera‑controlled rerendering, long‑horizon revisiting, and human‑motion transfer without additional training. The method employs a correspondence router and evidence‑wise attention to align token identities and regulate auxiliary channel contributions during a single denoising pass.

By Chenxi Song, Yanming Yang, Chi Zhang
arXiv Computer Vision
Sep 23

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models

arXiv:2605.30263v2 Announce Type: replace Abstract: Recent video diffusion foundation models have achieved remarkable progress in high-quality video generation, yet turning them into real-time intera...

By Min Zhao, Hongzhou Zhu, Bokai Yan, Zihan Zhou, Yimin Chen, Honglie Wang, Wenqiang Sun, Zhengwei Fang, Zizheng Xun, Zihao Li, Kaiwen Zheng, Guande He, Xiao Yang, Chongxuan Li, Fan Bao, Jun Zhu