arXiv Machine Learning
Sep 22

FIRM-WM: State-factorized factual-interventional recurrent modeling for reward-free visual planning

FIRM-WM is a compact pixel world model that separates a goal‑comparable configuration from a 128‑dimensional dynamic fiber, enabling reward‑free visual planning from offline videos. It addresses two key mismatches: aligning planning states with goal images and reconciling factual trajectories with interventional sampling. In experiments, FIRM‑WM achieves high success rates on TwoRoom, Reacher, and OGBench‑Cube while using fewer parameters and faster planning times than prior models.

By Yilun Wu, Yunjian Zhang, Aobo Li, Mujiangshan Wang, Haitao Wu, Aqiang Zhang
arXiv Computer Vision
2d ago

EVO-WAM: Evolving World Action Models through Video-Action Verification

arXiv:2609.38057v1 Announce Type: new Abstract: Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action m...

By Shiyang Zhou, Xionghao Wu, Wenbo Li, Shenghe Zheng, Jiyao Zhang, Songsong Yu, Yijun Yang, Jianhui Liu, Haoze Sun, Senqiao Yang, Li Jiang, Jingyong Su, Haoyang Huang, Zhuotao Tian
arXiv AI
Sep 18

JEPA-WAM: Connecting Generated Visual Instructions to World Action Models through JEPA Latent Representations

JEPA-WAM enhances World Action Models (WAMs) by pairing text instructions with stochastically generated visual cues, using a text-to-image generator and a frozen V‑JEPA encoder to create dense goal representations. These representations are compressed into goal tokens that condition both video and action experts via cross‑attention, enabling the model to better ground instructions. On a new real‑robot benchmark, JEPA‑WAM attains 87.3%, 74.5%, and 80.9% success rates across in‑distribution, out‑of‑distribution scenes, and out‑of‑distribution instructions, outperforming prior methods by significant margins.

By Tianbin Liu, Jian Zhu, Taiyi Su, Jianjun Zhang, Chong Ma, Zitai Huang, Yi Xu