arXiv Computer Vision
Sep 3

SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

SolarWM is an open foundation for building interactive video world models, offering a reconfigurable multi‑source data engine that unifies 1.43 million clips from 10 datasets into a consistent, frame‑aligned format. It provides a backbone‑native adaptation framework that preserves native representations of models ranging from 5 B to 33 B parameters, and a three‑stage training recipe combining bidirectional adaptation, teacher‑forced autoregressive initialization, and distribution‑matching distillation. The resulting causal models can interact in real‑time over rollouts from minutes to hours, trained only on 5‑second sequences, and the project releases data, pipeline, recipes, weights, and framework for reproducible research.

By Junchao Huang, Guian Fang, Shengju Qian, Xianghao Kong, Zhuoran Zhao, Wei Huang, Yihua Du, Zixin Zhang, Justin Cui, Yuchao Gu, Yukang Chen, Xinting Hu, Tianyu He, Shaoshuai Shi, Zhuotao Tian, Xin Wang, Mike Zheng Shou, Li Jiang
arXiv Computer Vision
Sep 7

Weather-Conditioned Depth Anything

Weather-Conditioned Depth Anything (DA‑W) is a new framework that enhances monocular depth estimation models, like the Depth Anything series, to perform robustly under adverse weather conditions such as fog, rain, snow, and low‑light. It achieves this by disentangling style from content: a Style Filter extracts weather‑specific embeddings from a curated mix of real and synthetic degradation data, which are then injected into the backbone via a lightweight, zero‑initialized adapter. The adapter is trained with pseudo‑label distillation and alignment, enabling a single unified model to adapt to diverse weather scenarios while preserving its generalization on clean data, and it achieves state‑of‑the‑art performance with an average 3.7% improvement in AbsRel on weather benchmarks.

By Zhaoming Xu, Chan-Wei Hu, Kuan-Ru Huang, Zihao Zhu, Renjie Li, Yang Zhou, Zhengzhong Tu