WAVE introduces a multi-level discrete wavelet transform (ML‑DWT) to reverse the typical fine‑to‑coarse bias in guided depth super‑resolution. By consuming wavelet sub‑bands and semantic tokens in reverse order, it separates structure and detail reconstruction, applies semantic gating to high‑frequency bands, and fuses modalities via an invertible coupling mechanism. Experiments on multiple benchmarks show that WAVE matches or outperforms existing methods, especially at high upsampling factors where low‑resolution depth has minimal structure.
By Tayyab Nasir, Daochang Liu, Ajmal Mian
arXiv:2601. 22054v2 Announce Type: replace-cross Abstract: Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camera-dependent biases, and metric ambiguity in noisy cross-source 3D data.
By Baorui Ma, Jiahui Yang, Donglin Di, Xuancheng Zhang, Jianxun Cui, Hao Li, Yan Xie, Wei Chen
arXiv:2503. 19947v2 Announce Type: replace-cross Abstract: Generalized metric depth understanding is critical for precise vision-guided robotics, which current state-of-the-art (SOTA) vision-encoders do not support.
By Paul Koch, J\"org Kr\"uger
arXiv:2505.16157v3 Announce Type: replace
Abstract: Transformer-based models have made remarkable progress in image restoration (IR) tasks. However, the quadratic complexity of self-attention in Tran...
By Yuang Ai
arXiv:2608. 08676v1 Announce Type: cross Abstract: Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation.
By Jinbo Yan, Limeng Qiao, Jie Qin, Junyan He, Feize Wu, Guanglu Wan
arXiv:2605. 18714v2 Announce Type: replace-cross Abstract: Unified multimodal models (UMMs) strive to consolidate visual understanding and visual generation within a single architecture.
By Songsong Yu, Yuxin Chen, Ying Shan, Yanwei Li