GeoMAD is a multi‑view anomaly detection framework that fuses multiple camera viewpoints while maintaining geometric awareness and scalability to multi‑class industrial settings. It introduces a Cross‑view Deformable Fusion Module (CDFM) that learns view‑pair‑specific sampling offsets on 2D feature maps, enabling hierarchical cross‑view correspondence without camera calibration or voxel construction. Additionally, Distributional View Alignment (DVA) provides a self‑supervised loss that aligns bottleneck distributions across views, ensuring global consistency without pixel‑level correspondence. Together, CDFM and DVA achieve geometry‑aware, distribution‑consistent fusion and demonstrate strong detection and localization performance on Real‑IAD and MANTA‑Tiny datasets.
By Shang-Fu Chen, Jhih-Ciang Wu, Kuan-Chuan Peng, Wen-Huang Cheng, Kai-Lung Hua
arXiv:2609.14419v1 Announce Type: cross
Abstract: Person re-identification (ReID) is essential for multi-camera surveillance and tracking, yet remains difficult due to viewpoint and illumination chan...
By Leon Fernando, C Dombawala, P. Hettigoda, Vanodhya G. Warnasooriya, Ishara Neranjana, Rashmika Nawaratne
arXiv:2604.02583v3 Announce Type: replace
Abstract: We propose FusionBERT, a novel multi-view visual fusion framework for image--3D multimodal retrieval. Existing image--3D representation learning me...
By Wei Li, Yufan Ren, Hanqing Jiang, Jianhui Ding, Zhen Peng, Leman Feng, Yichun Shentu, Guoqiang Xu, Baigui Sun
arXiv:2609.39486v1 Announce Type: new
Abstract: Estimating accident mechanics from real-world crashes is important for vehicle-safety analysis, injury modeling, crash-severity prediction, and operati...
By Ond\v{r}ej Valach, V\'aclav Divi\v{s}, Ivan Gruber
The paper introduces MVLGeo, a unified framework for cross-view object geo-localization that combines multiple viewpoints into a single model. It employs Vision‑Language Reranking to use contextual text from the query view, a multi‑view Mixture‑of‑Experts architecture to share knowledge and reduce redundancy, and an adaptive elliptical prior for positional encoding. Experiments on CVOGL benchmarks show that MVLGeo achieves state‑of‑the‑art performance and robustness to input degradation.
By Xuyu Fan, Qi Ming, Zhu Han, Liuqian Wang, Siyuan Cao, Xiaohan Zhang, Xudong Zhao, Mingjing Zhao, Yuhan Zhang
arXiv:2607. 22068v1 Announce Type: cross Abstract: Multi-branch architectures and CNN-Transformer fusion have long been regarded as effective ways to improve vehicle re-identification (Re-ID) by combining complementary representations.
By Yu Wang, Hongyu Yang