GeoMAD is a multi‑view anomaly detection framework that fuses multiple camera viewpoints while maintaining geometric awareness and scalability to multi‑class industrial settings. It introduces a Cross‑view Deformable Fusion Module (CDFM) that learns view‑pair‑specific sampling offsets on 2D feature maps, enabling hierarchical cross‑view correspondence without camera calibration or voxel construction. Additionally, Distributional View Alignment (DVA) provides a self‑supervised loss that aligns bottleneck distributions across views, ensuring global consistency without pixel‑level correspondence. Together, CDFM and DVA achieve geometry‑aware, distribution‑consistent fusion and demonstrate strong detection and localization performance on Real‑IAD and MANTA‑Tiny datasets.
By Shang-Fu Chen, Jhih-Ciang Wu, Kuan-Chuan Peng, Wen-Huang Cheng, Kai-Lung Hua
arXiv:2609.14419v1 Announce Type: cross
Abstract: Person re-identification (ReID) is essential for multi-camera surveillance and tracking, yet remains difficult due to viewpoint and illumination chan...
By Leon Fernando, C Dombawala, P. Hettigoda, Vanodhya G. Warnasooriya, Ishara Neranjana, Rashmika Nawaratne
arXiv:2604.02583v3 Announce Type: replace
Abstract: We propose FusionBERT, a novel multi-view visual fusion framework for image--3D multimodal retrieval. Existing image--3D representation learning me...
By Wei Li, Yufan Ren, Hanqing Jiang, Jianhui Ding, Zhen Peng, Leman Feng, Yichun Shentu, Guoqiang Xu, Baigui Sun
arXiv:2609.39486v1 Announce Type: new
Abstract: Estimating accident mechanics from real-world crashes is important for vehicle-safety analysis, injury modeling, crash-severity prediction, and operati...
By Ond\v{r}ej Valach, V\'aclav Divi\v{s}, Ivan Gruber
The paper introduces MVLGeo, a unified framework for cross-view object geo-localization that combines multiple viewpoints into a single model. It employs Vision‑Language Reranking to use contextual text from the query view, a multi‑view Mixture‑of‑Experts architecture to share knowledge and reduce redundancy, and an adaptive elliptical prior for positional encoding. Experiments on CVOGL benchmarks show that MVLGeo achieves state‑of‑the‑art performance and robustness to input degradation.
By Xuyu Fan, Qi Ming, Zhu Han, Liuqian Wang, Siyuan Cao, Xiaohan Zhang, Xudong Zhao, Mingjing Zhao, Yuhan Zhang
arXiv:2607. 22068v1 Announce Type: cross Abstract: Multi-branch architectures and CNN-Transformer fusion have long been regarded as effective ways to improve vehicle re-identification (Re-ID) by combining complementary representations.
By Yu Wang, Hongyu Yang
arXiv:2609.36894v1 Announce Type: new
Abstract: As a fundamental image processing task, object Re-Identification (ReID) aims to retrieve objects across non-overlapping cameras. Recently, with the dev...
By Yingquan Wang, Pingping Zhang, Dong Wang, Huchuan Lu
arXiv:2504. 11500v3 Announce Type: replace-cross Abstract: Transit Origin-Destination (OD) data are fundamental for optimizing public transit services, yet current collection methods, such as manual surveys, Bluetooth/WiFi tracking, and Automated Passenger Counters, are often costly, device-dependent, or unable to support individual-level matching.
By Kaicong Huang, Talha Azfar, Jack Reilly, Ruimin Ke
Multi-modality data from different sensors provides rich complementary information for 3D perception, becoming an essential component in reliable autonomous driving systems. Current research typically designs intricate and complex fusion strategies to integrate information from multimodal data on a unified bird's-eye-view (BEV) feature map for the joint learning of multiple perception tasks.
The paper introduces the concept of protocol divergence, showing that identical nominal missing rates can lead to vastly different learning regimes in incomplete multi‑view clustering. It critiques existing evaluation practices that ignore observation structure and proposes CRAFT, a train‑once framework that fuses observed views with mask‑aware attention, enabling efficient deployment across multiple missing‑view protocols. Experiments on CUB, MultiFashion, and other benchmarks demonstrate CRAFT’s superior performance and significant computational savings through checkpoint reuse.
By Haolu Liu, Xiyue Wang, Xuanting Xie, Liangjian Wen, Zhao Kang
arXiv:2606.03406v2 Announce Type: replace
Abstract: Reliable correspondence estimation supports image processing and 3D vision tasks, including Structure from Motion, visual localization, and image r...
By Xu Pan, Zhen Pang, Qiyuan Ma, Wei Ji, Shuhan Shen, Xianwei Zheng
arXiv:2607. 10391v1 Announce Type: cross Abstract: Despite exposing rich intermediate representations, Vision Transformers (ViTs) are almost exclusively utilized as black-box feature extractors, where only the last layer is considered for downstream tasks.
By Francesco Di Salvo, Shyam Nandan Rai, Hamed Damirchi, Ignacio Meza De la Jara, Sebastian Doerrich, Marco Lents, Christian Ledig