The paper identifies a problem in multi‑view anomaly detection called cross‑view information leakage, where fusing multiple inspection views can cause normal features to mask anomalies during reconstruction. To address this, the authors propose GLAD, a framework that uses a Global‑Local Attention Driven approach, combining vision foundation model features with two fusion modules: Multi‑view Merging Attention for local, weighted fusion and Object‑Guided Attention for global context aggregation. Experiments on Real‑IAD and MANTA‑Tiny demonstrate that GLAD outperforms existing methods across various metrics, underscoring the importance of restricting information flow to preserve the reconstruction gap.
By Shang-Fu Chen, Kuan-Chuan Peng, Jhih-Ciang Wu, Wen-Huang Cheng, Kai-Lung Hua
arXiv:2608.03423v2 Announce Type: replace
Abstract: Local feature matching is a fundamental component of photogrammetry, enabling accurate image correspondence critical for tasks such as 3D reconstru...
By Zhihua Xu, Runyu Zhu, Rongjun Qin
arXiv:2608. 07559v1 Announce Type: cross Abstract: In 3D reconstruction, mirrors introduce significant challenges by creating distorted and fragmented spaces, resulting in inaccurate and unreliable 3D models.
By Yidan Shen, Yu Wen, Chen Zhang, Xin Fu, Renjie Hu
The paper introduces S2A, a semantic-to-spatial alignment framework designed for alignment‑free RGB‑T salient object detection. It employs a global‑guided hierarchical fusion module to refine intra‑modal features, an alignment‑free cross‑modal channel attention module to exchange semantic information, and a spatial deformable cross‑attention module to recover local spatial correspondence. These components collectively reduce misalignment‑induced feature contamination and achieve competitive performance on public benchmarks without additional bells and whistles.
By Qiangqiang Zhou, Yang Luo, Yong Chen, Jiawei Xu
arXiv:2607. 00491v1 Announce Type: cross Abstract: Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input.
By Leyuan Yu, Xiao Tang, Minghao Liu, Xinyuan Li, Xiaokai Bai, Sheng Zhou, Qunshu Lin, Weihao Xuan, Naoto Yokoya
arXiv:2607. 08970v1 Announce Type: cross Abstract: Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observations across viewpoints into a coherent, world-centric (allocentric) 3D mental model.
By Hantao Zhang, Jinru Sui, Ed Li, Dirk Bergemann, Zhuoran Yang
The paper introduces a saliency-depth conditioning approach for zero‑shot segmentation of communication‑tower components in cluttered UAV imagery. By combining appearance‑based saliency with monocular relative depth, the method creates a coarse tower prior that suppresses irrelevant background, and integrates this module with Grounded‑SAM and SAM 3 to produce SD‑Grounded‑SAM and SD‑SAM 3. Experiments on the TOW‑300 dataset show that SD‑SAM 3 achieves the best instance‑segmentation performance while SD‑Grounded‑SAM reduces false positives, with ablations confirming the complementary benefits of saliency, depth, and box refinement.
By Ali Lesani, Chul Min Yeum, Su-Min Kang
arXiv:2606. 11683v1 Announce Type: cross Abstract: Spatial reasoning from egocentric videos is inherently challenging because the observable evidence is constrained by the camera trajectory.
By Chaofan Ma, Zhenjie Mao, Yuhuan Yang, Fanqin Zeng, Yue Shi, Yingjie Zhou, Xiaofeng Cao, Jiangchao Yao
Vision-Language Models (VLMs) are promising for construction-site monitoring, and recent construction-tailored VLMs have primarily adapted pretrained VLMs through direct QA-style fine-tuning from a single global image. We argue that this direct paradigm remains limited for in-the-wild deployment in terms of operational range, reliability under reduced-resolution inputs, and inference efficiency.
3DGS-HPC is a framework that improves 3D Gaussian Splatting for novel view synthesis by mitigating transient distractors such as moving objects and varying shadows. It combines a patch‑wise classification strategy that uses local spatial consistency for robust region‑level decisions with a hybrid classification metric that adaptively integrates photometric and perceptual cues. Experiments show that this approach outperforms existing methods in reducing distractor effects and enhancing 3DGS quality.
By Jiahao Chen, Yipeng Qin, Ganlong Zhao, Xin Li, Wenping Wang, Guanbin Li
The paper introduces a new hyperspectral image dataset for benchmarking salient object detection, comprising 60 hyperspectral images, their ground‑truth binary masks, and corresponding sRGB renderings. The dataset was curated to include diverse object sizes, counts, contrasts, and positions, addressing the lack of dedicated hyperspectral data for this task. The authors also evaluate existing hyperspectral saliency models using the AUC metric and provide the dataset on GitHub and Hugging Face.
By Nevrez Imamoglu, Yu Oishi, Xiaoqiang Zhang, Guanqun Ding, Yuming Fang, Toru Kouyama, Ryosuke Nakamura
Local feature matching is a fundamental component of photogrammetry, enabling accurate image correspondence critical for tasks such as 3D reconstruction, stereo mapping, and visual localization. While recent detector-free matching methods, like LoFTR, have advanced the field, the global features obtained by leveraging the global-range modeling capacity of the unconstrained attention mechanism compromise the model's attention to the salient structures in certain scenarios.