arXiv AI

Collaborative Space Object Detection with Multi-Satellite Viewpoints in LEO Constellations

arXiv:2606. 01895v1 Announce Type: cross Abstract: With the growing number of satellites in low Earth orbit (LEO) constellations, the near-Earth space environment has become increasingly congested, making space object detection (SOD) a pressing challenge for space safety and sustainability.

arXiv Computer Vision
Aug 28

GeoMAD: Geometry-Aware Multi-View Anomaly Detection via Deformable Fusion and Distributional Alignment

GeoMAD is a multi‑view anomaly detection framework that fuses multiple camera viewpoints while maintaining geometric awareness and scalability to multi‑class industrial settings. It introduces a Cross‑view Deformable Fusion Module (CDFM) that learns view‑pair‑specific sampling offsets on 2D feature maps, enabling hierarchical cross‑view correspondence without camera calibration or voxel construction. Additionally, Distributional View Alignment (DVA) provides a self‑supervised loss that aligns bottleneck distributions across views, ensuring global consistency without pixel‑level correspondence. Together, CDFM and DVA achieve geometry‑aware, distribution‑consistent fusion and demonstrate strong detection and localization performance on Real‑IAD and MANTA‑Tiny datasets.

By Shang-Fu Chen, Jhih-Ciang Wu, Kuan-Chuan Peng, Wen-Huang Cheng, Kai-Lung Hua
arXiv AI
Jul 7

BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations

arXiv:2603. 06576v2 Announce Type: replace-cross Abstract: The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and semantic understanding abilities, which are essential for handling complex decision-making and long-tail scenarios.

By Thomas Monninger, Shaoyuan Xie, Qi Alfred Chen, Sihao Ding
arXiv Computer Vision
Sep 17

Multi-View Mixture-of-Experts with Vision-Language Reranking for Cross-View Object Geo-Localization

The paper introduces MVLGeo, a unified framework for cross-view object geo-localization that combines multiple viewpoints into a single model. It employs Vision‑Language Reranking to use contextual text from the query view, a multi‑view Mixture‑of‑Experts architecture to share knowledge and reduce redundancy, and an adaptive elliptical prior for positional encoding. Experiments on CVOGL benchmarks show that MVLGeo achieves state‑of‑the‑art performance and robustness to input degradation.

By Xuyu Fan, Qi Ming, Zhu Han, Liuqian Wang, Siyuan Cao, Xiaohan Zhang, Xudong Zhao, Mingjing Zhao, Yuhan Zhang
arXiv AI
Sep 25

SARFusion: Scene-Aware Routing Fusion for Robust Camera-LiDAR 3D Object Detection

SARFusion introduces a scene-aware routing approach for camera‑LiDAR 3D object detection, decoupling object‑query decoding into separate camera, LiDAR, and fusion branches. By estimating a global scene reliability prior and incorporating object‑level evidence, each query is routed to the most suitable branch, reducing cross‑modal interference. The method achieves strong performance on the nuScenes test set (72.5 mAP, 74.4 NDS) and demonstrates robustness to sensor corruptions and environmental changes.

By Yuting Zhao, Ziyi Zheng, Shuxiao Li
arXiv AI
Sep 10

TriCCOT: Tri-part Convolutional Conformal Transformer for Onboard Space Object Detection

TriCCOT is a tri-part architecture designed for onboard space object detection that balances computational efficiency with robust performance. It combines a convolutional region proposal network, a conformal prediction stage that enlarges bounding boxes with distribution‑free probabilistic coverage, and Aper‑GATES—a hardware‑friendly attention‑based classifier that replaces standard transformer operations with convolutional projections and gating. Experiments on DIOR and VDVRaw datasets show competitive detection accuracy and improved robustness to blur and noise, and the model was fully deployed on a Xilinx Versal VCK190 FPGA without altering the underlying DPU architecture.

By Adrien Dorise, Marjorie Bellizzi, Julia Cohen, St\'ephane May
arXiv AI
Aug 20

One-Stage Object Detectors in Autonomous Driving

The paper surveys one‑stage object detectors for autonomous driving, covering the evolution from early models like YOLOv1 and SSD to recent real‑time architectures such as YOLOv10 and anchor‑free detectors like FCOS and CenterNet. It compares these methods on design choices, feature‑fusion strategies, loss functions, deployment trade‑offs, and benchmark performance, while also summarizing datasets, evaluation metrics, open challenges, and future research directions. The survey emphasizes how one‑stage detectors balance speed, accuracy, efficiency, and robustness, noting the gap between benchmark results and dependable real‑world performance.

By Jonel Roman, Ryan Sirjue, Peter Nguyen, Daniel Krutky, Juan Jesus, Sudip Dhakal