Vision-Wireless Fusion for Multi-User Localization: A Cross-Modal Transformer Approach
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2607. 15713v1 Announce Type: cross Abstract: Accurate and robust wireless localization is a critical enabler for emerging 5G/6G applications, including autonomous driving, extended reality, and smart manufacturing.
The paper introduces the Multi-Context Fusion Transformer (MFT), a model that predicts pedestrian crossing intentions in urban settings by integrating four types of contextual information—pedestrian behavior, environment, localization, and vehicle motion—through a progressive fusion strategy. MFT uses intra-context attention for reciprocal interactions within each context, cross-context attention to combine these contexts into a global representation, and guided attention mechanisms to refine both context tokens and the global token. Experiments on JAADbeh, JAADall, and PIE datasets show MFT outperforms existing methods with accuracies of 73%, 93%, and 90% respectively, and ablation studies confirm the importance of each network component and input context.
arXiv:2510. 13774v2 Announce Type: replace Abstract: Forecasting urban phenomena such as housing prices and public health indicators requires the effective integration of various geospatial data.
CIG-MAE is a self‑supervised framework for WiFi‑based human action recognition that uses a cross‑modal masked autoencoder to reconstruct both amplitude and phase of Channel State Information. It introduces an adaptive, information‑guided masking strategy that focuses on high‑density time‑frequency regions and employs a Barlow Twins regularizer to align cross‑modal representations without negative samples. Experiments on three public datasets show that CIG‑MAE outperforms state‑of‑the‑art SSL methods and even surpasses a fully supervised baseline, highlighting its data efficiency, robustness, and generalization.
arXiv:2606. 01899v1 Announce Type: cross Abstract: Wireless localization is a fundamental capability of sixth-generation (6G) networks.
arXiv:2608. 02092v2 Announce Type: replace Abstract: Deep multimodal fusion for object detection has demonstrated good performance through mining modal characteristics.