arXiv:2607. 15713v1 Announce Type: cross Abstract: Accurate and robust wireless localization is a critical enabler for emerging 5G/6G applications, including autonomous driving, extended reality, and smart manufacturing.
By Yong Chu, Xun Zhou, Zenglin Xu, Hui Wang, Yue Yu
The paper introduces the Multi-Context Fusion Transformer (MFT), a model that predicts pedestrian crossing intentions in urban settings by integrating four types of contextual information—pedestrian behavior, environment, localization, and vehicle motion—through a progressive fusion strategy. MFT uses intra-context attention for reciprocal interactions within each context, cross-context attention to combine these contexts into a global representation, and guided attention mechanisms to refine both context tokens and the global token. Experiments on JAADbeh, JAADall, and PIE datasets show MFT outperforms existing methods with accuracies of 73%, 93%, and 90% respectively, and ablation studies confirm the importance of each network component and input context.
By Yuanzhe Li, Hang Zhong, Steffen M\"uller
arXiv:2510. 13774v2 Announce Type: replace Abstract: Forecasting urban phenomena such as housing prices and public health indicators requires the effective integration of various geospatial data.
By Dominik J. M\"uhlematter, Lin Che, Ye Hong, Martin Raubal, Nina Wiedemann
CIG-MAE is a self‑supervised framework for WiFi‑based human action recognition that uses a cross‑modal masked autoencoder to reconstruct both amplitude and phase of Channel State Information. It introduces an adaptive, information‑guided masking strategy that focuses on high‑density time‑frequency regions and employs a Barlow Twins regularizer to align cross‑modal representations without negative samples. Experiments on three public datasets show that CIG‑MAE outperforms state‑of‑the‑art SSL methods and even surpasses a fully supervised baseline, highlighting its data efficiency, robustness, and generalization.
By Gang Liu, Yanling Hao, Yixuan Zou
arXiv:2606. 01899v1 Announce Type: cross Abstract: Wireless localization is a fundamental capability of sixth-generation (6G) networks.
By Guangjin Pan, Hui Chen, Hei Victor Cheng, Henk Wymeersch
arXiv:2608. 02092v2 Announce Type: replace Abstract: Deep multimodal fusion for object detection has demonstrated good performance through mining modal characteristics.
By Guandi Wang, Ming Li, Yunsen Xing, Junle Liu
The paper introduces an Attention-Driven Complementarity Resampling framework to enhance cross-modality object detection. It employs a shared channel spatial attention mechanism that exchanges semantic masks between modalities, encouraging the backbone to learn generalized features. Additionally, a learnable channel competition module samples and aggregates features channel‑wise, improving robustness and achieving competitive results on multiple datasets.
By Guandi Wang, Ming Li, Yunsen Xing, Junle Liu
arXiv:2609.23507v1 Announce Type: new
Abstract: The increasing reliance on mobile phones has made phone-induced pedestrian distraction increasingly prevalent. Activities such as texting, watching vid...
By Yuanzhe Li, Hounian Liu, Xiaotong Chang, Yidi Huang
arXiv:2606. 13509v1 Announce Type: cross Abstract: Indoor vision-based localization systems are affected by detection noise, occlusions, and limited camera coverage, leading to uncertainty at multiple stages of the pipeline.
By Mateo Toro Diz, Jonathan Hoss, Noah Klarmann
arXiv:2505.07622v2 Announce Type: replace
Abstract: Cross-view geo-localization is a promising solution for large-scale localization problems, requiring the sequential execution of retrieval and metr...
By Zhuo Song, Ye Zhang, Kunhong Li, Longguang Wang, Yulan Guo
The paper introduces a self‑localizing MIMO beam‑mapping framework that builds a hierarchical wireless memory using sparse channel state information (CSI) without explicit location labels. It employs beam‑domain RSS as compact inputs, a dual‑scale extractor for angular and temporal dependencies, and a hybrid temporal encoder to infer physical anchors that index a structured radio map. The radio‑map embedding enables continuous updates and full‑CSI reconstruction, yielding over 30% better anchor recovery and more than 20% channel‑capacity gains in NLOS beam tracking compared to Kalman‑filter methods.
By Wangqian Chen, Junting Chen, Shuguang Cui
The paper introduces Mobility Stream-Structure Synergy (MoSS), a method that fuses two complementary views of mobility data—an hourly inflow/outflow Sequence view and a Structure view derived from zigzag persistence diagrams—to capture temporal dynamics and evolving regional connectivity. MoSS employs a synergy module that extracts higher‑order representations from the co‑occurrence of these views, moving beyond additive fusion. Experiments on New York City and Chicago demonstrate that MoSS outperforms existing baselines on three downstream tasks using only mobility data.
By Namwoo Kim, Jeeyun Chang, Kanghoon Lee, Yoonjin Yoon