arXiv:2608. 15972v1 Announce Type: cross Abstract: Synchronized camera and wireless measurements observe the same scene through different physical channels.
By Yubo Zhang, Yiyao Liu
arXiv:2608. 09285v1 Announce Type: cross Abstract: Learning-based wireless localizers often fail to utilize geometric information about the propagation environment, limiting their ability to exploit non-line-of-sight (NLoS) propagation and generalize across scenes.
By Chenghong Bian, Chaozheng Wen, Hongze Chen, Jun Zhang
The paper surveys wireless foundation models (WFMs), highlighting their role in learning reusable representations from large-scale wireless data for physical-layer tasks. It systematically reviews WFM design components—pretraining, backbone architectures, and downstream adaptation—and categorizes the literature into five task families: signal recognition and demodulation, channel representation learning, RF sensing and localization, beam management, and spectrum sensing and monitoring, including multi-task models. The analysis reveals that while WFMs show promise, evidence of transferability varies across tasks and evaluation settings, and differences in datasets, modalities, architectures, and distribution shifts hinder clear conclusions about effective design choices.
By Alonso M. Pacheco Huachaca, Juan J. Rodriguez Rodriguez, Ahmed Aboulfotouh, Nelson L. S. da Fonseca, Carlos A. Astudillo, Hatem Abou-Zeid
arXiv:2607. 28994v1 Announce Type: cross Abstract: High-fidelity radio fields are typically simulated for every scene--transmitter configuration or fitted separately to each scene, failing to exploit propagation structures shared across environments.
By Chaozheng Wen, Chenghong Bian, Hongze Chen, Jun Zhang
arXiv:2607. 15713v1 Announce Type: cross Abstract: Accurate and robust wireless localization is a critical enabler for emerging 5G/6G applications, including autonomous driving, extended reality, and smart manufacturing.
By Yong Chu, Xun Zhou, Zenglin Xu, Hui Wang, Yue Yu
The paper introduces RSMEB, a unified benchmark for remote‑sensing multimodal retrieval that evaluates both cross‑modal and interleaved retrieval across 21 tasks under a single ranking protocol. It also presents VLM2GeoVec, an instruction‑conditioned single‑encoder model that embeds image, text, bounding‑box, and geo‑coordinate tokens into one sequence and achieves state‑of‑the‑art performance on region‑caption, referring‑expression, and semantic geo‑aware retrieval while remaining competitive on conventional tasks. The authors provide code, checkpoints, and data on GitHub to facilitate reproducibility.
By Emanuel S\'anchez Aimar, Gulnaz Zhambulova, Fahad Shahbaz Khan, Yonghao Xu, Michael Felsberg
arXiv:2505.07622v2 Announce Type: replace
Abstract: Cross-view geo-localization is a promising solution for large-scale localization problems, requiring the sequential execution of retrieval and metr...
By Zhuo Song, Ye Zhang, Kunhong Li, Longguang Wang, Yulan Guo
arXiv:2607. 11970v1 Announce Type: cross Abstract: We develop an enhanced in-context learning (ICL) framework to improve the performance of pilot-based beamforming in multi-user multiple-input single-output (MU-MISO) systems.
By Yubo Zhang, Xiaodong Wang
arXiv:2606.10277v2 Announce Type: replace
Abstract: Mobile systems increasingly rely on heterogeneous learning-enabled wireless functions, for which separate taskspecific models incur redundant train...
By Yuxuan Shi, Tingting Yang, Li Sun, Liwen Jing, Kangning Ma, Yuwei Wang, Mengfan Zheng
arXiv:2608. 15647v1 Announce Type: cross Abstract: Semantic segmentation of very-high-resolution (VHR) remote sensing imagery increasingly benefits from strong pretrained hierarchical encoders, yet exploiting their multi-stage representations remains difficult.
By Shuaishuai Cao, Meng Tang, Shuwei Peng, Xuan Liu, Min Huang, Jie Chen, Jiacheng Niu, Yong Chen, Edore Akpokodje, Hui Lin
arXiv:2606. 30576v1 Announce Type: cross Abstract: Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.
By Liyao Wang, Ruipu Wu, Haojun Xu, Lei Shi, Linjiang Huang, Si Liu
Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, acquisition-time and imaging-platform differences between UAV and reference imagery induce substantial cross-domain appearance and viewpoint shifts, challenging robust six-degree-of-freedom (6-DoF) pose estimation.