arXiv:2608. 15972v1 Announce Type: cross Abstract: Synchronized camera and wireless measurements observe the same scene through different physical channels.
By Yubo Zhang, Yiyao Liu
arXiv:2608. 09285v1 Announce Type: cross Abstract: Learning-based wireless localizers often fail to utilize geometric information about the propagation environment, limiting their ability to exploit non-line-of-sight (NLoS) propagation and generalize across scenes.
By Chenghong Bian, Chaozheng Wen, Hongze Chen, Jun Zhang
The paper surveys wireless foundation models (WFMs), highlighting their role in learning reusable representations from large-scale wireless data for physical-layer tasks. It systematically reviews WFM design components—pretraining, backbone architectures, and downstream adaptation—and categorizes the literature into five task families: signal recognition and demodulation, channel representation learning, RF sensing and localization, beam management, and spectrum sensing and monitoring, including multi-task models. The analysis reveals that while WFMs show promise, evidence of transferability varies across tasks and evaluation settings, and differences in datasets, modalities, architectures, and distribution shifts hinder clear conclusions about effective design choices.
By Alonso M. Pacheco Huachaca, Juan J. Rodriguez Rodriguez, Ahmed Aboulfotouh, Nelson L. S. da Fonseca, Carlos A. Astudillo, Hatem Abou-Zeid
arXiv:2607. 28994v1 Announce Type: cross Abstract: High-fidelity radio fields are typically simulated for every scene--transmitter configuration or fitted separately to each scene, failing to exploit propagation structures shared across environments.
By Chaozheng Wen, Chenghong Bian, Hongze Chen, Jun Zhang
arXiv:2607. 15713v1 Announce Type: cross Abstract: Accurate and robust wireless localization is a critical enabler for emerging 5G/6G applications, including autonomous driving, extended reality, and smart manufacturing.
By Yong Chu, Xun Zhou, Zenglin Xu, Hui Wang, Yue Yu
The paper introduces RSMEB, a unified benchmark for remote‑sensing multimodal retrieval that evaluates both cross‑modal and interleaved retrieval across 21 tasks under a single ranking protocol. It also presents VLM2GeoVec, an instruction‑conditioned single‑encoder model that embeds image, text, bounding‑box, and geo‑coordinate tokens into one sequence and achieves state‑of‑the‑art performance on region‑caption, referring‑expression, and semantic geo‑aware retrieval while remaining competitive on conventional tasks. The authors provide code, checkpoints, and data on GitHub to facilitate reproducibility.
By Emanuel S\'anchez Aimar, Gulnaz Zhambulova, Fahad Shahbaz Khan, Yonghao Xu, Michael Felsberg