arXiv:2606.10277v2 Announce Type: replace
Abstract: Mobile systems increasingly rely on heterogeneous learning-enabled wireless functions, for which separate taskspecific models incur redundant train...
By Yuxuan Shi, Tingting Yang, Li Sun, Liwen Jing, Kangning Ma, Yuwei Wang, Mengfan Zheng
The paper surveys wireless foundation models (WFMs), highlighting their role in learning reusable representations from large-scale wireless data for physical-layer tasks. It systematically reviews WFM design components—pretraining, backbone architectures, and downstream adaptation—and categorizes the literature into five task families: signal recognition and demodulation, channel representation learning, RF sensing and localization, beam management, and spectrum sensing and monitoring, including multi-task models. The analysis reveals that while WFMs show promise, evidence of transferability varies across tasks and evaluation settings, and differences in datasets, modalities, architectures, and distribution shifts hinder clear conclusions about effective design choices.
By Alonso M. Pacheco Huachaca, Juan J. Rodriguez Rodriguez, Ahmed Aboulfotouh, Nelson L. S. da Fonseca, Carlos A. Astudillo, Hatem Abou-Zeid
arXiv:2606. 10277v1 Announce Type: new Abstract: Though wireless foundation models (WFMs) have shown strong potential in learning universal channel representations, their adaptation to various downstream tasks remains constrained by existing paradigms.
By Yuxuan Shi, Tingting Yang, Kangning Ma, Liwen Jing, Yuwei Wang, Mengfan Zheng, Li Sun
arXiv:2606. 06373v1 Announce Type: cross Abstract: Wireless foundation models have emerged as a promising alternative to building separate models for each wireless task.
By Ahmed Mohamed, Ahmed Aboulfotouh, Hatem Abou-Zeid
The paper introduces a physics‑informed autoencoder that separates reflector speed from sensing geometry in WiFi spectrograms, enabling joint recovery of speed, geometry factor, relative amplitude, and ridge width for each Doppler ridge. It builds a compact parametric representation of spectrograms validated on a large human‑activity dataset and employs a synthetic‑to‑real training framework to avoid real‑data collection. Experiments on synthetic and 31 real WiFi scenarios show the method outperforms existing baselines in accurately extracting speed and geometry information.
By Mert Torun, Darius Cuenca, Yasamin Mostofi
arXiv:2609.22139v1 Announce Type: cross
Abstract: Automatic modulation classification (AMC) of received radio signals is prudent for further signal processing tasks such as communication monitoring,...
By Qamar Ijaz, Nayyer Aafaq
WiNeRF is a neural field framework that learns a spatially continuous, complex-valued wireless channel representation from sparse channel state information collected by commodity WiFi devices. It incorporates system constraints such as antenna geometry, limited spatial resolution, and phase uncertainty through a 3D conical wave sampling model, a multi-resolution implicit scene representation, and a differentiable optimization framework. In diverse indoor environments with non‑line‑of‑sight regions, WiNeRF achieves a median prediction SNR of 5.3 dB, outperforming prior neural baselines by 4.9 dB on average, and produces a task‑agnostic channel representation that can be reused in standard signal‑processing pipelines without hardware or protocol changes.
By Saif Ur Rahman, Rafid Umayer Murshed, Anton Dmitriev, Cagri Tanriover, Rahul C. Shah, Elah\'e Soltanaghai
The paper presents a deep learning‑enhanced Wi‑Fi sensing system that uses only a single transceiver pair to achieve real‑time human pose estimation and localization. By leveraging prior information and temporal correlation as side information, the system reduces estimation error under hardware constraints. Experimental results show an average pose error of 0.2189 m and a localization error of 0.6124 m while running at 42 fps on commodity hardware.
By Yuxuan Liu, Chiya Zhang, Yifeng Yuan, Chunlong He, Weizheng Zhang, Gaojie Chen
arXiv:2607. 09760v1 Announce Type: cross Abstract: Radio frequency fingerprint identification (RFFI) uses transmitter-specific hardware imperfections as a physicallayer identity cue for Internet of Things (IoT) devices, but deep RFFI models often degrade when the acquisition environment changes.
By Fengchong Yao, Jianbing Li, Qing Liu, Qikun Liu, Kefeng Song, Haitao Li, Song Wang
The paper introduces a self‑localizing MIMO beam‑mapping framework that builds a hierarchical wireless memory using sparse channel state information (CSI) without explicit location labels. It employs beam‑domain RSS as compact inputs, a dual‑scale extractor for angular and temporal dependencies, and a hybrid temporal encoder to infer physical anchors that index a structured radio map. The radio‑map embedding enables continuous updates and full‑CSI reconstruction, yielding over 30% better anchor recovery and more than 20% channel‑capacity gains in NLOS beam tracking compared to Kalman‑filter methods.
By Wangqian Chen, Junting Chen, Shuguang Cui
CIG-MAE is a self‑supervised framework for WiFi‑based human action recognition that uses a cross‑modal masked autoencoder to reconstruct both amplitude and phase of Channel State Information. It introduces an adaptive, information‑guided masking strategy that focuses on high‑density time‑frequency regions and employs a Barlow Twins regularizer to align cross‑modal representations without negative samples. Experiments on three public datasets show that CIG‑MAE outperforms state‑of‑the‑art SSL methods and even surpasses a fully supervised baseline, highlighting its data efficiency, robustness, and generalization.
By Gang Liu, Yanling Hao, Yixuan Zou
arXiv:2607. 09727v1 Announce Type: cross Abstract: WiFi sensing based on Channel State Information (CSI) promises ubiquitous, device-free perception, yet current research remains trapped in a Tower of Babel - fragmented into isolated silos where models are tailored to specific hardware dialects, fixed environments, and narrow tasks.
By Jiayi Chen, Weiting Ou, Guangxu Zhu