The paper introduces Embodied Semantic Communication (ESC), a new paradigm that redefines information transmission for autonomous agents by embedding multimodal perceptual states, hardware capabilities, and collaborative intents into unified, action‑oriented semantic representations. ESC enables heterogeneous agents to parse, align, and ground shared information directly into local motor control, addressing the limitations of traditional communication approaches that focus solely on bit delivery or single‑task optimization. The tutorial outlines ESC’s conceptual boundaries, system characteristics, and technical pathways, mapping relevant mathematical tools such as semantic information theory, world models, and multi‑agent decision theory, and concludes with a roadmap of open challenges like semantic reliability, dynamic interaction, and bandwidth‑adaptive transmission.
By Yizheng Huang, Wensheng Lin, Lixin Li, Qinghe Du, Wenchi Cheng, Zhu Han
arXiv:2607. 06651v1 Announce Type: new Abstract: Federated learning (FL) over mobile and edge devices increasingly involves multimodal models in which clients differ in both sensing capability and computational capacity.
By Quoc Bao Phan, Tuy Tan Nguyen
The paper investigates the environmental impact of running large language models (LLMs) on mobile devices. It evaluates 18 different LLM configurations on two smartphones and a server, measuring energy per token, latency, accuracy, and battery-cycle consumption. Findings reveal that on-device inference is about three times less energy‑efficient than batched server inference, that energy consumption varies non‑monotonically with quantization bit‑width, and that most models are not on the Pareto front of accuracy and energy efficiency. The study concludes that local AI is not inherently more sustainable than cloud inference, with the majority of environmental impact stemming from device embodied carbon.
By \'Edouard Gu\'egain, Tristan Coignion
The paper presents an autonomous agent that designs machine learning algorithms for wireless power control, eliminating manual specification of architecture, loss, and training details. Using an autoresearch protocol, the agent iteratively edits a training script, runs experiments, and evaluates changes against a single metric, ultimately achieving 99.5% of a reference solution with vastly reduced inference cost. The agent’s discovered output parameterization matches the exact max‑min‑optimal allocation at the minimum percentile for all trained weights, demonstrating a principled, scalable approach to a complex, NP‑hard problem.
By Ahmad Khan, Akram Bin Sediq, Sara Azadegi Naeini, Raviraj S. Adve
Sixth-generation (6G) wireless networks are expected to provide a key infrastructure for large-scale embodied intelligence, where heterogeneous robots collaborate through low-latency connectivity, edg...
arXiv:2607. 16133v1 Announce Type: cross Abstract: LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks.
By Wendi Yu, Lianhao Zhou, Xiangjue Dong, Sai Sudarshan Barath, Declan Staunton, Byung-Jun Yoon, Xiaoning Qian, James Caverlee, Shuiwang Ji
arXiv:2501. 16726v2 Announce Type: replace-cross Abstract: Semantic communications aim to enhance transmission efficiency by jointly optimizing source coding, channel coding, and modulation.
By Hanju Yoo, Dongha Choi, Yonghwi Kim, Yoontae Kim, Songkuk Kim, Chan-Byoung Chae, Robert W. Heath Jr
arXiv:2607. 13160v1 Announce Type: cross Abstract: Active beyond-diagonal reconfigurable intelligent surfaces (BD-RISs) enables hybrid transmitting and reflecting mode to achieve effective signal amplification and full-space coverage, thus providing a promising solution for blockage-aware uplink offloading in heterogeneous mobile edge computing (MEC) systems.
By Tianyu Pang, Hongyu Li
arXiv:2608. 13863v1 Announce Type: new Abstract: Deep neural network (DNN) inference on mobile devices often incurs high latency and energy consumption due to limited computing and memory resources.
By Yunchu Han, Zhaojun Nan, Sheng Zhou, Zhisheng Niu
arXiv:2608. 15502v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have emerged as a promising foundation for Embodied AI, but their high inference cost poses significant challenges for deployment in robotic systems.
By Ao Zhou, Bo Dai, Le Yu, Xingyu Liu, Zeyu Hao, Lingkun Long, Chunming Hu, Jianlei Yang
arXiv:2609.00284v1 Announce Type: cross
Abstract: Emerging 6G wireless networks are expected to operate across diverse deployment scenarios, where variations in network topology, user mobility, traff...
By Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci
The paper introduces FedMVLA, a modality‑decoupled federated learning framework designed for privacy‑preserving embodied intelligence in 6G networks. It addresses the unique challenges of vision‑language‑action models by employing modality‑aware federated aggregation, privacy allocation, and communication compression, along with a precision‑critical action transport slice. A case study on federated robotic manipulation demonstrates significant gains in task success, scalability, and uplink payload reduction compared to standard FedAvg.
By Zhuodong Liu, Xiangyu Li, Chunhong Yuan, Hongyang Du, Bodong Shang, Qingqing Wu, Tony Q. S. Quek, Mohsen Guizani