arXiv:2609.10714v1 Announce Type: cross
Abstract: The ambitious requirements of sixth-generation (6G) networks are driving communication systems from reliable bit delivery toward meaning-aware and ta...
By Yu Ma, Zhen Gao, Li Qiao, Xiaoyuan Zhang, Mahdi Boloursaz Mashhadi, Yin Xu, Wenjun Xu, Xiaodong Xu, Kaibin Huang, Jiangzhou Wang, Rahim Tafazolli, Sheng Chen, Tony Q. S. Quek, Ping Zhang
The paper introduces Embodied Semantic Communication (ESC), a new paradigm that redefines information transmission for autonomous agents by embedding multimodal perceptual states, hardware capabilities, and collaborative intents into unified, action‑oriented semantic representations. ESC enables heterogeneous agents to parse, align, and ground shared information directly into local motor control, addressing the limitations of traditional communication approaches that focus solely on bit delivery or single‑task optimization. The tutorial outlines ESC’s conceptual boundaries, system characteristics, and technical pathways, mapping relevant mathematical tools such as semantic information theory, world models, and multi‑agent decision theory, and concludes with a roadmap of open challenges like semantic reliability, dynamic interaction, and bandwidth‑adaptive transmission.
By Yizheng Huang, Wensheng Lin, Lixin Li, Qinghe Du, Wenchi Cheng, Zhu Han
arXiv:2607. 03283v1 Announce Type: new Abstract: Embodied intelligence systems require not only end-to-end policy models, but also reusable functional modules that transform multimodal observations, robot states, human demonstrations, and task contexts into structured representations, decisions, trajectories, control references, and system services.
By Junwu Xiong, Jiaxuan Gao, Wei Chai, Renxing Chen, Yuzhen Li, Yu Guo, Yucheng Guo, Mingxi Luo, Wenyang Ma, Yiyun Mou, Yifei Zhang, Chen Zhou, Yongjian Guo
Building capable embodied agents requires not only multimodal perception and understanding, but also agentic capabilities for reasoning about actions, adapting to evolving situations, and interacting with the physical world. In this report, we introduce Hy-Embodied-VLM-1.
arXiv:2608. 19794v1 Announce Type: new Abstract: The convergence of large language models (LLMs), structured knowledge bases (KBs), and reasoning ability (RA) presents a promising trajectory toward general embodied intelligence (GEI).
By Fujiang Yuan, Xia Huang, Lusheng Wang, Jun Ding, Zhen Tian, Yuxin Wang, Shaojie Gu, Yuki Funabora, Yanhong Peng, Zebing Mao
arXiv:2609.24526v2 Announce Type: replace
Abstract: Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints...
By Foundation Model, Li Auto Inc
arXiv:2609.08977v3 Announce Type: replace-cross
Abstract: In this work, we present Gander, a native multimodal duplex interaction model that builds on MiniCPM-o 4.5 and is further adapted for realtim...
By Orantqing, Shengpeng Ji, Junlong Tong, Jialong Zuo, Dongjie Fu, Di Cao, Yangzhuo Li, Shangda Wu, Franz, Evan, Theron Veyra, Changhao Pan, Jingyu Lu, Dongchao Yang, Zhifei Xie, Yang Tan, Xiaoyu Shen, Xiaoda Yang, Wenfu Wang, Teddy Sun, Steve Yves, Zhou Zhao
arXiv:2609.24526v1 Announce Type: new
Abstract: Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints and...
By Foundation Model, Li Auto Inc
The paper introduces a token‑oriented semantic communication framework that transmits only task‑relevant image latents instead of full token embeddings, reducing communication cost and improving interoperability. It leverages a spatial alignment between vision transformer patch tokens and learned image compression latents, enabling token‑level relevance estimation and selective transmission. Experiments on ImageNet demonstrate a superior rate–accuracy trade‑off compared to existing semantic communication methods and hand‑crafted codecs.
By Jiwoong Im, Minwoo Kim, Jaeho Lee, Yo-Seb Jeon, Yongjune Kim
arXiv:2505. 10946v3 Announce Type: replace-cross Abstract: Token communications (TokenCom) is an emerging generative semantic communication paradigm, where tokens serve as compact representation units across modalities.
By Li Qiao, Mahdi Boloursaz Mashhadi, Zhen Gao, Robert Schober, Deniz G\"und\"uz
arXiv:2606. 11324v1 Announce Type: cross Abstract: We introduce Embodied-R1.
By Yifu Yuan, Yaoting Huang, Xianze Yao, Yutong Li, Shuoheng Zhang, Linqi Han, Pengyi Li, Jiangeng Sun, Wenting Jia, Zhao Zhang, Yuhao Liu, Ruihao Liao, Yucheng Hu, Qiyu Wu, Yuxiao Li, Zibin Dong, Fei Ni, Yan Zheng, Shuyang Gu, Yi Ma, Hongyao Tang, Han Hu, Jianye Hao
arXiv:2606. 18363v1 Announce Type: cross Abstract: Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents.
By Haowen Liu, Xirui Li, Shaoxiong Yao, Peng Shi, Tianyi Zhou, Jia-Bin Huang, Furong Huang, Jiayuan Mao