The paper introduces a byte‑constrained cooperative perception framework that balances dense coverage with sparse refinement. Each vehicle sends a highly compressed coarse Bird’s‑Eye‑View (BEV) layer covering the entire map and uses the remaining bandwidth to transmit high‑resolution patches selected by a Task‑Aware Benefit Selector. Experiments on DAIR‑V2X and OPV2V demonstrate that this coverage‑refinement strategy achieves superior accuracy‑payload trade‑offs, reaching 0.60 AP@0.7 with only 1.87 KB per non‑ego agent.
By Melih Yazgan, Timon M\"uller, J. Marius Z\"ollner
arXiv:2608. 10198v1 Announce Type: new Abstract: Latent-space communication allows heterogeneous vision-language model agents to exchange continuous representations without serializing visual and reasoning states into text.
By Di Wu, Xiaohui Zhu
The paper introduces VLA-ULAP, a lightweight local action predictor that interleaves remote vision–language–action (VLA) calls with on‑edge inference. ULAP, with only 7.4 M parameters, predicts action chunks in a single pass using current views, proprioception, and action history, eliminating the need for VLA hidden states or server round‑trips. Experiments on Jetson Orin Nano and simulated benchmarks show that VLA-ULAP can remove 48.8–76.7 % of VLA calls while preserving 95–97.5 % of baseline success, and it outperforms local VLA‑acceleration alternatives in both inference time and energy consumption.
By Deyu Cao, Ryuji Oi, Kosuke Matsushima, Yuxuan Pan, Ziheng Wang, Daichi Fujiki, Atsutake Kosuge
arXiv:2609.18084v1 Announce Type: cross
Abstract: Fine-tuning a Vision-Language-Action (VLA) model for a new deployment environment is expensive, yet most methods apply uniform-capacity adapters to e...
By Shahram Najam Syed, Arthur Jakobsson, Prayuj Sachdev, Jeffrey Ichnowski
arXiv:2607. 09520v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are the perceptual backbone of embodied AI, but their energy footprint on edge hardware remains poorly understood.
By Junfei Zhan, Haoxun Shen, Mingang Guo, Zixuan Huang, Tengjiao He
The paper introduces NOSTRAdAMUS, a predictive link‑adaptation framework for 5G NR that forecasts retransmissions in the next radio frame using recent HARQ history and adjusts the Modulation and Coding Scheme accordingly. Gradient Boosting models achieve 82.9% overall accuracy, with high‑confidence predictions correct 94.2% of the time and a 5.5 µs inference latency. Evaluated OTA on the X5G testbed and various channel emulators, the approach boosts goodput by up to 71.5% and cuts retransmissions by up to 71.8% without retraining across diverse scenarios.
By Tamerlan Aghayev, Maxime Elkael, Michele Polese, Reshma Prasad, Salvatore D'Oro, Yunseong Lee, Koichiro Furueda, Tommaso Melodia
arXiv:2609.23974v1 Announce Type: new
Abstract: Foundation models are endowing autonomous systems with greater intelligence, enabling a more comprehensive understanding of the environment through vis...
By Boxun Hu, Jiawei Ge, Axel Krieger, Peng Wang, Tinoosh Mohsenin
arXiv:2609.39296v1 Announce Type: new
Abstract: Video semantic communication has attracted increasing attention as a promising approach to improving video transmission efficiency. However, most exist...
By Xiangben Zhu, Caili Guo, Yang Yang, Chuanhong Liu, Meiyi Zhu
The paper evaluates the claim that vision‑language models (VLMs) outperform task‑specific vision backbones for UAV power‑line defect assessment using the ElecVQA‑Bench benchmark. Across various evaluation settings—partitioning, item sets, label spaces, replication, resolution, and side information—the performance gap between VLMs and traditional backbones is minimal or even reversed when controlling for resolution and token budget. The study concludes that VLM superiority is not universally supported and emphasizes the importance of rigorous benchmark audits.
By Linghao Zhang, Siyu Xiang, Junwei Kuang, Peiyu Yi
Foundation models are endowing autonomous systems with greater intelligence, enabling a more comprehensive understanding of the environment through visual perception. A representative example is Human...
arXiv:2609.13947v1 Announce Type: cross
Abstract: In-sensor computing reduces the cost of transmitting high-resolution image data by performing early-stage processing near the sensor. However, the lo...
By Chengwei Zhou, Abu Masum, Xuming Chen, Mehran Moghadam, Sreetama Sarkar, Arnab Sanyal, Md Abdullah-Al Kaiser, M. Hassan Najafi, Sercan Aygun, Gourav Datta
WiSP (Working‑Set Paging) is a routing‑aware expert pager that allows Mixture‑of‑Experts models to run on GPUs that cannot hold the entire expert pool by paging experts in and out of VRAM while preserving byte‑identical outputs. On a 24 GiB RTX 3090, WiSP doubles decode throughput compared to static offload when the model does not fit, and its companion policy MV‑WSA allocates VRAM between resident experts and KV cache based on marginal latency benefit, reducing end‑to‑end time by up to 1.19× without altering model outputs.
By Jiamu Zhang, Liang Wu, Mayank Darbari, Liangjie Hong