arXiv Computer Vision

FRESHLATENT: Channel-Aware Latent Adaptation for Resource-Constrained Embodied VLM Perception

arXiv Computer Vision
Sep 25

Dense Coverage, Sparse Refinement: Byte-Constrained Cooperative Perception

The paper introduces a byte‑constrained cooperative perception framework that balances dense coverage with sparse refinement. Each vehicle sends a highly compressed coarse Bird’s‑Eye‑View (BEV) layer covering the entire map and uses the remaining bandwidth to transmit high‑resolution patches selected by a Task‑Aware Benefit Selector. Experiments on DAIR‑V2X and OPV2V demonstrate that this coverage‑refinement strategy achieves superior accuracy‑payload trade‑offs, reaching 0.60 AP@0.7 with only 1.87 KB per non‑ego agent.

By Melih Yazgan, Timon M\"uller, J. Marius Z\"ollner
arXiv Machine Learning
Sep 17

VLA-ULAP: Interleaving Cloud VLA Calls with Ultra-Lightweight Local Action Prediction at the Edge

The paper introduces VLA-ULAP, a lightweight local action predictor that interleaves remote vision–language–action (VLA) calls with on‑edge inference. ULAP, with only 7.4 M parameters, predicts action chunks in a single pass using current views, proprioception, and action history, eliminating the need for VLA hidden states or server round‑trips. Experiments on Jetson Orin Nano and simulated benchmarks show that VLA-ULAP can remove 48.8–76.7 % of VLA calls while preserving 95–97.5 % of baseline success, and it outperforms local VLA‑acceleration alternatives in both inference time and energy consumption.

By Deyu Cao, Ryuji Oi, Kosuke Matsushima, Yuxuan Pan, Ziheng Wang, Daichi Fujiki, Atsutake Kosuge
arXiv AI
Sep 11

Improving 5G AI-RAN MCS Selection by Predicting Retransmissions

The paper introduces NOSTRAdAMUS, a predictive link‑adaptation framework for 5G NR that forecasts retransmissions in the next radio frame using recent HARQ history and adjusts the Modulation and Coding Scheme accordingly. Gradient Boosting models achieve 82.9% overall accuracy, with high‑confidence predictions correct 94.2% of the time and a 5.5 µs inference latency. Evaluated OTA on the X5G testbed and various channel emulators, the approach boosts goodput by up to 71.5% and cuts retransmissions by up to 71.8% without retraining across diverse scenarios.

By Tamerlan Aghayev, Maxime Elkael, Michele Polese, Reshma Prasad, Salvatore D'Oro, Yunseong Lee, Koichiro Furueda, Tommaso Melodia
arXiv Computer Vision
Sep 24

Beyond Balanced Accuracy: A Resolution and Parity-Controlled Benchmark for Vision-Language and Vision-Only Defect Assessment in UAV Power-Line Inspection

The paper evaluates the claim that vision‑language models (VLMs) outperform task‑specific vision backbones for UAV power‑line defect assessment using the ElecVQA‑Bench benchmark. Across various evaluation settings—partitioning, item sets, label spaces, replication, resolution, and side information—the performance gap between VLMs and traditional backbones is minimal or even reversed when controlling for resolution and token budget. The study concludes that VLM superiority is not universally supported and emphasizes the importance of rigorous benchmark audits.

By Linghao Zhang, Siyu Xiang, Junwei Kuang, Peiyu Yi
arXiv Machine Learning
Sep 1

WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware

WiSP (Working‑Set Paging) is a routing‑aware expert pager that allows Mixture‑of‑Experts models to run on GPUs that cannot hold the entire expert pool by paging experts in and out of VRAM while preserving byte‑identical outputs. On a 24 GiB RTX 3090, WiSP doubles decode throughput compared to static offload when the model does not fit, and its companion policy MV‑WSA allocates VRAM between resident experts and KV cache based on marginal latency benefit, reducing end‑to‑end time by up to 1.19× without altering model outputs.

By Jiamu Zhang, Liang Wu, Mayank Darbari, Liangjie Hong