arXiv AI

SlackDrive: Reclaiming Runtime Slack for Adaptive Driving Inference

SlackDrive is a pre‑inference compute allocator that dynamically selects the compute budget for each driving control step by reusing the realized latency from previous inferences. By profiling a small set of discrete budgets once, it estimates the current compute state online and chooses the highest‑utility budget that stays within the admissible latency envelope. On the NAVSIM v2 benchmark with DriveDreamer‑Policy, SlackDrive boosts latency‑constrained EPDMS performance by 21.7% compared to the best baseline, while full‑budget and token‑pruning approaches exceed the latency limits under runtime contention.

arXiv AI
Aug 18

Efficient Block-Layer Parallel Inference for Vision-Language-Action on Hybrid Architectures

arXiv:2608. 14586v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are becoming a promising paradigm for autonomous driving, but their deployment on existing vehicle platforms remains difficult because they introduce both high inference latency and strong GPU-side resource pressure.

By Haibo HU, Lianming Huang, Qiao Li, Nan Guan, Chun Jason Xue
Hugging Face Trending Papers
Aug 3

TALSC: Timeliness-Aware Large-Small VLM Collaboration for Infrastructure-Assisted Autonomous Driving

The deployment of Vision-Language Models (VLMs) in autonomous driving (AD) systems is constrained by on-board computing power, restricting vehicles to small VLMs (SVLMs) with limited perception and reasoning capabilities. Infrastructure-assisted AD alleviates this resource constraint by enabling collaboration with large VLMs (LVLMs) at edge servers.

arXiv AI
6d ago

Guiding End-to-End Driving Models with Endpoint-Constrained Trajectory Optimization

The paper introduces Endpoint-Constrained Optimization (ECO), a lightweight postprocessing layer that corrects intermediate waypoints of end-to-end driving policies while preserving the predicted endpoint. ECO does not require maps, privileged simulator state, or additional training, and can be applied to a wide range of waypoint-emitting policies. Experiments on two closed-loop simulators show that ECO significantly improves closed-loop performance, achieving top results in the HUGSIM Closed-Loop Driving Challenge and boosting scene scores on AlpaSim.

By Brayden Zhang, Mahsa Golchoubian, Igor Gilitschenski, Boris Ivanovic, Kashyap Chitta
arXiv AI
Jun 29

End-to-End Dynamic Sparsity for Resource-Adaptive LLM Inference

arXiv:2606. 27743v1 Announce Type: cross Abstract: Large Language Models (LLMs) inference is typically deployed under a static resource assumption, where models execute a fixed computational graph regardless of the runtime environment.

By Yuhang Chen, Jinhao Duan, Ruichen Zhang, Mingfu Liang, Xiaohan Wei, Yunchen Pu, Fei Tian, Chonglin Sun, Parish Aggarwal, Frank Shyu, Luke Simon, Sandeep Pandey, Tianlong Chen, Xi Liu
arXiv AI
Jun 3

AUGUSTE: Online-Learning dApp for Predictive URLLC Scheduling

arXiv:2606. 03664v1 Announce Type: cross Abstract: Ultra Reliable and Low Latency Communications (URLLC) was one of the main motivations behind 5G, with 3GPP advertising 1-10 ms latency targets for applications such as industrial automation, Vehicle-To-Everything (V2X), tactical edge networking, and unmanned-system control.

By Maxime Elkael, Michele Polese, Yunseong Lee, Koichiro Furueda, Tommaso Melodia