arXiv:2606. 26922v1 Announce Type: cross Abstract: Continuous driver monitoring in automated vehicles requires low-latency inference while avoiding unsafe decisions under uncertain driver states.
By Daosheng Qiu, Haozhuang Chi, Hao Su, Shu Long, Xinyue Miao, Yongle Dong, Wei Zhang
arXiv:2605. 20301v2 Announce Type: replace-cross Abstract: In autonomous driving, 3D object detection is essential for accurate perception and reliable decision-making.
By Wenxuan Li, Qin Zou, Shoubing Chen, Chi Chen, Yingyi Yang, Qingxiang Meng
arXiv:2607. 23822v1 Announce Type: new Abstract: Driving style captures stable, driver-specific patterns in how a vehicle is driven.
By Yuhang Wang, Lingyao Li, Hao Zhou
arXiv:2604. 08435v2 Announce Type: replace-cross Abstract: It remains challenging to assess driver fatigue from untrimmed videos under constrained computational budgets, due to the difficulty of modeling long-range temporal dependencies in subtle facial expressions.
By Changdao Chen, Qinqiuhong Ye, Hao Chen, Jinyu Wang
The paper introduces LaPla, a Vision‑Language‑Action framework that uses a latent‑aligned planning approach to convert discrete semantic reasoning into continuous, physics‑constrained driving actions. It employs a residual VQ‑VAE to encode vehicle kinematics into a structured latent space, then projects multimodal inputs—images, past actions, and text—directly into this latent space, allowing a frozen decoder to generate physically plausible trajectories without quantization errors. Experiments on nuScenes and NVIDIA AlpaSim show LaPla reduces long‑horizon L2 error by 15.52% and improves closed‑loop success rates by 33.34 percentage points while cutting inference latency.
By Ruoyu Yao, Yusen Xie, Qingzhao Liu, Pei Liu, Zewei Yang, Yipeng Zhu, Xiaolong Wang, Jun Ma
arXiv:2512. 05277v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as the perception and reasoning backbone of autonomous agents acting in the wild, with autonomous driving (AD) being one of the most safety-critical instances.
By Kevin Cannons, Saeed Ranjbar Alvar, Mohammad Asiful Hossain, Ahmad Rezaei, Mohsen Gholami, Alireza Heidarikhazaei, Zhou Weimin, Yong Zhang, Mohammad Akbari
arXiv:2607. 20988v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-end autonomous driving.
By Quanfu Yu, Xian Wu, Hao Xu, Liulong Ma
arXiv:2607. 00296v1 Announce Type: cross Abstract: Human motion forecasting in unconstrained real-world videos remains challenging due to the ambiguity of future behaviors and the presence of noisy multimodal observations.
By Jingni Huang
arXiv:2512. 05277v4 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as the perception and reasoning backbone of autonomous agents acting in the wild, with autonomous driving (AD) being one of the most safety-critical instances.
By Kevin Cannons, Saeed Ranjbar Alvar, Mohammad Asiful Hossain, Ahmad Rezaei, Mohsen Gholami, Alireza Heidarikhazaei, Zhou Weimin, Yong Zhang, Mohammad Akbari
arXiv:2606. 24759v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision.
By Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye
UNWIND is a facial‑video framework that detects stress by treating an entire recording as a single input, avoiding the need for temporal windowing or segmentation. It folds the video’s temporal dimension into the channel dimension of a 2‑D spatial representation and processes it with an asymmetric‑attention architecture. Experiments on a 58‑subject stress dataset show that using all 3,600 frames (stride τ = 1) yields a 69.73 % accuracy, comparable to the best 70.02 % accuracy at τ = 15, while computational cost varies from 12.48 to 348.78 GFLOPs.
By Stefanos Gkikas, Christian Arzate Cruz, Eric Nichols, Giorgos Giannakakis, Randy Gomez
The paper introduces MiRA, a plug‑in framework that reweights framewise attention in Vision Transformer video models to better capture subtle facial dynamics for expression recognition. MiRA computes frame‑level confidence and intra‑frame concentration from self‑attention maps, redistributing attention toward localized facial cues without adding trainable parameters. Two modes—an exact post‑softmax redistribution and a lightweight flashLite pre‑softmax approximation—are proposed, and experiments on facial expression recognition benchmarks show consistent gains over strong ViT baselines.
By Seongro Yoon, Donghyeon Cho, Jinsun Park, Fran\c{c}ois Br\'emond