arXiv AI

Exploring Context-aware and LLM-driven Locomotion for Immersive Virtual Reality

arXiv:2504. 17331v3 Announce Type: replace-cross Abstract: Locomotion plays a crucial role in shaping the user experience within virtual reality environments.

arXiv Machine Learning
Sep 15

Speak to the City: Multimodal Resolution for Outside-the-Vehicle References

arXiv:2609.14691v1 Announce Type: cross Abstract: As autonomous vehicles and Extended Reality (XR) headsets enable novel in-car interactions, seamlessly querying physical landmarks, known as Outside-...

By Alireza Parchami (Mercedes-Benz Tech Innovation GmbH, Saarland University), Artin Saberpour (Saarland University), Robin Connor Schramm (Mercedes-Benz Tech Innovation GmbH, RheinMain University of Applied Sciences), J\"urgen Steimle (Saarland University), Ulrich Schwanecke (RheinMain University of Applied Sciences)
arXiv Computer Vision
Sep 7

EyeMakeYou: Identity-, Task-, and Subjective-State-Conditioned Diffusion for High-Frequency Gaze Synthesis

EyeMakeYou is a multi‑conditional denoising diffusion model that synthesizes high‑frequency, subject‑specific gaze velocity sequences. It conditions on identity, task, and self‑reported subjective states (difficulty, mental tiredness, eye tiredness) to generate realistic 5‑second, 1000‑Hz bivariate gaze data from a reference trajectory. Experiments on the GazeBase dataset show that EyeMakeYou outperforms existing generative methods in spatial accuracy and real‑synthetic similarity while preserving task‑dependent associations with subjective reports.

By Kamrul Hasan, Mehedi Hasan Raju, Oleg V. Komogortsev
arXiv AI
Sep 21

Touvigation: Embodied Adaptive Object Acquisition for Blind and Low-Vision Users in Unfamiliar Indoor Environments

arXiv:2609.21828v1 Announce Type: cross Abstract: Blind and low-vision users often face challenges when locating and physically acquiring objects in unfamiliar indoor environments. Existing vision-la...

By George Xi Wang, Xiangyu Li, Shaoyue Wen, Jiaqian Hu, Junan Xie, Yupeng Wang, Ziyue Shi, Qijun Chen, Maaike Bouwmeester, Yuhua Jin, Jing Qian
arXiv AI
Jun 18

Signals of Provenance: Practices & Challenges of Navigating Indicators in AI-Generated Media for Sighted and Blind Individuals

arXiv:2505. 16057v2 Announce Type: replace-cross Abstract: AI-Generated (AIG) content has become increasingly widespread by recent advances in generative models and the easy-to-use tools that have significantly lowered the technical barriers for producing highly realistic audio, images, and videos through simple natural language prompts.

By Ayae Ide, Tory Park, Jaron Mink, Tanusree Sharma
arXiv AI
Sep 2

RecalibrateGPT: AI Fatigue Resilient Conversational Interfaces

RecalibrateGPT is a system designed to reduce AI fatigue in conversational interfaces by introducing five cross-turn operators—Anchor, Replay, Delta, Scope, and Steer—that target specific fatigue types. Users can apply these operators with a single click via an AssistiveButton in one of three layout options (Vertical, Arc, Tablet). Pilot studies with advanced LLM users showed that the system cuts perceived cognitive workload by half while maintaining high usability.

By Nikhil Wani
arXiv Computer Vision
Aug 27

Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans

The paper introduces a real‑time framework for generating co‑speech gestures for digital humans, coupling a streaming speech response module with a causal multimodal autoregressive gesture generator that uses only current speech and motion history. It also presents an offline data synthesis pipeline for virtual companion dialogues and a self‑evolving training loop that incorporates user feedback to continually adapt the model. Experiments show the system achieves a better latency‑quality trade‑off, stronger speech‑motion synchronization, and higher user preference than existing baselines.

By Wentao Jiang, Youchen Xie, Haidi Fan, Yajing Chen, Xin Wang, Ye Shi, Jingya Wang