arXiv Computer Vision By Hao Wang, Yanyu Qian, Pengcheng Weng, Zixuan Xia, William Dan, Yangxin Xu, Fei Wang

COMPASS: Fusion-Matched Supervision for Missing-Modality Human Sensing

Read the original on arXiv Computer Vision →

COMPASS is a completion-and-fusion framework designed for multimodal human activity recognition (HAR) and human pose estimation (HPE) when some modalities are missing at inference. It assigns each modality to a fixed slot, filling it with either an observed representation or a completion inferred from available inputs, and uses fusion‑matched supervision to train completions against real targets at the readout level. Experiments on XRF55 and MM‑Fi datasets show that COMPASS outperforms strong baselines and alternative matching strategies for both HAR and HPE.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Aug 25

Multi-Context Fusion Transformer for Pedestrian Crossing Intention Prediction in Urban Environments

The paper introduces the Multi-Context Fusion Transformer (MFT), a model that predicts pedestrian crossing intentions in urban settings by integrating four types of contextual information—pedestrian behavior, environment, localization, and vehicle motion—through a progressive fusion strategy. MFT uses intra-context attention for reciprocal interactions within each context, cross-context attention to combine these contexts into a global representation, and guided attention mechanisms to refine both context tokens and the global token. Experiments on JAADbeh, JAADall, and PIE datasets show MFT outperforms existing methods with accuracies of 73%, 93%, and 90% respectively, and ablation studies confirm the importance of each network component and input context.

By Yuanzhe Li, Hang Zhong, Steffen M\"uller
arXiv Computer Vision
Sep 14

RoES: Rotational Equivariant Selective-frequency Fusion for Multimodal Images

RoES is a Rotational Equivariant Selective-frequency fusion network that dynamically separates low- and high-frequency components of infrared-visible images. It uses a trainable rotation-enhanced updater to decouple frequencies, then fuses them with a dual-branch module: a rotation-equivariant Mamba for low-frequency structural dependencies and a polar spectral attention Dual-Fourier block for high-frequency detail refinement. Experiments show RoES outperforms existing methods in fusion quality and downstream object detection, offering a robust multimodal fusion solution.

By Jiabao Wang, Wenjian Liu, Yaoming Cai, Gengyu Zhang, Boyan Zhao, Zijia Zhang, Yao Ding, Xiaobo Liu
arXiv AI
Jun 2

Interpretable Multimodal Gesture Recognition for Drone and Mobile Robot Teleoperation via Log-Likelihood Ratio Fusion

arXiv:2602. 23694v3 Announce Type: replace-cross Abstract: Human operators are still frequently exposed to hazardous environments such as disaster zones and industrial facilities, where intuitive and reliable teleoperation of mobile robots and Unmanned Aerial Vehicles (UAVs) is essential.

By Seungyeol Baek, Jaspreet Singh, Lala Shakti Swarup Ray, Hymalai Bello, Paul Lukowicz, Sungho Suh
arXiv Computer Vision
Aug 31

A-PAIR: A Benchmark and Identity-Consistent Grounding Framework for Air-Ground Cross-View Referring Person Detection

arXiv:2608.27997v1 Announce Type: new Abstract: Air-ground cross-view referring person detection is a necessary component in the language-to-perception-to-control chain of collective embodied intelli...

By Zhoupeng Guo, Xinjie Yao, Yunqi Zhu, Zhihe Fan, Siqi Zhao, Jianjun Chen, Yichen Dong, Yan Fan, Pengfei Zhu
arXiv AI
Sep 16

Coverage-Aware Virtual IMU Augmentation for Low-Resource Human Activity Recognition

The paper introduces a coverage-aware virtual IMU augmentation framework for human activity recognition. It selects diverse and scarce data points in a learned sensor embedding space, generates virtual IMU samples as prompts, ranks them by proximity and label consistency, and incorporates them into training with reliability-based weights. Experiments on public benchmarks demonstrate consistent performance gains over existing baselines, with ablation studies confirming the framework’s effectiveness.

By Jiayuan Gao, Yingwei Zhang, Ziyao Tang, Yuejia Ma, Yuanzhe Chen, Shuchao Song, Boshi Tang