arXiv Computer Vision

COMPASS: Fusion-Matched Supervision for Missing-Modality Human Sensing

COMPASS is a completion-and-fusion framework designed for multimodal human activity recognition (HAR) and human pose estimation (HPE) when some modalities are missing at inference. It assigns each modality to a fixed slot, filling it with either an observed representation or a completion inferred from available inputs, and uses fusion‑matched supervision to train completions against real targets at the readout level. Experiments on XRF55 and MM‑Fi datasets show that COMPASS outperforms strong baselines and alternative matching strategies for both HAR and HPE.

arXiv AI
Aug 25

Multi-Context Fusion Transformer for Pedestrian Crossing Intention Prediction in Urban Environments

The paper introduces the Multi-Context Fusion Transformer (MFT), a model that predicts pedestrian crossing intentions in urban settings by integrating four types of contextual information—pedestrian behavior, environment, localization, and vehicle motion—through a progressive fusion strategy. MFT uses intra-context attention for reciprocal interactions within each context, cross-context attention to combine these contexts into a global representation, and guided attention mechanisms to refine both context tokens and the global token. Experiments on JAADbeh, JAADall, and PIE datasets show MFT outperforms existing methods with accuracies of 73%, 93%, and 90% respectively, and ablation studies confirm the importance of each network component and input context.

By Yuanzhe Li, Hang Zhong, Steffen M\"uller
arXiv Computer Vision
Sep 14

RoES: Rotational Equivariant Selective-frequency Fusion for Multimodal Images

RoES is a Rotational Equivariant Selective-frequency fusion network that dynamically separates low- and high-frequency components of infrared-visible images. It uses a trainable rotation-enhanced updater to decouple frequencies, then fuses them with a dual-branch module: a rotation-equivariant Mamba for low-frequency structural dependencies and a polar spectral attention Dual-Fourier block for high-frequency detail refinement. Experiments show RoES outperforms existing methods in fusion quality and downstream object detection, offering a robust multimodal fusion solution.

By Jiabao Wang, Wenjian Liu, Yaoming Cai, Gengyu Zhang, Boyan Zhao, Zijia Zhang, Yao Ding, Xiaobo Liu
arXiv AI
Jun 2

Interpretable Multimodal Gesture Recognition for Drone and Mobile Robot Teleoperation via Log-Likelihood Ratio Fusion

arXiv:2602. 23694v3 Announce Type: replace-cross Abstract: Human operators are still frequently exposed to hazardous environments such as disaster zones and industrial facilities, where intuitive and reliable teleoperation of mobile robots and Unmanned Aerial Vehicles (UAVs) is essential.

By Seungyeol Baek, Jaspreet Singh, Lala Shakti Swarup Ray, Hymalai Bello, Paul Lukowicz, Sungho Suh
arXiv Computer Vision
Aug 31

A-PAIR: A Benchmark and Identity-Consistent Grounding Framework for Air-Ground Cross-View Referring Person Detection

arXiv:2608.27997v1 Announce Type: new Abstract: Air-ground cross-view referring person detection is a necessary component in the language-to-perception-to-control chain of collective embodied intelli...

By Zhoupeng Guo, Xinjie Yao, Yunqi Zhu, Zhihe Fan, Siqi Zhao, Jianjun Chen, Yichen Dong, Yan Fan, Pengfei Zhu
arXiv AI
Sep 16

Coverage-Aware Virtual IMU Augmentation for Low-Resource Human Activity Recognition

The paper introduces a coverage-aware virtual IMU augmentation framework for human activity recognition. It selects diverse and scarce data points in a learned sensor embedding space, generates virtual IMU samples as prompts, ranks them by proximity and label consistency, and incorporates them into training with reliability-based weights. Experiments on public benchmarks demonstrate consistent performance gains over existing baselines, with ablation studies confirming the framework’s effectiveness.

By Jiayuan Gao, Yingwei Zhang, Ziyao Tang, Yuejia Ma, Yuanzhe Chen, Shuchao Song, Boshi Tang
arXiv Computer Vision
Sep 18

MTF-Net: Multi-Modal Temporal Feature Fusion Network for Pedestrian Intention Prediction

MTF‑Net is a Multi‑Modal Temporal Feature Fusion Network that jointly models kinematic, appearance, and contextual cues for pedestrian intention prediction. It fuses four modalities—bounding‑box dynamics, human pose keypoints, local context, and scene‑level semantics—within a recurrent framework enhanced by gated linear units (GLUs) and an attention‑guided fusion head. Evaluations on the PIE and JAAD benchmarks show that MTF‑Net outperforms recent transformer‑ and graph‑based models, achieving up to 0.95 AUC on PIE and 0.94 AUC on JAAD while maintaining real‑time performance.

By Md Mahfuzur Rahman, Pengzhan Zhou, A. F. M. Abdun Noor, Md Imam Ahasan, Md Mustafizur Rahman, Fang Qu
arXiv AI
Sep 25

WildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model

WildHSR introduces a lightweight adaptation of 3D foundation models to jointly recover metric cameras, scene geometry, and persistent person identities from monocular video. By generating pseudo‑scale labels from curated web footage and fine‑tuning a Scale Readout, the method predicts metric scale directly from foundation‑model tokens. It also exploits intermediate query‑key features to associate per‑frame bodies, enabling feed‑forward reconstruction that outperforms state‑of‑the‑art optimization‑based methods on several benchmarks while running at 10.1 fps.

By Jerrin Bright, John Zelek