arXiv AI

MAND: Modality-Aware Novelty Detection for Open-World Egocentric Activity Recognition

arXiv:2603. 16970v2 Announce Type: replace-cross Abstract: Multimodal egocentric activity recognition integrates visual and inertial cues for robust first-person behavior understanding.

arXiv Computer Vision
Aug 27

Moving Beyond More Views: Redundancy-Aware Ego-Exo Fusion for Proficiency Estimation

The paper introduces a redundancy-aware fusion framework for EgoExo proficiency estimation, which integrates fine-grained motion cues from egocentric views with spatial context from exocentric views. It identifies multiview redundancy and overfitting as key challenges and proposes two modules—AdaMVS for adaptive view selection and VIB-GB for compressing redundant signals—to address them. Experiments on EgoExo-4D and EgoExo-Fitness show that the method learns to select informative views and fuse them effectively, achieving state‑of‑the‑art results.

By Xu Dong, Wanqing Li, Anthony Adeyemi-Ejeye, Andrew Gilbert
arXiv Machine Learning
Aug 28

HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition

HALO is a heterogeneity‑aware, language‑aligned foundation model for inertial measurement unit (IMU) based human activity recognition. It uses a two‑stage training process: first, a self‑supervised encoder learns to handle diverse sensor configurations and natural‑language sensor descriptions; second, the encoder is aligned with text embeddings through synonym‑aware contrastive learning, enabling open‑set recognition via cosine similarity. Trained on ten public HAR datasets, HALO outperforms five state‑of‑the‑art baselines across eight metrics while using only ~35 M parameters, and improves zero‑shot open‑set accuracy by 13.7 percentage points over 87 training labels.

By Zihan Ding, Liyu Zhang, Xiaomin Ouyang
arXiv AI
Aug 20

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding

EgoMemReason is a new benchmark for week‑long egocentric video understanding that focuses on memory‑driven reasoning rather than simple perception tasks. It tests three memory types—entity, event, and behavior—across 500 questions, each requiring evidence from an average of 5.1 video segments and 25.9 hours of backtracking. Evaluation of 17 models shows that even the best achieves only 39.6% accuracy, highlighting the difficulty of long‑horizon memory in multimodal systems.

By Ziyang Wang, Yue Zhang, Shoubin Yu, Ce Zhang, Zengqi Zhao, Jaehong Yoon, Hyunji Lee, Gedas Bertasius, Mohit Bansal
arXiv Computer Vision
Sep 18

BinoGen: Scaling egocentric binocular data for embodied visual perception and learning

BinoGen is an automated framework that generates large-scale, embodiment-aware egocentric binocular visual experiences in indoor environments. It models environmental and observer variation through generative scene synthesis, probabilistic object instantiation, appearance randomization, stochastic trajectory generation, and configurable binocular camera setups, producing synchronized videos with dense multimodal supervision such as depth maps, optical flow, surface normals, semantic maps, object coordinates, and camera poses. Using BinoGen, the authors created a dataset of over 20 million annotated images, demonstrating that incorporating this data improves real-world visual perception tasks like depth estimation, object detection, and video object tracking, and that embodiment-specific adaptation enhances performance while joint training enables a single model to perform competitively across different observer embodiments.

By Chunpeng Li, Ya-tang Li
arXiv AI
Sep 10

RevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion Mode Recognition in Older Adults and Clinical Cohorts

RevalExo is a new benchmark for locomotion mode recognition that focuses on functional daily activities performed by older adults and clinical cohorts. It includes 27 participants from three groups—healthy older adults, stroke survivors, and older adults with probable sarcopenia—recorded with lower-body IMUs and, for a subset, synchronized egocentric video. The dataset offers 10.1 hours of frame‑level annotations across 11 locomotion modes, and the authors evaluate unimodal, multimodal, cross‑population, and cross‑modal recognition challenges, finding that sensor fusion improves performance but transitions and generalization remain difficult.

By Diwas Lamsal, Juha Carlon, Reinhard Claeys, Maxim Yudayev, Louis Flynn, Tom Verstraten, David Beckw\'ee, Eva Swinnen, Mihai B\^ace, Bart Vanrumste, Benjamin Filtjens