Hugging Face Trending Papers

Closing the Modality Gap in Zero-Shot HAR: Contrastive Training and Separability-Optimized Prototypes on IMU Data

Zero-shot learning (ZSL) for inertial measurement unit (IMU)-based human activity recognition (HAR) faces a central challenge: bridging the gap between sensor embeddings and semantic class representations. We systematically evaluate seven configurations combining three inference methods with two training pipelines on the PAMAP2 dataset, using 14 seen and 4 unseen activity classes with subjects 108 and 109 held out for testing.

arXiv Machine Learning
Aug 28

HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition

HALO is a heterogeneity‑aware, language‑aligned foundation model for inertial measurement unit (IMU) based human activity recognition. It uses a two‑stage training process: first, a self‑supervised encoder learns to handle diverse sensor configurations and natural‑language sensor descriptions; second, the encoder is aligned with text embeddings through synonym‑aware contrastive learning, enabling open‑set recognition via cosine similarity. Trained on ten public HAR datasets, HALO outperforms five state‑of‑the‑art baselines across eight metrics while using only ~35 M parameters, and improves zero‑shot open‑set accuracy by 13.7 percentage points over 87 training labels.

By Zihan Ding, Liyu Zhang, Xiaomin Ouyang
arXiv Machine Learning
Sep 24

ChronoSteer: Bridging Large Language Model and Time Series Foundation Model via Synthetic Cross-Modal Alignment Dataset

ChronoSteer is a decoupled agentic framework that bridges large language models and time series foundation models by learning cross‑modal alignment from synthetic paired supervision. It converts textual events into revision instructions that steer a frozen time‑series model, discretizes these instructions into a compact codebook to reduce semantic divergence, and then refines the predictions with a two‑stage training strategy. The authors also release a leakage‑controlled multimodal benchmark and report a 25.8% improvement in zero‑shot prediction accuracy over the unimodal backbone.

By Chengsen Wang, Qi Qi, Zhongwen Rao, Lujia Pan, Jingyu Wang
arXiv AI
Sep 18

Lens: Bringing the Right Semantic Perspective into Focus for Training-Free Multimodal Representation Learning

The paper introduces Lens, a training‑free framework that aligns multimodal representations with the semantic perspective required by downstream tasks. Lens uses a task‑specific readout phrase to anchor the perspective and then aggregates token states after the full input, ensuring the extracted representation reflects task‑conditioned evidence integration rather than generic salient content. The method achieves a Precision@1 of 63.9 across 36 MMEB datasets, outperforming the nearest training‑free baseline by 10.2 points.

By Xinran Liu, Shouqian Shi, Yixian Chen, Ruizhi Chen, Xin-Wei Yao, Sheng Zhong
arXiv AI
Aug 28

Subspace Alignment for Vision-Language Model Test-time Adaptation

The paper introduces SubTTA, a test-time adaptation method for vision‑language models that aligns the semantic subspaces of visual and textual modalities to improve zero‑shot predictions. It addresses two issues: the modality gap caused by distribution shifts and visual nuisance that masks task‑specific semantics. By minimizing chordal distance between principal subspaces and projecting visual features onto a task‑specific textual subspace, SubTTA refines decision boundaries and achieves an average 2.24% improvement over existing TTA methods.

By Zhichen Zeng, Wenxuan Bao, Xiao Lin, Ruizhong Qiu, Tianxin Wei, Xuying Ning, Yuchen Yan, Chen Luo, Monica Xiao Cheng, Jingrui He, Hanghang Tong
arXiv AI
Aug 20

From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model

arXiv:2608. 18339v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference.

By Qi Yu, Zhichen Zeng, Katherine Tieu, Xiyuan Yang, Ruizhong Qiu, Yuchen Yan, Lihui Liu, Yanjun Zhao, Lingjie Chen, Jingrui He, Hanghang Tong
arXiv Computer Vision
3d ago

Project and Mix: Task-Semantic Prototypes for Few-Shot Image Classification

arXiv:2603.24528v2 Announce Type: replace Abstract: Vision-language models like CLIP are trained with the objective of aligning text and image pairs. Beyond text prompts alone, recent works show that...

By Dipam Goswami, Simone Magistri, Gido M. van de Ven, Bart{\l}omiej Twardowski, Andrew D. Bagdanov, Tinne Tuytelaars, Joost van de Weijer
arXiv AI
Jul 3

Neuron-Aware Active Few-Shot Learning for LLMs

arXiv:2607. 02423v1 Announce Type: cross Abstract: Active Few-Shot Learning (AFSL) adapts LLMs to specialized domains by identifying the most valuable unlabeled samples for annotation and use as few-shot demonstrations, effectively reducing human annotation costs while promoting high performance.

By Zhuowei Chen, Liwei Chen, Christian Schunn, Raquel Coelho, Xiang Lorraine Li
arXiv AI
4d ago

One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models

The paper investigates how the modality gap— the separation between image and text representations in contrastive vision‑language models—affects different downstream tasks. By showing that a single dominant direction accounts for most of the image‑text mean separation, the authors explain why reducing or removing this gap can improve zero‑shot classification, degrade retrieval, or restore performance depending on the task. The study provides a geometric framework that clarifies when and why gap interventions should be applied in vision‑language systems.

By Aditya Sharma, Divya Saxena