arXiv AI

Parameter-Efficient Vision-Language Adaptation with Continuous Metadata Conditioning for Animal Re-Identification

arXiv:2607. 09443v1 Announce Type: cross Abstract: Long-term animal re-identification (ReID) must remain robust to gradual morphological evolution and seasonal appearance shifts.

arXiv Computer Vision
Sep 15

GorillaWatch: An Automated System for In-the-Wild Gorilla Re-Identification and Population Monitoring

arXiv:2512.07776v2 Announce Type: replace Abstract: Monitoring critically endangered western lowland gorillas is currently hampered by the immense manual effort required to re-identify individuals fr...

By Maximilian Schall, Felix Leonard Kn\"ofel, Noah Elias K\"onig, Jan Jonas Kubeler, Maximilian von Klinski, Joan Wilhelm Linnemann, Xiaoshi Liu, Iven Jelle Schlegelmilch, Ole Woyciniuk, Alexandra Schild, Dante Wasmuht, Magdalena Bermejo Espinet, German Illera Basas, Gerard de Melo
arXiv AI
Sep 10

What Does Animal Re-Identification Learn? Linear Biological Concepts and Their Origins in Visual Representations

The study investigates whether Vision Transformer (ViT)-based animal re-identification models learn biologically meaningful concepts. Using a DINOv3 backbone fine‑tuned on Western lowland gorilla images, the authors find that sex and age emerge as linear directions in the model’s representations, generalizing to unseen individuals with high AUROC scores. They demonstrate that the sex direction is causally used by the model, that fine‑tuning relocates these concepts within the network, and that the representations reflect a graded biological axis encoded redundantly across the population.

By Robert Nolting, Alexandra Schild, Moritz Weckbecker, Maximilian Schall, Gerard de Melo
arXiv AI
Aug 20

From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model

arXiv:2608. 18339v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference.

By Qi Yu, Zhichen Zeng, Katherine Tieu, Xiyuan Yang, Ruizhong Qiu, Yuchen Yan, Lihui Liu, Yanjun Zhao, Lingjie Chen, Jingrui He, Hanghang Tong
Hugging Face Trending Papers
Aug 5

Promptable Animal Pose Tracking Across Species

Animal pose estimation and tracking is important for wildlife monitoring and conservation research, and with limited expert time for labelling automated approaches are imperative. While human pose estimation and tracking has seen rapid progress thanks to large annotated datasets, animal pose remain challenging, due to large morphological and behavioural differences between species and limited annotated data.

arXiv Computer Vision
Sep 3

Test-Time Logit Prompting for Source-Free Missing Modality Adaptation

The paper introduces Test-Time Logit Prompting (TLP), a lightweight framework that adapts vision-language models to missing-modality inputs without accessing source training data. TLP optimizes logit prompts using uncertainty-aware adjustments and modality-complete consistency regularization, thereby maintaining prediction confidence and semantic consistency. Experiments on various benchmarks show that TLP improves recognition performance by up to 8% while requiring only a few hundred tunable parameters and minimal test-time optimization steps.

By Taixi Chen, Nancy Guo
arXiv Computer Vision
Aug 24

ES-VP : Energy-Shaped Dynamic Visual Prompting for Efficient Model Adaptation

ES‑VP introduces Energy‑Shaped Visual Prompting, a method that generates image‑specific prompts through low‑rank initialization and energy‑guided dynamic adaptation. It achieves higher performance than existing single‑prompt and diverse‑prompt approaches while using far fewer parameters. Experiments on five architectures and fifteen datasets show consistent superiority, including a 2.6% accuracy gain over DAM‑VP on CLIP with 590× fewer prompt parameters.

By Can Jin, Ying Li, Jingchen Sun, Hongwu Peng, Jiahui Zhao, Yang Zhou, Lei Li, Dimitris N. Metaxas
arXiv AI
Aug 28

Subspace Alignment for Vision-Language Model Test-time Adaptation

The paper introduces SubTTA, a test-time adaptation method for vision‑language models that aligns the semantic subspaces of visual and textual modalities to improve zero‑shot predictions. It addresses two issues: the modality gap caused by distribution shifts and visual nuisance that masks task‑specific semantics. By minimizing chordal distance between principal subspaces and projecting visual features onto a task‑specific textual subspace, SubTTA refines decision boundaries and achieves an average 2.24% improvement over existing TTA methods.

By Zhichen Zeng, Wenxuan Bao, Xiao Lin, Ruizhong Qiu, Tianxin Wei, Xuying Ning, Yuchen Yan, Chen Luo, Monica Xiao Cheng, Jingrui He, Hanghang Tong