arXiv Machine Learning By Erfan Loweimi, Mengjie Qian, Kate Knill, Guanfeng Wu, Chi-Ho Chan, Abbas Haider, Muhammad Awan, Josef Kittler, Hui Wang, Mark Gales

To Be Multimodal or Not to Be: Query-Adaptive Audio-Visual Person Retrieval via Active Modality Detection

Read the original on arXiv Machine Learning →

arXiv:2606. 05931v1 Announce Type: cross Abstract: When retrieving a person from a video archive by voice and face, should the system be multimodal or not?

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
Sep 15

Query-Conditioned Spherical Centroid Aggregation for Multimodal Retrieval

The paper introduces SCALAR, a query‑conditioned spherical centroid aggregator that assigns relevance‑based weights to each available modality before computing a spherical centroid. SCALAR supports arbitrary modality subsets, is trained with rank‑8 LoRA adapters on masked views, and achieves positive aggregation gains on four of five benchmarks, outperforming prior symmetric aggregators. With only 4.8 million trainable parameters, SCALAR attains the highest text‑to‑video R@1 on three benchmarks and surpasses the released GRAM checkpoint under test‑time modality dropout by 3.2 to 10.9 R@1.

By Ambuj Mehrish, Anindya Nag, Sebastiano Vascon