Hugging Face Trending Papers

Fine-grained CLIP fine-tuning with self-annotated region alignment

Contrastive Language-Image Pre-training (CLIP) has been shown to have limitations in its fine-grained dense feature representation, due to its pre-training focusing on matching the whole image to a text description. Considering the large data and computational burden in pre-training a vision-language model from scratch, a series of works aim to enhance the fine-grained ability of CLIP through a fine-tuning scheme.

arXiv Computer Vision
Aug 21

ID-VTG: Image-Disambiguated Video Temporal Grounding

arXiv:2608. 20127v1 Announce Type: new Abstract: Video Temporal Grounding (VTG) faces significant challenges when natural language queries must distinguish between multiple events involving visually similar entities, particularly when relying on fine-grained visual attributes that are difficult to describe accurately in words alone.

By Minghang Zheng, Jingli Wei, Hongyi Yang, Yang Liu
arXiv Computation and Language
Sep 17

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

PANORAMA introduces a new panoptic grounded captioning framework that jointly generates detailed image captions and associates each phrase with precise pixel-level masks. The authors create PanoCaps, a human‑annotated benchmark with dense captions and near‑complete pixel coverage, and propose a phrase‑mask matching protocol with a generalized Panoptic Quality metric. PANORAMA conditions a pretrained segmenter on contextualized phrase representations, learns to select appropriate masks, and achieves state‑of‑the‑art grounding performance on PanoCaps and other pixel‑level tasks.

By Sara Pieri, Evangelos Kazakos, Shizhe Chen, Josef Sivic, Cordelia Schmid