From Alignment to Fusion in 3D Vision-Language
arXiv:2609.28222v1 Announce Type: new Abstract: Unified 3D vision-language systems must combine complementary geometry, scale, and appearance cues while supporting tasks from instance segmentation to...
arXiv:2609.28222v1 Announce Type: new Abstract: Unified 3D vision-language systems must combine complementary geometry, scale, and appearance cues while supporting tasks from instance segmentation to...
Vision Foundation Models (VFMs) have significantly advanced dense feature matching, yet severe in-plane rotation remains a critical challenge. Existing solutions face a fundamental dilemma: data-driven methods require inefficient parameter scaling to implicitly learn rotations, whereas strictly equivariant networks lack the semantic capacity of modern VFMs.
arXiv:2609.16664v1 Announce Type: new Abstract: Leveraging Large Vision-Language Models like CLIP has recently set new benchmarks for No-Reference Image Quality Assessment (NR-IQA). However, the cont...
arXiv:2609.23717v1 Announce Type: new Abstract: Global vision--language similarities compress an image and a caption into one vector, preserving semantics but not which word corresponds to which regi...
The paper demonstrates that sharing a deep encoder alone does not eliminate the confounding effects in task-comparison scores. By introducing a conditional two‑discriminator discrepancy within the embedding space, the authors achieve robust detection of task changes, maintaining stability under input rotations and accurately tracking label‑permutation drift. This approach, integrated into a mixture‑of‑heads framework, outperforms traditional novelty triggers and generalizes across multiple backbones and datasets, including ImageNet‑21k ViT‑B/16, DINOv2, and CIFAR‑100.
Leveraging Large Vision-Language Models like CLIP has recently set new benchmarks for No-Reference Image Quality Assessment (NR-IQA). However, the contrastive pretraining of CLIP inherently prioritize...
Multimodal models often build on architectures designed for generative vision-language modeling, typically combining separately pretrained vision encoders with causal language models. Visual document...
arXiv:2511. 01390v2 Announce Type: replace-cross Abstract: Fine-grained cross-modal alignment aims to establish precise local correspondences between vision and language, forming a cornerstone for visual question answering and related multimodal applications.
arXiv:2608. 08676v1 Announce Type: cross Abstract: Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation.
The paper investigates drift detection in deep learning models, showing that sharing a deep encoder alone does not eliminate confounding in task-comparison scores. By introducing a conditional two‑discriminator discrepancy into the embedding space, the authors create a two‑axis gate that remains stable under input rotations and accurately tracks label‑permutation drift. This approach outperforms traditional exchange or novelty triggers, achieving high AUROC in distinguishing semantic novelty from photometric shift across multiple backbones and datasets.
Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features can discard text, local attributes, and spatial relationships, while high-resolution inputs substantially increase context length and inference latency.
arXiv:2507.00754v3 Announce Type: replace Abstract: The integration of Large Language Model (LLMs) blocks with Vision Transformers (ViTs) holds immense promise for vision-only tasks by leveraging the...