arXiv AI By Chengzhen Yu, Canran Xiao, Siyuan Ma, Yang Liu

Text as Partial Constraint: Core-Residual Alignment for Robust Vision-Language Learning

Read the original on arXiv AI →

arXiv:2607. 03143v1 Announce Type: cross Abstract: Vision-language alignment powers open-vocabulary recognition, retrieval, and LVLM grounding, yet natural captions are often underspecified, making similarity brittle and overly confident under paraphrase and omitted details.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.