arXiv:2609.23717v1 Announce Type: new
Abstract: Global vision--language similarities compress an image and a caption into one vector, preserving semantics but not which word corresponds to which regi...
By Liuyang Song, Yi Zhang, Zhongyi Deng, Daqian Yang, Hongbo Zhang
arXiv:2606. 03879v1 Announce Type: cross Abstract: As foundation models scale toward fusing more heterogeneous visual streams, understanding how diverse encoders interact under joint training becomes a prerequisite for principled design.
By Wei Ding, Yudong Zhang, Ruobing Xie, Xingwu Sun, Jiansheng Chen, Yu Wang
The paper argues that meaning identity—whether two sentences convey the same idea after wording changes—is not encoded in the geometry of independently produced sentence embeddings. Experiments on frozen off‑the‑shelf encoders and language models show that identity can only be reliably computed when both sentences are processed together in a single forward pass, yielding high accuracy (0.90–0.96) on PAWS‑X, whereas independent embeddings or simple fusion methods perform near chance. Even advanced bi‑encoder fine‑tuning improves performance on PAWS but fails to generalize to other similarity tasks, underscoring that identity is a cheap computed operator rather than a property of individual sentence vectors.
By Jiaqi Deng
arXiv:2608.30725v1 Announce Type: new
Abstract: Recent multilingual vision--language encoders cover hundreds of languages in a single model, yet on two state-of-the-art instances retrieval on low-res...
By Donghoon Han, SungHyun Moon, Aidyn Zhakatayev, Junghun Cha, SeungJae Lee
arXiv:2605. 25889v4 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models reach high success rates on clean inputs but collapse under small adversarial perturbations: a $16/255$ PGD attack drops OpenVLA-7B's LIBERO success from $95\%$ to under $5\%$.
By Jianwei Tai
Global vision--language similarities compress an image and a caption into one vector, preserving semantics but not which word corresponds to which region or how those regions are arranged; a model can...
arXiv:2607. 22771v1 Announce Type: cross Abstract: Picking the frozen image encoder for a 3D~CT vision--language model (VLM), together with the token-compression scheme on top of it, is a search over many candidates.
By Renjie Liang
arXiv:2607. 04926v1 Announce Type: cross Abstract: How does the way information reaches a transformer -- as symbolic tokens, a clean per-factor "oracle" code, or an entangled perceptual vector -- shape whether it binds that information compositionally?
By Yoshiyuki Ootani
arXiv:2608. 11661v1 Announce Type: cross Abstract: A multiplicative dual-encoder network computes a real-valued output for a pair of inputs as the inner product of their separate encodings.
By Zijian Zhao, Sen Li
The paper introduces Auto-Comp, a fully automated, concept-driven pipeline that generates photorealistic compositional benchmarks for vision‑language models. Auto‑Comp creates paired Minimal and Contextual samples for each concept, enabling isolation of core binding abilities from visio‑linguistic complexity. Evaluations across 25 models reveal consistent failures in attribute and relational binding, with context helping relational tasks but hindering attribute tasks due to visual clutter.
By Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci
arXiv:2608.29924v1 Announce Type: cross
Abstract: Large Vision-Language Models (LVLMs) are prone to hallucinations: they fluently describe objects, attributes, and scenes that are not in the image. W...
By Aditi Sarker, Rafi Ibn Sultan, Hui Zhu, Dongxiao Zhu, Prashant Khanduri
arXiv:2607. 23050v1 Announce Type: new Abstract: Neural scaling laws describe how loss decreases as models, data, and compute grow, but they do not answer a prior question: for a fixed task, what is the minimum model capacity required to solve it?
By Byeong Hoon Yoon