arXiv Machine Learning By Changhao Xiang, Shangyu Xing, Zhen Wu, Jianbing Zhang, Xinyu Dai

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment

Read the original on arXiv Machine Learning →

arXiv:2608. 11167v1 Announce Type: cross Abstract: Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
Sep 22

MinCU: A Fine-Grained Benchmark for Grounded Minimal-Change Understanding in Image Pairs

MinCU is a new benchmark for grounded minimal‑change understanding that presents pairs of near‑identical images differing by a single atomic variation in object category, attribute, count, or spatial position. Models are evaluated on their ability to describe the change, localize the changed region, and identify the changed entity. The authors also introduce SG‑ISA, a structured autoregressive method that decomposes the task into a Think‑Locate‑Describe sequence, showing that fine‑tuning with SG‑ISA improves both grounding accuracy and description quality while reducing reasoning‑token overhead.

By Chaoqian Mu, Wenhao Wu, Zichen Liang, Jiaxu Li, Lijun Wang, Yifan Wang, Huchuan Lu
arXiv AI
Jun 8

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models

arXiv:2602. 07026v3 Announce Type: replace-cross Abstract: Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the Modality Gap, remains: embeddings of distinct modalities expressing identical semantics occupy systematically offset regions.

By Xiaomin Yu, Yi Xin, Yuhui Zhang, Wenjie Zhang, Chonghan Liu, Hanzhen Zhao, Chen Liu, Xiaoxing Hu, Ziyue Qiao, Hao Tang, Xiaobin Hu, Chengwei Qin, Hui Xiong, Yu Qiao, Shuicheng Yan