The paper introduces Lens, a training‑free framework that aligns multimodal representations with the semantic perspective required by downstream tasks. Lens uses a task‑specific readout phrase to anchor the perspective and then aggregates token states after the full input, ensuring the extracted representation reflects task‑conditioned evidence integration rather than generic salient content. The method achieves a Precision@1 of 63.9 across 36 MMEB datasets, outperforming the nearest training‑free baseline by 10.2 points.
By Xinran Liu, Shouqian Shi, Yixian Chen, Ruizhi Chen, Xin-Wei Yao, Sheng Zhong
JEPAMatch introduces a new semi‑supervised learning framework that replaces traditional output‑thresholding with explicit geometric shaping of latent representations. By combining the FlexMatch loss with a latent‑space regularization inspired by LeJEPA, the method encourages isotropic Gaussian structure in the embedding space, mitigating class imbalance and noisy pseudo‑labels. Experiments on CIFAR‑100, STL‑10, and Tiny‑ImageNet show consistent performance gains and faster convergence compared to existing FixMatch‑based baselines.
By Ali Aghababaei-Harandi, Aude Sportisse, Massih-Reza Amini
arXiv:2603. 01471v2 Announce Type: replace-cross Abstract: Multimodal embedding models, rooted in multimodal large language models (MLLMs), have yielded significant performance improvements across diverse tasks such as retrieval and classification.
By Jiahan Chen, Da Li, Hengran Zhang, Yinqiong Cai, Lixin Su, Jiafeng Guo, Daiting Shi, Dawei Yin, Keping Bi
arXiv:2512. 00088v2 Announce Type: replace-cross Abstract: We propose SemImage, a novel method for representing a text document as a two-dimensional semantic image to be processed by convolutional neural networks (CNNs).
By Mohammad Zare
arXiv:2512. 10092v2 Announce Type: replace Abstract: Analyzing large-scale text corpora is a core challenge in machine learning, crucial for tasks like identifying undesirable model behaviors or biases in training data.
By Nick Jiang, Xiaoqing Sun, Lisa Dunlap, Lewis Smith, Neel Nanda
arXiv:2603. 01471v3 Announce Type: replace-cross Abstract: Multimodal embedding models, rooted in multimodal large language models (MLLMs), have yielded significant performance improvements across diverse tasks such as retrieval and classification.
By Jiahan Chen, Da Li, Hengran Zhang, Yinqiong Cai, Lixin Su, Jiafeng Guo, Daiting Shi, Dawei Yin, Keping Bi
arXiv:2605. 16739v2 Announce Type: replace-cross Abstract: Decoding visual experience from brain activity has advanced substantially, but current brain-to-text systems largely recover semantic content while discarding affect.
By Bilal A. Mohammed, Lin Gu, Ruogu Fang
arXiv:2511. 16527v2 Announce Type: replace-cross Abstract: Contrastive vision-language models continue to be the dominant approach for image-text retrieval.
By Kwun Ho Ngan, Saman Sadeghi Afgeh, Joe Townsend, Artur d'Avila Garcez
arXiv:2606. 10789v1 Announce Type: new Abstract: Zero-shot learning (ZSL) for inertial measurement unit (IMU)-based human activity recognition (HAR) faces a central challenge: bridging the gap between sensor embeddings and semantic class representations.
By Anik Ghosh
arXiv:2609.15152v1 Announce Type: cross
Abstract: Multimodal embedding models encode heterogeneous inputs into a shared embedding space, enabling efficient similarity computation across modalities an...
By Yanping Li, Wei Zhou, Yawen Liu, Yibo Wang, Ke Zhu, Guangda Huzhang, Qing-Guo Chen, Zhao Xu, Jun Zhang, Wei Wei
arXiv:2605.28190v2 Announce Type: replace
Abstract: Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property. We argue that embeddin...
By Manuel Frank, Haithem Afli
The paper introduces DAN, a training‑free inference‑time framework that improves affective reasoning in multimodal large language models. It combines a Hierarchical Emotional Reasoning Chain (HERC) to better capture fine‑grained visual cues and a Contrastive Discriminative Visual Pruning (CDVP) module to isolate discriminative tokens for semantically similar emotions. Experiments show significant gains, notably a +10.47% improvement on the WebEmo25 benchmark with Qwen3‑VL‑8B‑Instruct.
By Cheng Ye, Weidong Chen, Zhaobo Qi, Beier Zhu, Zhendong Mao