arXiv Machine Learning By Niloufar Alipour Talemi, Hossein Kashiani, Fatemeh Afghah

Hyper-ICL: Attention Calibration with Hyperbolic Anchor Distillation for Multimodal In-Context Learning

Read the original on arXiv Machine Learning →

arXiv:2606. 04434v1 Announce Type: cross Abstract: Multimodal In-Context Learning (ICL) has emerged as a practical inference paradigm for Multimodal Large Language Models, where a small set of interleaved image-text In-Context Demonstrations (ICDs) conditions the model to solve new tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 11

Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning

The paper introduces a multimodal in‑context learning framework that uses contrastive demonstration modeling to align large language models’ responses with the required reasoning paths. By contrasting suboptimal and better responses and incorporating a response‑conditioned retrieval mechanism, the method explicitly guides models beyond surface imitation. Experiments on various multimodal tasks, especially visual question answering, show consistent performance gains.

By Mingbo Yang, Wenqiang Wang, Zhaolu Kang, Peng Chen, Yannan Chen, Sunshang Wang, Yan Xiao
arXiv Machine Learning
Sep 15

TwinICL: Diagnosing Multimodal In-Context Learning through Paired Counterfactuals

TwinICL is a procedurally generated benchmark that pairs matched text and image versions of tasks to enable controlled comparison of in‑context learning (ICL) across modalities. Experiments on six open‑weight models and 38 tasks show that multimodal ICL consistently underperforms text‑only ICL, with varying gaps by task family. Interventions targeting visual access, task framing, and reasoning can recover strong multimodal performance on a diagnostic subset, yet a modality gap remains even when explicit task instructions are provided, highlighting the dual role of demonstrations as context and evidence.

By Zihan Xue, Po-Yi Lu, Serhii Honcharenko, Zih-Ching Chen, Hsuan-Tien Lin, Nanyun Peng, I-Hung Hsu, Kuan-Hao Huang