arXiv AI By Shao-Jun Xia, Huixin Zhang, Zhengzhong Tu

T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs

Read the original on arXiv AI →

arXiv:2511. 16107v3 Announce Type: replace-cross Abstract: Visual in-context learning (VICL) solves visual tasks by conditioning on a few input-output demonstrations without any model training.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 7

Cross-Task Generalization Between Understanding and Generation in Unified Vision-Language Models: A Controlled Study

The study investigates how unified vision‑language models (VLMs) can simultaneously support visual understanding and generation. Using controlled benchmarks (SmartWatch and modified CelebA) that pair VQA, captioning, and text‑to‑image tasks, the authors evaluate several LLM‑based architectures built on SigLIP and VQ‑VAE visual spaces. Results show that mixed training can improve both understanding and generation, but the gains depend on how well the visual input and output spaces are aligned; misaligned or distorted visual spaces can weaken or reverse these benefits. The paper also demonstrates that balancing data across tasks and controlling attribute frequencies can help recover underrepresented visual concepts, and that the transfer is driven more by the base language model’s learned relationships than by visual adapters.

By Jihai Zhang, Tianle Li, Linjie Li, Zhengyuan Yang, Yu Cheng
arXiv AI
Jun 16

Retrieve, Don't Retrain: Extending Vision Language Action Models to New Tasks at Test Time

arXiv:2606. 15631v1 Announce Type: cross Abstract: Extending a vision-language-action (VLA) policy to a new task typically requires task-specific teleoperated demonstrations and per-task fine-tuning, making adaptation costly in both data collection and compute.

By Jeongeun Park, Juhan Park, Taekyung Kim, Sungjoon Choi, Dongyoon Han, Sangdoo Yun
arXiv Machine Learning
Sep 15

TwinICL: Diagnosing Multimodal In-Context Learning through Paired Counterfactuals

TwinICL is a procedurally generated benchmark that pairs matched text and image versions of tasks to enable controlled comparison of in‑context learning (ICL) across modalities. Experiments on six open‑weight models and 38 tasks show that multimodal ICL consistently underperforms text‑only ICL, with varying gaps by task family. Interventions targeting visual access, task framing, and reasoning can recover strong multimodal performance on a diagnostic subset, yet a modality gap remains even when explicit task instructions are provided, highlighting the dual role of demonstrations as context and evidence.

By Zihan Xue, Po-Yi Lu, Serhii Honcharenko, Zih-Ching Chen, Hsuan-Tien Lin, Nanyun Peng, I-Hung Hsu, Kuan-Hao Huang
arXiv AI
Jun 16

When RAG Hurts: Diagnosing and Mitigating Attention Distraction in Retrieval-Augmented LVLMs

arXiv:2602. 00344v2 Announce Type: replace-cross Abstract: While Retrieval-Augmented Generation (RAG) is one of the dominant paradigms for enhancing Large Vision-Language Models (LVLMs) on knowledge-based VQA tasks, recent work attributes RAG failures to insufficient attention towards the retrieved context, proposing to reduce the attention allocated to image tokens.

By Beidi Zhao, Wenlong Deng, Xinting Liao, Yushu Li, Nazim Shaikh, Yao Nie, Xiaoxiao Li
arXiv Machine Learning
Jun 4

Hyper-ICL: Attention Calibration with Hyperbolic Anchor Distillation for Multimodal In-Context Learning

arXiv:2606. 04434v1 Announce Type: cross Abstract: Multimodal In-Context Learning (ICL) has emerged as a practical inference paradigm for Multimodal Large Language Models, where a small set of interleaved image-text In-Context Demonstrations (ICDs) conditions the model to solve new tasks.

By Niloufar Alipour Talemi, Hossein Kashiani, Fatemeh Afghah