arXiv:2510. 02528v2 Announce Type: replace Abstract: Large Multimodal Models (LMMs) demonstrate impressive in-context learning abilities from few multimodal demonstrations, yet the internal mechanisms supporting such task learning remain opaque.
By Shuhao Fu, Esther Goldberg, Ying Nian Wu, Hongjing Lu
arXiv:2602. 23197v2 Announce Type: replace-cross Abstract: Transformer-based large language models exhibit in-context learning, enabling adaptation to downstream tasks via few-shot prompting with demonstrations.
By Chungpa Lee, Jy-yong Sohn, Kangwook Lee
arXiv:2610.01054v1 Announce Type: cross
Abstract: In-context learning (ICL) enables language models to perform new tasks from demonstrations without weight updates. However, every ICL inference requi...
By Guangzhi Xiong, Zhenghao He, Bohan Liu, Sanchit Sinha, Wenqian Ye, Aidong Zhang
arXiv:2606. 29407v1 Announce Type: cross Abstract: There has been increasing interest in exploring the capabilities of advanced large language models (LLMs) in the field of information extraction (IE), specifically focusing on tasks related to named entity recognition (NER) and relation extraction (RE).
By Xiao You, Tianwei Yan, Shan Zhao
arXiv:2606. 11745v1 Announce Type: cross Abstract: Visual causal reasoning is essential for understanding and intervening in the physical world, requiring identification of causal variables from visual inputs and reasoning over intervention effects.
By Haoping Yu, Yuanxi Li, Jing Ma
The paper investigates how large language models learn new tasks in-context, comparing rule-based instruction following to example-based few-shot prompting across five diverse tasks. Results show that models generally learn more reliably from rule descriptions than from examples alone, and adding more examples does not consistently improve performance. Instruction tuning further enhances rule-based learning while preserving example-based capabilities, with rule advantages being strongest for algebraic tasks and weaker for tasks requiring distributional sensitivity or parametric knowledge.
By Xiang Fu, Seungmin Cho, Yukyung Lee, Najoung Kim
The paper investigates how transformer language models perform few‑shot learning for a simple addition task, showing that the ability is concentrated in a handful of attention heads. Using dimensionality reduction, the authors identify low‑dimensional subspaces—three heads with six‑dimensional spaces in Llama‑3‑8B‑Instruct—where specific dimensions encode the units digit via trigonometric patterns and magnitude via low‑frequency components. They also derive a mathematical identity linking aggregator and extractor subspaces, enabling tracking of information flow from examples to the final prediction.
By Xinyan Hu, Kayo Yin, Michael I. Jordan, Jacob Steinhardt, Lijie Chen
arXiv:2606. 01810v1 Announce Type: new Abstract: Current benchmarks for embodied vision-language planning often favor linguistic next-token prediction over physically grounded next-state reasoning.
By Zheng Lu, Mingqi Gao, Qinlei Xie, Wanqi Zhong, Hanwen Cui, Heng Cao, Zirui Song, Yifan Yang, Chong Luo, Bei Liu, Yiming Li
The paper introduces Causal Shortcut Learning (CSL), a framework that identifies token chains—called causal shortcuts—that guide Diffusion Language Models (DLMs) toward correct reasoning paths. By extracting these shortcuts and applying parallel prioritized masking during training, CSL improves both convergence speed and generation accuracy. Experiments on several reasoning benchmarks and two base models show CSL outperforms existing SFT-variant baselines, achieving an average 1.92% improvement over SFT-only models and up to 4.20% on MATH-500.
By Dian Jin, Kairong Han, Baohong Li, Xinpeng Dong, Zijing Hu, Nuanqiao Shan, Fei Wu, Kun Kuang
Compositional Zero-Shot Learning (CZSL) aims to combine known attributes and objects as primitives for recognizing previously unseen attribute-object pairs. Prior works either predict attributes and objects independently, missing their strong contextual dependency, or use unidirectional conditional modeling (e.
arXiv:2608. 02830v1 Announce Type: cross Abstract: Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and is widely assumed to improve as more demonstrations are supplied.
By Mohammad Rostami
arXiv:2605. 13511v3 Announce Type: replace-cross Abstract: While many-shot ICL achieves remarkable performance, prior studies of its scaling behavior have mainly focused on non-reasoning tasks.
By Tsz Ting Chung, Lemao Liu, Mo Yu, Dit-Yan Yeung