arXiv:2602. 23197v2 Announce Type: replace-cross Abstract: Transformer-based large language models exhibit in-context learning, enabling adaptation to downstream tasks via few-shot prompting with demonstrations.
By Chungpa Lee, Jy-yong Sohn, Kangwook Lee
arXiv:2610.01054v1 Announce Type: cross
Abstract: In-context learning (ICL) enables language models to perform new tasks from demonstrations without weight updates. However, every ICL inference requi...
By Guangzhi Xiong, Zhenghao He, Bohan Liu, Sanchit Sinha, Wenqian Ye, Aidong Zhang
The paper introduces Declarative Attention (DA), a protocol that lets language models explicitly declare which parts of their context to focus on during generation. By partitioning decoding into full-context, region-specific, and recent-output-only modes, the inference engine can skip large portions of the KV cache, dramatically reducing attended tokens. Experiments on 15 long-context tasks with off-the-shelf models show significant savings (52.0% and 31.1% reductions) with only modest accuracy drops that diminish as model size increases.
By Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos
arXiv:2609.15990v1 Announce Type: cross
Abstract: Few-shot prompting sometimes degrades language models instead of helping them, but why this happens is unknown. We evaluate 12 open-weight models on...
By Volodymyr Ovcharov
The paper investigates how large language models use activation memory (KV caches) and parametric memory (updated parameters) during few‑shot learning. Experiments show activation memory excels at factual recall, while parametric memory does not consistently outperform it for task learning. The composite task Conditional Arithmetic requires both memory types, with neuron‑level analysis revealing distinct neuron sets activated by each memory, and their combined use is essential for success.
By Miaohe Niu, Runsong Zhao, Xinyu Liu, Bo Jin, Yucheng Qiao, Chunliang Zhang, Jingbo Zhu, Tong Xiao
arXiv:2511. 21338v2 Announce Type: replace Abstract: Masked Diffusion Language Models (MDLMs) have recently emerged as a promising alternative to Autoregressive Language Models (ARLMs), leveraging a denoising objective that, in principle, should enable more uniform context utilisation.
By Julianna Piskorz, Cristina Pinneri, Alvaro Correia, Motasem Alfarra, Risheek Garrepalli, Christos Louizos
The paper introduces PromptCCZSL, a framework that enables vision‑language models to continually learn new attributes, objects, and their unique compositions while avoiding forgetting. It uses a frozen VLM backbone with prompt‑based techniques, recency‑weighted multi‑teacher distillation, and several loss functions (CAL, OPL, IDL) to maintain prior knowledge and promote diverse, distinct embeddings. Experiments on UT‑Zappos and C‑GQA show significant performance gains over existing VLM‑based and non‑VLM baselines, establishing a new benchmark for continual compositional zero‑shot learning.
By Sauda Maryam, Sara Nadeem, Faisal Qureshi, Mohsen Ali
arXiv:2610.00526v1 Announce Type: cross
Abstract: In-context learning (ICL) can be amortized into latent objects (task vectors, function vectors, context vectors) that recover few-shot behavior at ze...
By Gunmay Jhingran
The paper investigates how transformer language models perform few‑shot learning for a simple addition task, showing that the ability is concentrated in a handful of attention heads. Using dimensionality reduction, the authors identify low‑dimensional subspaces—three heads with six‑dimensional spaces in Llama‑3‑8B‑Instruct—where specific dimensions encode the units digit via trigonometric patterns and magnitude via low‑frequency components. They also derive a mathematical identity linking aggregator and extractor subspaces, enabling tracking of information flow from examples to the final prediction.
By Xinyan Hu, Kayo Yin, Michael I. Jordan, Jacob Steinhardt, Lijie Chen
arXiv:2608.21487v1 Announce Type: cross
Abstract: Vision-Language Models (VLMs) exhibit strong zero-shot capabilities, making them an attractive solution for continual learning across diverse tasks....
By Chang Sun, Francesco Barbato, Matteo Caligiuri, Pietro Zanuttigh
arXiv:2605. 28854v2 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit remarkable flexibility in adapting to novel tasks from in-context examples without parameter updates, a capability known as in-context learning (ICL).
By Hua-Dong Xiong, Li Ji-An, Robert C. Wilson, Kwonjoon Lee, Xue-Xin Wei
The paper investigates how the modality gap— the separation between image and text representations in contrastive vision‑language models—affects different downstream tasks. By showing that a single dominant direction accounts for most of the image‑text mean separation, the authors explain why reducing or removing this gap can improve zero‑shot classification, degrade retrieval, or restore performance depending on the task. The study provides a geometric framework that clarifies when and why gap interventions should be applied in vision‑language systems.
By Aditya Sharma, Divya Saxena