arXiv Machine Learning

Activation-Based Active Learning for In-Context Learning: Challenges and Insights

arXiv:2606. 05134v1 Announce Type: cross Abstract: Deep active learning has previously been explored for LLM in-context sample selection, but not with methods that utilise recent advances in understanding of transformer activations.

arXiv Machine Learning
Aug 31

Generalized Context in Cross Attention for Transfer Learning of Disjoint Tabular Data

The paper introduces CATTLE, a transfer learning framework for disjoint tabular datasets that eliminates the need for shared features by leveraging generalized context learned through transformer projection weights. By using key, value, and query weights from source and target domains, CATTLE performs cross‑domain attention transfer in a data‑agnostic manner. Experiments on ten source‑target pairs demonstrate that CATTLE outperforms nine state‑of‑the‑art baselines, achieving the best average rank (2.9) and a 3.7% AUROC improvement.

By Kazi F. Akhter, Ibna Kowsar, Manar D. Samad
arXiv AI
Jul 28

Extracting Algorithms in Pre-trained LLMs: A Case on Hidden Markov Models

arXiv:2607. 22646v1 Announce Type: new Abstract: Large language models (LLMs) display a striking ability to predict next observations from Hidden Markov Models (HMMs) via in-context learning (ICL), but the algorithm underlying this capability remains undetermined: prior work has proposed several candidates without consensus, and none has been grounded in the model's internal activations.

By Yijia Dai, Zhaolin Gao, Yahya Sattar, Jennifer J. Sun, Sarah Dean
arXiv AI
Sep 3

Language Models Can Control Their Own Attention

The paper introduces Declarative Attention (DA), a protocol that lets language models explicitly declare which parts of their context to focus on during generation. By partitioning decoding into full-context, region-specific, and recent-output-only modes, the inference engine can skip large portions of the KV cache, dramatically reducing attended tokens. Experiments on 15 long-context tasks with off-the-shelf models show significant savings (52.0% and 31.1% reductions) with only modest accuracy drops that diminish as model size increases.

By Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos
arXiv AI
Jun 9

ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning

arXiv:2505. 21457v2 Announce Type: replace-cross Abstract: Active vision, also known as active perception, refers to actively selecting where and how to look in order to gather task-relevant information.

By Muzhi Zhu, Hao Zhong, Canyu Zhao, Zongze Du, Mingyu Liu, Zheng Huang, Anzhou Li, Hao Chen, Cheng Zou, Jingdong Chen, Ming Yang, Chunhua Shen