arXiv AI

Neuron-Aware Active Few-Shot Learning for LLMs

arXiv:2607. 02423v1 Announce Type: cross Abstract: Active Few-Shot Learning (AFSL) adapts LLMs to specialized domains by identifying the most valuable unlabeled samples for annotation and use as few-shot demonstrations, effectively reducing human annotation costs while promoting high performance.

arXiv Machine Learning
Sep 11

Are We Really Doing Few-Shot Learning? A Critical Examination of Pre-Training Assumptions

The paper critically evaluates common few‑shot learning protocols that rely on pre‑training a model on a large auxiliary set with classes disjoint from the target but drawn from the same visual domain. By comparing no pre‑training, class‑disjoint in‑domain pre‑training, supervised out‑of‑domain pre‑training, and label‑free out‑of‑domain pre‑training across eight datasets and three architectures, the authors find that in‑domain pre‑training yields a 33.41‑point average improvement, while out‑of‑domain pre‑training offers a 23.75‑point gain, revealing a 9.66‑point optimistic bias due to domain overlap. They also demonstrate that a label‑free augmentation strategy can match supervised out‑of‑domain performance and propose a descriptor‑based source‑selection method that closely approximates oracle selection, underscoring the need to move beyond in‑domain pre‑training as the default evaluation protocol.

By Alejandro Galan-Cuenca, Marcelo Saval-Calvo, Antonio Javier Gallego
Hugging Face Trending Papers
Jun 9

Closing the Modality Gap in Zero-Shot HAR: Contrastive Training and Separability-Optimized Prototypes on IMU Data

Zero-shot learning (ZSL) for inertial measurement unit (IMU)-based human activity recognition (HAR) faces a central challenge: bridging the gap between sensor embeddings and semantic class representations. We systematically evaluate seven configurations combining three inference methods with two training pipelines on the PAMAP2 dataset, using 14 seen and 4 unseen activity classes with subjects 108 and 109 held out for testing.

arXiv AI
Aug 5

From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model

arXiv:2508. 00955v3 Announce Type: replace-cross Abstract: Adapting generative Multimodal Large Language Models (MLLMs) into universal embedding models typically demands resource-intensive contrastive pre-training, while traditional hard negative mining methods suffer from severe false negative contamination.

By Yeong-Joon Ju, Seong-Whan Lee
arXiv Computation and Language
Sep 24

Complementary Roles of Activation and Parametric Memory in Few-Shot Learning

The paper investigates how large language models use activation memory (KV caches) and parametric memory (updated parameters) during few‑shot learning. Experiments show activation memory excels at factual recall, while parametric memory does not consistently outperform it for task learning. The composite task Conditional Arithmetic requires both memory types, with neuron‑level analysis revealing distinct neuron sets activated by each memory, and their combined use is essential for success.

By Miaohe Niu, Runsong Zhao, Xinyu Liu, Bo Jin, Yucheng Qiao, Chunliang Zhang, Jingbo Zhu, Tong Xiao
arXiv AI
Sep 15

Convergent Emergence of In-Context Learning Across Modalities

The paper investigates whether few-shot in-context learning (ICL) emerges similarly across different data modalities. Using a controlled cross-modality framework, the authors test the Convergent Emergence Hypothesis, which posits that tasks benefiting from ICL in one modality will also benefit in others. They find that paired-mapping ICL appears in six modalities—language, genome, integer sequences, time series, images, and proteins—outperforming baselines and showing correlated task effects in five of them, supporting the hypothesis in some but not all cases.

By Nathan Breslow, Seungwook Han, Daniel Hyunsoo Lee, Aayush Mishra, Anqi Liu, Daniel Khashabi