SetFit: Efficient Few-Shot Learning Without Prompts
Related stories
Language models are few-shot learners
Are We Really Doing Few-Shot Learning? A Critical Examination of Pre-Training Assumptions
The paper critically evaluates common fewâshot learning protocols that rely on preâtraining a model on a large auxiliary set with classes disjoint from the target but drawn from the same visual domain. By comparing no preâtraining, classâdisjoint inâdomain preâtraining, supervised outâofâdomain preâtraining, and labelâfree outâofâdomain preâtraining across eight datasets and three architectures, the authors find that inâdomain preâtraining yields a 33.41âpoint average improvement, while outâofâdomain preâtraining offers a 23.75âpoint gain, revealing a 9.66âpoint optimistic bias due to domain overlap. They also demonstrate that a labelâfree augmentation strategy can match supervised outâofâdomain performance and propose a descriptorâbased sourceâselection method that closely approximates oracle selection, underscoring the need to move beyond inâdomain preâtraining as the default evaluation protocol.
Training-free LLM Verification via Recycling Few-shot Examples
arXiv:2506.17251v3 Announce Type: replace-cross Abstract: Although large language models (LLMs) have achieved remarkable performance, the inherent stochasticity of their reasoning processes and varyi...
Re-Evaluating Continual Learning with Few-Shot Adaptation
arXiv:2606. 03843v1 Announce Type: cross Abstract: Continual learning methods aim to maximize the stability and plasticity of machine learning models that are trained on a sequence of tasks.
One-shot imitation learning
Guided Prompt Evolution for Vision-Language Models Adaptation
The paper introduces EvoPrompt, a framework for adapting visionâlanguage models to new tasks with limited data while preventing catastrophic forgetting. EvoPrompt uses a ModalityâShared Prompt Projector to create hierarchical prompts and an evolutionary training strategy that separates lowârank updates into directional and magnitude components, preserving learned semantic directions. Experiments show that EvoPrompt achieves stateâofâtheâart fewâshot performance while maintaining the original zeroâshot capabilities of the preâtrained models.
From Prompts to Trees: Effective LLM-Guided Tree Generation for Few-Shot Tabular Classification
While Large Language Models (LLMs) possess rich world knowledge and impressive generalization capabilities, their direct application to tabular data classification is hindered by high inference costs...
ARASH: Adaptive Retrieval And Shot Selection for Tabular Prediction
ARASH is a method that improves the efficiency of Tabular Foundation Models by selecting optimal few-shot prompts based on local neighborhood analysis within the training set. It reduces the prompt length and memory usage of TabPFN by 1261.5Ă and 2.56Ă, respectively, while maintaining comparable accuracy. This approach addresses the challenge of identifying relevant rows for in-context learning in tabular data.
Few-Medoids: An Embarrassingly Simple Coreset Selection Method for Few-Shot Knowledge Distillation
arXiv:2607. 05891v1 Announce Type: cross Abstract: Coreset selection aims to identify a small and highly representative subset of a massive dataset for efficient model training.
Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
arXiv:2602. 23197v2 Announce Type: replace-cross Abstract: Transformer-based large language models exhibit in-context learning, enabling adaptation to downstream tasks via few-shot prompting with demonstrations.
How Much Prompt Is Enough? A Blackbox Minimization of Few-Shots in LLMs
The paper introduces ramework, a blackbox promptâminimization framework that identifies the minimal subset of fewâshot prompts necessary for large language models (LLMs). In a case study, the framework reduces fewâshot exemplars by an average of 65.3% in character count while maintaining full propositional output fidelity, revealing that models tend to keep logical identifiers and constraint declarations while discarding natural language prose. The analysis further distinguishes between universal encoder and decoder models, offering insights into prompt compression and structural analysis.