Language models are few-shot learners
Related stories
Few-shot learning in practice: GPT-Neo and the 🤗 Accelerated Inference API
How Much Prompt Is Enough? A Blackbox Minimization of Few-Shots in LLMs
The paper introduces ramework, a blackbox prompt‑minimization framework that identifies the minimal subset of few‑shot prompts necessary for large language models (LLMs). In a case study, the framework reduces few‑shot exemplars by an average of 65.3% in character count while maintaining full propositional output fidelity, revealing that models tend to keep logical identifiers and constraint declarations while discarding natural language prose. The analysis further distinguishes between universal encoder and decoder models, offering insights into prompt compression and structural analysis.
FSA-GRPO: Teaching Auditory LLMs to Use Few-Shot Demonstrations
arXiv:2606.02615v2 Announce Type: replace-cross Abstract: Few-shot prompting provides an effective way to adapt auditory large language models to low-resource tasks such as children's speech recognit...
A Dive into Vision-Language Models
FSA-GRPO: Teaching Auditory LLMs to Use Few-shot Demonstrations
arXiv:2606. 02615v1 Announce Type: cross Abstract: Few-shot prompting provides an effective way to adapt auditory large language models to low-resource tasks such as children's speech recognition.
Training-free LLM Verification via Recycling Few-shot Examples
arXiv:2506.17251v3 Announce Type: replace-cross Abstract: Although large language models (LLMs) have achieved remarkable performance, the inherent stochasticity of their reasoning processes and varyi...
Efficient training of language models to fill in the middle
In-Context Learning Amplifies a Latent Symbolic Circuit
The paper investigates how large language models activate a latent symbolic reasoning circuit—comprising abstraction, induction, and retrieval—when presented with in-context examples. By tracking this circuit across different shot counts and model families, the authors show that its components become detectable and functional long before the model reaches high accuracy. They further demonstrate that per-head causal contributions can increase eightfold from 1- to 10-shot, and that interventions such as cross-shot activation patching or function vector injection can dramatically improve accuracy, even at 0-shot, by leveraging the pre‑existing circuit in the model weights.
Convergent Emergence of In-Context Learning Across Modalities
The paper investigates whether few-shot in-context learning (ICL) emerges similarly across different data modalities. Using a controlled cross-modality framework, the authors test the Convergent Emergence Hypothesis, which posits that tasks benefiting from ICL in one modality will also benefit in others. They find that paired-mapping ICL appears in six modalities—language, genome, integer sequences, time series, images, and proteins—outperforming baselines and showing correlated task effects in five of them, supporting the hypothesis in some but not all cases.
Neuron-Aware Active Few-Shot Learning for LLMs
arXiv:2607. 02423v1 Announce Type: cross Abstract: Active Few-Shot Learning (AFSL) adapts LLMs to specialized domains by identifying the most valuable unlabeled samples for annotation and use as few-shot demonstrations, effectively reducing human annotation costs while promoting high performance.
Shared Doubt: Zero-Shot Cross-Lingual Confidence Estimation for Language Models
arXiv:2605. 31220v2 Announce Type: replace-cross Abstract: Confidence estimation (CE), i.