Red-Teaming Large Language Models
Related stories
Understanding the capabilities, limitations, and societal impact of large language models
Evaluating large language models trained on code
Introducing The World's Largest Open Multilingual Language Model: BLOOM
The Curse of Multilinguality in Lexical Normalization
The paper investigates how many languages should be jointly trained in a single lexical normalization model. Using a fixed-capacity character-level model across twelve languages, it finds that accuracy peaks when a language is trained with only a few others—typically one to four—and then declines sharply as more languages are added, dropping about forty percent. A control experiment keeping total training data constant shows the decline is due to competition for model capacity rather than data scarcity, and no reliable typological rule predicts the optimal number of co‑training languages.
One Model to Translate Them All? A Journey to Mount Doom for Multilingual Model Merging
The paper investigates weight‑space merging of independently fine‑tuned multilingual machine translation models. Experiments show that merging is more successful when models share a target language, yet it still cannot match the peak performance of language‑specific checkpoints. When target languages differ, performance drops sharply, and analysis reveals that overlapping neuron activation and incompatible upper‑layer geometries cause these failures.
The Reformer - Pushing the limits of language modeling
Foundations of Large Language Models
Foundations of Large Language Models is a book that focuses on core concepts of large language models rather than exhaustive coverage of the latest technologies. It is organized into six chapters covering pre‑training, generative models, prompting, alignment, inference, and reasoning. The book targets college students, professionals, and practitioners in NLP and related fields, serving as a reference for anyone interested in large language models.
Efficient training of language models to fill in the middle
Learning diverse attacks on large language models for robust red-teaming and safety tuning
arXiv:2405.18540v3 Announce Type: replace-cross Abstract: Red-teaming, or identifying prompts that elicit harmful responses, is a critical step in ensuring the safe and responsible deployment of larg...
Advancing red teaming with people and AI
Advancing red teaming with people and AI
English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training
arXiv:2604.13286v2 Announce Type: replace Abstract: Despite the widespread multilingual deployment of large language models, post-training pipelines remain predominantly English-centric, contributing...