Aligning language models to follow instructions
Related stories
The Reformer - Pushing the limits of language modeling
Efficient training of language models to fill in the middle
Very Large Language Models and How to Evaluate Them
Cross-Lingual Alignment Without Joint Training: Do Monolingual Language Models Converge on Universal Representations?
The study investigates whether monolingual language models, trained without joint multilingual objectives, develop cross-lingual alignment. By evaluating models such as Goldfish and independently built monolingual systems, the authors find that alignable representational geometry emerges across layers, strengthening with larger data, larger models, or closer linguistic proximity. A single Procrustes rotation on parallel sentences can map hidden states between models, and applying this rotation to a German model’s residuals swaps factual predictions to those of the donor English model, demonstrating functional transfer.
Foundations of Large Language Models
Foundations of Large Language Models is a book that focuses on core concepts of large language models rather than exhaustive coverage of the latest technologies. It is organized into six chapters covering pre‑training, generative models, prompting, alignment, inference, and reasoning. The book targets college students, professionals, and practitioners in NLP and related fields, serving as a reference for anyone interested in large language models.
Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin
Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties. We address this by training text-dependent and text-independent aligners for Chengdu Mandarin using a 17-hour corpus and a custom G2P dictionary.
Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin
arXiv:2607. 21332v1 Announce Type: cross Abstract: Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties.
Evaluating large language models trained on code
How to train a new language model from scratch using Transformers and Tokenizers
Best practices for deploying language models
Cohere, OpenAI, and AI21 Labs have developed a preliminary set of best practices applicable to any organization developing or deploying large language models.
Deliberative alignment: reasoning enables safer language models
Deliberative alignment: reasoning enables safer language models Introducing our new alignment strategy for o1 models, which are directly taught safety specifications and how to reason over them.