OpenAI Blog

Better language models and their implications

We’ve trained a large-scale unsupervised language model which generates coherent paragraphs of text, achieves state-of-the-art performance on many language modeling benchmarks, and performs rudimentary reading comprehension, machine translation, question answering, and summarization—all without task-specific training.

arXiv AI
Aug 24

PSK at WMT 2026 MIST: Task-Specialized QLoRA Adapters for Multilingual Summarization and Question Answering

The PSK submission to the WMT 2026 Multilingual Instruction Shared Task employs a 3.35B‑parameter Tiny Aya Global model enhanced with three QLoRA adapters, each dedicated to a specific task: multilingual summarization, passage‑based question answering, and filtered standalone question answering. The summarization adapter is trained on multilingual document‑summary pairs, including scientific papers with author‑written abstracts, and outperforms a multitask adapter trained solely on organizer data on a held‑out split. For open question answering, results vary with answer length and evaluation method, prompting the submission of three systems that share the same context and summarization adapters but differ in their open‑QA adapters.

By Srikar Kashyap Pulipaka
Hugging Face Trending Papers
Sep 2

Unifying Conformal Language Tasks with In-Context Ensembles

The paper introduces the Conformal Relevance framework, which leverages in-context learning example curation and ensembling to generate a score function that preserves coverage while enhancing conciseness for NLP tasks such as summarization and extractive question answering. Unlike traditional methods that rely on labor-intensive, hand-engineered LLM prompts to rate content importance, this approach requires minimal manual input. The authors validate the framework across seven NLP tasks and provide theoretical insights into how diversity in ensembled conformal scores can improve worst-case sentence scores, including a saturation bound on ensemble gains.

arXiv Machine Learning
Sep 4

Unifying Conformal Language Tasks with In-Context Ensembles

The paper introduces the Conformal Relevance framework, which employs in-context learning example curation and ensembling to generate a score function that preserves coverage while enhancing conciseness for NLP tasks such as summarization and extractive question answering. Unlike previous methods that rely on labor-intensive, task‑specific prompt engineering, this approach requires minimal manual input. The authors validate the framework across seven NLP tasks and provide a theoretical analysis of how diversity in ensembled conformal scores can improve worst‑case sentence scores, including a saturation bound on ensemble gains.

By Xiao Shi Huang, Chen-Yuan Lin, Bruce Kuwahara, Kin Kwan Leung, Jesse C. Cresswell
arXiv Machine Learning
Sep 10

Retrieval-augmented Decoding for Improving Truthfulness in Open-ended Generation

The paper introduces Retrieval-Augmented Decoding (RAD), a decoding-time method that improves the truthfulness of large language models without retraining. RAD uses a small reference set of up to ten annotated examples to build a grounding space of context embeddings and next-token logits, which it retrieves and aggregates during inference to shape the model’s output. Experiments on four open-ended generation benchmarks and four different LLMs show that RAD consistently outperforms strong baselines and generalizes well across tasks.

By Manh Nguyen, Sunil Gupta, Hung Le
arXiv Computation and Language
Sep 4

FrameBench:A Language Understanding Benchmark Based on Frame Semantics

FrameBench is a new benchmark that evaluates language models on their ability to distinguish semantic frames evoked by the same verb in different contexts, using multiple-choice questions grounded in FrameNet-style resources for English and Japanese. The dataset is generated and verified through a pipeline that incorporates native-speaker judgments, and the authors provide both the data and the code for construction and evaluation. Experiments show that small models struggle with this task, while several large models outperform human reference scores.

By Chihiro Yano, Ryohei Sasano
arXiv Computation and Language
Sep 15

To Each Language Its Tokenizer: Modular Tokenizers for Efficient Multilingual LLMs

The paper proposes a modular tokenizer framework for multilingual large language models, allowing the creation of language‑specific subtokenizers that match monolingual compression quality. It introduces a pretraining strategy that samples these subtokenizers to limit predictions to relevant vocabularies, enabling efficient training and inference. This approach reduces memory usage and speeds up inference without compromising performance.

By Franck Signe, Hippolyte Pilchen, Fran\c{c}ois Yvon, \'Edouard Grave