Introducing gpt-oss
We’re releasing gpt-oss-120b and gpt-oss-20b—two state-of-the-art open-weight language models that deliver strong real-world performance at low cost. Available under the flexible Apache 2.
We’ve fine-tuned GPT-3 to more accurately answer open-ended questions using a text-based web browser.
We’re releasing gpt-oss-120b and gpt-oss-20b—two state-of-the-art open-weight language models that deliver strong real-world performance at low cost. Available under the flexible Apache 2.
A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions.
arXiv:2504.11582v3 Announce Type: replace Abstract: How can a monolingual English speaker determine whether an automatic translation in French is good enough to be shared? Existing MT error detection...
arXiv:2601.21225v4 Announce Type: replace-cross Abstract: Large language models have made substantial progress in mathematical reasoning. However, benchmark development for multilingual evaluation ha...
The paper investigates whether language models can reason across languages by introducing a two‑hop question answering task that requires inference over two multilingual documents. Results show that models are more sensitive to language variation in answer‑span documents than in bridging documents, and that up to 33% of multilingual cases involve correct final answers despite failing to infer bridging information in the first step. The study also reveals an 18% composition failure rate and proposes a three‑stage SUBQ prompting method that improves accuracy from 10.1% to 66.5%.
Mathematical Word Problems (MWPs) are an important benchmark for evaluating natural language understanding and quantitative reasoning. Despite recent progress in high resource languages, Bengali remains underexplored due to the limited availability of large scale annotated datasets.
arXiv:2505.16227v4 Announce Type: replace-cross Abstract: Personalizing jargon detection and explanation is essential for making technical documents accessible to readers with diverse disciplinary ba...
VākQA is a newly introduced benchmark for Telugu spoken factoid question answering, comprising 2,001 question‑answer pairs across six domains, 2.53 hours of speech audio, bilingual transcriptions, and human‑verified reference answers. The study validates automatic evaluation methods against human judgments, finding that Gemini‑as‑a‑judge best approximates human ratings but is inconsistently strict, while open‑weight judges tend to penalize correct Telugu answers that differ in surface form. Using this validated setup, the authors benchmark proprietary and open‑weight models, highlighting challenges such as cultural specificity loss in translation, phonetic confusions from speech input, and compounded errors from cascaded ASR‑MT pipelines.
arXiv:2606. 14817v1 Announce Type: cross Abstract: This work presents the design, implementation, and evaluation of a system for generating personalized reading content using Large Language Models (LLMs) combined with Retrieval-Augmented Generation (RAG).
arXiv:2510.18368v2 Announce Type: replace Abstract: We present $\textbf{Korean SimpleQA (KoSimpleQA)}$, a benchmark for evaluating factuality in large language models (LLMs) with a focus on Korean cu...
Systematically evaluating the factuality of large language models with the FACTS Benchmark Suite.
While large language models have been dominating the research landscape recently, small language models remain highly relevant across various domains; yet, they receive far less attention. In this study, we investigate how smaller language models perform during the generation stage within a Retrieval-Augmented Generation (RAG) system.