OpenAI Blog

WebGPT: Improving the factual accuracy of language models through web browsing

We’ve fine-tuned GPT-3 to more accurately answer open-ended questions using a text-based web browser.

arXiv AI
4d ago

MGSM-Pro: A Simple Strategy for Robust Multilingual Mathematical Reasoning Evaluation

arXiv:2601.21225v4 Announce Type: replace-cross Abstract: Large language models have made substantial progress in mathematical reasoning. However, benchmark development for multilingual evaluation ha...

By Tianyi Xu, Kosei Uemura, Alfred Malengo Kondoro, Tadesse Destaw Belay, Catherine Nana Nyaah Essuman, Ifeoma Okoh, Ganiyat Afolabi, Ayodele Awokoya, David Ifeoluwa Adelani
arXiv AI
Sep 1

Do Language Models Reason Across Languages?

The paper investigates whether language models can reason across languages by introducing a two‑hop question answering task that requires inference over two multilingual documents. Results show that models are more sensitive to language variation in answer‑span documents than in bridging documents, and that up to 33% of multilingual cases involve correct final answers despite failing to infer bridging information in the first step. The study also reveals an 18% composition failure rate and proposes a three‑stage SUBQ prompting method that improves accuracy from 10.1% to 66.5%.

By Yan Meng, Wafaa Mohammed, Christof Monz
Hugging Face Trending Papers
Sep 17

VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering

VākQA is a newly introduced benchmark for Telugu spoken factoid question answering, comprising 2,001 question‑answer pairs across six domains, 2.53 hours of speech audio, bilingual transcriptions, and human‑verified reference answers. The study validates automatic evaluation methods against human judgments, finding that Gemini‑as‑a‑judge best approximates human ratings but is inconsistently strict, while open‑weight judges tend to penalize correct Telugu answers that differ in surface form. Using this validated setup, the authors benchmark proprietary and open‑weight models, highlighting challenges such as cultural specificity loss in translation, phonetic confusions from speech input, and compounded errors from cascaded ASR‑MT pipelines.

Hugging Face Trending Papers
Jun 29

Little Brains, Big Feats: Exploring Compact Language Models

While large language models have been dominating the research landscape recently, small language models remain highly relevant across various domains; yet, they receive far less attention. In this study, we investigate how smaller language models perform during the generation stage within a Retrieval-Augmented Generation (RAG) system.