arXiv Machine Learning

Shared Doubt: Zero-Shot Cross-Lingual Confidence Estimation for Language Models

arXiv:2605. 31220v2 Announce Type: replace-cross Abstract: Confidence estimation (CE), i.

arXiv Computation and Language
Sep 1

Beyond the Final Layer: Intermediate Representations for Better Multilingual Calibration in Large Language Models

The paper investigates multilingual confidence calibration in large language models, revealing that non‑English languages are systematically less well calibrated than English. By analyzing internal representations, the authors find that late‑intermediate layers provide a more reliable confidence signal than the final layer, which is biased by English‑centric training. They propose training‑free methods such as Language‑Aware Confidence Ensemble (LACE) to adaptively select optimal layers per language, aiming to improve global equity and trustworthiness of LLMs.

By Ej Zhou, Caiqi Zhang, Tiancheng Hu, Chengzu Li, Nigel Collier, Ivan Vuli\'c, Anna Korhonen
arXiv AI
Sep 7

A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models

The paper presents a unified evaluation of cross‑lingual consistency (CLC) enhancement methods for multilingual language models, covering inference‑time interventions and post‑training approaches across three model families and three closed‑form benchmarks. Results indicate that post‑training methods, especially direct distribution alignment, consistently improve CLC across all model‑dataset combinations, while other methods are more sensitive to answer format and language coverage. The study also examines the impact of CLC enhancement on culturally diverse question answering, finding no systematic degradation in controlled settings but occasional accuracy drops in open‑ended generation, particularly for non‑English responses.

By Jirui Qi, Mingyang Wang, Hinrich Sch\"utze, Raquel Fern\'andez, Arianna Bisazza
arXiv Computation and Language
Aug 27

Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation

The paper investigates how to fairly compare language models across languages, noting that current evaluation methods vary widely and lack empirical validation. By training controlled monolingual models on parallel data and testing multilingual LLMs, the authors find that many normalized metrics suffer from biases due to tokenization, encoding, and orthographic differences. Instead, they recommend using sentence‑level negative log‑likelihood over semantically equivalent sequences for more reliable cross‑lingual comparisons.

By Xiulin Yang, Ethan Gotlieb Wilcox, Catherine Arnett
arXiv Computation and Language
Sep 7

Choosing the Right Language Mode at Inference Time for Multilingual Reliability

The paper investigates how multilingual large language models can be guided to reason more reliably in low- to mid-resource languages by selecting appropriate language modes during inference. Experiments with LLaMA and Qwen models show that using English context can correct errors from non‑English comprehension, but adding redundant bilingual context can cause interference. To balance this trade‑off, the authors propose Reliability‑Aware Adaptive Inference (RAAI), a training‑free test‑time framework that routes prompts based on Expected Calibration Error and gates reasoning with a mid‑layer Risk Index, achieving up to 37.7% accuracy gains and reduced calibration error on low‑resource languages.

By Ekata Mitra, Ameeta Agrawal