arXiv Computation and Language

Quality-Aware Self-Correcting Speech Translation on an Edge Device

arXiv Machine Learning
Aug 5

M-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language Models

arXiv:2608. 03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency.

By Tom\'a\v{s} Burkert, Angelika Peljak-{\L}api\'nska, David Zelen\'y
arXiv AI
2d ago

Refinement Buys Intelligibility, Search Buys Identity: What Test-Time Compute Buys in Masked-Diffusion TTS

The study investigates how two computational dimensions—model depth and refinement steps—affect intelligibility and speaker identity in masked-diffusion text‑to‑speech systems. Experiments with 15 models (19–133 M parameters) and up to 16 refinement steps show that refinement improves intelligibility more than identity, with a 1.86× asymmetry that persists even after retraining. Best‑of‑K search can recover identity when refinement fails, and analysis indicates that depth and steps target distinct bottlenecks, requiring separate optimization.

By Nityanand Mathur, Hamees Sayed, Ayush Pratap Singh
arXiv Computation and Language
2d ago

Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation

The paper presents a method for improving simultaneous speech translation by adapting a full‑utterance speech language model with prefix supervision derived from its own complete and partial waveform translations, eliminating the need for transcripts or human translations. Experiments on FLEURS and CoVoST2 across three language directions show that prefix training enhances quality–latency trade‑offs, especially when combined with multi‑turn append‑only decoding, and that a confidence threshold effectively controls the inference‑time quality–latency balance. The study also explores the impact of synthesis margin on translation quality and calibration, finding a non‑monotonic relationship with latency.

By Hieu Hoang, Amittai Axelrod
arXiv Computation and Language
Sep 25

Confident but Wrong: A Constrained Decoding Diagnostic for Low-Resource Automatic Post-Editing

The paper introduces a black‑box, inference‑time diagnostic for low‑resource Automatic Post‑Editing (APE) that distinguishes whether poor performance is due to insufficient training data or inconsistent training signals. By varying an edit‑distance penalty and analyzing the resulting TER‑vs‑λ curve and confidence‑based constraint ordering, the authors identify two failure modes—Binary Collapse and Confident Miscalibration—across multiple language pairs. The diagnostic also suggests practical next steps, such as applying a static constraint for immediate accuracy gains, and the authors release new English‑Sinhala and English‑Tamil APE datasets with accompanying code.

By Isuru Wijesiri, Nisansa de Silva, Kavindu Warnakulasuriya, Aloka Fernando, Surangika Ranathunga
arXiv Machine Learning
Jun 18

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs

arXiv:2606. 18323v1 Announce Type: cross Abstract: Open autoregressive neural-codec text-to-speech (TTS) models sound excellent on typical inputs yet suffer stochastic catastrophic failures: on a meaningful fraction of utterances they emit silence, terminate early, or collapse into repetitive or hallucinated content.

By Ali Asaria, Tony Salomone, Deep Gandhi
arXiv Computation and Language
Sep 4

Breaking the Likelihood Trap: Variance-Calibrated Modulation for Large Language Model Decoding

The paper introduces Variance‑Calibrated Modulation (VCM), a training‑free pre‑decoding technique that reshapes language model probability distributions before truncation. VCM uses two dynamic mechanisms: a Contextual Searchlight via PMI to suppress stopwords and highlight context‑relevant tokens, and an Adaptive Self‑Debiasing that applies scale‑invariant penalization based on real‑time logit standard deviation. Experiments on open‑ended generation, factual QA, and mathematical reasoning show that VCM consistently reduces the likelihood trap, improving diversity, coherence, and reasoning accuracy with minimal computational cost.

By Yuanhao Ding, Meimingwei Li, Esteban Garces Arias, Matthias A{\ss}enmacher, Christian Heumann, Chongsheng Zhang
Hugging Face Trending Papers
Aug 17

IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages

Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and evaluation across tasks and language pairs on one footing. We consolidate the WMT 2020--2024 shared-task lineage with an extended English--Malayalam resource into \indicqe: $126{,}754$ instances over nine directional pairs, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level OK/BAD tags and an error explanation, and a test set stratified over four difficulty axes.

arXiv Computation and Language
Sep 1

IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages

IndicQE-APE is a consolidated benchmark that unifies quality estimation (QE) and automatic post‑editing (APE) data for nine Indic language pairs, comprising 126,754 instances with multiple aligned labels such as direct assessment, human post‑edit, word‑level OK/BAD tags, and error explanations. The dataset includes a stratified test set across four difficulty axes and supports training and evaluation of six prompted large language models, three COMET metrics, and three APE systems. Experiments reveal that segments with conflicting holistic and token‑level quality signals are consistently ranked lower, while annotator disagreement shows no effect when controlled for score distribution. whyItMatters":"The benchmark provides a unified resource for training and evaluating QE and APE across Indic languages, enabling consistent comparison of models and metrics on a shared dataset."

By Diptesh Kanojia, Archchana Sindhujan, Sourabh Deoghare, Daria Sokova, Shenbin Qian, Girish Koushik, Tharindu Ranasinghe, Constantin Or\u{a}san, Chrysoula Zerva, Ricardo Rei, Fr\'ed\'eric Blain, Andr\'e F. T. Martins, Marco Turchi, Matteo Negri, Anoop Kunchukuttan, Mitesh M. Khapra, Pushpak Bhattacharyya
arXiv Computation and Language
Sep 24

Text Scores Can Miss Waveform Use: A Qwen2-Audio Quantization Case Study

The paper presents a new evaluation protocol for post‑training quantization of speech language models that separates lexical output, transcript‑insufficient endpoints, and packed implementations. In a Qwen2‑Audio case study, a 6‑bit allocation selected for translation improves chrF scores but degrades emotion recognition, while uniform and front‑layer controls perform better on emotion tasks. Similar patterns hold at 7 bits, and a 4‑bit study shows consistent emotion deficits across all low‑bit allocations, with no advantage for the selected scheme. The study highlights a precision‑dependent mismatch between lexical output, waveform‑dependent behavior, and nominal precision, without claiming a general failure of low‑bit models or a deployment benefit for the selected allocation.

By Mengzhe Geng, Jinxi Jin, Junhao Xu
arXiv Computation and Language
Sep 21

Dynamic Lagging using Stable-Prefix Training for Simultaneous Translation

The paper introduces a training strategy for cascaded simultaneous speech translation that allows the system to dynamically decide how much of the source prefix to translate. By fine‑tuning a large language model (Qwen3‑8B) on stable prefixes—pairs of source prefixes and the longest shared translation with the full sentence—the authors enable contextual read‑write decisions beyond fixed wait‑k or target‑suffix deletion. Experiments on English‑to‑German, Japanese, and Chinese demonstrate that stable prefixes improve the quality‑latency tradeoff across various test sets.

By Hieu Hoang, Amittai Axelrod, Matt Post