arXiv AI By Thilo Tamme, Anton Hantel, Bijan Khosrawi-Rad

Whose Voice Survives the Summary? A Voice-Retention Audit of LLM Employee Listening

Read the original on arXiv AI →

The paper introduces a Voice Retention / Representation Ratio metric to assess bias in large language model (LLM) summaries of employee feedback. Using a bilingual corpus of 2,586 responses from a global professional services firm, the study finds that criticism is reported more reliably than praise, and that LLM summaries tend to filter by popularity rather than sentiment—criticism often survives while single-mention concerns, short or German-only content are frequently omitted. The authors argue that prevalence, not sentiment, drives the bias, and provide a metric, field evidence, and a disaggregated voice‑retention card for future audits.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English

The paper introduces BASSE, a multilingual meta‑evaluation dataset containing 2,040 human‑rated abstractive summaries produced manually or by five LLMs with four prompts. Annotators scored each summary on coherence, consistency, fluency, relevance, and 5W1H using a 5‑point Likert scale. Benchmarking shows proprietary LLM‑judge models best align with human judgments, followed by criteria‑specific automatic metrics, while open‑source judge LLMs perform poorly.

By Jeremy Barnes, Naiara Perez, Alba Bonet-Jover, Bego\~na Altuna
arXiv AI
Jun 9

Summarization is Not Dead Yet

arXiv:2606. 08000v1 Announce Type: cross Abstract: The progress of large language models (LLMs) has fueled claims that model-generated summaries rival or even surpass human-written references, raising questions about whether summarization remains an open research problem.

By Dongqi Liu, Chenxi Whitehouse, Zheng Zhao, Zhuchen Cao, Jian Li, Yabiao Wang
Hugging Face Trending Papers
Sep 8

EviSI: An Evaluation Agent for Simultaneous Interpreting

EviSI is a large language model evaluation agent designed for simultaneous speech-to-speech translation. It adapts Multidimensional Quality Metrics to assess semantic fidelity and oral expression, using shared source evidence and deterministic scoring. In English‑to‑Chinese, EviSI achieves a mean Kendall agreement of 0.707 with human system rankings, outperforming baseline metrics, and shows positive concordance with COMET across five translation directions.

arXiv Computation and Language
Aug 31

A Shaky Voice Is Not Always a Dodge: Benchmarking Textual and Vocal Evasion Detection in Earnings Calls

The paper introduces DualEvasion, a benchmark that evaluates evasion detection in earnings call Q&A using both textual transcripts and vocal cues. It contains 505 annotated question‑answer pairs from 60 calls, each labeled for textual evasion (direct vs. evasive) and speaker confidence (confident vs. unconfident). Experiments show that current multimodal models struggle to detect vocal confidence, especially in unconfident responses, and that providing speaker‑level references only modestly improves performance, leaving a significant gap compared to humans.

By Mirae Kim, Seonghun Jeong, Youngjun Kwak
arXiv AI
6d ago

Same Text, Different Numbers: The Divergence of LLM-Based Measures

Researchers investigated how different large language models (LLMs) convert corporate text into empirical variables, focusing on thirteen measures such as sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Using seven LLMs to score earnings call transcripts of S&P 500 companies, they found low cross-model rank correlations (average 0.52) and that transcript-level differences across providers explained only 34% of total score variation. The study shows that model choice significantly alters downstream inference, with varying coefficient magnitudes, signs, and statistical significance, and that averaging across providers stabilizes rankings but not score levels.

By Hamid Boustanifar, Sasan Mansouri