The paper introduces BASSE, a multilingual meta‑evaluation dataset containing 2,040 human‑rated abstractive summaries produced manually or by five LLMs with four prompts. Annotators scored each summary on coherence, consistency, fluency, relevance, and 5W1H using a 5‑point Likert scale. Benchmarking shows proprietary LLM‑judge models best align with human judgments, followed by criteria‑specific automatic metrics, while open‑source judge LLMs perform poorly.
By Jeremy Barnes, Naiara Perez, Alba Bonet-Jover, Bego\~na Altuna
arXiv:2606. 08000v1 Announce Type: cross Abstract: The progress of large language models (LLMs) has fueled claims that model-generated summaries rival or even surpass human-written references, raising questions about whether summarization remains an open research problem.
By Dongqi Liu, Chenxi Whitehouse, Zheng Zhao, Zhuchen Cao, Jian Li, Yabiao Wang
EviSI is a large language model evaluation agent designed for simultaneous speech-to-speech translation. It adapts Multidimensional Quality Metrics to assess semantic fidelity and oral expression, using shared source evidence and deterministic scoring. In English‑to‑Chinese, EviSI achieves a mean Kendall agreement of 0.707 with human system rankings, outperforming baseline metrics, and shows positive concordance with COMET across five translation directions.
The paper introduces DualEvasion, a benchmark that evaluates evasion detection in earnings call Q&A using both textual transcripts and vocal cues. It contains 505 annotated question‑answer pairs from 60 calls, each labeled for textual evasion (direct vs. evasive) and speaker confidence (confident vs. unconfident). Experiments show that current multimodal models struggle to detect vocal confidence, especially in unconfident responses, and that providing speaker‑level references only modestly improves performance, leaving a significant gap compared to humans.
By Mirae Kim, Seonghun Jeong, Youngjun Kwak
arXiv:2609.13893v1 Announce Type: new
Abstract: Earnings conference calls are a primary channel through which managers disclose information under analyst scrutiny. Prior work has linked vocal and lex...
By Huizhong Chen, Huan Zhang
Researchers investigated how different large language models (LLMs) convert corporate text into empirical variables, focusing on thirteen measures such as sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Using seven LLMs to score earnings call transcripts of S&P 500 companies, they found low cross-model rank correlations (average 0.52) and that transcript-level differences across providers explained only 34% of total score variation. The study shows that model choice significantly alters downstream inference, with varying coefficient magnitudes, signs, and statistical significance, and that averaging across providers stabilizes rankings but not score levels.
By Hamid Boustanifar, Sasan Mansouri
arXiv:2607. 18983v1 Announce Type: cross Abstract: We present AutoJourn, a demonstration system for multi-perspective news generation and bias-aware evaluation using large language models (LLMs).
By Himel Ghosh, Ahmed Mosharafa, Georg Groh
We present AutoJourn, a demonstration system for multi-perspective news generation and bias-aware evaluation using large language models (LLMs). The system tackles three core challenges in responsible automated journalism: extracting diverse perspectives from unstructured social media discussions, generating summaries that preserve viewpoint diversity, and detecting or mitigating bias in AI-generated news.
The paper introduces the concept of summarization bias in large language models (LLMs), describing a systematic tendency for LLMs to represent narrative meaning as an abstract summary label rather than the reconstructable inferential structure that produces it. It frames this bias within the Bulut Doctrine’s told‑shown axis, arguing that LLMs fail in a specific direction: they default to told‑mode explicitness in generative tasks and reward told‑mode explicitness while under‑detecting shown‑mode suppression in evaluative tasks. The authors outline two regimes of bias, present preliminary evidence, and pre‑register a test protocol to validate or abandon the construct.
By Levent Bulut
The paper investigates language-of-study (LoS) bias in NLP peer reviews, defining and distinguishing negative and positive forms of bias. Using a new dataset, LOBSTER, and an LLM-based detection pipeline, the authors analyze 15,645 reviews and find that non‑English papers experience significantly higher bias rates, with negative bias outweighing positive bias. They further identify four subcategories of negative bias, noting that demanding unjustified cross‑lingual generalization is the most common.
By Ehsan Barkhordar, Abdulfattah Safa, Verena Blaschke, Erika Lombart, Marie-Catherine de Marneffe, G\"ozde G\"ul \c{S}ahin
Large language models (LLMs) are increasingly used to assess social bias in text, but the passages they evaluate often contain surface noise such as typos and broken punctuation. This study applied five realistic noise conditions at varying intensities to 3,822 stereotype‑related responses and compared bias judgments on noisy versus original text. The findings show that noise disproportionately turns neutral judgments into biased ones—up to 120 times more likely—while rarely converting biased judgments into neutral ones, and that the most fragile LLM judge exhibits the greatest distortion at mild noise levels. As LLMs become more robust, the bias distortion tends toward parity rather than reversal, meaning bias measured on noisy text is systematically overestimated, especially in fairness‑critical categories.
By DongHyun Ryu, Jaehyeok Lee, YeongJun Hwang, JinYeong Bak
The paper introduces a bias depth score to differentiate between stable model preferences (Deep biases) and prompt‑dependent responses (Shallow biases) in large language models. By analyzing 4,442 opinion prompts across four models, it finds that only about a quarter of concentrated preferences persist after scenario reframing, indicating that most are shallow. The study shows Deep biases are more often inherited from pretraining and harder to remove through fine‑tuning or prompt‑based debiasing, highlighting the need to distinguish learned biases from prompt artifacts.
By An Vo, Vy Tuong Dang, Khai-Nguyen Nguyen, Emilio Villa-Cueva, Thamar Solorio, Anh Totti Nguyen, Daeyoung Kim