arXiv AI By Peng Sun, Xiangyu Zhang, Duan Wu, Lu Tan, Jian Lin, He Yang, Qi Qian, Yikai Wang

BoRP: Bootstrapped Regression Probing for Scalable and Human-Aligned LLM Evaluation

Read the original on arXiv AI →

arXiv:2601. 18253v2 Announce Type: replace-cross Abstract: Accurate evaluation of user satisfaction is critical for iterative development of conversational AI.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 11

Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting

The paper introduces a new automatic evaluation metric for human simultaneous interpreting that aligns with traditional analytic rubrics. A corpus of 1,101 professionally scored SI segments is created, covering meaning transfer, delivery quality, and perceived latency. Using a LoRA‑adapted COMET‑KIWI encoder with dual regression heads, the model achieves Pearson correlations of 0.388 for meaning transfer and 0.301 for delivery quality, outperforming the frozen baseline while acknowledging low rater agreement.

By Ziyu Zhang, Satoshi Nakamura
arXiv Computation and Language
Aug 28

Evaluating Language Models in Realistic Conversational Contexts

The paper introduces UPHELD, a large benchmark of human-to-human dialogues written by professional script writers, featuring realistic turn densities and over 36,000 per-turn human annotations. It evaluates existing automatic metrics and LLM-as-a-judge methods, finding them unreliable against expert human judgment. Using UPHELD, the authors develop a Mixture-of-Judges framework that improves correlation with human assessments by about 30%.

By Ilija Subasic, Andrew Rabinovich, Zhao Chen
arXiv AI
Sep 4

Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation

The paper examines LLM-as-a-Judge systems used to assess AI-generated text, questioning the assumption that judgments are derived from reasoning over responses and rubrics. It finds that classifiers trained solely on rubric text can predict judge outputs, indicating that rubrics contain recoverable evaluative signals independent of the responses. Counterfactual experiments show judges often fail to adjust decisions when either the response or rubric criterion is reversed, raising doubts about the reliability of rubric-based LLM evaluation.

By Anshul Bagaria, Sowmya S Sundaram, Gokul S Krishnan, Balaraman Ravindran
arXiv AI
Jul 15

Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking

arXiv:2607. 12085v1 Announce Type: new Abstract: Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality.

By Niranjan Kumar M, Balaji Nagarajan, Karthik Nair, Faysal Satter, Nithin Surendran