arXiv AI By Peng Sun, Xiangyu Zhang, Duan Wu, Lu Tan, Jian Lin, He Yang, Qi Qian, Yikai Wang

BoRP: Bootstrapped Regression Probing for Scalable and Human-Aligned LLM Evaluation

Read the original on arXiv AI →

arXiv:2601. 18253v2 Announce Type: replace-cross Abstract: Accurate evaluation of user satisfaction is critical for iterative development of conversational AI.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jul 15

Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking

arXiv:2607. 12085v1 Announce Type: new Abstract: Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality.

By Niranjan Kumar M, Balaji Nagarajan, Karthik Nair, Faysal Satter, Nithin Surendran