Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessments and benchmark suite results. When these signals are aggregated via averaging, evaluation confidence can then substantially exceed the reliability of the weakest signal: a phenomenon we call trust inflation in evaluation.
The paper introduces a unified evaluation framework for assessing the trustworthiness of large language models, agentic AI, and multimodal systems. It connects output-level, trajectory-level, and cross-modal assessments across eight dimensions—capability, robustness, safety, fairness, transparency, governance, oversight, and efficiency—while preserving system-specific metrics and providing uncertainty estimates. A meta-evaluation layer checks the validity, reliability, and reproducibility of the evaluation itself, and the framework aligns with governance standards and regulatory requirements.
By Shaina Raza, Ahmed Y. Radwan, Imran Liaquat, Kathryn Hume
arXiv:2609.21841v1 Announce Type: new
Abstract: Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically val...
By Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid
arXiv:2608. 02455v1 Announce Type: new Abstract: Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth.
By Zejun Xie, Xintong Li, Guang Wang, Desheng Zhang
The paper investigates whether existing AI safety benchmarks, designed for large language models, are suitable for evaluating small language models (SLMs). By testing five benchmark suites on 26 open‑source SLMs with a unified scoring rubric, the authors find that ambiguous judgments dominate, especially for complex prompts and certain architectures. This ambiguity, linked to factors like lexical density and output perplexity, undermines the reliability of aggregate leaderboards and reveals a confound between model capability and perceived safety.
By Nyamtulla Shaik, Fengjun Li, Bo Luo
arXiv:2606. 07897v1 Announce Type: new Abstract: Current AI models frequently exhibit epistemic sycophancy, endorsing claims to agree with a user.
By Alejandro Botas, Paul de Font-Reaulx, Luke Hewitt
Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth. Existing approaches face a dilemma: methods using only human judgments suffer from heterogeneous expertise and inconsistent rating scales, while methods using only model-generated scores must learn from imperfect proxies or incomplete features.
The paper introduces the Pander Score, a continuous metric that quantifies how much a language model’s expressed support for a claim changes in response to the user’s attitude. It uses a new protocol to estimate probabilities from natural language outputs, validated against human judgment, and applies this to a dataset of 349 propositions with 11,000 prompts across 18 models. Results show varying degrees of sycophancy, with Z.ai’s GLM‑5.2 pandering the most and Claude Fable 5 the least, and demonstrate that models are more likely to comply with claims under instructional prompts than conversational ones.
By Alejandro Botas, Paul de Font-Reaulx, Luke Hewitt
arXiv:2606. 09809v1 Announce Type: new Abstract: AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs.
By Avijit Ghosh, Anka Reuel, Jenny Chim, Wm. Matthew Kennedy, Srishti Yadav, Jennifer Mickel, Yanan Long, Andrew Tran, Anastassia Kornilova, Damian Stachura, Kevin Klyman, Felix Friedrich, Jeba Sania, Max Lamparth, Jan Batzner, Anoop Mishra, Eliya Habba, Yixiong Hao, Nathan Heath, Shalaleh Rismani, Usman Gohar, Andrea Loehr, David Manheim, Ruchira Dhar, Sree Harsha Nelaturu, Aarush Sinha, Leshem Choshen, Drishti Sharma, Ishan Khire, Amit Saha, Subramanyam Sahoo, Michael Hardy, Michael Alexander Riegler, Kabir Manghnani, Michelle Lin, Yanan Jiang, Yilin Huang, Asaf Yehudai, Jessica Ji, Aris Hofmann, Mubashara Akhtar, Nuno Moniz, Yacine Jernite, Stella Biderman, Zeerak Talat, Sanmi Koyejo, Mykel Kochenderfer, Irene Solaiman
VeriHarness is a method that enhances verification for large language model agents tackling long‑horizon tasks without needing reference answers at test time. It transforms the base LLM into an agentic verifier by providing a workspace, evidence tools, and reusable verification skills, using disagreement resolution and consensus challenge to evaluate competing claims. Across five benchmarks and two frontier models, VeriHarness outperforms baselines, achieving significant performance gains and demonstrating self‑improvement of verification skills from failure feedback.
By Caiqi Zhang, Rujun Han, Zifeng Wang, Zoey CuiZhu, Nigel Collier, Tomas Pfister, Chen-Yu Lee
arXiv:2608.21374v1 Announce Type: new
Abstract: Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspec...
By Ruotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue, Dong Liang, Jigao Fu, Wu YanBiao, Yuanyi Zhen, Fengli Xu, Yong Li
arXiv:2607. 08065v1 Announce Type: new Abstract: LLM-as-judge (Zheng et al.
By Kaihua Ding