arXiv Machine Learning By Sankalp Gilda, Shlok Gilda

Position: Evaluation Scores Are Perishable Knowledge Claims

Read the original on arXiv Machine Learning →

arXiv:2607. 26191v1 Announce Type: cross Abstract: Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessments and benchmark suite results.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 28

Position: Evaluation Scores Are Perishable Knowledge Claims

Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessments and benchmark suite results. When these signals are aggregated via averaging, evaluation confidence can then substantially exceed the reliability of the weakest signal: a phenomenon we call trust inflation in evaluation.

arXiv AI
Sep 18

A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems

The paper introduces a unified evaluation framework for assessing the trustworthiness of large language models, agentic AI, and multimodal systems. It connects output-level, trajectory-level, and cross-modal assessments across eight dimensions—capability, robustness, safety, fairness, transparency, governance, oversight, and efficiency—while preserving system-specific metrics and providing uncertainty estimates. A meta-evaluation layer checks the validity, reliability, and reproducibility of the evaluation itself, and the framework aligns with governance standards and regulatory requirements.

By Shaina Raza, Ahmed Y. Radwan, Imran Liaquat, Kathryn Hume
arXiv AI
Aug 19

Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models

The paper investigates whether existing AI safety benchmarks, designed for large language models, are suitable for evaluating small language models (SLMs). By testing five benchmark suites on 26 open‑source SLMs with a unified scoring rubric, the authors find that ambiguous judgments dominate, especially for complex prompts and certain architectures. This ambiguity, linked to factors like lexical density and output perplexity, undermines the reliability of aggregate leaderboards and reveals a confound between model capability and perceived safety.

By Nyamtulla Shaik, Fengjun Li, Bo Luo