The paper investigates whether existing AI safety benchmarks, designed for large language models, are suitable for evaluating small language models (SLMs). By testing five benchmark suites on 26 open‑source SLMs with a unified scoring rubric, the authors find that ambiguous judgments dominate, especially for complex prompts and certain architectures. This ambiguity, linked to factors like lexical density and output perplexity, undermines the reliability of aggregate leaderboards and reveals a confound between model capability and perceived safety.
By Nyamtulla Shaik, Fengjun Li, Bo Luo
arXiv:2608. 00794v2 Announce Type: replace Abstract: Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims.
By William Caban
arXiv:2607. 26191v1 Announce Type: cross Abstract: Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessments and benchmark suite results.
By Sankalp Gilda, Shlok Gilda
EvalDetectBench is an open pipeline and benchmark designed to measure evaluation awareness in frontier large language models, enabling practitioners to test models against any Inspect-compatible evaluation. It includes a curated transcript suite from current frontier system-card evaluations and diverse deployment sources, and it assesses both how reliably models recognize they are being evaluated and how detectable individual benchmarks are. The benchmark addresses systematic bias by calibrating probes per model and harmonizing generator selection to correct for variance caused by model identity and prompt choice.
By Xinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj, Alexandra Souly, Robert Kirk
arXiv:2607. 01153v1 Announce Type: cross Abstract: Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded command, or misreported progress in an agentic task.
By Brett Reynolds
arXiv:2608. 06202v1 Announce Type: cross Abstract: Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness.
By Ro Encarnaci\'on, Tina Behzad, Emma Lurie, Dana\'e Metaxa
The paper discusses the trustworthiness of agentic AI systems built on large language models, highlighting new security and operational risks such as indirect prompt injection, memory contamination, and cross‑session data leakage. It categorizes failure modes, reviews mitigation strategies—including instruction hierarchies, context isolation, and constrained tool use—and introduces the Trustworthy Agent Development Lifecycle (TADL), a six‑phase framework for specification, design, training, evaluation, deployment, and monitoring. The authors note that TADL has not yet been empirically validated but offers a structured foundation for developing more secure and accountable agentic systems, and they call for improved benchmarks and future research priorities.
By Fayeq Jeelani Syed, Rehan Ahmad, Ali Al Bataineh, Aakriti Adhikari
Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessments and benchmark suite results. When these signals are aggregated via averaging, evaluation confidence can then substantially exceed the reliability of the weakest signal: a phenomenon we call trust inflation in evaluation.
arXiv:2605. 28591v2 Announce Type: replace-cross Abstract: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings.
By Katharina Deckenbach, Haritz Puerto, Jonas Geiping, Sahar Abdelnabi
arXiv:2602.24055v5 Announce Type: replace
Abstract: This study proposes CIRCLE, a six-stage, lifecycle-based framework to bridge the reality gap between model-centric performance metrics and AI syste...
By Reva Schwartz, Carina Westling, Morgan Briggs, Marzieh Fadaee, Isar Nejadgholi, Matthew Holmes, Fariza Rashid, Maya Carlyle, Afaf Ta\"ik, Kyra Wilson, Peter Douglas, Theodora Skeadas, Gabriella Waters, Rumman Chowdhury, Thiago Lacerda
The paper proposes rethinking bias in AI as a diagnostic tool rather than merely a flaw to be minimized. It introduces a multidimensional framework that examines bias across origin, lifecycle emergence, technical causes, and validation methods, covering 30 bias types, 16 verification methods, and 20 countermeasures for both traditional and generative AI. The authors present a hierarchical evidence framework distinguishing internal and external validity, and advocate for Ethics by Design principles to embed bias verification throughout the AI development lifecycle.
By Samira Maghool, Paolo Ceravolo
The paper reviews how large language models have evolved into agents that can influence external environments through tool use, interface operation, delegation, state retention, virtual world inhabitation, and robotic control. It critiques the narrative of a single march toward autonomy, distinguishing model competence from system integration, persistence, and safe authority. The authors find that action-interface expansion is well documented, while robust completion, recovery, authorization, and independent verification remain less proven, and they propose a framework of justified delegation to guide future research.
By Linsen Zhu, Mengqing Cai