arXiv AI By Viviana Crescitelli, Generoso Immediato, Fabio Persia, Stefania Costantini

AI Evaluation Should Measure Verification Cost, Not Correctness Alone

Read the original on arXiv AI →

arXiv:2608. 08709v1 Announce Type: new Abstract: The reliability of AI generative models is typically measured by output correctness, yet in practice it depends on the effort required to verify those outputs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 15

The geometry of AI validation: From structural blindness to reusable audits

The paper investigates how AI systems that perform best‑of‑n search require different validation strategies as the search width changes. It shows that auditing only small search widths leaves a gap in reliability estimates for larger widths, and proposes retaining candidate ranks and truth labels to estimate reliability across all widths up to N. The authors derive theoretical bounds on the minimax mean‑squared error, design procedures that achieve these bounds, and demonstrate that a shared audit can significantly reduce maximum error across many widths in practical CodeRM pools.

By Ricardo Fitas
arXiv AI
Sep 25

The Gold in Bias: Maturing the AI Design Process through Verification

The paper proposes rethinking bias in AI as a diagnostic tool rather than merely a flaw to be minimized. It introduces a multidimensional framework that examines bias across origin, lifecycle emergence, technical causes, and validation methods, covering 30 bias types, 16 verification methods, and 20 countermeasures for both traditional and generative AI. The authors present a hierarchical evidence framework distinguishing internal and external validity, and advocate for Ethics by Design principles to embed bias verification throughout the AI development lifecycle.

By Samira Maghool, Paolo Ceravolo