arXiv AI By Jun Wang, Ninglun Gu, Kailai Zhang, Pengyong Li, Yelun Bao, Jin Yang, Xu Yin, Liwei Liu, Zijiao Zhang, Yihuan Liu, Gary G. Yen, Junchi Yan

Beyond Benchmarks: LLM Evaluation with an Anthropomorphic and Lifecycle-oriented Roadmap

Read the original on arXiv AI →

The paper argues that large language models (LLMs) are evaluated too narrowly, focusing on isolated technical metrics rather than holistic, developmental, and societal aspects. It proposes a diagnostic ontology that links evaluation dimensions to the LLM training pipeline, turning evaluation into a root‑cause analysis tool. The authors introduce an anthropomorphic framework—IQ, PQ, EQ, and VQ—to assess LLM capabilities, operationalize it with a modular architecture, and validate it through meta‑analysis of over 200 benchmarks, outlining key challenges and future directions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 20

MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned Evaluators

MAVEN is a hierarchical framework for evaluating whether multimodal content aligns with macro‑societal values such as peace, justice, and freedom. It organizes values into six primary dimensions and 72 secondary indicators, enabling multi‑level quantitative scoring. The authors build a human‑verified multimodal benchmark, a soft‑match metric, and propose efficient evaluator optimization techniques, demonstrating that a compact 2B evaluator performs comparably to larger models and approaches state‑of‑the‑art closed‑source VLMs.

By Zijuan Zhao, Zheren Fu, Hou Xia, Licheng Zhang, Yi Liu, Zhendong Mao
arXiv AI
Jun 6

SAGE: Scalable AI Governance & Evaluation

arXiv:2602. 07840v3 Announce Type: replace-cross Abstract: Evaluating relevance in large-scale search systems is fundamentally constrained by the governance gap between nuanced, resource-constrained human oversight and the high-throughput requirements of production systems.

By Benjamin Le, Xueying Lu, Nick Stern, Wenqiong Liu, Igor Lapchuk, Xiang Li, Baofen Zheng, Kevin Rosenberg, Jiewen Huang, Zhe Zhang, Abraham Cabangbang, Satej Milind Wagle, Jianqiang Shen, Raghavan Muthuregunathan, Abhinav Gupta, Mathew Teoh, Andrew Kirk, Thomas Kwan, Jingwei Wu, Wenjing Zhang