arXiv Machine Learning

MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned Evaluators

MAVEN is a hierarchical framework for evaluating whether multimodal content aligns with macro‑societal values such as peace, justice, and freedom. It organizes values into six primary dimensions and 72 secondary indicators, enabling multi‑level quantitative scoring. The authors build a human‑verified multimodal benchmark, a soft‑match metric, and propose efficient evaluator optimization techniques, demonstrating that a compact 2B evaluator performs comparably to larger models and approaches state‑of‑the‑art closed‑source VLMs.

arXiv AI
Sep 3

Multimodal Language Models as Text-to-Image Model Evaluators

Multimodal Language Models as Text-to-Image Model Evaluators presents MT2IE, a framework where a multimodal large language model generates evaluation prompts and scores images, achieving higher correlation with human judgment than prior metrics. MT2IE recovers official T2I model rankings using only 20 prompts—far fewer than traditional benchmarks—and adapts prompts to each model’s performance, maintaining informative scoring ranges. The approach demonstrates that dynamic, interactive evaluation can replace static benchmarks as T2I models improve.

By Jiahui Chen, Candace Ross, Reyhane Askari-Hemmat, Koustuv Sinha, Melissa Hall, Amy Zhang, Michal Drozdzal, Adriana Romero-Soriano
arXiv AI
Aug 25

Beyond Benchmarks: LLM Evaluation with an Anthropomorphic and Lifecycle-oriented Roadmap

The paper argues that large language models (LLMs) are evaluated too narrowly, focusing on isolated technical metrics rather than holistic, developmental, and societal aspects. It proposes a diagnostic ontology that links evaluation dimensions to the LLM training pipeline, turning evaluation into a root‑cause analysis tool. The authors introduce an anthropomorphic framework—IQ, PQ, EQ, and VQ—to assess LLM capabilities, operationalize it with a modular architecture, and validate it through meta‑analysis of over 200 benchmarks, outlining key challenges and future directions.

By Jun Wang, Ninglun Gu, Kailai Zhang, Pengyong Li, Yelun Bao, Jin Yang, Xu Yin, Liwei Liu, Zijiao Zhang, Yihuan Liu, Gary G. Yen, Junchi Yan
arXiv AI
Aug 28

Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models

The Modality Maturity Index (MMI) is a new benchmark that evaluates large language models on their ability to handle five different modalities—text, image, audio, video, and document—across up to three-input and three-output combinations. It contains 893 self‑contained questions, each with human‑authored rubric criteria for the expected output modalities, and measures performance via an MMI Value and a Modality Presence Score (MPS). Experiments on five frontier multimodal models show low MPS scores, indicating limited modality availability, and confirm that LLM judges can reliably assess output correctness against human‑blind rubric scoring on 70.8% of cases.

By Rohit Patel, Dieuwke Hupkes, Sloan Strader
Hugging Face Trending Papers
Jul 9

PLURAL: A Global Dataset for Value Alignment

Large language models (LLMs) are used worldwide, yet disproportionately reflect Western values, limiting their ability to represent diverse value systems. We introduce PLURAL, a large-scale, value-focused preference dataset grounded in the Integrated Values Survey (IVS), a nationally representative survey spanning 92 countries.

arXiv AI
Aug 21

DiverValue-Bench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values

arXiv:2509. 08022v3 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with diverse human values is essential for safe and effective deployment, yet existing benchmarks often overlook cultural and demographic variation.

By Yao Liang, Dongcheng Zhao, Feifei Zhao, Guobin Shen, Yuwei Wang, Dongqi Liang, Yi Zeng
arXiv AI
Aug 3

M3MAD-Bench: Multi-Dimensional Evaluation of Multi-Agent Debate Across Domains and Modalities

arXiv:2601. 02854v2 Announce Type: replace Abstract: As an agent-level reasoning and coordination paradigm, Multi-Agent Debate (MAD) orchestrates multiple agents through structured debate to improve answer quality and support complex reasoning.

By Ao Li, Jinghui Zhang, Luyu Li, Yuxiang Duan, Lang Gao, Mingcai Chen, Weijun Qin, Shaopeng Li, Fengxian Ji, Ning Liu, Lizhen Cui, Xiuying Chen, Yuntao Du