arXiv Machine Learning By Haji Gul, Ajaz Ahmad Bhat

When Metrics Disagree: A Meta-Analysis of Knowledge-Graph-Completion Model Benchmarking

Read the original on arXiv Machine Learning →

arXiv:2606. 10287v1 Announce Type: new Abstract: Evaluating Knowledge Graph Completion (KGC) models remains challenging because standard assessment relies on isolated rank-based metrics such as MRR, Hits$@$k, and Mean Rank, which often produce conflicting model orderings across datasets.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 9

Generalized Rank-based Evaluation for Knowledge Graph Completion: Perspectives, Framework, and Analyses

arXiv:2606. 08921v1 Announce Type: new Abstract: Knowledge graph completion (KGC) aims to predict missing facts from an observed knowledge graph (KG), playing a crucial role in a wide range of real-world applications such as drug discovery, recommender systems, and retrieval-augmented generation (RAG).

By Sooho Moon, Jian Kang, Yunyong Ko
arXiv Machine Learning
5d ago

One Capability or Many? Structural and Predictive Tests of Benchmark Validity Disagree About Economic Benchmarks for Frontier AI

The paper examines whether economic benchmarks used in frontier AI leaderboards measure a distinct capability or merely reflect general test-taking ability. Using a structural factor analysis and a predictive leave-one-benchmark-out test on a snapshot of 421 model configurations, the authors find that economic benchmarks do not form a separate factor but are better predicted by a multi‑factor representation than by a single general index, especially for linear learners. They argue that construct validity should be evaluated with both structural and predictive tests and provide a two‑test protocol along with data and code.

By Louis Yiven Zhu
arXiv AI
Jul 14

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

arXiv:2607. 09739v1 Announce Type: new Abstract: We study LLM benchmark coreset selection: selecting a small subset of prompts over multiple benchmarks whose induced model scores and rankings approximate those obtained from the full benchmark suite.

By Jihan Yao, Gantavya Bhatt, Arnav Das, Peter Jin, Ke Bao, Qiaolin Yu, Khushi Bhardwaj, Chang Su, Jialei Wang, Yikai Zhu, Sugam Devare, Damon Mosk-Aoyama, Zhen Dong, Venkat Krishna Srinivasan, Yineng Zhang, Oleksii Kuchaiev, Jiantao Jiao, Banghua Zhu, Jeff Bilmes