arXiv:2606. 08921v1 Announce Type: new Abstract: Knowledge graph completion (KGC) aims to predict missing facts from an observed knowledge graph (KG), playing a crucial role in a wide range of real-world applications such as drug discovery, recommender systems, and retrieval-augmented generation (RAG).
By Sooho Moon, Jian Kang, Yunyong Ko
arXiv:2606. 08926v1 Announce Type: new Abstract: Knowledge graph completion (KGC) models are commonly evaluated using rank-based metrics such as MRR and Hits@K, despite different users often requiring different evaluation perspectives.
By Sooho Moon, Yunyong Ko
The paper examines whether economic benchmarks used in frontier AI leaderboards measure a distinct capability or merely reflect general test-taking ability. Using a structural factor analysis and a predictive leave-one-benchmark-out test on a snapshot of 421 model configurations, the authors find that economic benchmarks do not form a separate factor but are better predicted by a multi‑factor representation than by a single general index, especially for linear learners. They argue that construct validity should be evaluated with both structural and predictive tests and provide a two‑test protocol along with data and code.
By Louis Yiven Zhu
arXiv:2604. 24827v2 Announce Type: replace-cross Abstract: Closed-source frontier labs do not disclose parameter counts.
By Bojie Li
arXiv:2607. 09739v1 Announce Type: new Abstract: We study LLM benchmark coreset selection: selecting a small subset of prompts over multiple benchmarks whose induced model scores and rankings approximate those obtained from the full benchmark suite.
By Jihan Yao, Gantavya Bhatt, Arnav Das, Peter Jin, Ke Bao, Qiaolin Yu, Khushi Bhardwaj, Chang Su, Jialei Wang, Yikai Zhu, Sugam Devare, Damon Mosk-Aoyama, Zhen Dong, Venkat Krishna Srinivasan, Yineng Zhang, Oleksii Kuchaiev, Jiantao Jiao, Banghua Zhu, Jeff Bilmes
arXiv:2606. 10358v1 Announce Type: cross Abstract: Learning Bayesian network (BN) structure from sparse discrete data is hard: when each instance records only a few variables, most variable pairs lack the joint observations needed for reliable scoring, and data-only methods recover little structure.
By Guoliang Xu, James E. Corter