arXiv Machine Learning By Sooho Moon, Yunyong Ko

PROBE-Web: An Interactive System for Probing Evaluation Landscapes of Knowledge Graph Completion Models

Read the original on arXiv Machine Learning →

arXiv:2606. 08926v1 Announce Type: new Abstract: Knowledge graph completion (KGC) models are commonly evaluated using rank-based metrics such as MRR and Hits@K, despite different users often requiring different evaluation perspectives.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 9

Generalized Rank-based Evaluation for Knowledge Graph Completion: Perspectives, Framework, and Analyses

arXiv:2606. 08921v1 Announce Type: new Abstract: Knowledge graph completion (KGC) aims to predict missing facts from an observed knowledge graph (KG), playing a crucial role in a wide range of real-world applications such as drug discovery, recommender systems, and retrieval-augmented generation (RAG).

By Sooho Moon, Jian Kang, Yunyong Ko
arXiv AI
Sep 23

WebCraftBench: Evaluating Web Application Generation from a Software Testing Perspective

arXiv:2609.15387v3 Announce Type: replace-cross Abstract: Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automat...

By Chenxu Liu, Zilu Zou, Peizhong Gao, Jiawen Tao, Zhexin Zhang, Guang Chen, Haowei Lin, Ying Zhou, Tianyi Bai, Dolly Deng, Suncong Zheng, Maxm Pan
arXiv AI
Jun 16

Unifying Post-hoc Explanations of Knowledge Graph Completions

arXiv:2507. 22951v2 Announce Type: replace Abstract: Knowledge Graphs organize information as entity-relation-entity triples, enabling machine learning models to predict plausible missing triples in a task known as Knowledge Graph Completion (KGC).

By Alessandro Lonardi, Samy Badreddine, Tarek R. Besold, Pablo Sanchez Martin
arXiv AI
Aug 11

HugSelect: An Explainable Multi-Criteria Decision-Support Framework for foundation-model selection

arXiv:2608. 08069v1 Announce Type: cross Abstract: Foundation models are increasingly reused as software components, making model selection a critical software-engineering decision.

By Alireza Joonbakhsh (Shiraz University), Arda Canser Adal{\i} (Utrecht University), Slinger Jansen (Utrecht University), Farshad Khunjush (Shiraz University), Siamak Farshidi (Wageningen University,Research)
arXiv Machine Learning
Jun 26

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation

arXiv:2606. 26429v1 Announce Type: new Abstract: Current LLM evaluation relies on two complementary but often disconnected signals: static benchmarks with objective correctness labels and arena-style preference data that better reflect open-ended user interactions.

By Aaron J. Li, Hao Huang, Youngmin Park, Yitong Ma, Wei-Lin Chiang, Li Chen, Cho-Jui Hsieh, Bin Yu, Ion Stoica