arXiv Machine Learning By Guillaume M\'erou\'e, Fabien Gandon, Pierre Monnin

Oracle, will I ever learn? A study of prediction convergence and complementarity across link prediction models

Read the original on arXiv Machine Learning →

The paper investigates how different link prediction models for knowledge graphs produce varying predictions and explores the extent of their complementary knowledge. By evaluating an oracle that selects the best prediction from a set of models, the authors show a significant performance gap between individual models and the oracle, indicating substantial complementarity. However, this complementarity quickly saturates as more models are added, leaving many queries unsolved even with many models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 1

Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs

The paper evaluates how modern large language models use internal web search to answer factual questions. Using 783 static queries and 288 dynamic queries, the authors find that enabling retrieval improves accuracy on static questions but hurts confidence calibration. On dynamic queries, models often retrieve but still achieve less than 70% accuracy, mainly due to poor query formulation and source selection, indicating that internal web search works better as a quick verification tool than a full information‑retrieval system.

By Sahil Kale
arXiv Machine Learning
Sep 23

TailSpec-EASE: Knowledge-Graph-Regularized Linear Recommendation for Web Long-Tail Discovery

TailSpec-EASE is a lightweight linear recommender that incorporates a relation‑aware spectral knowledge‑graph prior into a local closed‑form reconstruction objective. By adapting the prior strength to item popularity, it provides stronger semantic guidance for long‑tail items. Across four public benchmarks, it achieves a favorable balance of overall accuracy, long‑tail performance, and training cost, improving NDCG@20 by up to 24% over a no‑KG baseline and training in just 37 seconds on CPU compared to thousands of seconds for GPU‑based KGAT and CPU LightGCN.

By Jianru Shen
arXiv Machine Learning
Aug 27

A Storage-Retrieval Gap in Parametric Knowledge Graph Memory

The paper investigates a parametric approach to knowledge graph memory by compiling each entity into a LoRA adapter, enabling zero‑cost query-time retrieval via weight injection. On the MetaQA dataset, these adapters encode context‑free factual knowledge, improving exact‑match scores by up to +0.243 over a base model and achieving an oracle gap of +0.283. However, the stored knowledge is not recoverable through similarity or embedding‑based methods, indicating that knowledge is stored locally and does not transfer across semantically neighboring entities.

By Martino M. L. Pulici, Cuong Xuan Chu, Evgeny Kharlamov, Volker Tresp
arXiv AI
Aug 25

What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"

The paper introduces Knowledge-Weighted Fine‑Tuning, a method that estimates an instance‑level knowledge score through multi‑sampled inference and uses it to scale the learning signal. This approach encourages large language models to explicitly say "I don't know" on out‑of‑scope queries while preserving accuracy on known questions. The authors also propose new evaluation metrics for uncertainty, demonstrating that better discrimination between known and unknown instances improves overall performance.

By Joosung Lee, Hwiyeol Jo, Donghyeon Ko, Kyubyung Chae, Cheonbok Park, Jeonghoon Kim