arXiv AI

KG-SoftMAP: Soft Knowledge-Graph Priors for Bayesian Network Structure Learning from Sparse Discrete Data

arXiv:2606. 10358v1 Announce Type: cross Abstract: Learning Bayesian network (BN) structure from sparse discrete data is hard: when each instance records only a few variables, most variable pairs lack the joint observations needed for reliable scoring, and data-only methods recover little structure.

arXiv Computation and Language
Sep 7

KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human Difficulty

KCSAT-ML is a new benchmark built from 664 Korean College Scholastic Ability Test mathematics problems, including a 339‑item core set with official per‑item error rates from nationwide cohorts of hundreds of thousands of examinees. The benchmark introduces the Difficulty‑aligned Reasoning Gain (DRG) metric, which evaluates whether a model’s mistakes align with items humans find hard or easy, revealing distinct patterns in how vision‑language and large language models perform across difficulty levels. The dataset and code are publicly available at https://github.com/naver-ai/KCSAT-ML.

By Sanghee Park, Geewook Kim, Kee-Eung Kim
arXiv Machine Learning
Jul 27

Neural Feature Governance: Extending Atom Prevalence

arXiv:2607. 21671v1 Announce Type: new Abstract: Neural network compression and interpretability remain open challenges in modern deep learn- ing, where billion-parameter architectures deliver impressive accuracy at the cost of trans- parency, computational efficiency, and reliable uncertainty quantification.

By Idris Karel Seunda Ekwe, Patrick Tenga Shako, Ernest Parfait Fokou\'e
arXiv Machine Learning
Aug 24

When Graph-JEPA Learns the Wrong Thing: Diagnosing and Repairing Category-Conditional Collapse

The paper investigates a failure mode in Graph-JEPA, a joint‑embedding predictive model trained on a large scientific‑reasoning graph. Despite achieving high linear‑probe accuracy and effective rank, the learned representation contains almost no usable instance information, as shown by retrieval metrics. The authors diagnose the issue to variance allocation in the objective, propose a repair that restores near‑perfect information recovery, and demonstrate that the problem persists even after repair, highlighting limitations in the evaluation metrics used.

By Gollam Rabby, S\"oren Auer
arXiv Machine Learning
Sep 25

BLADE: Distilled LLM Regularization for Calibrated Knowledge Graph Completion

BLADE is a variational model for knowledge graph completion that separates latent truth from graph recording and uses distilled offline language‑model judgments as a frozen teacher regularizer. The model provides calibrated probabilities and epistemic uncertainty through posterior samples, while the teacher is only an optional triage factor during inference. Across five benchmarks, BLADE matches ranking performance and significantly reduces expected calibration error, improving ECE, Brier score, and NLL over several baselines, and shows strong performance under controlled missingness and leakage stress tests.

By Ibne Farabi Shihab, Rabeya Bosri Tamanna, Abdo El Karaky, Sanjeda Akter, Anuj Sharma