Hierarchical data is ubiquitous in the empirical sciences and is most commonly analyzed with generalized linear mixed-effects models (GLMMs). Bayesian inference for GLMMs yields calibrated uncertainty...
arXiv:2609.24422v1 Announce Type: new
Abstract: Hierarchical data is ubiquitous in the empirical sciences and is most commonly analyzed with generalized linear mixed-effects models (GLMMs). Bayesian...
By Alex Kipnis, Marcel Binz, Eric Schulz
arXiv:2610.01269v1 Announce Type: cross
Abstract: Bayesian Optimisation (BO) is a powerful framework for the optimisation of expensive black-box functions, but typically requires refitting a surrogat...
By Luca Geminiani, Nadja Klein
arXiv:2606. 10287v1 Announce Type: new Abstract: Evaluating Knowledge Graph Completion (KGC) models remains challenging because standard assessment relies on isolated rank-based metrics such as MRR, Hits$@$k, and Mean Rank, which often produce conflicting model orderings across datasets.
By Haji Gul, Ajaz Ahmad Bhat
arXiv:2609.06941v1 Announce Type: new
Abstract: Causal effect estimation asks how an outcome would change under an intervention, and medicine, economics, and public policy all treat it as a foundatio...
By Haohao Zhou
arXiv:2607. 11956v1 Announce Type: cross Abstract: Data Shapley is the standard principled answer to which training points are worth what, and its k-nearest-neighbor (KNN) specialization is the version deployed in practice: the exact estimator shipped by toolkits such as pyDVL and OpenDataVal.
By Zongye Lyu
arXiv:2608. 10441v1 Announce Type: new Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expensive measurement -- and then must decide when the acquired signal is worth using.
By Ying Yuan
KCSAT-ML is a new benchmark built from 664 Korean College Scholastic Ability Test mathematics problems, including a 339‑item core set with official per‑item error rates from nationwide cohorts of hundreds of thousands of examinees. The benchmark introduces the Difficulty‑aligned Reasoning Gain (DRG) metric, which evaluates whether a model’s mistakes align with items humans find hard or easy, revealing distinct patterns in how vision‑language and large language models perform across difficulty levels. The dataset and code are publicly available at https://github.com/naver-ai/KCSAT-ML.
By Sanghee Park, Geewook Kim, Kee-Eung Kim
arXiv:2607. 21671v1 Announce Type: new Abstract: Neural network compression and interpretability remain open challenges in modern deep learn- ing, where billion-parameter architectures deliver impressive accuracy at the cost of trans- parency, computational efficiency, and reliable uncertainty quantification.
By Idris Karel Seunda Ekwe, Patrick Tenga Shako, Ernest Parfait Fokou\'e
The paper investigates a failure mode in Graph-JEPA, a joint‑embedding predictive model trained on a large scientific‑reasoning graph. Despite achieving high linear‑probe accuracy and effective rank, the learned representation contains almost no usable instance information, as shown by retrieval metrics. The authors diagnose the issue to variance allocation in the objective, propose a repair that restores near‑perfect information recovery, and demonstrate that the problem persists even after repair, highlighting limitations in the evaluation metrics used.
By Gollam Rabby, S\"oren Auer
arXiv:2608. 15565v1 Announce Type: new Abstract: Experience-learning agents for optimization modeling improve by storing verified skills, but existing learners admit knowledge by checking against known answers, which real ticket streams do not provide.
By Junbo Jacob Lian, Huiling Chen, Hanzhang Qin, Chung-Piaw Teo
BLADE is a variational model for knowledge graph completion that separates latent truth from graph recording and uses distilled offline language‑model judgments as a frozen teacher regularizer. The model provides calibrated probabilities and epistemic uncertainty through posterior samples, while the teacher is only an optional triage factor during inference. Across five benchmarks, BLADE matches ranking performance and significantly reduces expected calibration error, improving ECE, Brier score, and NLL over several baselines, and shows strong performance under controlled missingness and leakage stress tests.
By Ibne Farabi Shihab, Rabeya Bosri Tamanna, Abdo El Karaky, Sanjeda Akter, Anuj Sharma