arXiv AI

LLM Evaluation as Tensor Completion: Low Rank Structure and Semiparametric Efficiency

The paper treats large language model (LLM) evaluation as a tensor completion problem, modeling noisy, sparse, and non‑uniform pairwise human judgments through a low‑rank latent score tensor under Bradley‑Terry‑Luce‑type models. It derives the efficient influence function and semiparametric efficiency bound for smooth functionals of the true tensor, and proposes a one‑step debiased estimator with asymptotic normality. A key innovation is a score‑whitening technique that equalizes local Fisher information, overcoming anisotropy in the information operator and enabling stable inference at optimal sample‑complexity.

arXiv Statistics ML
2d ago

Low Rank for Rank: Uncertainty-Aware Task-Specific LLM Ranking under Sparse Pairwise Comparisons

The paper introduces a low‑rank framework for ranking large language models (LLMs) on task‑specific benchmarks using sparse pairwise comparisons. By modeling the task‑by‑model ability matrix as low rank, the method shares information across related tasks while preserving task‑specific differences, and it provides uncertainty‑aware ranking through debiased estimators and simultaneous confidence sets. Experiments on synthetic data and the Chatbot Arena benchmark demonstrate improved sample efficiency and tighter, better‑calibrated ranking certificates, especially in the sparse comparison regime typical of real LLM evaluations.

By Jiachun Li, David Simchi-Levi, Will Wei Sun
arXiv Machine Learning
Jun 4

Low-rank Distributional Matrix Completion

arXiv:2606. 04176v1 Announce Type: new Abstract: We study a distributional generalization of the matrix completion problem in which each entry of the target matrix is a probability distribution rather than a scalar.

By Jiayi Wang, Raymond K. W. Wong
arXiv Machine Learning
Jun 30

BaRA: Bayesian Adaptive Rank Allocation for Parameter-Efficient Fine-Tuning

arXiv:2606. 29184v1 Announce Type: new Abstract: While Low-rank adaptation (LoRA) enables highly efficient fine-tuning by constraining task-specific updates to fixed low-rank subspaces, this rigid design limits representational flexibility and often results in overconfident predictions and miscalibrated uncertainty, especially in low-data regimes.

By Zhibin Duan, Yuhong Wang, Jiahong Fu, Zongsheng Yue, Bo Chen, Zongben Xu
arXiv AI
Jul 28

Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models

arXiv:2601. 21003v3 Announce Type: replace Abstract: Large Language Models usually put more emphasis on accuracy and therefore, will guess even when not certain about the prediction, which is especially severe when fine-tuned on small datasets due to the inherent tendency toward miscalibration.

By Moule Lin, Shuhao Guan, Andrea Patane, David Gregg, Goetz Botterweck
arXiv AI
Jul 21

Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals

arXiv:2602. 03061v2 Announce Type: replace-cross Abstract: Evaluating mathematical reasoning in LLMs is constrained by limited benchmark sizes and inherent model stochasticity, yielding high-variance accuracy estimates and unstable rankings across platforms.

By Zihan Dong, Zhixian Zhang, Yang Zhou, Can Jin, Ruijia Wu, Linjun Zhang