The paper treats large language model (LLM) evaluation as a tensor completion problem, modeling noisy, sparse, and non‑uniform pairwise human judgments through a low‑rank latent score tensor under Bradley‑Terry‑Luce‑type models. It derives the efficient influence function and semiparametric efficiency bound for smooth functionals of the true tensor, and proposes a one‑step debiased estimator with asymptotic normality. A key innovation is a score‑whitening technique that equalizes local Fisher information, overcoming anisotropy in the information operator and enabling stable inference at optimal sample‑complexity.
By Jiachun Li, David Simchi-Levi, Will Wei Sun
arXiv:2606. 08679v1 Announce Type: cross Abstract: Pretrained models are often evaluated on multi-task leaderboards to measure their applicability in diverse contexts.
By Bitya Neuhof, Yuval Benjamini
arXiv:2607. 20205v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) has become a widely used parameter-efficient fine-tuning method for large language models.
By Yihang Gao, Vincent Y. F. Tan
Low-rank adaptation (LoRA) has become a widely used parameter-efficient fine-tuning method for large language models. Since different modules and layers may contribute unequally to downstream adaptation, allocating rank resources under a fixed parameter budget is an important problem for balancing efficiency, expressiveness, and generalization.
arXiv:2606. 29184v1 Announce Type: new Abstract: While Low-rank adaptation (LoRA) enables highly efficient fine-tuning by constraining task-specific updates to fixed low-rank subspaces, this rigid design limits representational flexibility and often results in overconfident predictions and miscalibrated uncertainty, especially in low-data regimes.
By Zhibin Duan, Yuhong Wang, Jiahong Fu, Zongsheng Yue, Bo Chen, Zongben Xu
arXiv:2601. 21817v2 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on open-ended tasks without ground-truth labels is increasingly done via the LLM-as-a-judge paradigm.
By Mingyuan Xu, Xinzi Tan, Jiawei Wu, Doudou Zhou
arXiv:2607. 25257v1 Announce Type: cross Abstract: Item Response Theory (IRT) has recently been proposed as a framework for evaluating large language model (LLM) benchmarks by separating a model's latent ability from the properties of individual benchmark items.
By Juan Francisco, Mandujano Reyes
arXiv:2607. 02182v1 Announce Type: new Abstract: Large language models (LLMs) exhibit remarkable reasoning capabilities, but their task-specific fine-tuning is notoriously plagued by overconfidence, severely hindering trustworthy deployment.
By Jijie Zhang, Zhe Ren, Quan Zhang, Dandan Guo
arXiv:2608.30044v1 Announce Type: new
Abstract: Language models are commonly compared by averaging scores across a benchmark list with equal weight. Such lists grow through publication outside an exp...
By Jhen-Ke Lin
arXiv:2606. 27997v1 Announce Type: new Abstract: Benchmarks of machine learning models often include many datasets, making evaluation expensive.
By Rostislav Gusev, Alexey Zaytsev
arXiv:2608.22432v1 Announce Type: cross
Abstract: Multilingual LLM judges produce different evaluator-backbone rankings depending on the prompt language: on an eight-language Agent-as-a-Judge benchma...
By Alhasan Mahmood, Samir Abdaljalil, Hasan Kurban
arXiv:2608. 02455v1 Announce Type: new Abstract: Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth.
By Zejun Xie, Xintong Li, Guang Wang, Desheng Zhang