arXiv:2608. 06609v1 Announce Type: new Abstract: Automated item evaluation (AIE) refers to the use of computational methods to assess item quality without requiring manual expert review or field testing of the items under evaluation.
By Hotaka Maeda, Yikai Lu
arXiv:2608.20385v1 Announce Type: new
Abstract: Systematic reviews rely on quality appraisal of included studies, a process that is time-consuming and sensitive to ambiguity in checklist criteria. Al...
By Timo van der Kuil (Methodology and Statistics Utrecht University), Bruno Messina Coimbra (Methodology and Statistics Utrecht University), Mirjam van Zuiden (Clinical Psychology Utrecht University), Robert A. Bagheri (Methodology and Statistics Utrecht University), Rens van de Schoot (Methodology and Statistics Utrecht University), Klaas Dieleman (Methodology and Statistics Utrecht University), Berend Greijn (Methodology and Statistics Utrecht University), Stefan Houkes (Methodology and Statistics Utrecht University), Sebastiaan Rodenhuis (Methodology and Statistics Utrecht University), Elizabeth M. Grandfield (Methodology and Statistics Utrecht University)
arXiv:2608.21374v1 Announce Type: new
Abstract: Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspec...
By Ruotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue, Dong Liang, Jigao Fu, Wu YanBiao, Yuanyi Zhen, Fengli Xu, Yong Li
arXiv:2607. 16989v1 Announce Type: cross Abstract: Introduction.
By Mohammad Arvan, Amber E. Osterholt, Bailee Rue, Yuvaneswaren Ramakrishnan Sureshbabu, Krishna Riteshkumar Patel, Rebecca T. Feinstein, Bethany C. Bray, Niranjan S. Karnik
arXiv:2608. 03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting.
By Abraham Camelo-Guerrero, Jairo Diaz-Rodriguez
arXiv:2608. 10385v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about how assessor framing affects judgment reliability and downstream system comparison.
By Samaneh Mohtadi, Pietro Bernardelle, Joel Mackenzie, Gianluca Demartini
The paper introduces a dual‑dimensional framework called Automated Item Similarity Analysis (AISA) that uses Large Language Models to assess incidental content similarity in large‑scale assessments. It combines Structured Decomposition and Semantic Relatedness to capture both structural and semantic nuances that traditional metrics miss. Psychometric validation shows that LLM‑derived metrics better align with construct‑irrelevant local dependence and produce more coherent item groupings, and simulations in Computerized Adaptive Testing demonstrate improved estimation stability and reduced bias with minimal efficiency loss.
By Jing Huang, Jihong Zhang, Hua-Hua Chang
arXiv:2606. 15887v1 Announce Type: cross Abstract: Large language model (LLM) systems are increasingly proposed to assist peer review, yet most evaluations judge the prose of machine-generated review text, not the validity of the numeric score a system assigns.
By Costa Georgantas
arXiv:2510. 02027v2 Announce Type: replace Abstract: Scholarly publishing requires scalable scrutiny supported by auditable evidence.
By Khalid M. Saqr
arXiv:2606. 27383v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as research assistants, yet it remains unclear whether they can calibrate research takeaways to the strength and scope of the supporting evidence.
By Yu Fu, Yongqi Kang, Yong Zhao
arXiv:2606. 30256v1 Announce Type: new Abstract: Safety benchmarks often buy scalability by fixing the prompt, the language, and the turn structure.
By Camilo Chac\'on Sartori
arXiv:2608. 04549v1 Announce Type: cross Abstract: Frontier LLMs are increasingly put to use on open-ended complex questions, different in nature from the ones they are typically evaluated on.
By Pau Arnal, Khaled Denfir, Danylo Smahliuk, Amrut Avhad, Marcus A. Castro