arXiv AI

RE-AD: Real-Time Requirement Adherence for Data Labeling

arXiv:2607. 20455v1 Announce Type: cross Abstract: Human-annotated data remains fundamental to training frontier Large Language Models (LLMs).

arXiv Computation and Language
Sep 23

Calibration as a First-Class Criterion in LLM Evaluation

The paper argues that calibration—how well a language model’s confidence aligns with its actual correctness—should be a standard evaluation metric for large language models (LLMs). It notes that while calibration metrics exist, they are rarely applied outside specialized NLP subfields, leading to unverified confidence scores in new models, datasets, and benchmarks. The authors highlight the risks of miscalibration both at deployment (overconfident errors causing harm) and in research workflows (affecting LLM-as-a-judge, synthetic data generation, and active learning). They call for every NLP subfield to pair its primary performance metric with a calibration score, treating calibration as an essential property of every model.

By Mario Sanz-Guerrero, Katharina von der Wense
arXiv AI
Sep 3

EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models

EvalDetectBench is an open pipeline and benchmark designed to measure evaluation awareness in frontier large language models, enabling practitioners to test models against any Inspect-compatible evaluation. It includes a curated transcript suite from current frontier system-card evaluations and diverse deployment sources, and it assesses both how reliably models recognize they are being evaluated and how detectable individual benchmarks are. The benchmark addresses systematic bias by calibrating probes per model and harmonizing generator selection to correct for variance caused by model identity and prompt choice.

By Xinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj, Alexandra Souly, Robert Kirk
arXiv AI
Sep 2

Towards a Reliable and Practical Eval Pipeline

The paper introduces an end-to-end evaluation pipeline for large language model (LLM) software systems that integrates checklist creation, learned aggregation of checklist responses, and additional features such as self‑consistency, explanations, and prediction uncertainty. This pipeline aims to enhance agreement among LLM judges and improve alignment with human judgments, addressing practical reliability concerns that previous work has only partially covered. Empirical results demonstrate the effectiveness of the proposed framework.

By Emma Thuong Nguyen, Abhishek Ghose
arXiv AI
Jun 3

Fixing FOLIO and MALLS: Verified Annotations and an LLM-assisted Framework to Focus Human Relabeling

arXiv:2606. 02837v1 Announce Type: cross Abstract: Accurate translation from Natural Language to First-Order Logic (NL-to-FOL) underpins neurosymbolic AI systems and Natural Language Inference (NLI), making the quality of NL-to-FOL benchmarks essential -- yet these datasets have never been rigorously audited.

By Andrea Brunello, Cristian Curaba, Luca Geatti, Michele Mignani, Angelo Montanari, Nicola Saccomanno
arXiv Computation and Language
Aug 31

Select, Label, Evaluate: Active Testing in NLP

The paper introduces Active Testing, a framework that selects the most informative test samples for annotation in NLP, aiming to reduce human effort while accurately estimating model performance. Experiments across 18 datasets and 4 embedding strategies show up to 95% annotation savings with less than 1% loss in performance estimation accuracy. The authors also propose an adaptive stopping criterion to determine the optimal number of samples without a predefined budget.

By Antonio Purificato, Maria Sofia Bucarelli, Andrea Bacciu, Fabrizio Silvestri, Amin Mantrach
arXiv Machine Learning
Jun 5

Domain-Adapted Small Language Models with Hybrid Post-Processing: Achieving Cost-Efficient, Low-Latency Multi-Label Structured Prediction via LoRA Fine-Tuning on Scarce Data

arXiv:2606. 05781v1 Announce Type: new Abstract: Deploying frontier large language models (LLMs) for domain-specific structured evaluation tasks often incurs substantial latency, cost, and data privacy overhead.

By Srinivasan Manoharan, Dilipkumar Nallusamy, Sachin Kumar, Haifeng Wu