arXiv Machine Learning By Shengwei Xu, Yuxuan Lu, Yifan Wu, Jason Hartline, Grant Schoenebeck

Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics

Read the original on arXiv Machine Learning →

arXiv:2608. 01423v1 Announce Type: cross Abstract: Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
4d ago

PADM\'E: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators

PADM'E is a method for synthesizing preference‑aligned data to meta‑evaluate language‑model (LM) evaluators of agentic behaviors. It reframes meta‑evaluation as a preference judgment problem, generating criterion‑based data with small LMs and no human involvement. In a prototype, PADM'E produced 1,000 samples across four domains and three criteria, and human validation showed agreement with human judgment rising from 73% to 85% compared to a naive baseline.

By Cheng Chang, Yining Mao, Peng Qi
arXiv AI
Sep 18

Form Over Content In Gradient-Based Data Attribution Methods

The paper investigates what gradient similarity measures in data attribution for large language models. By independently varying task and answer format in supervised fine‑tuning benchmarks, the authors show that gradient alignment is driven by answer format rather than task semantics, with strong alignment for shared formats and none for differing formats. This pattern persists across training stages, model sizes, and families, and is evident in the selections of the LESS data‑selection method, which over‑represents its own answer format.

By Sunwoo Kim, Seokwon Jung, Sohyung Kim, Seong Joon Oh, Alice Oh