arXiv AI

Evaluating Nonuniform Dependability Across Response Conditions: A Conditional Generalizability Framework Illustrated in Automated Essay Scoring

arXiv:2607. 11981v1 Announce Type: cross Abstract: Aggregate reliability estimates can obscure heterogeneity in measurement-design burden across response conditions, so a single G- or D-study may mischaracterize a design's adequacy for particular strata.

arXiv Machine Learning
Sep 18

The Complexity Kink: A Prompt-Side Structural Complexity Index for Code-Generation Reliability

The paper introduces a six‑dimension prompt‑side structural‑complexity index to assess code‑generation reliability before a model generates output. Using 5,000 Python prompts and 21 large language models, the authors find that pass rates exhibit a non‑monotonic breakpoint around a composite score of 13.75, with task‑type and construction‑frame adjustments shifting this threshold. The study also reports high inter‑rater reliability (ICC = 0.872) and demonstrates that the index can predict failure likelihood without relying on output correctness.

By Michael Hernandez, Tian Zhao
Hugging Face Trending Papers
Jul 8

From Text to Parameters: Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings

Newly developed items must ordinarily be field tested before their psychometric properties are known, creating a cold start problem for item calibration. Predicting item parameters from features is a long standing measurement problem dating back to the Linear Logistic Test Model; modern text embeddings now automate the design matrices traditionally specified by hand.

arXiv Machine Learning
Sep 25

Three Ways Classical Test Theory Misleads for LLM Judges

The article examines how classical test theory reliability statistics misrepresent the performance of large language model (LLM) judges. It shows that internal‑consistency coefficients, the dependability index, and Livingston‑Lewis accuracy each conflate judge error with item design or criterion validity, making it impossible to attribute a single reliability value to the judge alone. The authors argue that such misattribution can influence deployment decisions and documentation.

By Louis Yiven Zhu
Hugging Face Trending Papers
Aug 11

Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations

Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent-trace benchmarks (TheAgentCompany, $τ^2$-bench, and AppWorld), that the agent main effect accounts for less than 3% of total variance in every dataset and check type, while the agent-by-task interaction accounts for 7-23%.