Hugging Face Trending Papers

Evaluating Nonuniform Dependability Across Response Conditions: A Conditional Generalizability Framework Illustrated in Automated Essay Scoring

Aggregate reliability estimates can obscure heterogeneity in measurement-design burden across response conditions, so a single G- or D-study may mischaracterize a design's adequacy for particular strata. This study introduces a conditional generalizability framework with three components.

arXiv Machine Learning
Sep 18

The Complexity Kink: A Prompt-Side Structural Complexity Index for Code-Generation Reliability

The paper introduces a six‑dimension prompt‑side structural‑complexity index to assess code‑generation reliability before a model generates output. Using 5,000 Python prompts and 21 large language models, the authors find that pass rates exhibit a non‑monotonic breakpoint around a composite score of 13.75, with task‑type and construction‑frame adjustments shifting this threshold. The study also reports high inter‑rater reliability (ICC = 0.872) and demonstrates that the index can predict failure likelihood without relying on output correctness.

By Michael Hernandez, Tian Zhao
Hugging Face Trending Papers
Jul 8

From Text to Parameters: Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings

Newly developed items must ordinarily be field tested before their psychometric properties are known, creating a cold start problem for item calibration. Predicting item parameters from features is a long standing measurement problem dating back to the Linear Logistic Test Model; modern text embeddings now automate the design matrices traditionally specified by hand.

arXiv AI
Jul 7

When Rubrics Fail: Error Enumeration as Reward in Reference-Free RL Post-Training for Virtual Try-On

arXiv:2603. 05659v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) and Rubrics as Rewards (RaR) have driven strong gains in domains with clear correctness signals and even in subjective domains by synthesizing evaluation criteria from ideal reference answers.

By Wisdom Ikezogwo, Mehmet Saygin Seyfioglu, Ranjay Krishna, Karim Bouyarmane
Hugging Face Trending Papers
Aug 11

Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations

Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent-trace benchmarks (TheAgentCompany, $τ^2$-bench, and AppWorld), that the agent main effect accounts for less than 3% of total variance in every dataset and check type, while the agent-by-task interaction accounts for 7-23%.