Hugging Face Trending Papers

Evaluating Nonuniform Dependability Across Response Conditions: A Conditional Generalizability Framework Illustrated in Automated Essay Scoring

Read the original on Hugging Face Trending Papers →

Aggregate reliability estimates can obscure heterogeneity in measurement-design burden across response conditions, so a single G- or D-study may mischaracterize a design's adequacy for particular strata. This study introduces a conditional generalizability framework with three components.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
Sep 18

The Complexity Kink: A Prompt-Side Structural Complexity Index for Code-Generation Reliability

The paper introduces a six‑dimension prompt‑side structural‑complexity index to assess code‑generation reliability before a model generates output. Using 5,000 Python prompts and 21 large language models, the authors find that pass rates exhibit a non‑monotonic breakpoint around a composite score of 13.75, with task‑type and construction‑frame adjustments shifting this threshold. The study also reports high inter‑rater reliability (ICC = 0.872) and demonstrates that the index can predict failure likelihood without relying on output correctness.

By Michael Hernandez, Tian Zhao