The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure
Read the original on arXiv Machine Learning →The paper critiques a recent NLI benchmark that tests the imperfective paradox, arguing that the benchmark suffers from conceptual and evaluation mis-specifications, notably Aspectual Reduction and a lack of strict NLI standards. The authors re-evaluate the benchmark, identify mis-specifications, and construct lexically matched minimal pairs to control for lexical variation. Their experiments reveal that models often exhibit a Sufficiency Bias, accept simple‑past hypotheses without affirming culmination, and that prompting interventions shift label decisions without improving true semantic understanding, highlighting additional failure modes such as compositional aspectual classification errors and surface‑form attraction.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.