arXiv Computation and Language By Sanghee Park, Geewook Kim, Kee-Eung Kim

KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human Difficulty

Read the original on arXiv Computation and Language →

KCSAT-ML is a new benchmark built from 664 Korean College Scholastic Ability Test mathematics problems, including a 339‑item core set with official per‑item error rates from nationwide cohorts of hundreds of thousands of examinees. The benchmark introduces the Difficulty‑aligned Reasoning Gain (DRG) metric, which evaluates whether a model’s mistakes align with items humans find hard or easy, revealing distinct patterns in how vision‑language and large language models perform across difficulty levels. The dataset and code are publicly available at https://github.com/naver-ai/KCSAT-ML.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Aug 19

What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations

The paper investigates how aggregate benchmark scores can obscure item-level changes when commercial large language model APIs are upgraded. By querying 900 benchmark items across three GPT-5.4 to GPT-5.6 upgrades, the authors classify each item as reliably improved, reliably regressed, practically equivalent, or inconclusive, revealing that both improvements and regressions coexist within the same upgrade. The study shows that even large aggregate gains can hide up to 8.3% of reliably regressed items, and that strict versus loose scoring can dramatically alter perceived performance changes.

By Xiaonan Xu, Wenjing Wu
arXiv AI
Aug 7

Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning

arXiv:2608. 05643v1 Announce Type: new Abstract: Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing answer patterns instead of adding useful reasoning diversity.

By Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Lena Trigg, Ali Subhan, Muhammad Ali, Dean F. Hougen
arXiv AI
Aug 25

GIM: Evaluating models via tasks that integrate multiple cognitive domains

The paper introduces the Grounded Integration Measure (GIM), a benchmark of 820 expert‑authored problems designed to test models on tasks that integrate multiple cognitive operations such as constraint satisfaction, state tracking, epistemic vigilance, and audience calibration. GIM emphasizes realistic, broadly accessible knowledge rather than specialized expertise, and uses a judge‑aware 2‑parameter logistic IRT model to produce robust ability estimates across 53 model‑thinking‑level configurations. The authors provide a comprehensive leaderboard of 22 models and 47 test configurations, and conduct an extensive study on how test‑time compute affects model capability, finding that configuration choices like thinking budget and quantization can be as influential as model selection itself. whyItMatters":"By focusing on integration of multiple cognitive domains, GIM offers a more realistic assessment of model reasoning capabilities than benchmarks that either overemphasize memorization or abstract reasoning alone."

By Rohit Patel, Alexandre Rezende, Steven McClain