arXiv Computation and Language

KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human Difficulty

KCSAT-ML is a new benchmark built from 664 Korean College Scholastic Ability Test mathematics problems, including a 339‑item core set with official per‑item error rates from nationwide cohorts of hundreds of thousands of examinees. The benchmark introduces the Difficulty‑aligned Reasoning Gain (DRG) metric, which evaluates whether a model’s mistakes align with items humans find hard or easy, revealing distinct patterns in how vision‑language and large language models perform across difficulty levels. The dataset and code are publicly available at https://github.com/naver-ai/KCSAT-ML.

arXiv AI
Aug 19

What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations

The paper investigates how aggregate benchmark scores can obscure item-level changes when commercial large language model APIs are upgraded. By querying 900 benchmark items across three GPT-5.4 to GPT-5.6 upgrades, the authors classify each item as reliably improved, reliably regressed, practically equivalent, or inconclusive, revealing that both improvements and regressions coexist within the same upgrade. The study shows that even large aggregate gains can hide up to 8.3% of reliably regressed items, and that strict versus loose scoring can dramatically alter perceived performance changes.

By Xiaonan Xu, Wenjing Wu
arXiv AI
Aug 7

Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning

arXiv:2608. 05643v1 Announce Type: new Abstract: Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing answer patterns instead of adding useful reasoning diversity.

By Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Lena Trigg, Ali Subhan, Muhammad Ali, Dean F. Hougen
arXiv AI
Aug 25

GIM: Evaluating models via tasks that integrate multiple cognitive domains

The paper introduces the Grounded Integration Measure (GIM), a benchmark of 820 expert‑authored problems designed to test models on tasks that integrate multiple cognitive operations such as constraint satisfaction, state tracking, epistemic vigilance, and audience calibration. GIM emphasizes realistic, broadly accessible knowledge rather than specialized expertise, and uses a judge‑aware 2‑parameter logistic IRT model to produce robust ability estimates across 53 model‑thinking‑level configurations. The authors provide a comprehensive leaderboard of 22 models and 47 test configurations, and conduct an extensive study on how test‑time compute affects model capability, finding that configuration choices like thinking budget and quantization can be as influential as model selection itself. whyItMatters":"By focusing on integration of multiple cognitive domains, GIM offers a more realistic assessment of model reasoning capabilities than benchmarks that either overemphasize memorization or abstract reasoning alone."

By Rohit Patel, Alexandre Rezende, Steven McClain
arXiv AI
Jul 14

LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning

arXiv:2607. 10139v1 Announce Type: cross Abstract: Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution.

By Ning Liu
arXiv Machine Learning
Jul 7

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs

arXiv:2602. 20629v3 Announce Type: replace Abstract: As Large Language Models (LLMs) saturate elementary benchmarks, the research frontier has shifted from generation to the reliability of automated evaluation.

By Santiago Gonzalez, Alireza Amiri Bavandpour, Peter Ye, Edward Zhang, Ruslans Aleksejevs, Todor Anti\'c, Polina Baron, Sujeet Bhalerao, Shubhrajit Bhattacharya, Zachary Burton, John Byrne, Hyungjun Choi, Nujhat Ahmed Disha, Koppany Istv\'an Encz, Yuchen Fang, Robert Joseph George, Ebrahim Ghorbani, Alan Goldfarb, Jing Guo, Meghal Gupta, Stefano Huber, Annika Kanckos, Minjung Kang, Hyun Jong Kim, Dino Lorenzini, Levi Lorenzo, Tianyi Mao, Giovanni Marzenta, Ariane M. Masuda, Lukas Mauth, Ana Mickovic, Andres Miniguano-Trujillo, Antoine Moulin, Wenqi Ni, Tomos Parry, Kevin Ren, Hossein Roodbarani, Mathieu Rundstr\"om, Manjil Saikia, Detchat Samart, Rebecca Steiner, Connor Stewart, Dhara Thakkar, Jeffrey Tse, Vasiliki Velona, Yunhai Xiang, Sibel Yal\c{c}{\i}n, Jun Yan, Ji Zeng, Arman Cohan, Quanquan C. Liu
arXiv AI
Aug 17

The Metacognitive Bottleneck: Japanese Riddles Reveal Fundamental Limits of Machine Insight and Self-Evaluation in Reasoning AI

arXiv:2509. 14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems.

By Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Xingfu Li, Yo Nakawake, Le Minh Nguyen
arXiv AI
Sep 2

Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random

The paper investigates item-sensitivity—whether a language model’s choice depends on the specific input—in a forced-choice signalling task derived from the board game Deception: Murder in Hong Kong. Across seven models, two families, a post‑training ablation, and three scoring rules, every tested cell shows item‑sensitivity, yet many are statistically indistinguishable from random choice and some perform worse than random. The authors term this phenomenon "consistency without alignment" and argue it undermines evaluations that rely solely on item‑sensitivity, permutation consistency, or self‑consistency without an independent reference.

By Cris Huynh