arXiv:2607. 14109v1 Announce Type: cross Abstract: Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain central challenges in natural language understanding.
By Inder Preet, Shuxin Lin, Dhaval Patel
The paper introduces Evaluation-as-Search (EaS), a feedback‑driven method that adaptively probes LLM‑powered meeting assistants by focusing on natural questions likely to reveal grounding failures. Using EaS, the authors build MeetingProbe, a benchmark of over 3,000 annotated question‑answer pairs from 20 transcripts across three meeting genres and three assistants. Ablation studies show that adaptive search uncovers 2.5× more failures than random probing, revealing a capability gradient and eight recurring failure categories dominated by discourse‑pragmatic challenges.
By Sami Khairy, Yasaman Hosseinkashi, Vishak Gopal, Ross Cutler
arXiv:2607.14109v2 Announce Type: replace
Abstract: Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain central...
By Inder Preet, Shuxin Lin, Dhaval Patel
arXiv:2509. 14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems.
By Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Xingfu Li, Yo Nakawake, Le Minh Nguyen
GRACE is a step‑level benchmark for evaluating the faithfulness of chain‑of‑thought reasoning over context. It provides human annotations for each step in CoT traces from 10 models across 4 datasets, labeling faithfulness, error category, and natural‑language explanations. The benchmark introduces a data‑driven taxonomy that splits errors into GRACE‑Inference (deductive) and GRACE‑Grounding (factual) tracks, each with four categories, and demonstrates that incorporating step‑level faithfulness signals can improve downstream accuracy and reasoning reliability.
By Hoang Pham, Dong Le, Anh Tuan Luu
arXiv:2603. 03824v2 Announce Type: replace Abstract: Humans often become more self-aware under threat, yet can lose self-awareness when absorbed in a task; we hypothesize that language models exhibit environment-dependent \textit{evaluation awareness}.
By Maheep Chaudhary