arXiv AI

InvestPhilBench: A Multi-Layer Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy

arXiv:2606. 25984v2 Announce Type: replace Abstract: Large language models are increasingly deployed as investment research assistants, yet no benchmark tests whether they can accurately reconstruct and apply the specific procedural decision frameworks of expert investors.

arXiv AI
Aug 25

GIM: Evaluating models via tasks that integrate multiple cognitive domains

The paper introduces the Grounded Integration Measure (GIM), a benchmark of 820 expert‑authored problems designed to test models on tasks that integrate multiple cognitive operations such as constraint satisfaction, state tracking, epistemic vigilance, and audience calibration. GIM emphasizes realistic, broadly accessible knowledge rather than specialized expertise, and uses a judge‑aware 2‑parameter logistic IRT model to produce robust ability estimates across 53 model‑thinking‑level configurations. The authors provide a comprehensive leaderboard of 22 models and 47 test configurations, and conduct an extensive study on how test‑time compute affects model capability, finding that configuration choices like thinking budget and quantization can be as influential as model selection itself. whyItMatters":"By focusing on integration of multiple cognitive domains, GIM offers a more realistic assessment of model reasoning capabilities than benchmarks that either overemphasize memorization or abstract reasoning alone."

By Rohit Patel, Alexandre Rezende, Steven McClain
arXiv AI
Aug 11

WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management

arXiv:2608. 07529v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains difficult to assess because existing benchmarks emphasize general knowledge rather than professional decisions under engineering, environmental, and policy constraints.

By Yi Zhang, Hongyang Wang, Zheng Hao Leong, Zihao Wu, Kaijun Lin, Zhixing Pan, Qixun Huangfu, Wei Ren, Wenyan Wu, Fangyun Wang, Wenting Yu, Hengyu Lin, Muling Yang, Zongguo Wen
arXiv Computation and Language
Aug 31

FinExam-10K: When Retrieval Helps Financial Reasoning?

FinExam-10K is a new English benchmark for financial reasoning, comprising 10,198 expert‑reannotated questions covering CFA Levels I‑III and FRM Parts I‑II. The dataset is split into a 5,110‑question release and a 5,088‑question held‑out set for a quarterly leaderboard, with separate Full‑Coverage and Context‑Complete Reasoning tracks. Across 17 models, the best overall accuracy is 85.29 %, but performance drops on harder subsets, and retrieval‑augmented methods like Function‑Graph‑RAG provide modest gains when gated appropriately.

By Yan Lin, Jingyu Sun, Zhongliang Guo, Qing Li, Zhuohan Xie, Yuxia Wang
arXiv AI
Aug 17

The Metacognitive Bottleneck: Japanese Riddles Reveal Fundamental Limits of Machine Insight and Self-Evaluation in Reasoning AI

arXiv:2509. 14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems.

By Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Xingfu Li, Yo Nakawake, Le Minh Nguyen