BENCHCOMPASS is a new payment‑domain benchmark that transforms typed evidence packs into scenario‑grounded tasks, applies LLM‑based quality checks, generates attack variants, and reserves final item admission for domain experts. It includes an expert‑reviewed Pro benchmark covering payment knowledge, context‑grounded scenario reasoning, and attacked open robustness, plus a lower‑assurance Normal pool. Across 16 model variants, BENCHCOMPASS reveals distinct failure modes—missing payment knowledge, incomplete reasoning, and failure to reject invalid workflows—while the best model scores 89.6% on open context‑grounded reasoning and 81.7% under attacked inputs.
"whyItMatters":"The benchmark provides a structured way to isolate and evaluate specific weaknesses in LLMs for payment operations, a critical financial infrastructure where rules change rapidly and decisions depend on complex contextual factors."
By Sijie Dong, Wei Ren, Xuanwei Hu, Jiawei Luo, Zifan Wang, Xiaoyun Feng, Hui Cai, Lyuxin Xue, Peng Lu, Jianshe Li, Xin Zhang, Wei Wu
arXiv:2607. 23123v1 Announce Type: new Abstract: Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable deliverable produced within a constrained workflow as the unit of evaluation.
By Summer Sun (Shaqiu Community)
arXiv:2607. 06411v1 Announce Type: cross Abstract: Developers increasingly delegate real maintenance work to product-grade coding agents, and many state tasks in their native language, in the style of a customer request rather than a curated English issue.
By Evgeny Shilov (Independent Researcher)
arXiv:2608. 09072v1 Announce Type: cross Abstract: Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or fixing bugs.
By Xin Zhou, Chun Yong Chong, Kisub Kim, Yun Peng, Rui Shu, Zihan Wu, Xu Han, Guowen Yuan, Zeyang Zhuang, Jounghoon Kim, Jeongjin Ju, Seongmin Ju, Taein Yoon, David Lo
Developers increasingly delegate real maintenance work to product-grade coding agents, and many state tasks in their native language, in the style of a customer request rather than a curated English issue. Existing repository-level agentic benchmarks do not measure this setting: their task statements are English by design.
arXiv:2609.06059v1 Announce Type: new
Abstract: As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to a...
By Yu Liu, Zhilin Liu, Zhiwei Yang, Shaojie Zhang, Zheyuan Deng, Tingwei Huang, Zhenbo Luo, Lei Jiang, Yanbing Liu, Pei Fu