arXiv AI By Shiyu Ying, Xuejie Cao, Yingfan Ma, Yuanhao Dong, Wenyu Chen, Bowen Song, Lin Zhu

Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents

Read the original on arXiv AI →

arXiv:2607. 14573v1 Announce Type: new Abstract: Payment integration is a demanding repository-level software task: agents must select a suitable product, implement coordinated client-server flows, verify payment outcomes, and preserve consistency between transaction and business states.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 17

BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs

BENCHCOMPASS is a new payment‑domain benchmark that transforms typed evidence packs into scenario‑grounded tasks, applies LLM‑based quality checks, generates attack variants, and reserves final item admission for domain experts. It includes an expert‑reviewed Pro benchmark covering payment knowledge, context‑grounded scenario reasoning, and attacked open robustness, plus a lower‑assurance Normal pool. Across 16 model variants, BENCHCOMPASS reveals distinct failure modes—missing payment knowledge, incomplete reasoning, and failure to reject invalid workflows—while the best model scores 89.6% on open context‑grounded reasoning and 81.7% under attacked inputs. "whyItMatters":"The benchmark provides a structured way to isolate and evaluate specific weaknesses in LLMs for payment operations, a critical financial infrastructure where rules change rapidly and decisions depend on complex contextual factors."

By Sijie Dong, Wei Ren, Xuanwei Hu, Jiawei Luo, Zifan Wang, Xiaoyun Feng, Hui Cai, Lyuxin Xue, Peng Lu, Jianshe Li, Xin Zhang, Wei Wu
arXiv AI
Aug 11

A Unified Issue Resolution Benchmark for Requirement Clarification, Planning, and Code Generation for Coding Agents

arXiv:2608. 09072v1 Announce Type: cross Abstract: Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or fixing bugs.

By Xin Zhou, Chun Yong Chong, Kisub Kim, Yun Peng, Rui Shu, Zihan Wu, Xu Han, Guowen Yuan, Zeyang Zhuang, Jounghoon Kim, Jeongjin Ju, Seongmin Ju, Taein Yoon, David Lo
Hugging Face Trending Papers
Jul 7

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications

Developers increasingly delegate real maintenance work to product-grade coding agents, and many state tasks in their native language, in the style of a customer request rather than a curated English issue. Existing repository-level agentic benchmarks do not measure this setting: their task statements are English by design.